Method and system for modeling the reverberation of speech signals
The convolutional prediction approach using DNNs to estimate RIRs effectively addresses speech reverberation challenges, enhancing speech quality and recognition by accurately separating direct path signals in noisy and multi-speaker environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MITSUBISHI ELECTRIC CORP
- Filing Date
- 2023-06-02
- Publication Date
- 2026-04-20
AI Technical Summary
Existing methods struggle to effectively mitigate speech reverberation in enclosed environments, particularly in the presence of noise and multiple speakers, leading to degraded speech quality and inaccurate automatic speech recognition, due to challenges in identifying and separating the direct path signal from attenuated and delayed copies.
A convolutional prediction approach using deep neural networks (DNNs) to estimate room impulse responses (RIRs) and apply filters that model reverberation, allowing for both early and late reverberation removal, and incorporating amplitude and phase information for improved accuracy.
The method achieves effective reverberation reduction, enhancing speech quality and improving automatic speech recognition performance by accurately separating direct path signals from reverberant mixtures, even in noisy and multi-speaker scenarios.
Smart Images

Figure 0007848410000026 
Figure 0007848410000027 
Figure 0007848410000028
Abstract
Description
[Technical Field]
[0001] This disclosure relates to audio signal processing in general, and more specifically to methods and systems for modeling the reverberation of speech signals. [Background technology]
[0002] In general, in enclosed rooms, reverberation occurs in speech signals (e.g., speech) during modern hands-free speech communication, such as remote conferencing and interactions with smart devices like smart speaker microphones. In such enclosed rooms, speech signals propagate through the air and may be reflected by walls, floors, ceilings, and other objects in the room before being picked up by a microphone. Reverberation is the multipath propagation of a speech signal from the source or speaker to the receiving end, such as a microphone. Such speech reverberation occurs when sound reflects off surfaces in the environment. Some of the sound may be absorbed by those surfaces, resulting in multiple attenuation of the speech signal. The reflection and absorption of sound by those surfaces can produce multiple attenuated and delayed copies of the speech signal. These multiple attenuated and delayed copies can degrade the quality of speech and interfere with the performance of automatic speech recognition (ASR) systems or any speech / voice processing systems. For example, an ASR may produce inaccurate output due to a degraded speech input.
[0003] Speech reverberation can be mitigated by removing the reverberation effect from the sound. Such removal of reverberation effects is known as reverberation rejection. Reverberation rejection may involve identifying the direct path signal and distinguishing it from attenuated and delayed copies. The direct path signal corresponds to the signal that the sound follows when the source and microphone are in line of sight. However, identifying the direct path signal and distinguishing it from copies can be difficult, especially when there is significant reverberation and noise from non-stationary sources. For example, environments such as enclosed rooms with non-stationary sources such as air conditioning systems can have significant room reverberation. Reducing reverberation can be challenging due to noise from air conditioning systems or any multi-source environmental noise. Multi-source environmental noise can also be present in scenarios where multiple people are speaking in that environment.
[0004] Therefore, modeling reverberation in such enclosed environments is advantageous for use in multiple applications, including speech reverberation rejection, room impulse response modeling, and acoustic modeling.
[0005] Therefore, the above problems need to be overcome. More specifically, methods and systems for modeling the reverberation of speech signals need to be developed while overcoming reverberation conditions and transient noise in reverberant environments. [Overview of the Initiative]
[0006] The objective of some embodiments is to develop methods and systems for modeling the reverberation of speech signals. Another objective of some embodiments is to perform deroguerization of speech signals using deep learning techniques. Deroguerization of speech signals can be extended to tasks such as reverberation reduction, speech enhancement, and speaker separation.
[0007] Some embodiments are based on the understanding that clean speech exhibits spectral-temporal patterns. Such spectral-temporal patterns are inherent patterns exhibited in the time-frequency domain and can provide useful clues for reverberation reduction. Some of these patterns originate from the structure of the speech signal itself, but Several patterns may also correspond to a linear filter structure of reverberation (i.e., reflection of sound waves) specific to the space in which recording takes place, including all objects, structures, or entities present in the physical space, as well as the positions of the source speech signal and receivers such as microphones that record the signal. This linear filter structure can be used to explain the signal emanating from the source signal at the microphone position and the reflection of that signal from the walls and surfaces of objects or people in the space, and the linear filter structure expresses the effect of reverberation on the input signal as a linear convolution of the input signal and a Room Impulse Response (RIR). The input signal is the original source signal, also known as the dry source signal. The room impulse response represents the influence of space and everything within that space on the input signal. For example, by playing a short-duration time-domain signal, such as an impulse sound (e.g., a blank shot or a balloon burst), at the source position in a room and recording the resulting signal at the receiver position, an estimate of the RIR between the source and receiver positions can be recorded in a physical space such as a room. The impulse excites the room, producing a reverberation impulse signal, which can be used to estimate the RIR. The reverberation of the dry source sound signal, which will then be played back at the same source position and recorded at the same receiver position, can be modeled by convolving the dry source signal and the estimated RIR. For that purpose, the objective of some embodiments is to estimate a basic filter for approximating or modeling the RIR. In some exemplary embodiments, the RIR can be estimated based on a linear regression problem solved frequency by frequency in the time-frequency domain. The filter estimate for modeling the RIR can be used to identify delayed and attenuated copies of the input signal for reverberation rejection of speech signals.
[0008] Furthermore, such linear filters can be utilized as regularization to improve the de-reverberation process. For example, a linear filter as a regularization prevents the de-reverberation process model from overfitting to the training data. Several embodiments are based on the understanding that linear filter structures can be utilized in combination with linear prediction and deep learning for single-channel and multi-channel reverberation speaker separation and de-reverberation tasks. For this purpose, deep learning techniques supported by convolutional prediction can be used for de-reverberation in environments with noise signals, speech signal reverberation, etc. Convolutional prediction is a linear prediction method for speech de-reverberation in reverberant situations, and is a source estimation obtained by a Deep Neural Network (DNN). It relies on the value and utilizes a linear filter structure between the source estimate and the reverberation version of the source signal in the observed input signal.
[0009] To obtain source estimation, the DNN is trained in the time-frequency domain or time domain to predict the target speech from reverberant speech. The target speech corresponds to the target direct path signal between the source and the receiver, such as a microphone. This approach can leverage prior knowledge of speech patterns.
[0010] Previous studies have also attempted to utilize some form of linear filter structure to perform reverberation removal. For example, Weighted Prediction Error (WPE) is sometimes used for reverberation removal of speech signals. The WPE method computes an inverse linear filter based on variance-normalized delayed linear predictions. The computed linear filter is applied to past observations of a mixture input signal containing reverberation and possibly noise, and for reverberation removal, the late reverberation of the target source signal in the mixture input signal is estimated from past observations of reverberation. The estimated late reverberation is subtracted from the acoustic signal mixture received from various sources to estimate the target speech signal in the acoustic signal mixture. In some embodiments, the filter may also be estimated using the time-varying power spectral density (PSD) of the target speech signal. The PSD is the distribution of the signal's power across the frequency domain. Such linear filters can be iteratively estimated using WPE in an unsupervised manner. However, the iterative procedure of WPE for filter estimation can lead to suboptimal results and can be computationally expensive.
[0011] To overcome the shortcomings of WPE described above, the iterative procedure for filter estimation can be replaced with a DNN-based WPE (DNN-WPE) approach. DNN-WPE uses the amplitude estimated by the DNN as the PSD of the target speech signal for filter estimation. However, DNN-WPE cannot reduce early reflections because it requires a strict non-zero frame delay to avoid trivial solutions and lacks a mechanism to utilize the phase estimated by the DNN for filter estimation. Furthermore, DNN-WPE may lack robustness to interference caused by noise signals. For example, DNN-WPE may estimate a filter that associates past observations containing noise with current observations containing noise, thereby limiting the accuracy of the filter estimation. Additionally, DNN-WPE directly uses linear prediction results as its output, which may result in partial or minimal reverberation reduction.
[0012] To that end, another objective of some embodiments is to remove both early and late reverberations for reverberation removal. Early and late reverberations can be removed using a convolutional prediction approach. The convolutional prediction approach utilizes both amplitude and phase estimated by the DNN for filter estimation. Furthermore, the convolutional prediction approach (similar to the DNN-WPE approach described above) provides closed-form solutions for linear filters, and these closed-form solutions are suitable for online real-time processing applications and can be trained in conjunction with other DNN modules, such as acoustic models.
[0013] In some embodiments, a filter is determined using a convolutional prediction approach with respect to a first estimate of the target direct path signal, such that the application of the filter to the target direct path estimate is as close as possible under a weighted distance function to the residuals between the acoustic signal mixture and the first estimate of the target direct path signal. The filter models the reverberation of the target direct path signal in the mixture. Alternatively, functionally equivalent, the filter can be such that the application of the filter to the target direct path estimate is as close as possible to the acoustic signal mixture under a weighted distance function, in which case the filter models the reverberation including the direct path contribution that gives rise to the target direct path signal. These filters are considered equivalent since each filter can be readily obtained from the others. Furthermore, the filter is applied to the first estimate of the target direct path signal in the time-frequency domain. When the filter is applied to the first estimate of the target direct path signal, the result is to identify delayed and attenuated copies of the estimated direct path signal from the acoustic signal mixture. These delayed and attenuated copies are, as herein it is, derived signals of the target direct path signal reflected in multiple paths due to reverberation. For example, a target direct path signal is reflected in various directions by various objects in the environment, such as a room. Such identified delayed and attenuated copies can be removed from the acoustic signal mixture for reverberation removal. Removal of delayed and attenuated copies produces a mixture with reduced reverberation.
[0014] The result obtained when the filter is applied to a first estimate of the target direct-path signal is, by the above configuration, closest to the residual between the acoustic signal mixture and the first estimate of the target direct-path signal according to a distance function. The distance function is a weighted distance between the filtered direct-path signal and the residual obtained by subtracting the direct-path estimate from the mixture, where the weight at each time-frequency point in the time-frequency domain is determined by one or a combination of the received acoustic signal mixture and the first estimate of the target direct-path signal. In some embodiments, the distance function is based on the least-squares distance. Furthermore, the result of applying the filter to the first estimate of the target direct-path signal is removed from the acoustic signal mixture to obtain a mixture with reduced reverberation of the target direct-path signal.
[0015] In some embodiments, this reverberation-reduced mixture is input to a second DNN. The second DNN outputs a second estimate of the target direct path signal, which may be an improved estimate of the target direct path signal compared to the previous estimate of the target direct path signal. The second DNN can also perform steps similar to those of the first DNN; however, in some embodiments, the second DNN may take a set of different signals as input, such as one or a combination of the acoustic signal mixture, the reverberation-reduced mixture, and the direct path signal estimate. By similarly using the second estimate of the target direct path signal, an improved filter can be obtained using a convolutional prediction approach, such that the application of the improved filter to the improved estimate of the target direct path signal is as close as possible to the residual between the acoustic signal mixture and the improved estimate of the target direct path signal under some weighted distance function. Alternatively, and functionally equivalent, the improved filter can be applied to an improved estimate of the target direct-path signal so that it is as close as possible under some weighted distance function to the acoustic signal mixture, in which case the improved filter models the reverberation, including the direct-path contribution that gives rise to the target direct-path signal.
[0016] In some embodiments, the downstream DNN uses one or more of the filtered estimates output by the first estimate or the improved filtered estimates output by the second DNN to perform tasks that benefit from RIR estimation. For example, the downstream DNN is used to perform localization functions.
[0017] Furthermore, some embodiments are based on the understanding that individual speakers of multiple speakers, or each speaker, are convolved with a different RIR. Acquiring each RIR individually may also be desirable for downstream applications or to perform data augmentation. The WPE method estimates a single filter to reduce reverberation from all sources. However, calculating a single filter for dereverberation of a mixture may not be feasible if noise or competing speakers are louder than the target source. Such a calculated filter would be biased towards suppressing reverberation from higher-energy sources. Therefore, a dereverberation filter needs to be estimated for each source, because each source is convolved with a different RIR. While the DNN-WPE method can calculate different filters for each source, it can only do so by using the estimated PSD of each source as weights in the distance function that DNN-WPE uses to estimate the linear predictive filter, which may limit the accuracy and variety of these different filters.
[0018] In some embodiments, two DNNs, a first DNN and a second DNN, are trained on a convolutional prediction approach for de-reverberation of speech signals. First, the first DNN of the two DNNs outputs a first estimate of the direct path signal of the target source (hereinafter referred to as the speaker, the person speaking) from the input, i.e., an acoustic signal mixture including the speaker's utterance. The direct path signal of the target source is hereinafter referred to as the target direct path signal. The first estimate of the target direct path signal is used to determine a filter using the convolutional prediction approach. The filter is such that, under some weighted distance function, the application of the filter to the target direct path estimate is as close as possible to the residual obtained by subtracting the target direct path estimate from the mixture. Furthermore, the filter is applied to the first estimate of the target direct path signal in the time-frequency domain. When the filter is applied to the first estimate of the target direct path signal, the result is to identify delayed and attenuated copies of the estimated target direct path signal from the acoustic signal mixture. These delayed and attenuated copies are, in this specification, derived signals of the target direct path signal that are reflected in multiple paths due to reverberation. For example, the target direct path signal is reflected in various directions by various objects in the environment, such as a room. Such identified delayed and attenuated copies are removed from the acoustic signal mixture for reverberation removal. Removal of delayed and attenuated copies produces a mixture with reduced reverberation.
[0019] When the filter is applied to the first estimate of the target direct path signal, the result obtained, by virtue of the above configuration, comes closest to the residual between the acoustic signal mixture and the first estimate of the target direct path signal according to the distance function. In some embodiments, the distance function is based on the least squares distance. Further, the result of applying the filter to the first estimate of the target direct path signal is removed from the acoustic signal mixture, yielding a mixture with reduced reverberation of the target direct path signal. In some embodiments, this mixture with reduced reverberation is input to the second of the two DNNs. The second DNN outputs a second estimate of the target direct path signal, which can be an improved estimate of the target direct path signal compared to the first estimate of the target direct path signal.
[0020] In some embodiments, the first DNN can be trained for the purpose of speaker separation. For that purpose, the first DNN generates a plurality of outputs corresponding to the first estimate of the target direct path signal for a particular speaker out of a plurality of speakers. Further, the estimation of the filter and the acquisition of the mixture with reduced reverberation are repeated for each of the plurality of speakers, generating a corresponding filter and a corresponding mixture with reduced reverberation for each of the plurality of speakers. Next, the corresponding mixtures with reduced reverberation for each of the plurality of speakers are combined, and the combined mixtures with reduced reverberation for each of the plurality of speakers are input to the second DNN. Next, the second DNN generates a second estimate of the target direct path signal for each of the plurality of speakers.
[0021] Additionally or alternatively, reverberation-reduced mixtures, i.e., delayed copies and attenuated copies, may be used as additional features for the second DNN to determine a second estimate of the target direct-path signal, thereby improving reverberation removal. Additionally or alternatively, features corresponding to delayed and attenuated copies may also be used for speaker separation tasks. In some exemplary embodiments, delayed and attenuated copies may be identified based on a linear regression problem. In some embodiments, one or a combination of the acoustic signal mixture and the first estimate of the target direct-path signal may be provided as input to the second DNN to generate a second estimate of the target direct-path signal. In some embodiments, the acoustic signal mixture, the first estimate, and the reverberation-reduced mixture are provided as input to the second DNN to determine a second estimate of the target direct-path signal.
[0022] Some embodiments are based on the understanding that, when there are multiple speakers in a room, corresponding filters are estimated for each individual speaker for reverberation removal. In the case of multiple speakers, the acoustic signal mixture includes speech signals from multiple speakers. In such a case, the first DNN generates a corresponding first estimate of the target direct path signal for each of the multiple speakers. To generate a reverberation-reduced mixture for each of the multiple speakers, the steps of determining a first estimate for each speaker, determining a filter for each speaker, and feeding in one or a combination of the first estimate and the reverberation-reduced mixture for each speaker can be combined and fed into a second DNN to generate a second estimate of the target direct path signal for each of the multiple speakers.
[0023] In some cases, the acoustic signal mixture may be received from a single channel such as a single microphone, or may be received from multiple channels such as an array of microphones. Each different channel measures a different version of the acoustic signal mixture. The DNN can be trained to estimate the target direct-path signal in a reference channel or in each channel. The training can be based on composite spectral mapping in one or more channels. The DNN is trained to output an estimate of the target direct-path signal in the time-frequency domain in one or more channels such that the distance between the estimate and a reference in the time-frequency domain of the target direct-path signal in one or more channels is minimized. In the case of an array of microphones, a beamforming output can be obtained. The beamforming output can be obtained based on statistics calculated from a first estimate of the target direct-path signal at each microphone of the microphone array and one or a combination of these of the mixture with reduced reverberation of the target direct-path signal. The beamforming output can be input into a second DNN to generate a second estimate of the target direct-path signal for each of a plurality of speakers. Additionally or alternatively, the beamforming output and the reverberation removal result may be used as additional features for the second DNN to perform better separation and reverberation removal tasks.
[0024] In some embodiments, a first DNN may be pre-trained to obtain a first estimate of a target direct path signal from an observed acoustic signal mixture. Pre-training of the first DNN may be performed using a training dataset of acoustic signal mixtures and corresponding reference target direct path signals within this training dataset. In particular, pre-training of the first DNN may be performed by minimizing a loss function. The loss function may include one or a combination of distance functions defined based on the real and imaginary (RI) components of the first estimate of the target direct path signal in the complex time-frequency domain and the RI component of the corresponding reference target direct path signal. Alternatively, the distance function may be defined based on the magnitude obtained from the RI component of the first estimate of the target direct path signal in the complex time-frequency domain and the corresponding magnitude of the reference target direct path signal.
[0025] Additionally or alternatively, the distance function may be defined based on a reconstructed waveform obtained from the RI component of the first estimate of the target direct path signal by reconstruction in the time domain, and the corresponding waveform of the reference target direct path signal.
[0026] In some alternative embodiments, the distance function may be defined based on the RI component of the first estimate in a second complex time-frequency domain, obtained by further transforming the reconstructed waveform in a second time-frequency domain, and the corresponding RI component of the reference target direct path signal in the second time-frequency domain.
[0027] In some alternative embodiments, the distance function may be defined based on the magnitude of the RI component of the first estimate in a second complex time-frequency domain, obtained by further transforming the reconstructed waveform in a second time-frequency domain, and the corresponding magnitude of the reference target direct path signal in the second time-frequency domain.
[0028] In some exemplary embodiments, an updated first estimate of the target direct path signal can be obtained by replacing a first estimate of the target direct path signal with a second estimate of the target direct path signal. An updated second estimate of the target direct signal can be obtained by repeating the steps of obtaining the first estimate, obtaining a filter, and introducing the first estimate and a mixture with reduced reverberation for the updated first estimate of the target direct signal.
[0029] In some examples, in multi-speaker scenarios, the above steps are repeated for each of the speakers to generate a corresponding filter for each speaker. Furthermore, a portion of the received acoustic signal mixture corresponding to a particular speaker can be extracted by removing the reverberant speech of other speakers from the acoustic signal mixture. An estimate of the reverberant speech of another speaker is obtained by adding the first estimate of the target direct path signal for the other speaker to the result of applying the corresponding filter for the other speaker to the first estimate of the target direct path signal for the other speaker. After extraction, a filter for estimating the reverberation-reduced mixture for each speaker of the multi-speaker can be estimated based on that portion of the received mixture.
[0030] Some embodiments provide evaluation results for speech derafter and speaker separation that demonstrate the effectiveness of derafter removal of speech signals based on a convolutional prediction approach.
[0031] Accordingly, one embodiment of the present disclosure discloses a method performed by a computer to estimate a room impulse response (RIR). The method is performed by a processor coupled with stored instructions for performing the method. When the instructions are executed by the processor, they perform steps of the method. The method comprises the step of receiving an acoustic signal mixture including a target direct path signal propagating in a room and the reverberation of the target direct path signal in the room. The acoustic signal mixture is received on a wired or wireless communication channel. The method further comprises the step of feeding the received acoustic signal mixture into a first deep neural network (DNN) to generate an estimate of the direct path signal. The method further comprises the step of estimating a filter that models a room impulse response (RIR) representing the relationship between the direct path signal and the reverberation of the direct path signal, the filter which, when applied to the estimate of the direct path signal, produces a result that is closest to one of the acoustic signal mixture and the residuals between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function. The filter is then transmitted. The filter is used for further applications such as performing speech processing, generating augmented training data, and separating multiple speakers.
[0032] In some embodiments, a first DNN is trained to estimate the amplitude and phase of a direct-path signal, and the application of a filter to the direct-path signal estimate involves calculations on both the estimated amplitude and phase. Furthermore, a distance function between the result of applying the filter to the direct-path signal estimate and the acoustic signal mixture is measured in both the amplitude and phase domains.
[0033] In some embodiments, the filter coefficients are estimated in the time-frequency domain.
[0034] In various embodiments, the direct path signal is the signal transmitted from the source of the original signal to the sensor. To this end, the direct path signal represents the original signal at the source as it would be measured by the sensor if there were no reverberation, and the reverberation of the direct path signal represents one or more transmissions of the original signal from the source to the sensor along one or more paths longer than the shortest path, such that the RIR is modeled with respect to the direct path signal.
[0035] In some embodiments, the filter modeling the RIR is a linear filter estimated using a linear convolutional prediction module. The linear convolutional prediction enhances the linear structure of the filter to approximate the linear reflection of the direct path signal from the surface in the room.
[0036] Some embodiments provide a method for synthesizing a new acoustic mixture by applying a filter that models the RIR to an external signal measured outside the room, the new acoustic mixture resulting in the acoustic effect of the external signal propagation within the room. The first DNN can be updated using the new acoustic mixture.
[0037] Some embodiments provide updating the direct path signal estimate based on the application of a filter that models the RIR. The training data is further updated based on the acoustic signal mixture and the updated direct path signal estimate as a pseudo-label. The updated training data is used to retrain the first DNN based on the updated training data.
[0038] Various embodiments provide the creation of augmented data based on a filter that models the RIR, and the creation of an augmented training dataset, which is obtained by applying the filter that models the RIR to data that includes external data measured outside the room and one or more of the data obtained as a first measurement of the target direct path signal of another acoustic signal mixture in the training data, and one or more of the second measurements of the target direct path signal. Furthermore, one or more of the first and second DNNs can be trained to output a first or second estimate of the direct path signal based on the augmented training dataset. In some embodiments, the augmented training dataset can also be used to train a system for automatic speech recognition or a system for sound event detection.
[0039] In some embodiments, a filter that models the RIR is used to perform speech analysis for at least one or a combination thereof of room acoustic parameter analysis, room geometry reconstruction, speech enhancement, and speech signal de-reverberation. Room acoustic parameters include at least one of direct-to-reverberation ratio, reverberation time (RT60), early decay time (EDT), center time, clarity C80, and resolution D50.
[0040] Accordingly, another embodiment of the present disclosure discloses a system for estimating RIR. The system comprises an input interface configured to receive an acoustic signal mixture, including a direct path signal propagating through a room and the reverberation of the direct path signal in the room, over a wired communication channel or a wireless communication channel. The system further comprises a memory for storing at least a first DNN. The system further comprises a processor configured to feed the received acoustic signal mixture into the first DNN to generate an estimate of the direct path signal. The processor further comprises a filter that models an RIR representing the relationship between the direct path signal and the reverberation of the direct path signal, the filter, when applied to the estimate of the direct path signal, produces a result that is closest to one of the acoustic signal mixture and the residuals between the acoustic signal mixture and a first estimate of the target direct path signal, according to a distance function. The system further comprises an output interface configured to output a filter that models an RIR over a communication channel.
[0041] Further features and advantages will become more readily apparent from the detailed description below, when considered in conjunction with the attached drawings.
[0042] The Disclosure will be further described below in detail with reference to several drawings, which are non-limiting examples of exemplary embodiments of the Disclosure. In the drawings, similar reference numerals indicate similar parts across several drawings. The drawings shown are not necessarily to scale and are, as a whole, focused on illustrating the principles of the embodiments disclosed herein. [Brief explanation of the drawing]
[0043] [Figure 1A] This figure shows an exemplary representation for modeling the reverberation of a speech signal according to an embodiment of the present disclosure. [Figure 1B] This figure shows an exemplary representation for modeling the reverberation of a speech signal, relating to another embodiment of the present disclosure. [Figure 2A]This is a schematic block diagram showing a system for modeling the reverberation of a speech signal according to an embodiment of the present disclosure. [Figure 2B] This is a schematic block diagram illustrating another system for modeling the reverberation of a speech signal according to an embodiment of the present disclosure. [Figure 2C] This is a schematic block diagram illustrating another system for modeling the reverberation of a speech signal according to an embodiment of the present disclosure. [Figure 3A] This is a schematic block diagram illustrating a process for modeling the reverberation of a speech signal according to an embodiment of the present disclosure. [Figure 3B] This figure shows a representation of the intra-intra-horizon impulse response (RIR) in the time domain according to an embodiment of the present disclosure. [Figure 3C] This figure shows a representation of the application of a filter that models the RIR in the frequency bin, according to an embodiment of the present disclosure. [Figure 4] This is a schematic diagram showing an architecture for modeling the reverberation of a speech signal according to an embodiment of the present disclosure. [Figure 5] This is a schematic diagram illustrating an architecture for modeling the reverberation of speech signals for multiple speakers, according to some embodiments of the present disclosure. [Figure 6] This is a schematic diagram illustrating an architecture for modeling the reverberation of speech signals for multiple speakers, according to some other embodiments of the present disclosure. [Figure 7] This is a schematic diagram showing an architectural representation for improving speech signal reverberation modeling according to some embodiments of the present disclosure. [Figure 8A] This is a schematic diagram showing a network architecture for speech signal reverberation modeling according to some other embodiments of the present disclosure. [Figure 8B] This is a schematic diagram showing a network architecture for speech signal reverberation modeling according to some other embodiments of the present disclosure. [Figure 8C]This is a schematic diagram showing a network architecture for speech signal reverberation modeling according to some other embodiments of the present disclosure. [Figure 8D] This is a schematic diagram showing a network architecture for speech signal reverberation modeling according to some other embodiments of the present disclosure. [Figure 9A] This is a flowchart illustrating a method for estimating the intra-room impulse response (RIR) according to an embodiment of the present disclosure. [Figure 9B] This is a flowchart illustrating a method for modeling the reverberation of a speech signal according to an embodiment of the present disclosure. [Figure 10] This figure shows a tabular representation corresponding to a simulated test for reverberation removal of a speech signal according to an embodiment of the present disclosure. [Figure 11] This figure shows a tabular representation of evaluation results for speech signal reverberation removal using a test dataset, according to an embodiment of this disclosure. [Figure 12] This figure shows a tabular representation of evaluation results for speech signal reverberation removal using a test dataset, according to some other embodiments of the present disclosure. [Figure 13] This is a block diagram of an audio processing system according to an embodiment of the present disclosure. [Figure 14A] This block diagram shows a system for modeling the reverberation of a speech signal, according to some exemplary embodiments of the present disclosure. [Figure 14B] This block diagram shows a system for modeling the reverberation of a speech signal, according to some other exemplary embodiments of the present disclosure. [Figure 15] This figure shows use cases for speech signal reverberation modeling according to some exemplary embodiments of the present disclosure. [Figure 16] This figure shows a use case for speech signal reverberation modeling according to some other exemplary embodiments of the present disclosure. [Figure 17]This figure shows a use case for speech signal reverberation modeling according to some further exemplary embodiments of the present disclosure. [Figure 18] This figure shows a use case for speech signal reverberation modeling according to some further exemplary embodiments of the present disclosure. [Modes for carrying out the invention]
[0044] [Description of Embodiments] While the above drawings illustrate embodiments disclosed herein, other embodiments are also conceivable, as described herein. This disclosure uses exemplary embodiments as an expression, not an limitation. Many other modifications and embodiments, within the scope and spirit of the principles of the embodiments disclosed herein, can be devised by those skilled in the art.
[0045] For the sake of clarity, the following description includes numerous specific details to ensure a complete understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure can be implemented without these specific details. In other examples, the apparatus and methods are shown only in block diagram form to avoid obscuring this disclosure. Various modifications may be made to the function and arrangement of the elements without departing from the spirit and scope of the subject matter disclosed as described in the appended claims.
[0046] Where used herein and in the claims, the terms “for example,” “such as,” and “such,” as well as the verbs “equip,” “have,” and “include,” and their other verbal forms, when used with a list of one or more components or other items, are each to be interpreted as open-ended, meaning that the list should not be considered to exclude any other additional components or items. The term “based on” means based on at least in part. Furthermore, it should be understood that the expressions and terms used herein are for illustrative purposes only and should not be considered restrictive. Any headings used herein are for convenience only and have no legal or restrictive effect.
[0047] Specific details are given in the following description to provide a complete understanding of the embodiments. However, those skilled in the art will understand that embodiments can be carried out even without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in the form of block diagrams to avoid obscuring the embodiments with unnecessary details. In other examples, well-known processes, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments. Furthermore, similar reference numbers and names in different drawings refer to similar elements.
[0048] While the explanation primarily uses speech as the target audio source, the same method can be applied to other types of audio signals.
[0049] [System Overview] Figure 1A shows a representation of an environment 100A for speech signal reverberation modeling according to an embodiment of the present disclosure. Environment 100A may correspond to a closed environment having speakers 102 such as speaker 102A and speaker 102B. Figure 1A also shows a device 104 including at least a microphone or an array of microphones. In some exemplary embodiments, device 104 may correspond to an automatic speech recognition (ASR) system, a speech signal processing system, or any speech processing system.
[0050] In the illustrated illustrative scenario, when speaker 102 outputs speech, the corresponding acoustic speech signal may travel toward device 104 via different paths. The acoustic speech signal may be linearly distorted by object reflections such as wall and ceiling reflections, as shown in Figure 1A. In particular, speaker 102's acoustic speech signal is distorted in a multipath direction before reaching device 104, resulting in reverberation of the acoustic speech signal.
[0051] Therefore, device 104 receives such an acoustic speech signal from speaker 102A as an acoustic signal mixture. The acoustic signal mixture includes an anechoic speech signal and a reverberation speech signal. The anechoic speech signal is the target direct path signal 106A. The reverberation speech signal, collectively referred to below as reverberation 108A, includes a non-direct path signal or a multipath signal. There may be multiple speakers in environment 100, such as speaker 102A being together with another speaker 102B. In such cases, the acoustic signal mixture includes the target direct path signal 106B and the reverberation speech signal, collectively referred below as reverberation 108B corresponding to speaker 102B. The acoustic signal mixture may include reverberation noise signals 110A from non-target sources, such as an air conditioner 110 in environment 100A.
[0052] The speech signals from speaker 102A and / or speaker 102B may be interrupted before reaching device 104, as shown in Figure 1B.
[0053] Figure 1B shows an exemplary representation for speech signal reverberation modeling according to another embodiment of the present disclosure. As shown in environment 100B of Figure 1B, the speech signal of speaker 102A or speaker 102B is blocked by block 114 before reaching device 104. Block 114 may cause the speech signal of the corresponding speaker (e.g., speaker 102A or speaker 102B) to reverberate in a different direction. Such reverberation may increase attenuated and delayed copies (not shown in Figure 1B) of the speech signal of speaker 102A or speaker 102B. When device 104 is blocked by block 114, the speech signal of speaker 102A or speaker 102B may not have a corresponding target direct path signal. Instead, the speech signal may include shortest paths such as the shortest path 106C for the speech signal corresponding to speaker 102A and / or the shortest path 106D for the speech signal corresponding to speaker 102B. In this situation, for explanatory purposes, the shortest path signal is considered the target direct path signal, and the signal corresponding to a path longer than the shortest path is considered reverberation.
[0054] Device 104 may reduce such reverberations, for example, reverberations 108A and 108B, using a system 112 that may be integrated or built into device 104. System 112 will be further described with reference to Figures 2A and 2B.
[0055] Figure 2A is a schematic block diagram showing a system 200a for speech signal reverberation removal according to an embodiment of the present disclosure. System 200a corresponds to system 112 in Figures 1A and 1B.
[0056] In some exemplary embodiments, system 200a includes an input interface 202, a memory 204 for storing a first deep neural network (DNN1) (e.g., DNN1206A), a processor 208, and an output interface 210. The input interface 202 and the output interface 210 are further coupled to a wired communication channel or a wireless communication channel for communication between system 200a and external components of system 200a.
[0057] The input interface 202 is configured to receive an acoustic signal mixture, including a target direct path signal (e.g., target direct path signal 106A or target direct path signal 106B) and reverberations of the target direct path signal (e.g., reverberation 108A and / or reverberation 108B), over a wired or wireless communication channel. In some exemplary embodiments, the input interface 202 may be configured to connect to at least the microphones of device 104, or an array of microphones of device 104.
[0058] The processor 208 feeds an acoustic signal mixture, including the target direct path signal 106A and the reverberation 108A, into the DNN 1206A. The DNN 1206A generates a first estimate of the target direct path signal 106A. In a multi-speaker scenario including speakers 102A and 102B who generate sound signals in environment 100A or environment 100B, the DNN 1206A estimates the target direct path signals corresponding to each of speakers 102A and 102B. The DNN 1206A may generate corresponding estimates of the target direct path signal for each of speakers 102A and 102B, either individually or simultaneously. For example, the DNN 1206A simultaneously determines a first estimate of the target direct path signal 106A for speaker 102A and a first estimate of the target direct path signal 106B for speaker 102B.
[0059] A first estimate of the target direct path signal 106A is used together with the received acoustic signal mixture to estimate a filter that models the room impulse response (RIR) to the generated first estimate of the target direct path signal 106A. The RIR is the impulse response of the room between the sound source (e.g., speaker 102A and speaker 102B) and the microphone in device 104, e.g., environment 100A or environment 100B. When the filter that models the RIR is applied to the first estimate of the target direct path signal, it produces a result that is closest to at least one of the acoustic signal mixture and the residuals between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function. This estimated filter that models the RIR is transmitted over the communication channel. For this purpose, the filter that models the RIR may be output via the output interface 210 and then transmitted over a wired communication channel or a wireless communication channel connected to system 200a.
[0060] The filter that models RIR is a linear filter estimated using a linear convolutional prediction module. The linearity of this filter is modeled to reduce computational complexity in applications related to reverberation removal of sound signals. Furthermore, the linear filter provides a good approximation of the physical relationship between sound signals and their reflections. However, in reality, sound reflection is a nonlinear phenomenon. Nevertheless, those skilled in the art will consider the choice of whether to make the filter linear or nonlinear to be a design choice and well within the scope of this disclosure.
[0061] Linear filters are chosen for applications where simplicity and accuracy are required. This allows linear convolutional predictions to enhance the linear structure of the filter, approximating the linear reflection of the direct-path signal from the room's surface.
[0062] In some embodiments, the filter is applied to a first estimate of the target direct path signal in the time-frequency domain. This is functionally and computationally superior to applying the filter to the first estimate of the target direct path signal in the time domain, as the time-frequency domain allows for more accurate modeling of the filter. Furthermore, estimating the filter coefficients in the time-frequency domain provides a low-cost closed-form solution to the estimation problem. This may also be useful because other types of processing can be performed in the time-frequency domain, eliminating the need to switch back and forth between the time domain and the time-frequency domain.
[0063] When the estimated filter that models the RIR is applied to the first estimate of the target direct path signal 106A, a corresponding result is obtained. This result is closest to the acoustic signal mixture and one of the residuals between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function. The distance function provides the distance between the residual between the acoustic signal mixture and the first estimate of the target direct path signal and the result of applying the filter to the first estimate of the target direct path signal. In some embodiments, the distance function may correspond to a weighted distance having weights at each time-frequency point in the time-frequency domain. The weights may be determined by one or a combination of the received acoustic signal mixture and the estimate of the target direct path signal. In exemplary embodiments, the distance function may be based on the least-squares distance.
[0064] In some embodiments, the first DNN, namely DNN1206A, is trained to estimate both the amplitude and phase of the target direct-path signal. Therefore, the application of the filter estimate (which is also interchangeably and equivalently called the estimated filter modeling the RIR) to the first estimate of the target direct-path signal involves calculations on both the estimated amplitude and phase, and the distance function between the result of applying the filter to the direct-path signal estimate and the residual between the acoustic signal mixture and the direct-path signal is measured in both the amplitude and phase domains. Calculating both amplitude and phase allows for the use of information regarding the phase of the estimated direct-path signal in the filter estimation, leading to improved accuracy in estimating the filter.
[0065] In some embodiments, a first DNN, namely DNN1206A, is pre-trained to obtain a first estimate of a target direct path signal from an observed acoustic signal mixture. The pre-training of DNN1206A is performed using a training dataset of acoustic signal mixtures and the corresponding reference direct path signal within the training dataset, by minimizing a loss function.
[0066] When the result of applying the filtered estimate to the first estimate of the target direct path signal 106A is removed from the acoustic signal mixture, a mixture is obtained in which the reverberation of the target direct path signal 106A is reduced, achieving the objective of reverberation modeling of the speech signal, similar to the acoustic signal mixture.
[0067] Figure 2B is a schematic block diagram showing system 200b for speech signal reverberation modeling according to an embodiment of the present disclosure. System 200b corresponds to system 112 in Figures 1A and 1B. System 200b includes all the components of system 200a, plus an additional second DNN2206B.
[0068] To this end, system 200b includes an input interface 202, a memory 204 for storing a first DNN1 (e.g., DNN1206A) and a second DNN2 (e.g., DNN2206B), a processor 208, and an output interface 210.
[0069] When the result of applying the filter estimate to the first estimate of the target direct path signal 106A is removed from the acoustic signal mixture, a mixture of the target direct path signal 106A with reduced reverberation is obtained. This mixture of the target direct path signal 106A with reduced reverberation is provided as input to DNN2206B. DNN2206B generates a second estimate of the target direct path signal 106A. This second estimate of the target direct path signal is used with the received acoustic signal mixture to estimate an improved filter (or improved filter estimate) that models the RIR of the target direct path signal with higher accuracy than the filter obtained using the first estimate. The improved filter estimate is output via the output interface 210.
[0070] Similarly for speaker 102B, the first estimate of the target direct path signal 106B is used with the received acoustic signal mixture to estimate a filter that models the RIR of the target direct path signal 106B. The filter estimate is applied to the first estimate of the target direct path signal 106B to obtain the corresponding result. This result is removed from the acoustic signal mixture to obtain a mixture with reduced reverberation of the target direct path signal 106B. The mixture with reduced reverberation of the target direct path signal 106B is input to DNN2206B, which generates a second estimate of the target direct path signal 106B. The second estimate of the target direct path signal 106B for speaker 102B is used with the received acoustic signal mixture to estimate an improved filter that models the RIR of speaker 102B's target direct path signal with higher accuracy than the filter obtained using the first estimate. The improved filter that models the RIR is output via the output interface 210.
[0071] In some embodiments, a downstream DNN is used to apply one or more of the filter estimates output by the first DNN, i.e., DNN1206A, or the improved filter estimates output by the second DNN, i.e., DNN2206B, in order to perform different types of speech processing.
[0072] Figure 2C shows a downstream DNN, namely DNN3206C, used to perform multiple tasks not necessarily related to speech transcription and / or separation. For example, the downstream DNN, DNN3206C, is used in some embodiments for tasks such as speaker identification, sentiment recognition, event detection, and localization. For all these tasks, having information about the reverberation of the target direct path signal, modeled as the RIR of the target direct path signal, is beneficial for training or inference purposes. In various embodiments, the downstream DNN utilizes the RIR of a filter determined by the first DNN, or optionally by the second DNN.
[0073] For example, in a given application, the RIR (Resonance Indicator) might be modeled as a filter estimate for a location with different settings than the user's current location. For instance, in a film, the RIR might be determined for a church scene and then applied to a recording taken in a small conference room. In this way, RIRs for different settings can be combined with the RIR for the current location / setting to obtain new sound effects. Thus, the reverberation in the church setting is synthesized as if it were applied in the same room setting. In this manner, any number and type of sounds can be synthesized, providing diverse applications for the filter.
[0074] Furthermore, as mentioned above, a second estimate of the target direct path signal (such as the second estimate of target direct path signal 106A or the second estimate of target direct path signal 106B) is obtained as the de-reverberation speech signal of the corresponding speaker (such as speaker 102A or speaker 102B). The reverberation modeling of the speech signal by any of systems 200a, 200b, or 200c is described in more detail with reference to Figures 3A, 3B, and 3C.
[0075] Figure 3A is a schematic block diagram showing a process 300 for modeling the reverberation of a speech signal according to an embodiment of the present disclosure. The process 300 is performed by a system 200. In an exemplary embodiment, an acoustic signal mixture 302(Y) is received via an input interface 202 of the system 200. The acoustic signal mixture includes reverberation from other sources, such as a noise signal 110A from device 110, as well as a target direct path signal such as a target direct path signal 106A from speaker 102A and reverberation 108A of the target direct path signal 106A, or a target direct path signal 106B from speaker 102B and reverberation 108B of the target direct path signal 106B. The received acoustic signal mixture 302 is fed into a DNN 1206A.
[0076] The DNN1206A determines a first estimate 304 of the target direct path signal, such as target direct path signal 106A or target direct path signal 106B. Furthermore, a filter estimate 306 (hereinafter interchangeably referred to as filter 306) is determined to model the room impulse response (RIR) 308 for the first estimate 304 of the target direct path signal 106A. The RIR model 308, hereinafter referred to as RIR 308, can correspond to the impulse response of the environment, such as environment 100A or environment 100B, between a source, such as speaker 102A and / or speaker 102B, and a receiver, such as device 104. For this purpose, absolute delays and attenuations due to propagation from the source to the microphone are not modeled; only relative delays and attenuations with respect to the direct path signal are modeled. The impulse response is considered relative to the direct path signal received in the mixture, not to the actual dry source signal at the source location. For the sake of clarity, the application of filter estimate 306 to the direct-path signal is such that it includes only the early reflections and late reverberations of the direct-path signal, and does not include the direct-path signal itself. The associated full filter estimate is equivalently obtained by modifying filter estimate 306 to further include the direct-path signal. These two filter estimates are equivalent, and one can be directly obtained from the other.
[0077] In some exemplary embodiments, the acoustic signal mixture 302 may correspond to a monaural signal recorded in a noisy environment, such as environment 100A or environment 100B. Such a monaural signal can be formulated into a physical model in the time domain. This physical model represents the relationship between the acoustic signal mixture 302(y) and a reverberation target speech signal (x) (including both a target direct path signal such as target direct path signal 106A and reverberation such as reverberation 108A) and a non-target source (v) (e.g., device 110) including a reverberation noise signal (e.g., reverberation noise signal 110A) and a reverberation competing speaker (e.g., speaker 102B).
[0078]
Number
[0079] The term "r d ", "r e " and "r l " represent the direct part, the initial part and the late part of the RIR308 of the environment 100, respectively. The term "s" represents a target direct path signal (such as the target direct path signal 106A), and the target direct path signal is defined as s = a * r d . The term "h" represents an indirect path signal (for example, reverberation 108A), and the indirect path signal is the sum of the early reflection a * r e and the late reverberation a * r l , that is, h = a * r e + a * r l = a * r e+l is defined as. The part r d+e of the RIR308 corresponding to both the direct path and the early reflection can be defined as a set of impulses up to 50 milliseconds after the direct path peak of r, and the early reflection component re of the RIR can be defined as r e = r d+e - r d . The filter for modeling the RIR in this application is considered in relation to r d . That is, the starting point of the time of the filter is implicitly considered to be the time of the impulse of r d , and the scaling of the elements of the filter is considered based on the height of the impulse of r d .
[0080]
Number
[0081] The target direct path signal 106A, represented as S(t,f) in equation (2), is estimated from the STFT coefficient (Y(t,f)) of the acoustic signal mixture 302 using a DNN. The recovered target direct path signal 106A(S(t,f)) can be used as the first estimate 304 of the target direct path signal 106A.
[0082]
number
[0083]
number
[0084]
number
[0085]
number
[0086]
number
[0087]
number
[0088]
number
[0089]
number
[0090] The estimation of filter 306 can be improved by calculating a full-filtered estimate using equation (7), if an estimate of the reverberation target speech X can be obtained. In some embodiments, estimates of each speaker's reverberation speech are iteratively removed from the acoustic signal mixture 302 to refine the reverberation direct signal used in the estimation of filter 306.
[0091] In this embodiment, equation (4) of the FCP can remove reverberation associated with target speaker 102A. The ability to obtain the reverberation of target speaker 102A can be particularly useful in multi-speaker separation tasks because each target speaker can be convolved with a different RIR. For this purpose, in some embodiments, different filters may be calculated to remove reverberation from each speaker (as illustrated in Figure 6). The estimated filter, for example, filter 306, may focus on reducing the reverberation of target speaker 102A rather than the reverberation of a combination of another speaker (e.g., speaker 102B) and a non-target source (e.g., air conditioner 110). To remove reverberation from the speech signal even when a non-target source is present, the output of DNN 1206A, such as a first estimate 304 of the target direct path signal 106A, and a reverberation-reduced mixture obtained using filter 306 may be used for removing reverberation from the speech signal. To this end, the first estimate 304 and the reverberation-reduced mixture obtained using filter 306 may be input to DNN2206B to output a second estimate 314 of the target direct path signal 106A (or target direct path signal 106B). Outputs generated by DNN2206B, such as the second estimate 314, will likely be superior to the output of DNN1206A because the input to DNN2206B (i.e., the first estimate 304 and the reverberation-reduced mixture obtained using filter 306 output by DNN1206A) is more refined than the input to DNN1206A. For example, there may be less interference between the first estimate 304 and the reverberation-reduced mixture obtained using filter 306 output by DNN1206A. When the reverberation-reduced mixture obtained using these less interfering first estimates 304 and filter 306 is processed by DNN2206B, the corresponding output (i.e., second estimate 314) may be better than the output of DNN1206A (i.e., first estimate 304).Therefore, using the second estimate generated by DNN2206B, another iteration of the convolutional prediction can be performed to obtain a second filter and a second mixture with reduced reverberation, which can then be input to DNN2206B along with the second estimate to produce a refined output.
[0092] In some exemplary embodiments, the corresponding RIR for each speaker, such as RIR308, can be estimated by solving a linear regression problem for each frequency in the time-frequency domain or time domain. A filter 306 that models RIR308 can be used to identify delayed and attenuated copies of the target direct path signal for speaker 102A and / or speaker 102B. Delayed and attenuated copies, which are repeating patterns due to reverberation, can be removed from the received acoustic signal mixture 302. To this end, the filter 306 is applied to a first estimate 304 to output result 310. Result 310 may be closest to the residual between the acoustic signal mixture 302 and the first estimate 304 of the target direct path signal, based on a distance function such as a weighted least-squares distance function. When result 310 is removed from the acoustic signal mixture 302, a mixture 312 with reduced reverberation is obtained.
[0093]
number
[0094] Removal of result 310 reduces delayed and decayed copies from mixture 312, which has reduced reverberation. Delayed and decayed copies may correspond to late reverberation and early reflections of the target direct path signal. These early and late reverberations can be identified from RIR 308, which is modeled by the filtered estimate 306. RIR 308 with early and late reverberations is shown in Figure 3B.
[0095] Figure 3B shows a representation 316 of the Room Impulse Response (RIR) model 316A for a source of the original signal from a speaker such as speaker 102A, showing the impulse corresponding to the target direct path signal 320A, the impulse corresponding to the early reflection 320B, and the impulse corresponding to the late reverberation 320C. In this application, the target direct path signal is considered as the reference instead of the source of the original signal from the speaker. In other words, the application of the RIR to the target direct path signal results in the speaker's reverberation signal, which is the sum of the target direct path signal and the early reflection and late reverberation of the target direct path signal.
[0096] Figure 3C shows an application 326 of filter 316B that models RIR 316A in frequency bin f according to an embodiment of the present disclosure. RIR model 316A corresponds to RIR model 308, and filter estimate 316B corresponds to the full filter estimate associated with filter estimate 306.
[0097] The RIR model 316A has a structure that can be represented as a sequence of impulses in the time domain. For example, the RIR model 316A is represented as a graph plot having amplitude 318A and a tap number axis 318B representing time delay. The structure of the RIR model 316A is a target direct path signal 320A(r d The impulse corresponding to ) and the target direct path signal 320A(r) followed by the late reverberation 320C(r1) of the target direct path signal 320A(r d ) Discrete initial reflection 320B(r e ) may include multiple impulses corresponding to ). Target direct path signal 320A may correspond to target direct path signal 106A or target direct path signal 106B.
[0098] In some exemplary embodiments, early reflections 320B and late reverberations 320C are identified from the RIR model 316A. Assuming that the filter is modeled using K coefficients at each frequency f, the coefficients of the filter estimate 307 at frequency f are obtained such that the application of the filter 326 to the first estimate 304 of the target direct path signal, by summing the results of multiplying the time-frequency bin of the first estimate at time t-k+1 (all k=1,...,K) at the same frequency f by the k-th coefficient, best approximates the reverberation mixture 322 at the same frequency f at the current time t.
[0099] As shown in graph plot 316B, the acoustic signal mixture 322(Y) is approximated by applying a K-tap filter 324 to a first estimate 304 of the target direct path signal. Filter 324 is estimated by optimizing the forward filtering of the first estimate 304 of the target direct path signal 302A. Filter 324 is an example of filter 307. For example, the number of taps K of filter 324 may be set to 40, which corresponds to a filter length of ((40-1)×8+32) ms in the time domain.
[0100] There can be different scenarios for reverberation modeling, including speech signal de-reverberation, using DNN1206A and DNN2206B. For example, the acoustic signal mixture 302 may be received by a single microphone or array of microphones of device 104 from a single speaker (e.g., speaker 102A) or from multiple speakers (e.g., speakers 102A and 102B). In the case of multiple speakers, the first DNN1206A estimates a different first estimate of the target direct path signal for each of the multiple speakers. Speech signal de-reverberation for different scenarios will be further illustrated with reference to Figures 4, 5, and 6.
[0101] Figure 4 is a schematic diagram showing an architectural representation 400 for speech signal reverberation modeling according to an embodiment of the present disclosure. As shown in Figure 4, the architectural representation 400 includes DNN1402, DNN2406, and a convolutional prediction module 404 between DNN1402 and DNN1406. DNN1402 corresponds to DNN1206A, and DNN2406 corresponds to DNN2206B.
[0102]
number
[0103]
number
[0104]
number
[0105] Some embodiments are based on the understanding that the second estimate 410 is superior to the first estimate 408 because the DNN2406 processes a refined mixture of acoustic signals, which is a reverberation-reduced mixture 412. The second estimate 410 may be further improved to perform better than the first estimate 408. To this end, the DNN2406 may be input with one of the acoustic signal mixture 302 and the first estimate 408, or a combination thereof, to generate the second estimate 410. In some cases, the acoustic signal mixture 302, the first estimate 408, and the reverberation-reduced mixture 412 may be input to the DNN2406 to generate the second estimate 410. In other cases, the first estimate 408 and the reverberation-reduced mixture 412 may be input to the DNN2406 to generate the second estimate 410. Furthermore, the estimation of the filter, acquisition of the reverberation-reduced mixture 412, and input of the reverberation-reduced mixture 412 may be repeated to gradually refine the second estimate 410 of the target direct-path signal 106A to improve the de-reverberation of the speech signal for speaker 102A. This iteration may be terminated when a termination condition is met. The termination condition may correspond to a user-defined condition. Thus, the second estimate 410 will be better than the first estimate 408 because it is refined with the reverberation-reduced mixture 412. In some embodiments, the DNN 2406 may be trained using the acoustic signal mixture 302, the reverberation-reduced mixture 412, and the first estimate 408 to output a second estimate 410 that improves the de-reverberation of the speech signal.
[0106] In the case of multiple speakers, the received acoustic signal mixture 302 may contain speech signals from multiple speakers, such as speaker 102A and speaker 102B. In such cases, the DNN 1402 can generate multiple outputs, such as different first estimates of the target direct path signal, and from these multiple outputs, different filters can be obtained to model the corresponding RIR for the multiple speakers, which will be further explained with reference to Figure 5.
[0107] Figure 5 is a schematic diagram showing an architectural representation 500 for modeling the reverberation of a speech signal for multiple speakers (e.g., speakers 102A and 102B) according to some embodiments of the present disclosure. As shown in Figure 5, the architectural representation 500 corresponds to a multiple speaker scenario and includes multiple instances of convolutional prediction modules, such as DNN1502, DNN2506, and convolutional prediction modules 504A and 504B between DNN1502 and DNN2506. DNN1502 corresponds to DNN1206A, and DNN2506 corresponds to DNN2206B.
[0108]
number
[0109]
number
[0110] The reverberation-reduced mixture 510A and the reverberation-reduced mixture 510B are concatenated and provided as input to the DNN2506 to generate corresponding second estimates 512A and 512B for speakers 102A and 102B. In some exemplary embodiments, the DNN2506 may be input with the first estimate 508A along with the reverberation-reduced mixture 510A, with the first estimate 508B along with the reverberation-reduced mixture 510B, and the acoustic signal mixture 302 to output the second estimates 512A and 512B.
[0111] In some exemplary embodiments, the filtered and reverberation-reduced mixtures 510A and 510B for the first estimates 508A and 508B may be repeated by substituting the first estimates 508A and 508B with the second estimates 512A and 512B in order to generate a filtered and reverberation-reduced mixture for each of the multiple speakers 102A and 102B. This iteration terminates when a user-defined termination condition is met. This termination condition may include, for example, terminating after three iterations.
[0112] In some exemplary embodiments, reverberation-reduced mixtures 510A and 510B may be coupled to a tensor. The tensor is a dimensional data structure representing all reverberation-reduced mixtures for multiple speakers 102A and 102B. The tensor is fed into a DNN2506 to output corresponding second estimates 512A and 512B for each of the multiple speakers 102A and 102B.
[0113] It is also possible for each of the multiple speakers 102A and 102B to have one corresponding second estimate, which will be explained in Figure 6 below.
[0114] Figure 6 is a schematic diagram showing an architectural representation 600 for modeling the reverberation of speech signals of multiple speakers 102A and 102B according to some other embodiments of the present disclosure. As shown in Figure 6, the architectural representation 600 corresponds to a multi-speaker scenario and includes DNN1602, multiple instances of a second DNN such as DNN2606A and DNN2606B, and multiple instances of convolutional prediction modules such as convolutional prediction modules 604A and convolutional prediction modules 604B between DNN1602 and the multiple instances of the second DNN such as DNN2606A and DNN2606B. DNN1602 corresponds to DNN1206A, and each of DNN2606A and DNN2606B corresponds to DNN2206B.
[0115] The DNN1602 receives the acoustic signal mixture 302 and estimates the corresponding target direct path signal for each of the multiple speakers 102A and 102B. For example, the DNN1602 estimates a first estimate 608A of the target direct path signal 106A for speaker 102A. The DNN1602 estimates a first estimate 608B of the target direct path signal 106B for speaker 102B. The first estimate 608A is input to the convolutional prediction module 604A, and the first estimate 608B is input to the convolutional prediction module 604B.
[0116] The convolutional prediction module 604A estimates a filter that models the RIR of the first estimate 608A. The filter that models the RIR is applied to the first estimate 608A to obtain a mixture in which the reverberation 610A of the target direct path signal 106A is reduced. Similarly, the convolutional prediction module 604B estimates a filter that models the RIR of the first estimate 608B. The estimated filter that models the RIR, output by the convolutional prediction module 604B, is applied to the first estimate 608B to obtain a mixture in which the reverberation 610B of the target direct path signal 106B is reduced.
[0117]
number
[0118] Furthermore, each of the reverberation-reduced mixtures is fed into an instance of DNN2. Each of the reverberation-reduced mixtures 610A and 610B is fed into the corresponding instances DNN2606A and DNN2606B (essentially the same DNN2, but applied to different inputs), respectively. DNN2606A generates a second estimate 612A of the target direct path signal 106A for speaker 102B. DNN2606B generates a second estimate 612B of the target direct path signal 106B for speaker 102B. Multiple instances of the second DNN, such as DNN2606A and DNN2606B, which each generate second estimates 612A and 612B for the corresponding speakers 102A and 102B, can be used to obtain the clear speech of individual speakers from multiple speakers.
[0119] To improve the second estimates 612A and 612B, DNN2606A and DNN2606B may be input to one or a combination of the acoustic signal mixtures, the first estimates 608A and 608B, and the reverberation-reduced mixtures 610A and 610B.
[0120] In some exemplary embodiments, the first estimate 608A may be replaced with the second estimate 612A to generate an updated first estimate 608A of the target direct path signal 106A. Similarly, the first estimate 608B may be replaced with the second estimate 612B to generate an updated first estimate 608B of the target direct path signal 106B. Furthermore, by repeating the estimation of the filter by the DNN 1602, the estimation of the reverberation-reduced mixtures 610A and 610B, and the input of the reverberation-reduced mixtures 610A and 610B, an updated second estimate of the target direct path signal may be output for each of the multiple speakers 102A and 102B.
[0121] In some other exemplary embodiments, a portion of the acoustic signal mixture corresponding to a speaker (e.g., speaker 102A) may be extracted. This portion is extracted by removing the reverberant speech of another speaker, e.g., speaker 102B, from the acoustic signal mixture. An estimate of the reverberant speech of another speaker among multiple speakers is obtained by adding a first estimate of the target direct path signal of that other speaker to the result of applying the corresponding filter for that other speaker to the first estimate of the target direct path signal of that other speaker. After the extraction of the portion of the acoustic signal corresponding to speaker 102A, a filter is estimated for the first estimate of the extracted portion. This filter is used to estimate the mixture with speaker 102A's reverberation reduced based on the portion. Processing of the portion can improve the quality of the estimated filter for the speaker and the corresponding second estimate.
[0122] In some exemplary embodiments, the acoustic signal mixture of a single speaker 102A and / or multiple speakers 102A and 102B may be received from a single microphone or from an array of microphones. Therefore, DNNs such as DNN1602 and DNN2606A and DNN2606B can be trained based on spectral mappings corresponding to a single microphone and an array of microphones. The spectral mapping trains DNN1602 to predict the real and imaginary (RI) components (i.e., frequencies) of an estimate, such as a first estimate 608A of the target direct path signal 106A, from the RI component of the acoustic signal mixture 704. The RI component of the acoustic signal mixture 704 and the RI component of the first estimate 608A are input to DNN2606A, which may then predict a second estimate of the target direct path signal 106A. DNN1602 can be pre-trained using a training dataset of acoustic signal mixtures and a training dataset of the corresponding reference target direct path signal within the training dataset.
[0123] In some embodiments, pre-training of DNN1602 may be performed by minimizing a loss function. This loss function may include one or a combination of distance functions defined based on the RI component of the target direct path signal 106A in a first time-frequency domain and the RI component of a reference target direct path signal in a first time-frequency domain. The reference target direct path signal can be obtained from a training dataset of utterances, and the corresponding reverberation mixture can be obtained by convolving the reference target direct path signal with a recorded RIR or synthesized RIR and summing it with other interfering signals. The distance function may be defined based on the magnitude obtained from the RI component of the estimated target direct path signal in a first time-frequency domain and the corresponding magnitude of the reference target direct path signal.
[0124] In an alternative embodiment, the distance function may be defined based on the reconstructed waveform obtained from the RI component of the target direct path signal estimated in the first time-frequency domain by reconstruction in the time domain, and the waveform of the reference target direct path signal. Alternatively, the distance function may be defined based on the RI component in the complex time-frequency domain obtained by further transforming the reconstructed waveform in the second time-frequency domain, and the RI component of the reference target direct path signal in the second time-frequency domain. Alternatively, the distance function may be defined based on the magnitude obtained from the RI component in the second time-frequency domain obtained by transforming the reconstructed waveform in the second time-frequency domain, and the corresponding magnitude of the reference target direct path signal in the second time-frequency domain.
[0125]
number
[0126]
number
[0127]
number
[0128]
number
[0129]
number
[0130] In some exemplary embodiments, the acoustic signal mixture 302 may correspond to a multi-channel signal that may be received from an array of microphones. Beamforming is performed on such a multi-channel signal, which will be further explained with reference to Figure 7.
[0131] Figure 7 is a schematic diagram showing an architectural representation 700 for improving speech signal reverberation rejection according to some embodiments of the present disclosure. The architectural representation 700 is similar to the architectural representation in Figure 5, but further includes multiple instances of a Minimum Variance Distortionless Response (MVDR) beamforming module 704. In some exemplary embodiments, each instance of the MVDR module may output a beamforming output for a multi-channel signal. The beamforming filter can be obtained based on statistics derived from one or a combination thereof of first estimates, such as first estimate 508A (and / or first estimate 508B) output by a first DNN such as DNN1502, reverberation-reduced mixture 510A (and / or reverberation-reduced mixture 510B), and second estimates, such as second estimate 512A (and / or second estimate 512B) output by a second DNN such as DNN2506, the second estimates of which would have been obtained using the architectural representation of Figure 5, which includes only a convolutional predictive module between the two DNNs, or a previous iteration of the architectural representation of Figure 7, which includes MVDR beamforming. The beamforming output for the speaker can be obtained by applying the beamforming filter to reverberation-reduced mixture 510A or mixture 502. An MVDR beamforming module may be used between two DNNs, such as DNN1502 and DNN2506. The output of the MVDR beamforming module, such as beamforming output 514A (and / or beamforming output 514B), may be used as an input to a second DNN, such as DNN2506. In some exemplary embodiments, the output of the MVDR beamforming module, such as beamforming output 514A, may be combined with one or a combination of a first estimate, such as first estimate 508A, a reverberation-reduced mixture, such as reverberation-reduced mixture 510A, and mixture 502.In some exemplary embodiments, the beamforming output for all speakers is combined with a reverberation-reduced mixture for all speakers, a first estimate for all speakers, and the mixture, and used as input to the DNN2506. In some exemplary embodiments, an MVDR beamforming module may output beamforming using MVDR technology so that signals from multiple channels can be combined to derive a better estimate of the target direct path signal.
[0132] Therefore, MVDR beamforming can be applied to mixtures with reduced reverberation to further improve reverberation removal and separation tasks.
[0133]
number
[0134]
number
[0135] Furthermore, DNNs, such as DNN1602 and DNN2606A, can be easily replaced with size or time domain models and more advanced DNN architectures. One such model is further described with reference to Figures 8A, 8B, 8C, and 8D.
[0136] Figures 8A, 8B, 8C, and 8D are schematic diagrams showing a network architecture 800 for speech signal reverberation removal according to some other embodiments of the present disclosure. The network architecture 800 corresponds to DNNs such as DNN1206A and DNN2206B.
[0137] The network architecture 800 is a Temporal Convolutional Network (TCN) 806. The TCN 806 includes four layers, each of which contains six extended convolutional blocks, such as extended convolutional block 802A, extended convolutional block 802B, extended convolutional block 802C, extended convolutional block 802D, extended convolutional block 802E, and extended convolutional block 802F (hereinafter referred to as extended convolutional blocks 802A-802F). In each of the extended convolutional blocks 802A-802F, one one-dimensional (1D) depth-separable convolution 804 is used to reduce the number of parameters. For example, each of the extended convolutional blocks 802A-802F may contain approximately 6.9 million parameters for speech signal de-reverberation. These numerous parameters can be reduced by the 1D depth-separable convolution 804.
[0138] Furthermore, the TCN806 is sandwiched between U-Nets, which include an encoder 808 and a decoder 810. In each of the encoder 808 and decoder 810, DenseNet blocks are inserted at multiple frequency scales. A DenseNet block is an architecture that trains DNNs such as DNN1602 and DNN2606A using shorter connections between layers of the DNN. For example, the encoder 808 includes DenseNet blocks 808A, 808B, 808C, 808D, and 808E (hereinafter simply referred to as DenseNet blocks 808A-808E) at multiple frequency scales. Similarly, the U-Net decoder 810 includes DenseNet blocks 810A, 810B, 810C, 810D, and 810E (hereinafter simply referred to as DenseNet blocks 810A-810E) across multiple frequency scales. U-Net can maintain fine-grained local structure through skip connections and frequency-aligned model context information via downsampling and upsampling. TCN 806 leverages long-term information of the received acoustic signal mixture by using extended convolution along the time domain. DenseNet blocks 808A-808E enable feature reuse and improve the discriminability of speech signals from multiple speakers 102A and 102B in speaker separation tasks.
[0139] The encoder 808 includes one two-dimensional (2D) convolution 812 and seven convolution blocks, such as convolution blocks 814A, 814B, 814C, 814D, 814E, 814F, and 814G (hereinafter referred to as convolution blocks 814A-814G). Each of the convolution blocks 814A-814G includes a 2D convolution, exponential linear unit (ELU) nonlinearity, and instance normalization (IN) to downsample, i.e., reduce the sampling rate or sample size (bits per sample) of an input signal, such as an acoustic signal mixture 704. The 2D convolution forms an essential component of feature extraction corresponding to the estimate of the target direct-path signal. ELU is the activation function for DNNs (e.g., DNN1602 and DNN2606A), and IN is the normalization layer for stabilizing the hidden state dynamics in DNN1602 and DNN2606A.
[0140] Decoder 810 includes seven blocks of 2D deconvolutions, such as deconvolution 816A, deconvolution 816B, deconvolution 816C, deconvolution 816D, deconvolution 816E, deconvolution 816F, and deconvolution 816G (hereinafter referred to as deconvolution 816A-816G), along with ELU and IN and one 2D deconvolution 820, for upsampling by adding zero-value samples between the original samples to increase the sampling rate.
[0141] As mentioned above, reverberation-reduced mixtures of multiple speakers 102A and 102B (such as reverberation-reduced mixture 510A and reverberation-reduced mixture 510B) are represented by tensors. The tensors are of the form featureMapstimeStepsfrequencyChannels. Each of the convolutional blocks 814A-814G (i.e., Conv2D+ELU+IN) and deconvolutional blocks 816A-816G (i.e., Deconv2D+ELU+IN) is specified in the form kernelSizeTimekernelSizeFreq, (stridesTime,stridesFreq), (paddingsTime,paddingsFreq) and featureMaps.
[0142] Each of the DenseNet blocks 808A-808E, such as DenseBlock(g1,g2), contains five Conv2D+ELU+IN blocks, each having a growth rate g1 for the first four layers and a growth rate g2 for the last layer of the DenseNet block 808A-808E. The tensor shape after each TCN block is in the form featureMapstimeSteps. Each IN+ELU+Conv1D block is specified in the form kernelSizeTime, stridesTime, paddingsTime, dilationTime, featureMaps.
[0143] Figure 9A is a flowchart illustrating a method 900a for estimating RIR for reverberation modeling of a speech signal, according to an embodiment of the present disclosure. Method 900a is performed by system 200. Method 900a comprises the step in operation 902 of receiving an acoustic signal mixture (e.g., acoustic signal mixture 302) via a communication channel, which includes a target direct path signal (e.g., direct path signal 106A) and the reverberation of the target direct path signal. The acoustic signal mixture may include at least one of single-channel signals or multi-channel signals that can be received from a single microphone or an array of microphones connected to an input interface connected to the communication channel.
[0144] In operation 904, the received acoustic signal mixture is fed into a first DNN, such as DNN 1206, to generate a first estimate (e.g., first estimate 408) of the target direct path signal 106A. In a multi-speaker scenario, the first DNN determines a corresponding first estimate for each of the multiple speakers. The corresponding first estimates may be determined one by one for each of the multiple speakers, or they may be determined simultaneously for multiple speakers. In some embodiments, the first DNN may be pre-trained to generate a first estimate from an observed acoustic signal mixture based on a training dataset of acoustic signal mixtures and a corresponding reference target direct path signal in the training dataset. Pre-training of the first DNN may be performed by minimizing a loss function.
[0145] In operation 906, a filter (e.g., filter 306) that models the room impulse response (RIR) (e.g., RIR model 308) is estimated for a first estimate 408 of the target direct path signal 106A, and the filter is estimated such that the result of applying the filter to the first estimate of the target direct path signal is closest to the residual between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function (e.g., least squares distance function). In some embodiments, the filter corresponds to a linear filter structure estimated based on convolutional predictions. The first estimate is forward-filtered frequency by frequency in the time-frequency domain using a linear filter of convolutional predictions (described in Figures 3A, 3B, 4, 5, and 6). In some exemplary embodiments, the received acoustic signal mixture includes speech signals from multiple speakers. The first DNN generates multiple outputs, each output containing a first estimate of the target direct path signal for one speaker from the multiple speakers. In some embodiments, the early reflections (e.g., early reflection 320B) and late reverberations (e.g., late reverberation 320C) of the first estimate may be identified based on the RIR modeled by the filter. The identified early reflections and late reverberations may be removed from the first estimate to estimate a mixture with reduced reverberation.
[0146] In operation 908, the estimated filter that models the RIR is transmitted over a communication channel. The communication channel is any combination of wired or wireless communication channels required for each application of the filter after transmission. To this end, method 900a may be extended with additional processing steps, as shown in Figure 9B below.
[0147] Figure 9B is a flowchart illustrating another method 900b for de-reverberation of a speech signal according to an embodiment of the present disclosure. Method 900b is performed by system 200.
[0148] Method 900b comprises the step in operation 910 of receiving an acoustic signal mixture (e.g., acoustic signal mixture 302) via an input interface, which includes a target direct path signal (e.g., target direct path signal 106A) and the reverberation of the target direct path signal. The acoustic signal mixture may include at least one of single-channel signals or multi-channel signals that can be received from a single microphone or an array of microphones connected to the input interface.
[0149] In operation 912, the received acoustic signal mixture is fed into a first DNN, such as DNN 1206, to generate a first estimate (e.g., first estimate 408) of the target direct path signal 106A. In a multi-speaker scenario, the first DNN determines a corresponding first estimate for each of the multiple speakers. The corresponding first estimates may be determined one by one for each of the multiple speakers, or they may be determined simultaneously for multiple speakers. In some embodiments, the first DNN may be pre-trained to generate first estimates based on the observed acoustic signal mixture, or a training dataset of acoustic signal mixtures, and at least one of the corresponding reference target direct path signals in the training dataset. Pre-training of the first DNN may be performed by minimizing a loss function.
[0150] In some embodiments, the first estimate of the target direct path signal is updated based on the application of a filter that models the RIR. Next, the training data of the training dataset is updated using the updated first estimate and the acoustic signal mixture. The updated first estimate of the target direct path signal is used as a pseudo-label to identify the target direct path signal in the updated training dataset. Furthermore, the first DNN is then retrained using this updated training dataset.
[0151] In operation 914, a filter (e.g., filter 306) that models the room impulse response (RIR) (e.g., RIR model 308) is estimated for a first estimate of the target direct path signal 106A 408 such that the result of applying the filter that models the RIR to a first estimate of the target direct path signal is closest to the residual between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function (e.g., least squares distance function). In some embodiments, the filter corresponds to a linear filter structure estimated based on convolutional predictions. The first estimate is forward filtered frequency by frequency in the time-frequency domain using a linear filter of convolutional predictions (shown in Figures 3A, 3B, 4, 5, and 6). In some exemplary embodiments, the received acoustic signal mixture includes speech signals from multiple speakers. The first DNN generates multiple outputs, each output containing a first estimate of the target direct path signal for one of the multiple speakers. In some embodiments, early reflections (e.g., early reflection 320B) and late reverberations (e.g., late reverberation 320C) of the first estimate may be identified based on the RIR modeled by the filter. The identified early reflections and late reverberations may be removed from the first estimate to estimate the acoustic signal mixture.
[0152] In operation 916, the mixture with reduced reverberation of the target direct path signal 106A is obtained by removing from the received mixture the result of applying a filter to a first estimate 408 of the target direct path signal 106A. In some embodiments, the second DNN may be trained on a training dataset created from augmented data obtained using the estimated set of filters and the estimated set of target direct path signals to create the reverberation mixture. For example, the augmented data may include external data measured outside the room or one or more other acoustic signal mixtures in the training dataset. Next, reverberation removal processing based on the techniques described above is performed on this augmented data. Furthermore, the results of applying filters and reverberation removal processing include augmented data that is then used to further train the second DNN.
[0153] In operation 918, the reverberation-reduced mixture is fed into a second DNN (e.g., DNN2206B) to generate a second estimate of the target direct path signal. In some exemplary embodiments, one or a combination of the received acoustic signal mixture and the first estimate of the target direct path signal is fed into the second DNN to generate a second estimate of the target direct path signal. In some other exemplary embodiments, the received acoustic signal mixture, the first estimate of the target direct path signal, and the reverberation-reduced mixture are fed into the second DNN to generate a second estimate of the target direct path signal. In some yet other exemplary embodiments, the first estimate of the target direct path signal and the reverberation-reduced mixture are fed into the second DNN to generate a second estimate of the target direct path signal. In some embodiments, the second DNN may be trained on a training dataset created from augmented data acquired using an estimated set of filters and an estimated set of target direct path signals to create a reverberation mixture.
[0154] In operation 920, a second estimate of the target direct path signal is output via an output interface, such as output interface 210. To further improve speech signal de-reverberation, the steps of estimating a filter, obtaining a de-reverberated mixture, and feeding the de-reverberated mixture may be repeated for each of the multiple outputs of the first DNN. The output interface may also be configured to output the RIR modeled by the filter. The output RIR can be used to perform speech analysis for room acoustic parameter analysis, room geometry reconstruction, speech enhancement, and speech signal de-reverberation, or a combination thereof.
[0155] In some exemplary embodiments, speech signal dereverberation using estimates, i.e., first and second estimates of the target direct path signal, and a filter of the target direct path signal, is evaluated for three tasks: 1) speech dereverberation using weak stationary noise, 2) two-speaker separation in a reverberant situation using white noise, and 3) two-speaker separation in a reverberant situation using difficult transient noise. The evaluation results are shown in Figures 10, 11, and 12.
[0156] Figure 10 shows a tabular representation 1000 corresponding to a simulated test set for speech signal de-reverbing according to an embodiment of the present disclosure. The tabular representation 1000 shows the dataset used for de-reverbing, the reverberation speaker separation and speech enhancement tasks, the hyperparameter settings, and the baseline system for speech signal de-reverbing. The tabular representation 1000 also shows the results for the ASR task on the REVERB corpus.
[0157] For speech signal reverberation removal, DNNs such as DNN1206A and DNN2206B can be trained using simulated reverberation datasets under low air conditioning noise conditions. In addition to evaluating trained DNNs on simulated test sets, DNNs are directly applied to the Reverberant Voice Enhancement and Recognition Benchmark (REVERB) corpus to demonstrate their effectiveness in processing actually recorded noisy reverberation utterances. The REVERB corpus is a benchmark for evaluating automatic speech recognition technologies. The dataset also includes clean signals for simulation obtained from the WSJCAM0 corpus. The WSJCAM0 corpus contains 7,861 utterances in its training set, 742 in its validation set, and 1,088 in its test set, respectively. Using these utterances from the WSJCAM0 corpus, reverberation mixtures with 39,305 (7,861 × 5) noises, 2,968 (742 × 4) noises, and 3,264 (1,088 × 3) noises are simulated as training, validation, and test sets, respectively. Subsequently, a data spatialization process is performed, and for each utterance, the room is randomly sampled with random room features and speaker and microphone positions, using the estimated RIR for de-echoing of the speech signal. The distance between the speaker and microphone is sampled from the range [0.75, 2.5] m. The reverberation time (T60) is derived from the range [0.2, 1.3] seconds. For each utterance, diffuse air conditioning noise is sampled from the REVERB corpus and added to the speaker's reverberation speech. The signal-to-noise ratio between anechoic speech and noise is sampled from the range [5, 25] dB. The sampling rate is 16kHz.
[0158] The trained model is applied to practical reverberation recordings without retraining and to the REVERB ASR task. The test mixture is obtained from actual recordings made in a room (e.g., environment 100) with a reverberation time T60 of approximately 0.7 seconds and a distance of approximately 1 m between the speaker and the microphone for the near-field and 2.5 m for the far-field. The recorded noise is diffuse air conditioning noise and is weak.
[0159] In software such as Kaldi, an official reverb corpus is used to build a backend for ASR, which is trained using noisy reverberation speech and a clean source signal for reverb. In an exemplary embodiment, a plug-and-play approach is then performed for ASR, with the enhanced time-domain signal directly input to the backend for decoding.
[0160] For the reverberation speaker separation task, the 6-channel Spatialized Multi-Speaker Wall Street Journal (SMS-WSJ) dataset is used. The SMS-WSJ dataset contains simulated two-speaker mixtures in reverberation situations. Clean speech is sampled from the WSJ0 and WSJ1 datasets. The corpus contains 33,561 two-speaker mixtures, 982 two-speaker mixtures, and 1,332 two-speaker mixtures for training, validation, and testing, respectively. The distance between the speaker and the array is sampled from the range [1.0, 2.0] m, and T60 is derived from the range [0.2, 0.5] seconds. Weak white noise is added to simulate microphone noise. The energy level between the sum of the reverberation target speech signals and the noise is sampled from the range [20, 30] dB. The sampling rate is 8 kHz. The first channel of the 6-channel SMS-WSJ dataset is used for training and evaluation. Furthermore, direct sound is used as a training target, and both reverberation removal and separation tasks are performed.
[0161] For ASR, the default Kaldi-based backend acoustic model specified for the SMS-WSJ dataset is used. This model is trained using single-speaker noise reverberation speech as input and the state alignment of its corresponding direct-path signal as labels. Signals in the first, third, and fifth channels (i.e., more signals than the microphone) are used to train the acoustic model. A task-standard trigram language model is used for decoding.
[0162] The noisy reverberation speaker separation task is evaluated using the noisy reverberation WHAMR! (WSJ0 Hipster Ambient Mixtures) dataset. WHAMR! pairs two-speaker mixtures from the wsj0-2mix dataset with the noisy background scenes used for noisy reverberation binaural two-speaker separation. In this evaluation, clean two-speaker mixtures are reused in the WSJ0-2mix dataset, with each clean signal reverberated and transient ambient noise recorded in WHAM! added. Reverberation time T60 is randomly sampled from the range [0.2,1.0] seconds. The signal-to-noise ratio between the louder speaker and the noise is derived from the range [-6,3] dB. The energy level between the two speakers in each mixture is sampled from the range [-5,5] dB. The distance between the speakers and the array is sampled from the range [0.66,2.0] m. The training, validation, and test sets each contain 20,000 binaural mixtures, 5,000 binaural mixtures, and 3,000 binaural mixtures, respectively. The corpus used is the 1-minute and 8kHz versions.
[0163] For STFT, the window length is 32 milliseconds, the hop size is 8 milliseconds, and the analysis window is the square root of the Hann window. A 512-point FFT is applied to extract 257-dimensional STFT features when the sampling rate is 16 kHz, and a 256-point FFT is used to extract 129-dimensional features when the sampling rate is 8 kHz. Mean-variance normalization at the sentence level or global level is not performed on the input features. For each mixture, its sample variance is normalized to 1 before any processing. During training, the target signal should be scaled by the same coefficient used to scale the mixture.
[0164] For WPE and DNN-WPE, the number of filter taps K is set to 37 and the filter delay Δ is set to 3. The number of iterations in WPE is set to 3. No PSD context is used. Based on the validation set, K and Δ were adjusted to 40 and 0, 39 and 1, 38 and 2, 37 and 3, and 36 and 4, of which, setting the filter taps and filter delay to 37 and 3 worked best across the entire dataset. For convolutional prediction, K is set to 40, which gives the same amount of context as in WPE. This means the filter length in the time domain is 344 (=(40-1)×8+32) milliseconds. The filter tap K is increased up to 125, which corresponds to an RIR length up to 1.0 second. This results in an increase in the amount of computation spent on the linear regression step, but does not make a significant difference in terms of evaluation score. The RIR has energy mostly within a range of 0.35 seconds after the peak impulse. The floor value ε used to calculate the reverberation removal result is set to either 1.0, indicating that no weights are used, or 0.001. The PSD at each TF unit will be -30 dB lower than the TF unit with the highest energy.
[0165] For all tasks, the primary evaluation metric is the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR). SI-SDR measures the quality of time-domain sample-level predictions. Extended Short-Time Objective Intelligibility (eSTOI) and Perceptual Evaluation of Speech Quality (PESQ) scores are measured. For PESQ, the narrowband MOS-LQO score is reported using the Python-pesq toolkit, based on the ITU P.862.1 standard. The baseline for metric calculation is the target direct-path signal obtained by setting the reverberation time T60 parameter to zero in the RIR. The Word Error Rate (WER) of the ASR is also shown in tabular representation 1000.
[0166] In tabular format 1000, the target direct path signal is represented by "d", the target direct path signal with early reflections is represented by "d+e", and the target direct path signal with early reflections and noise is represented by "d+e+v".
[0167] As shown in tabular representation 1000, when the first estimate of the first DNN (DNN1) is considered to be the final prediction, the training target of DNN1 shows better performance than the other two (i.e., "d+e" and "d+e+v"). Compared to using various things as the training target of DNN1, there is no significant difference in DNN1-WPE, which improves WPE by applying the DNN1 output. However, it is found that training DNN1 with the target direct path signal shows an improvement in performance in DNN1+DNN2, which stacks two DNNs using an acoustic signal mixture and the output of DNN1, i.e., DNN1+DNN2, which trains a second DNN2 using the first estimate of the target direct path signal.
[0168] Furthermore, tabular representation 1000 includes a comparison of using Inverse Convolutive Prediction (ICP), Forward Convolutive Prediction (FCP), or Weighted Prediction Error (WPE) methods between two DNNs, namely DNN1 and DNN2, and shows that DNN1+FCP+DNN2, with a floor value ε set to 0.001, performs better than DNN1+WPE+DNN2 and DNN1+ICP+DNN2. As shown in tabular representation 1000, by running linear or convolutive prediction and DNN2 one or more times at runtime, DNN1+(WPE+DNN2)×2 and DNN1+(ICP+DNN2)×2 show slight improvements in SI-SDR and PESQ and slight decreases in Word Error Rate (WER). On the other hand, DNN1+(FCP+DNN2)×2 shows improvements across all metrics. These results indicate that the DNN1+FCP+DNN2 approach is more effective than the WPE and DNN1+WPE+DNN2 approaches.
[0169] In DNN1+ICP+DNN2, the SI-SDR and PESQ scores improved by setting the floor value ε to 1.0. The SI-SDR and PESQ scores in DNN1+FCP+DNN2 further improved when the floor value was set to 0.001. For example, with a floor value of 1.0, the SI-SDR score is 11.9 and the PESQ score is 3.15. With a floor value of 0.001, the SI-SDR score is 12.3 and the PESQ score is 3.18. The floor values of 1.0 and 0.001 are also used to evaluate the trained DNN1 using ICP and FCP. As shown in tabular representation 1000, for DNN1+ICP with a floor value of 1.0, the SI-SDR score is 3.2 and the PESQ score is 1.78; for DNN1+ICP with a floor value of 0.001, the SI-SDR score is 0.7 and the PESQ score is 1.77; for DNN1+FCP with a floor value of 1.0, the SI-SDR score is 3.6 and the PESQ score is 1.82; and for DNN1+FCP with a floor value of 0.001, the SI-SDR score is 3.0 and the PESQ score is 1.82. Therefore, DNN1+FCP+DNN2 shows better scores than training DNN1 using the ICP method and the FCP method.
[0170] Overall, for speech reverberation removal, the mixed SI-SDR and PESQ improve from -3.6dB and 1.64 to 8.2dB and 2.65 by using one DNN (i.e., DNN1), to 9.1dB and 2.82 by using two DNNs (i.e., DNN1 + DNN2), to 12.3dB and 3.18 by adding an FCP module between the two DNNs (DNN1 + FCP + DNN2), and to 12.8dB and 3.24 by using one additional repetition of FCP and DNN2 (DNN1 + (FCP + DNN2) × 2).
[0171] Finally, size domain loss is added during the training of the second DNN2. While improvements are obtained in word error rate (WER) and PESQ, the SI-SDR decreases by approximately 0.5 dB.
[0172] Figure 11 shows a tabular representation 1100 of the evaluation results for speech signal reverberation removal using a test dataset according to an embodiment of the present disclosure. The evaluation results show the performance on the SMS-WSJ dataset and the oracle results obtained by using a target direct path signal with or without early reflections, and an oracle mask such as a spectral intensity mask (|S| / |Y|) and a phase-sensitive mask (|S| / |Y|cos(∠S-∠Y)). As shown in tabular representation 1100, using an oracle target direct path signal for ASR yields a better WER than using a target direct path signal with early reflections (6.4% vs. 7.04%), demonstrating the potential benefits of removing early reflections.
[0173]
number
[0174] In the case of DNN-WPE, two variants are used for multi-speaker scenarios. The first variant calculates a different WPE for each speaker using the PSD of each estimated target speaker generated by DNN1. In tabular representation 1100, the DNN-WPE in a multi-speaker scenario is represented as DNN1 + mfWPE + DNN2, where "mf" indicates a multi-filter. The multi-filter sums all estimated target speakers provided by DNN1 and uses the PSD of the summed signals to calculate a single WPE filter to remove reverberation from the mixture. The second variant is represented as DNN1 + sfWPE + DNN2, where "sf" indicates a single filter.
[0175] As shown in tabular representation 1100, DNN1+sfWPE+DNN2 showed slightly better performance than DNN1+mfWPE+DNN2, suggesting that calculating separate filters for each target speaker is not effective for WPE.
[0176] In the scenario where all speakers provide speech signals, represented as "allSpks" in tabular representation 1100, DNN2 is trained to emphasize all target speakers simultaneously. As shown in tabular representation 1100, DNN1+FCP+DNN2 shows superior performance across all metrics compared to DNN1+sfWPE+DNN2 and DNN1+ICP+DNN2. This demonstrates that forward filtering of convolutional predictions (explained in Figures 5 and 6) is more effective than WPE in reverberation removal when competing speakers are present.
[0177] Further improvements are achieved when DNN2 is trained to emphasize each target speaker individually, as explained in Figure 6 (represented as "perSpk" in tabular representation 1100). This suggests that individually de-echoing each speaker can improve speaker speech emphasis. Steady improvements can be achieved by repeating convolutional prediction and DNN2 one or more times, as shown in tabular representation 1100. Additionally, DNN2 trained by including magnitude level loss improves PESQ, eSTOI, and WER, but decreases SI-SDR.
[0178] Tabular representation 1100 further shows that DNN1+(FCP+DNN2)×2 trained with a size-level loss function achieves SI-SDR scores of 12.2, 3.24, 89.0, and 12.77 for PESQ, eSTOI, and WER, respectively. DNN1+(FCP+DNN2)×2 trained with a size-level loss function can perform better than DNN1+(FCP+DNN2)×2 trained with a spectral mapping corresponding to a single microphone, such as a single-input single-output microphone (SISO1), or with a different complex spectral mapping (SI-SDR 12.5dB vs 5.1dB). DNN1+(FCP+DNN2)×2 trained with a size-level loss function can perform better than DNN1+(FCP+DNN2)×2 trained with DPRNN-TasNet (SI-SDR 12.5dB vs 6.5dB).
[0179] Furthermore, tabular representation 1100 shows the performance of DNN1 and DNN2 trained on spectral mapping corresponding to microphone arrays such as a 6-microphone SISO (SISO1-BF-SISO2) with beamforming of the microphone array, which combines monaural complex spectral mapping with beamforming and post-filtering. These results suggest that combining end-to-end DNNs with convolutional prediction may be effective in reducing reverberation in acoustic signal mixtures including speech signals from speakers (e.g., speakers 102A and 102B).
[0180] Figure 12 shows a tabular representation 1200 illustrating evaluation results for speech signal derafter removal using a test dataset, according to some other embodiments of the present disclosure. The tabular representation 1200 shows the SI-SDR for the WHAMR! dataset. As shown in the tabular representation 1200, DNN1+FCP+DNN2 produces better results than DNN1+mfWPE+DNN2 (SI-SDR of 7.4dB vs. 6.8dB). This indicates that DNN-FCP may be more robust than DNN-WPE in derafter removal in the presence of noise and competing speakers.
[0181] Furthermore, tabular representation 1200 shows a comparison with end-to-end speech separation systems such as Wavesplit. DNN1+(FCP+DNN2)×2 achieves an SI-SDR score of 7.5dB, which is higher than Wavesplit's SI-SDR score of 5.9dB. Wavesplit may use speaker identity as secondary information during training for target speaker extraction. DNN1+(FCP+DNN2)×2 does not rely on the availability of speaker identity information. Additionally, dynamic mixing may be applied for data augmentation, which further improves the SI-SDR (7.1dB). DNN1+(FCP+DNN2)×2 may be trained without data augmentation, which performs better than Wavesplit with dynamic mixing.
[0182] Figure 13 is a block diagram showing an audio signal processing system 1300 according to an embodiment of the present disclosure. The audio signal processing system 1300 uses system 200. In some exemplary embodiments, system 200, which has DNNs for de-reverberation of speech signals, such as DNN1206A and DNN2206B, may be implemented on a remote server or within a cloud network. In some embodiments, the audio signal processing system 1300 (hereinafter referred to as system 1300) may receive an RIR model, such as RIR model 316A, to the audio signal processing system 1300. System 1300 may process this RIR model to perform speech analysis for at least one or a combination thereof of room geometry reconstruction, speech enhancement, and de-reverberation of speech signals.
[0183] In some exemplary embodiments, the system 1300 includes one or more sensors 1302, such as an acoustic sensor, which collects data from the environment 1306, including an acoustic signal 1204. The environment 1306 corresponds to environment 100.
[0184] The acoustic signal 1304 may include one or more target direct path signals and their reverberations. For example, the acoustic signal 1304 may include multiple speakers with overlapping speech and their reverberations. Furthermore, the sensor 1302 can convert the acoustic input into the acoustic signal 1304.
[0185] The audio signal processing system 1300 includes a hardware processor 1308 that communicates with computer storage memory, such as memory 1310. Memory 1310 contains stored data, including algorithms, instructions, and other data that can be executed by the hardware processor 1308. Depending on the requirements of a particular application, the hardware processor 1308 may include two or more hardware processors. These two or more hardware processors may be internal or external. The audio signal processing system 1300 may be incorporated into other components of the device, particularly output interfaces and transceivers.
[0186] In some alternative embodiments, the hardware processor 1308 may be connected to a network 1312, which communicates with one or more data sources 1314, a computer device 1316, a mobile phone device 1318, and a storage device 1320. Network 1312 may, in non-limiting examples, include one or more local area networks (LANs) and / or wide area networks (WANs). Network 1312 may also include enterprise-scale computer networks, intranets, and the internet. The voice signal processing system 1300 may include one or more client devices, storage components, and data sources. Each of the one or more client devices, storage components, and data sources may include a single device or multiple devices cooperating in the distributed environment of network 1312.
[0187] In some other alternative embodiments, the hardware processor 1308 may be connected to a network-enabled server 1322 connected to a client device 1324. The hardware processor 1308 may also be connected to an external memory device 1326 and a transmitter 1328. Furthermore, output may be output for each target speaker according to a usage 1330 intended by a particular user. For example, a usage 1330 intended by a particular user may correspond to displaying the speech as text (such as speech commands) on one or more display devices such as a monitor or screen, or entering the text for each target speaker into a computer-related device for further analysis.
[0188] The data source 1314 may include data resources for training DNNs such as DNN1206A and DNN2206B for the speech separation task. For example, in one embodiment, the training data may include acoustic signals from multiple speakers, such as speaker 102A and speaker 102B speaking simultaneously. The training data may also include acoustic signals from a single speaker speaking alone, acoustic signals from one or more speakers speaking in a noisy environment, and acoustic signals from a noisy environment (e.g., environment 100 having reverberation noise signal 110A).
[0189] Furthermore, data source 1314 may include data resources for training DNN1206A and DNN2206B for a speech recognition task. The data provided by data source 1314 may include labeled and unlabeled data, such as transcribed data and untranscribed data. For example, in one embodiment, the data may include one or more sounds and corresponding transcribed information or labels that can be used to initialize a speech recognition task.
[0190] Furthermore, unlabeled data within data source 1314 may be provided by one or more feedback loops. For example, usage data from oral search queries performed on a search engine may be provided as untranscribed data. Other examples of data sources, not as an extension but as examples, may include a variety of oral language audio or image sources, including streaming sound or video, web queries, mobile device camera or audio information, webcam feeds, smart glasses and smartwatch feeds, customer care systems, security camera feeds, web documents, catalogs, user feeds, SMS logs, instant messaging logs, spoken word transcripts, voice commands or captured images (e.g., depth camera images), game system user interactions, tweets, chat or video call recordings, or social networking media. The specific data source 1314 used may be determined based on the application, including whether the data is of a particular class (e.g., data related only to a specific type of sound, including machine systems and entertainment systems) or is substantially general (not class-specific) data.
[0191] Furthermore, the voice signal processing system 1300 may include a third-party device which may consist of any type of computing device, such as an automatic speech recognition (ASR) system on a computing device. For example, the third-party device may include a computer device or a mobile device 1318. The mobile device 1318 may include a personal data assistant (PDA), smartphone, smartwatch, smart glasses (or other wearable smart device), augmented reality headset, virtual reality headset, laptop, tablet, remote control device, entertainment system, vehicle computer system, embedded system controller, appliance, home computer system, security system, consumer electronics, or other similar electronic devices. The mobile device 1318 may also include a microphone or line input terminal for receiving voice information, a camera for receiving video information or image information, or a communication component (e.g., Wi-Fi function) for receiving such information from the internet or another source such as data source 1314. In one exemplary embodiment, the mobile device 1318 may be capable of receiving input data such as voice information and image information. For example, the input data may include speaker queries to the microphone of the mobile device 1318 while multiple speakers are speaking in a room. To determine the content of the query, the input data may be processed by the ASR in the mobile device 1318 using system 200. System 200 enhances the input data by reducing noise in the speaker's environment, isolating the speaker from other speakers, or emphasizing the voice signal of the query, so that the ASR can output an accurate response to the query.
[0192] In some exemplary embodiments, storage 1320 may store information including data, computer instructions (e.g., software program instructions, routines, or services), and / or data related to DNNs such as DNN1206A and DNN2206B of system 200. For example, storage 1320 may store data from one or more data sources 1314, one or more deep neural network models, information for generating and training deep neural network models, and computer-available information output by one or more deep neural network models.
[0193] Figure 14A is a block diagram showing a system 1400A for speech signal reverberation removal according to some exemplary embodiments of the present disclosure. The system 1400A can be used to estimate a target speech signal from an input speech signal 1402 acquired from a sensor 1404 that monitors the environment 1406.
[0194] The input audio signal 1402 includes an acoustic signal mixture containing a target direct path signal (e.g., target direct path signal 106A) and a corresponding reverberation (e.g., reverberation 108A). The system 1400 processes the audio signal 1402 via a processor 1408 using a feature extraction module 1410. The feature extraction module 1410 calculates an audio feature sequence from the input audio signal 1402. A first target direct path signal estimation module 1412 processes the audio feature sequence and outputs a first estimate (e.g., a first estimate 408 of the target direct path signal 106A). The first estimate of the target direct path signal is processed by a filter estimation module 1414 to output a filter that models the room impulse response affecting the target direct path signal. For example, the target direct path signal may affect the target reverberation signal to change. The filter is applied to the first estimate to output a mixture with reduced reverberation. The filter and the first estimate are further processed by a target direct path reverberation reduction mixture estimation module 1416, which estimates a mixture with reduced target direct path reverberation. The mixture with reduced target direct path reverberation, the first estimate, and features are further processed by a second target direct path estimation module 1418 to calculate a signal estimate 1424 (e.g., a second estimate 410) of the target direct path signal. The signal estimate 1424 is output via an output interface 1422. In some embodiments, the room impulse response modeled by the filter may be output via the output interface 1422. The output room impulse response can be used in a speech analysis application to perform one or a combination of room geometry reconstruction, speech enhancement, and speech signal de-reverberation.
[0195] In some exemplary embodiments, the network parameters 1420 may be input to a first target direct path signal estimation module 1412, a filter estimation module 1414, a target direct path reverberation reduction mixture estimation module 1416, and a second target direct path estimation module 1418. The network parameters 1420 may include labeled and unlabeled data, such as transcribed and untranscribed data, for various sounds or utterances that can be used to initialize a speech recognition task.
[0196] Figure 14B is a block diagram showing a system 1400B for de-reverberation of speech signals, according to some other exemplary embodiments of the present disclosure.
[0197] System 1400B includes a processor 1426 configured to execute stored instructions, and memory 1428 that stores instructions for a neural network 1430, including a speech separation network 1432 with reverberation reduction, which enables speech separation and reverberation reduction. The processor 1426 may be a single-core processor, a multi-core processor, a graphics processing unit (GPU), a computing cluster, or any number of other configurations. The memory / storage 1428 may include random access memory (RAM), read-only memory (ROM), flash memory, or other suitable memory systems. The memory 1328 may also include a hard drive, an optical drive, a thumb drive, an array of drives, or any combination thereof. The processor 1426 is connected to one or more input and output interfaces / devices via a bus 1434. Furthermore, System 1400B may include one or more microphones 1438 connected via the bus 1434. System 1400B is configured to receive / acquire speech signals 1456 via one or more microphones 1438, or via network interfaces 1452 and network 1454 connected to the data source of speech signals 1456.
[0198] Memory 1428 stores a neural network 1430 trained to convert an acoustic signal mixture, including a speech signal mixture and corresponding reverberation, into a separated speech signal with reduced reverberation. A processor 1426, executing stored instructions, uses the neural network 1430 retrieved from memory 1428 to perform speech separation. The neural network 1430 is trained to convert an acoustic signal, including a speech signal mixture, into a separated speech signal. The neural network 1430 may include a speech separation network 1432 trained to estimate the separated signal from the acoustic features of the acoustic signal.
[0199] Figure 15 shows a use case 1500 for speech signal reverberation removal according to some exemplary embodiments of the present disclosure. Use case 1500 corresponds to a teleconference room including a group of speakers such as speaker 1502A, speaker 1502B, speaker 1502C, speaker 1502D, speaker 1502E, and speaker 1502F (the group of speakers 1502A-1502F). The speech signals of one or more speakers from the group of speakers 1502A-1502F are received by an audio receiver 1506 of device 1504. The audio receiver 1506 comprises a system 200 that receives acoustic speech signals of one or more speakers from the group of speakers 1502A-1502F.
[0200] The audio receiver 1506 may include a single microphone and / or an array of microphones for receiving a mixture of acoustic signals and noise signals from a group of speakers 1502A–1502F in a teleconference room. These mixtures of acoustic signals from the group of speakers 1502A–1502F can be processed using system 200. For example, system 200 may analyze an RIR model of the teleconference room. This RIR model can be used to generate the room geometry structure of the teleconference room. The room geometry structure can be used to arrange reflection boundaries within the teleconference room. For example, the corresponding room geometry structure can be used to determine speaker placement, seating arrangement for the group of speakers 1502A–1502F, etc., to offset noise and other disturbances within the teleconference room. Furthermore, the RIR model can be used to remove reflections and reverberations of speech signals from one or more speakers within the group of speakers 1502A–1502F.
[0201] In the illustrated exemplary scenario, multiple speakers within the group of speakers 1502A–1502F may simultaneously output speech signals. In such a scenario, system 200 reduces reverberation within the teleconference room to separate the speech signals of each speaker 1502A–1502F. System 200 may also perform beamforming of the acoustic signal mixture from the microphone array to highlight the speech signals of corresponding speakers within the group of speakers 1502A–1502F. The highlighted speech signals can be used for transcription of the speakers' utterances. For example, device 1504 may include an ASR module. The ASR module may receive the highlighted speech signals and output a transcription. The transcription may be displayed on the display screen of device 1504.
[0202] Figure 16 shows a use case 1600 for speech signal reverberation removal according to some other exemplary embodiments of the present disclosure. Use case 1600 corresponds to a factory site including one or more speakers, such as speaker 1602A and speaker 1602B. This factory site may have high reverberation signals and noise due to the operation of various industrial machines. The factory site may also include an audio device 1604 to facilitate communication between a control operator (not shown) of the factory site and one or more speakers 1602A and 1602B within the factory site. The audio device 1604 may comprise a system 200.
[0203] In the illustrated exemplary scenario, audio device 1604 may be transmitting a voice command that can be addressed to person 1602A, who manages the factory floor. This voice command may include "Please report the status of machine 1." Speaker 1602A may utter "Machine 1 is operating." However, the speech signal of speaker 1602A's utterance may be mixed with noise from the machine, noise from the background, and other utterances from speaker 1602B in the background.
[0204] Such noise and reverberation signals can be reduced by system 200. System 200 outputs clean speech from speaker 1602A. This clean speech is input to audio device 1604. Audio device 1604 receives this clean speech and takes a response to a voice command from the clean speech corresponding to speaker 1602A's utterance. System 200 enables the audio device to achieve improved communication with the intended speaker, such as speaker 1602A.
[0205] Figure 17 shows a use case 1700 for speech signal reverberation removal according to some yet other exemplary embodiments of the present disclosure. Use case 1700 corresponds to a driver assistance system 1702. The driver assistance system 1702 is implemented in a vehicle such as a manually operated vehicle, an automated vehicle, or a semi-automated vehicle. The vehicle is occupied by one or more people, such as person 1704A and person 1704B. The driver assistance system 1702 comprises a system 200. For example, the driver assistance system 1702 may be remotely connected to the system 1702 via a network such as network 1754. In some alternative exemplary embodiments, system 200 may be incorporated within the driver assistance system 1702.
[0206] Furthermore, the driver assistance system 1702 may include one or more microphones to receive an acoustic signal mixture. This acoustic signal mixture may include speech signals from people 1704A and 1704B and external noise signals such as the horns of other vehicles. In some cases, when person 1704A is sending a speech command to the driver assistance system 1702, person 1704B may speak louder than person 1704A. The utterance from person 1704B may interfere with person 1704A's speech command. For example, person 1704A's speech command might be "Find the nearest parking lot," while person 1704B's utterance might be "Find a shopping mall to park in." In such cases, the system 200 processes the utterances of person 1704A and person 1704B simultaneously or separately. The system 200 separates the utterances of person 1704A from those of person 1704B. The separated utterances are used by the driver assistance system 1702. The driver assistance system 1702 processes and executes speech commands from person 1704A and utterances from person 1704B, and may output a response to each utterance accordingly.
[0207] Figure 18 shows a use case 1800 for speech signal reverberation removal according to several other exemplary embodiments of the present disclosure. In some exemplary embodiments, systems 200a, 200b, 200c (illustrated in Figures 2A, 2B, and 2C) may process pre-recorded or live recordings of sound to determine estimates of the target direct path signal. Pre-recorded sound data may be accessed from a database via network 1808, which is an example of network 1312. Similarly, live recordings of sources may be streamed from corresponding sources at remote locations via network 1808.
[0208] Estimates of the target direct path signal can be filtered by system 200 to determine an RIR model. This RIR model can be analyzed by an audio signal processing system, such as an audio signal processing system 1300 connected to system 200. The audio signal processing system 1300 can process the RIR model for a room acoustic simulation 1802 of an environment such as a music concert hall 1806. The RIR model can be convolved with a recorded soundtrack source to imprint the acoustics of the music concert hall 1806 based on the room acoustic simulation 1802. Using the room acoustic simulation 1802, a simulated or virtual reality environment of the music concert hall 1806 can be created. The simulated environment of the music concert hall 1806 can allow performers to rehearse before actually performing in the music concert hall 1806.
[0209] In some cases, the room acoustic simulation 1802 can be used to model the room acoustic behavior for the room geometry reconstruction 1804. The room geometry reconstruction 1804 can provide architectural aspects for the design and structure to maximize the listening experience of the audience in a music concert hall, such as a music concert hall 1806.
[0210] In some embodiments, room acoustic simulation 1802 can be used to model room acoustic parameters such as at least one of the direct-to-reverberation ratio, reverberation time (RT60), initial decay time (EDT), center time, clarity C80, and resolution D50. These parameters are well known in the art.
[0211] By incorporating operations 902-920 in the manner described above, methods 900a and 900b, performed using the processor 208 located within system 200, can improve speech signal de-reverberation by enabling the estimation of a filter that includes both reverberation magnitude and phase. Since the filter is estimated based on a convolutional prediction approach, it allows the filter to reduce early reflections of the target direct path signal. Furthermore, since the filter models signal propagation in a room, i.e., RIR, the accuracy of the reverberation estimation can be improved. In addition, the use of two DNNs in system 200 can improve the performance of speech signal de-reverberation, as well as tasks such as speech enhancement and speaker separation. More specifically, the first DNN estimates a first estimate of the target direct path signal from an acoustic signal mixture including reverberation. The second DNN estimates a refined estimate of the target direct path signal using the first estimate along with other data such as the filter and the reverberation reduction estimated by the filter. Thus, the two DNNs enable the identification and differentiation of the target direct path signal from high reverberation and noise in an efficient and feasible manner.
[0212] Furthermore, individual embodiments may be described as processes depicted as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. While flowcharts can describe operations as sequential processes, many operations can be performed in parallel or concurrently. Moreover, the order of operations can be rearranged. A process can terminate when its operations are complete, but it may have additional steps that are not discussed or included in the diagrams. Furthermore, not all operations in a particularly described process occur in all embodiments. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. If a process corresponds to a function, the termination of the function may correspond to the function's return to the calling function or the main function.
[0213] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, manually or automatically. Examples of manual or automatic implementations may be performed, or at least assisted, by the use of a machine, hardware, software, firmware, middleware, microphone code, hardware description language, or any combination thereof. If implemented with software, firmware, middleware, or microphone code, program code or code segments for performing the required tasks may be stored in a machine-readable medium. A processor may perform the required tasks.
[0214] The embodiments described above in this disclosure may be implemented in any of a number of ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. If implemented in software, the software code may run on any suitable processor or array of processors, whether located on a single computer or distributed across multiple computers. Such a processor may be implemented as an integrated circuit, or may comprise one or more processors within an integrated circuit component. However, the processor may be implemented using any suitable form of circuitry.
[0215] The various methods or processes outlined herein can be coded as software executable on one or more processors employing any one of various operating systems or platforms. Furthermore, such software may be written using a number of suitable programming languages and / or programming tools or scripting tools, and may be compiled as executable machine code or intermediate code that runs on a framework or virtual machine. The functions of program modules may typically be combined or distributed as desired in various embodiments.
[0216] Embodiments of this disclosure may also be embodied as methods, and examples of such methods are provided. The actions performed as part of the method can be ordered in any suitable manner. Thus, even though they are shown as sequential actions in the descriptive embodiments, embodiments can be constructed in which the actions are performed in a different order than shown, including performing several actions simultaneously. Accordingly, it is the object of the appended claims to cover all variations and modifications that fall within the true spirit and scope of this disclosure.
[0217] While this disclosure has been described with reference to certain preferred embodiments, it should be understood that various other adaptations and modifications can be made within the spirit and scope of this disclosure. Therefore, it is the nature of the appended claims to cover all such variations and modifications that fall within the true spirit and scope of this disclosure.
Claims
1. A method for estimating a room impulse response (RIR), wherein the method uses a processor coupled with stored instructions for performing the method, and the instructions, when executed by the processor, perform the steps of the method, The steps include receiving an acoustic signal mixture, which includes a target direct path signal propagating within a room and the reverberation of the target direct path signal within the room, via a wired communication channel or a wireless communication channel, The steps include feeding the received acoustic signal mixture into a first deep neural network (DNN) to generate a first estimate of the target direct path signal, The process includes the step of estimating a filter that models the room impulse response (RIR) representing the relationship between the target direct path signal and the reverberation of the target direct path signal, When the filter is applied to the first estimate of the target direct path signal, it produces a result that is closest to at least one of the acoustic signal mixture and the residual between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function, and the method further, The steps include transmitting the filter that models the RIR over the wired communication channel or the wireless communication channel, The further step is to apply the filter that models the RIR to an external signal measured outside the room to synthesize a new acoustic mixture, A method for providing the new acoustic mixture to have an acoustic effect on the propagation of the external signal within the room.
2. The method further comprises the step of training the first DNN to estimate the amplitude and phase of the target direct path signal, Applying the filter that models the RIR to the first estimate of the target direct path signal includes calculations relating to both the estimated amplitude and the phase, The method according to claim 1, wherein the distance function between the result of applying the filter that models the RIR to the first estimate of the target direct path signal and the acoustic signal mixture is measured in both the amplitude domain and the phase domain.
3. The method according to claim 1, further comprising the step of pre-training the first DNN to obtain the first estimate of the target direct path signal from the observed acoustic signal mixture.
4. Pre-training the first DNN is A distance function defined based on the real and imaginary (RI) components of the first estimated value of the target direct path signal in the first time-frequency domain and the RI component of the corresponding reference direct path signal in the first time-frequency domain, A distance function defined based on the magnitude and phase obtained from the RI component of the first estimate of the target direct path signal in the first time-frequency domain, and the corresponding magnitude and phase of the reference direct path signal in the first time-frequency domain, A distance function defined based on the reconstructed waveform obtained from the RI component of the first estimated value of the target direct path signal in the first time-frequency domain, by reconstructing the reference direct path signal in the time domain and waveform, A distance function defined based on the RI component of the first estimate in the second time-frequency domain, obtained by further transforming the waveform reconstructed in the second time-frequency domain, and the RI component of the reference direct path signal in the second time-frequency domain, The method according to claim 3, performed using a training dataset of acoustic signal mixtures and a corresponding reference direct path signal in the training dataset, by minimizing a loss function that includes one or a combination thereof: a distance function defined based on the magnitude obtained from the RI component of the first estimate of the target direct path signal in the second time-frequency domain, obtained by further transforming the waveform reconstructed in the second time-frequency domain, and the corresponding magnitude of the reference direct path signal in the second time-frequency domain.
5. The method according to claim 1, wherein the coefficients of the filter that models the RIR are estimated in the time-frequency domain.
6. The aforementioned target direct path signal is a signal transmitted from the source of the original signal to the sensor. The target direct path signal represents the original signal at the source as measured by the sensor when there is no reverberation. The method according to claim 1, wherein the reverberation of the target direct path signal represents one or more transmissions of the original signal from the source to the sensor along one or more paths longer than the shortest path, such that the RIR is modeled with respect to the target direct path signal.
7. The method according to claim 1, wherein the filter for modeling the RIR is a linear filter estimated using a linear convolutional prediction module.
8. The method according to claim 7, wherein the linear convolution prediction enhances the linear structure of the filter that models the RIR to approximate the linear reflection of the target direct path signal from the surface in the room.
9. The filter that models the RIR is applied to the first estimate of the target direct path signal in the time-frequency domain, The distance function is a weighted distance having weights at each time-frequency point in the time-frequency domain determined by one or a combination of the first estimates of the received acoustic signal mixture and the target direct path signal, The method according to claim 1, wherein the distance function is based on the least squares distance.
10. The method according to claim 1, further comprising the step of updating the first DNN using the new acoustic mixture.
11. The steps include obtaining a mixture from the received acoustic signal mixture in which the reverberation of the target direct path signal has been reduced by removing the result of applying the filter that models the RIR to the first estimate of the target direct path signal, The steps include feeding the mixture with reduced reverberation into a second DNN to generate a second estimate of the target direct path signal, The method according to claim 1, further comprising the step of outputting the second estimated value of the target direct path signal via an output interface.
12. The steps include updating the first estimated value of the target direct path signal based on the application of the filter that models the RIR, The steps include updating the training data based on the acoustic signal mixture and the updated first estimate of the target direct path signal as a pseudo-label, The method according to claim 1, further comprising the step of retraining the first DNN based on the updated training data.
13. The steps include updating the first estimated value of the target direct path signal based on the application of the filter that models the RIR, The steps include updating the training data based on the acoustic signal mixture and the updated first estimate of the target direct path signal as a pseudo-label, The method according to claim 11, further comprising the step of retraining the first DNN based on the updated training data.
14. The step further comprises generating extended data based on the filter that models the RIR, The extended data is obtained by applying the filter that models the RIR to data which includes one or more of the following: external data measured outside the room and data obtained as one or more of the first estimate and second estimate of the target direct path signal of another acoustic signal mixture in the training data, and the method further, The method according to claim 13, further comprising the step of creating an augmented training dataset based on the augmented data.
15. The method according to claim 14, further comprising the step of training a system for automatic speech recognition using the extended training dataset.
16. The method according to claim 14, further comprising the step of training a system for sound event detection using the extended training dataset.
17. The further step involves processing the filter that models the RIR to perform speech analysis for at least one or a combination thereof of room acoustic parameter analysis, room geometry reconstruction, speech enhancement, and speech signal de-reverberation. The method according to claim 1, wherein the room acoustic parameters include at least one of the direct-to-reverberation ratio, reverberation time (RT60), early decay time (EDT), center time, clarity C80, and resolution D50.
18. The received acoustic signal mixture includes speech signals from multiple speakers. The first DNN generates multiple outputs, The method according to claim 1, wherein each output includes the first estimate of the target direct path signal for one of the plurality of speakers.
19. The method according to claim 18, further comprising the step of estimating a corresponding filter for modeling the corresponding RIR for each speaker among the plurality of speakers.
20. The step further comprises processing the filter that models the RIR, using together the first DNN, a downstream DNN trained to perform a task based on the RIR, The method according to claim 1, wherein the task includes one or a combination of speaker identification, sentiment recognition, event detection, and localization.
21. A system for estimating the room impulse response (RIR), An input interface configured to receive an acoustic signal mixture, including a target direct path signal propagating within a room and the reverberation of the target direct path signal within the room, via a wired communication channel or a wireless communication channel, Memory for storing the first Deep Neural Network (DNN), The system comprises a processor, and the processor is The received acoustic signal mixture is fed into the first DNN to generate a first estimate of the target direct path signal. The system is configured to estimate a filter that models the RIR representing the relationship between the target direct path signal and the reverberation of the target direct path signal, When the filter is applied to the first estimate of the target direct path signal, it produces a result that is closest to the residual between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function, and the system further, The filter that models the RIR is provided with an output interface configured to output over the wired communication channel or the wireless communication channel, The processor further applies the filter that models the RIR to an external signal measured outside the room to synthesize a new acoustic mixture. The system provides the new acoustic mixture for the acoustic effect of the propagation of the external signal within the room.
22. A method for estimating an indoor impulse response (RIR), the method using a processor coupled with stored instructions for performing the method, wherein the instructions, when executed by the processor, perform the steps of the method, The steps include receiving an acoustic signal mixture, which includes a target direct path signal propagating within a room and the reverberation of the target direct path signal within the room, via a wired communication channel or a wireless communication channel, The steps include feeding the received acoustic signal mixture into a first deep neural network (DNN) to generate a first estimate of the target direct path signal, The process includes the step of estimating a filter that models the room impulse response (RIR) representing the relationship between the target direct path signal and the reverberation of the target direct path signal, When the filter is applied to the first estimate of the target direct path signal, it produces a result that is closest to at least one of the acoustic signal mixture and the residual between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function, and the method further, The steps include transmitting the filter that models the RIR over the wired communication channel or the wireless communication channel, The steps include updating the first estimated value of the target direct path signal based on the application of the filter that models the RIR, The steps include updating the training data based on the acoustic signal mixture and the updated first estimate of the target direct path signal as a pseudo-label, A method comprising the step of retraining the first DNN based on the updated training data.
23. A method for estimating an indoor impulse response (RIR), the method using a processor coupled with stored instructions for performing the method, wherein the instructions, when executed by the processor, perform the steps of the method, The steps include receiving an acoustic signal mixture, which includes a target direct path signal propagating within a room and the reverberation of the target direct path signal within the room, via a wired communication channel or a wireless communication channel, The steps include feeding the received acoustic signal mixture into a first deep neural network (DNN) to generate a first estimate of the target direct path signal, The process includes the step of estimating a filter that models the room impulse response (RIR) representing the relationship between the target direct path signal and the reverberation of the target direct path signal, When the filter is applied to the first estimate of the target direct path signal, it produces a result that is closest to at least one of the acoustic signal mixture and the residual between the acoustic signal mixture and the first estimate of the target direct path signal, according to a distance function, and the method further, The steps include transmitting the filter that models the RIR over the wired communication channel or the wireless communication channel, The step of processing the filter that models the RIR, using together the first DNN, a downstream DNN trained to perform a task based on the RIR, The task is a method comprising one or a combination thereof of speaker identification, sentiment recognition, event detection, and localization.