Data augmentation for speech enhancement

Low-complexity machine learning models with augmented training sets using synthesized acoustic impulse responses effectively address dereverberation challenges, enhancing audio quality and intelligibility while reducing computational intensity and overfitting.

JP2026035611APending Publication Date: 2026-03-04DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing audio enhancement techniques introduce undesirable perceptual distortions such as changes in loudness or timbre, and training complex machine learning models for dereverberation is computationally intensive and prone to overfitting due to limited training sets.

Method used

Utilize low-complexity machine learning models combined with augmented training sets generated by modifying real acoustic impulse responses to create synthesized AIRs, incorporating convolutional neural networks with recurrent elements and loss functions to generate smooth enhancement masks.

Benefits of technology

Achieves efficient dereverberation with improved speech intelligibility and preserved perceptual quality by using a computationally efficient, robust machine learning model trained on a larger, more representative dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035611000001_ABST
    Figure 2026035611000001_ABST
Patent Text Reader

Abstract

To provide a method of dereverberating an audio signal to improve the intelligibility and clarity of speech.SOLUTION: The method includes obtaining an actual acoustic impulse response (AIR), identifying a first portion of the actual AIR corresponding to an early reflection of a direct sound and a second portion of the actual AIR corresponding to a late reflection of the direct sound, generating one or more synthesized AIRs by modifying the first portion of the actual AIR and / or the second portion of the actual AIR, and generating a plurality of training samples for training a machine learning model using the actual AIR and the one or more synthesized AIRs. Each training sample includes an input audio signal and a reverberated audio signal, wherein the reverberated audio signal is generated based on the input audio signal and one or more synthesized AIRs of a real AIR.SELECTED DRAWING: FIG. 5A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 260201, filed August 12, 2021, and International Application No. PCT / CN2021 / 106536, filed July 15, 2021, the contents of which are incorporated herein in their entireties.

[0002] Technical Field SUMMARY The present disclosure relates to systems, methods, and media for speech enhancement by attenuating distortion. [Background technology]

[0003] Audio devices such as headphones and speakers are widely deployed. People often listen to audio content (e.g., podcasts, radio programs, television programs, music videos, user-generated content, short videos, video conferences, teleconferences, panel discussions, interviews, etc.) that may contain distortions such as reverberation and / or noise. Furthermore, audio content may include far-field audio content such as background noise. Enhancements such as dereverberation and / or noise suppression may be performed on such audio content. However, enhancement techniques may introduce undesirable perceptual distortions such as changes in loudness or timbre.

[0004] Notation and Nomenclature Throughout this disclosure, including the claims, the terms "speaker," "loudspeaker," and "audio reproduction transducer" are used interchangeably to refer to any sound-emitting transducer (or set of transducers). A typical headphone set includes two speakers. A speaker may be implemented to include multiple transducers (e.g., woofers and tweeters) that may be driven by a single common speaker feed or multiple speaker feeds. In some examples, the speaker feeds may undergo different processing in different circuit branches coupled to different transducers.

[0005] Throughout this disclosure, including the claims, the expression "performing an operation on a signal or data" (e.g., filtering, scaling, transforming, or applying a gain to the signal or data) is used broadly to refer to performing an operation on the signal or data directly, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone pre-filtering or pre-processing before performing the operation).

[0006] Throughout this disclosure, including the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system that includes such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, where the subsystem generates M of the inputs and the other XM inputs are received from external sources) may also be referred to as a decoder system.

[0007] Throughout this disclosure, including the claims, the term "processor" is used broadly to refer to a system or device that is programmable or configurable (e.g., with software or firmware) to perform operations on data (e.g., audio or video or other image data). Examples of processors include field programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipeline processing on audio or other sound data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets. Summary of the Invention [Means for solving the problem]

[0008] At least some aspects of the present disclosure may be implemented via methods. Some methods may include, by a control system, acquiring a real acoustic impulse response (AIR). Some methods may include, by the control system, identifying a first portion of the real AIR corresponding to early reflections of direct sound and a second portion of the real AIR corresponding to late reflections of the direct sound. Some methods may include, by the control system, generating one or more synthesized AIRs by modifying the first portion of the real AIR and / or the second portion of the real AIR. Some methods may include, by the control system, generating a plurality of training samples using the real AIR and the one or more synthesized AIRs, each training sample including an input audio signal and a reverberant audio signal, the reverberant audio signal being generated at least in part based on the input audio signal and one of the real AIRs or one of the one or more synthesized AIRs, the plurality of training samples being used to train a machine learning model that takes a reverberant test audio signal as input and generates a dereverberated audio signal as output.

[0009] In some examples, identifying the first portion of the real AIR corresponding to an early reflection and the second portion of the real AIR corresponding to a late reflection includes selecting a random time value within a predetermined range, wherein the first portion includes a portion of the real AIR before the random time value and the second portion includes a portion of the real AIR after the random time value. In some examples, the predetermined range is from about 20 milliseconds to about 80 milliseconds.

[0010] In some examples, modifying the second portion of the actual AIR includes truncating the second portion of the actual AIR after a duration randomly selected from a predetermined range of late reflection durations.

[0011] In some examples, modifying the second portion of the real AIR includes modifying amplitudes of one or more responses included in the second portion of the real AIR. In some examples, modifying the amplitudes of the one or more responses included in the second portion of the real AIR includes: determining a target decay function associated with the second portion of the real AIR; and modifying the amplitudes of the one or more responses included in the second portion of the real AIR according to the target decay function.

[0012] In some examples, the reverberant audio signal is generated by convolving the input audio signal with the one of the real AIRs or the one of the one or more synthesized AIRs.

[0013] In some examples, the methods may further include adding noise to a convolution of the input audio signal with the one of the real AIRs or the one of the one or more synthesized AIRs to generate the reverberant audio signal.

[0014] In some examples, the methods may further include identifying an updated first portion of the real AIR and an updated second portion of the real AIR; and generating additional synthesized AIR by modifying the updated first portion of the real AIR and / or the updated second portion of the real AIR.

[0015] In some examples, the methods may further include providing the plurality of training samples to the machine learning model to generate a machine learning model that takes the test audio signal with reverberation as the input and generates the dereverberated audio signal as the output. In some examples, the test audio signal is a live-captured audio signal.

[0016] In some examples, the actual AIR is the measured AIR measured in a physical room.

[0017] The real AIR is generated using a room acoustic model.

[0018] In some examples, the input audio signal is associated with a particular audio content type. In some examples, the particular audio content type includes far-field noise. In some examples, the methods further include obtaining a training set of input audio signals, each associated with a particular audio content type, prior to generating the plurality of training samples.

[0019] Some or all of the operations, functions, and / or methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. Thus, some innovative aspects of the subject matter described in this disclosure may be implemented via one or more non-transitory media that store software.

[0020] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of at least partially performing the methods disclosed herein. In some implementations, the apparatus is or includes an audio processing system having an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof.

[0021] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. It should be noted that the relative dimensions of the following figures may not be drawn to scale. [Brief explanation of the drawings]

[0022] [Figure 1] 1 illustrates examples of audio signals in the time and frequency domains according to some implementations.

[0023] [Figure 2] 1 shows a block diagram of an example system for performing dereverberation of an audio signal, according to some implementations.

[0024] [Figure 3] 1 illustrates an example process for performing dereverberation of an audio signal, according to some implementations.

[0025] [Figure 4A] An example of an acoustic impulse response (AIR) is shown below. [Figure 4B] An example of an acoustic impulse response (AIR) is shown below.

[0026] [Figure 5A] 1 illustrates an example process for generating synthesized AIR according to some implementations.

[0027] [Figure 5B] Here is an example of a process for generating a training set using synthesized AIR, according to some implementations.

[0028] [Figure 6] 1 illustrates an example architecture of a machine learning model for dereverberating an audio signal, according to some implementations.

[0029] [Figure 7] 1 illustrates an exemplary process for training a machine learning model for dereverberating an audio signal, according to some implementations.

[0030] [Figure 8] 1 shows a block diagram of an example system for performing dereverberation of an audio signal, according to some implementations.

[0031] [Figure 9] FIG. 1 shows a block diagram illustrating example components of an apparatus capable of implementing various aspects of the present disclosure.

[0032] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION

[0033] An audio signal may contain various types of distortions, such as noise and / or reverberation. For example, reverberation occurs when an audio signal is distorted by various reflections from various surfaces (e.g., walls, ceilings, floors, furniture, etc.). Reverberation can have a substantial impact on sound quality and speech intelligibility. Thus, dereverberation of an audio signal may be performed, for example, to improve speech intelligibility and clarity.

[0034] Sound arriving at a receiver (e.g., a human listener, a microphone, etc.) consists of direct sound, which includes sound coming directly from the sound source without reflection, and reverberant sound, which includes sound reflected from various surfaces in the environment. Reverberant sound includes early and late reflections. Early reflections can arrive at the receiver immediately after or simultaneously with the direct sound and thus may be partially integrated into the direct sound. The integration of direct sound and early reflections produces a spectral coloring effect that contributes to the perceived sound quality. Late reflections arrive at the receiver after the early reflections (e.g., more than 50-80 milliseconds after the direct sound). Late reflections can have a detrimental effect on speech intelligibility. Therefore, dereverberation can be performed on an audio signal to reduce the effects of late reflections present in the audio signal, thereby improving speech intelligibility.

[0035] 1 shows an example of a time-domain input audio signal 100 and a corresponding spectrogram 102. As shown in spectrogram 102, early reflections can cause changes in spectrogram 104, as indicated by spectral coloration 106. Spectrogram 104 also shows late reflections 108, which can adversely affect speech intelligibility.

[0036] Performing enhancements (e.g., dereverberation and / or noise suppression) on an audio signal such that speech intelligibility is improved by the enhancements and the perceptual quality of the audio signal is preserved can be challenging. For example, machine learning models such as deep neural networks can be used to predict a dereverberation mask that, when applied to a reverberant audio signal, produces a dereverberated audio signal. However, training such machine learning models can be computationally intensive and inefficient. For example, such machine learning models may require a high degree of complexity to achieve a certain level of accuracy. As a more specific example, such machine learning models may include a large number of layers, thereby requiring a correspondingly large number of parameters to be optimized. Furthermore, such complex machine learning models may be prone to overfitting due to training on a limited training set and a large number of parameters to be optimized. In such cases, such machine learning models may be computationally intensive to train and may ultimately achieve poorer performance.

[0037] Disclosed herein are methods, systems, media, and techniques for enhancing audio signals using low-complexity machine learning models and / or using an augmented training set. As described herein (e.g., with respect to FIGS. 4A, 4B, 5A, and 5B), the augmented training set may be generated by generating a synthesized acoustic impulse response (AIR). The augmented training set may be able to better span potential combinations of room environment, noise, speaker type, etc., thereby allowing the machine learning model to be trained using a larger, more representative training set, thereby mitigating the problem of model overfitting. Furthermore, as described herein, low-complexity machine learning models may be used that utilize convolutional neural networks (CNNs) with relatively few layers (and therefore relatively few parameters to be optimized) in combination with recurrent elements. By combining a CNN with a recurrent element in parallel (e.g., as shown and described below in connection with FIG. 6), a low-complexity machine learning model can be trained that generates smooth enhancement masks in a computationally efficient manner. In particular, the recurrent element can inform portions of the CNN of the audio signal to be used in subsequent iterations of training, thereby resulting in a smoother predicted enhancement mask. Examples of recurrent elements that can be used include gated recurrent units (GRUs), long short-term memory (LSTM) networks, Elman recurrent neural networks (RNNs), and / or any other suitable recurrent elements. Additionally, loss functions are described herein that allow the machine learning model to both generate a predicted enhanced audio signal that is accurate with respect to a signal of interest in the input distorted audio signal and optimize to minimize the amount of reverberation in the predicted clean audio signal.In particular, such a loss function, as described in more detail in connection with FIG. 7, may incorporate a parameter that approximates the degree of reverberation in the predicted clean audio signal, thereby allowing a machine learning model to be trained based on the final parameter of interest, i.e., whether the output signal is substantially dereverberated compared to the input signal.

[0038] In some implementations, the input audio signal may be enhanced using a trained machine learning model. In some implementations, the input audio signal may be transformed into the frequency domain by extracting frequency-domain features. In some implementations, a perceptual transform based on processing by the human cochlea may be applied to the frequency-domain representation to obtain banded features. Examples of perceptual transforms that may be applied to the frequency-domain representation include a gamma tone filter, an equal-rectangular bandwidth filter, a mel-scale-based transform, etc. In some implementations, the frequency-domain representation may be provided as input to a trained machine learning model that generates a predicted enhancement mask as output. The predicted enhancement mask may be a frequency-domain representation of a mask that, when applied to the frequency-domain representation of the input audio signal, produces an enhanced audio signal. In some implementations, the inverse of the perceptual transform may be applied to the predicted enhancement mask to generate a modified predicted enhancement mask. A frequency-domain representation of the enhanced audio signal may then be generated by multiplying the frequency-domain representation of the input audio signal by the modified predicted enhancement mask. The enhanced audio signal may then be generated by transforming the frequency-domain representation of the enhanced audio signal into the time domain.

[0039] In other words, the trained machine learning model for enhancing an audio signal may be trained to generate, for a given frequency-domain input audio signal, a predicted enhancement mask that, when applied to the frequency-domain input audio signal, generates a frequency-domain representation of a corresponding enhanced audio signal. In some implementations, the predicted enhancement mask may be applied to the frequency-domain representation of the input audio signal by multiplying the frequency-domain representation of the input audio signal with the predicted enhancement mask. Alternatively, in some implementations, the logarithm of the frequency-domain representation of the input audio signal may be taken. In such implementations, the frequency-domain representation of the enhanced audio signal may be obtained by subtracting the logarithm of the predicted enhancement mask from the logarithm of the frequency-domain representation of the enhanced audio signal.

[0040] Note that in some implementations, training the machine learning model may include determining weights associated with one or more nodes and / or connections between nodes of the machine learning model. In some implementations, the machine learning model may be trained on a first device (e.g., a server, a desktop computer, a laptop computer, etc.). Once trained, weights associated with the trained machine learning model may then be provided (e.g., transmitted) to a second device (e.g., a server, a desktop computer, a laptop computer, a media device, a smart television, a mobile device, a wearable computer, etc.) for use by the second device in dereverberating the audio signal.

[0041] 2 and 3 show examples of systems and techniques for dereverberating an audio signal. While Figures 2 and 3 describe dereverberating an audio signal, it should be noted that the systems and techniques described in connection with Figures 2 and 3 may be applied to other types of enhancement, such as noise suppression, a combination of noise suppression and dereverberation, etc. In other words, rather than generating a predicted dereverberation mask and a predicted dereverberated audio signal, in some implementations a predicted enhancement mask may be generated, and the predicted enhancement mask may be used to generate a predicted enhanced audio signal, which is a denoised and / or dereverberated version of the distorted input audio signal.

[0042] FIG. 2 illustrates an example system 200 for dereverberating an audio signal, according to some implementations. As shown, a dereverberated audio component 206 takes an input audio signal 202 as an input and generates a dereverberated audio signal 204 as an output. In some implementations, the dereverberated audio component 206 includes a feature extractor 208. The feature extractor 208 may generate a frequency-domain representation of the input audio signal 202, which may be considered an input signal spectrum. The input signal spectrum may then be provided to a trained machine learning model 210. The trained machine learning model 210 may generate a predicted dereverberation mask as an output. The predicted dereverberation mask may be provided to a dereverberated signal spectrum generator 212. The dereverberated signal spectrum generator 212 may apply the predicted dereverberation mask to the input signal spectrum to generate a dereverberated signal spectrum (e.g., a frequency-domain representation of a dereverberated audio signal). The dereverberated signal spectrum may then be provided to a time domain transform component 214. The time domain transform component 214 may generate a dereverberated audio signal 204.

[0043] FIG. 3 shows an example process 300 for dereverberating an audio signal according to some implementations. In some implementations, the system shown in FIG. 2 and described above in connection therewith may implement the blocks of process 300 to generate a dereverberated audio signal. In some implementations, the blocks of process 300 may be implemented by a user device, such as a mobile phone, a tablet computer, a laptop computer, a wearable computer (e.g., a smart watch), a desktop computer, a game console, a smart television, etc. In some implementations, the blocks of process 300 may be performed in an order not shown in FIG. 3. In some implementations, one or more blocks of process 300 may be omitted. In some implementations, two or more blocks of process 300 may be performed substantially in parallel.

[0044] Process 300 may begin at 302 by receiving an input audio signal including reverberation. The input audio signal may be a live-captured audio signal, such as live-streamed content, an audio signal corresponding to an ongoing video or audio conference, etc. In some implementations, the input audio signal may be a pre-recorded audio signal, such as an audio signal associated with pre-recorded audio content (e.g., television content, video, movies, podcasts, etc.). In some implementations, the input audio signal may be received by a microphone of a user device. In some implementations, the input audio signal may be transmitted to the user device, such as from a server device, another user device, etc.

[0045] At 304, process 300 may extract features of the input audio signal by generating a frequency-domain representation of the input audio signal. For example, process 300 may generate the frequency-domain representation of the input audio signal using a transform such as a short-time Fourier transform (STFT), a modified discrete cosine transform (MDCT), or the like. In some implementations, the frequency-domain representation of the input audio signal is referred to herein as the “binned features” of the input audio signal. In some implementations, the frequency-domain representation of the input audio signal may be modified by applying a perceptually-based transform that mimics the filtering of the human cochlea. Examples of perceptually-based transforms include a gamma tone filter, an equivalent rectangular bandwidth filter, a Mel-scale filter, etc. The modified frequency-domain transform may be referred to herein as the “banded features” of the input audio signal.

[0046] At 306, process 300 can provide the extracted features (e.g., a frequency-domain representation of the input audio signal or a modified frequency-domain representation of the input audio signal) to a trained machine learning model. The machine learning model may be trained to generate a dereverberation mask that, when applied to the band representation of the input audio signal, generates a frequency-domain representation of a dereverberated audio signal. In some implementations, logarithms of the extracted features may be provided to the trained machine learning model.

[0047] A machine learning model may have any suitable architecture or topology. For example, in some implementations, a machine learning model may be or include a deep neural network, a convolutional neural network (CNN), a long short-term memory (LSTM) network, a recurrent neural network (RNN), etc. In some implementations, a machine learning model may combine two or more types of networks. For example, in some implementations, a machine learning model may combine a CNN with a recurrent element. Examples of recurrent elements that may be used include a GRU, an LSTM network, an Elman RNN, etc. An example machine learning model architecture that combines a CNN with a GRU is shown in and described below in connection with FIG. 6. Note that a technique for training a machine learning model is shown in and described below in connection with FIG. 7.

[0048] At 308, process 300 can obtain a predicted dereverberation mask from the output of the trained machine learning model. The predicted dereverberation mask, when applied to the frequency-domain representation of the input audio signal, produces a frequency-domain representation of the dereverberated audio signal. In some implementations, process 300 can modify the predicted dereverberation mask by applying an inverse perceptually based transform, such as an inverse gamma tone filter or an inverse equivalent rectangular bandwidth filter.

[0049] At 310, process 300 may generate a frequency-domain representation of the dereverberated audio signal based on the predicted dereverberation mask generated by the trained machine learning model and the frequency-domain representation of the input audio signal. For example, in some implementations, process 300 may multiply the predicted dereverberation mask by the frequency-domain representation of the input audio signal. If the logarithm of the frequency-domain representation of the input audio signal is provided to the trained machine learning model, process 300 may generate the frequency-domain representation of the dereverberated audio signal by subtracting the logarithm of the predicted reverberation mask from the logarithm of the frequency-domain representation of the input audio signal. Continuing with this example, process 300 may then raise the power of the difference between the logarithm of the predicted reverberation mask and the logarithm of the frequency-domain representation of the input audio signal to obtain the frequency-domain representation of the dereverberated audio signal.

[0050] At 312, process 300 may generate a time-domain representation of the dereverberated audio signal. For example, in some implementations, process 300 may generate the time-domain representation of the dereverberated audio signal by applying an inverse transform (e.g., an inverse STFT, an inverse MDCT, etc.) to the frequency-domain representation of the dereverberated audio signal.

[0051] The process 300 may end at 314.

[0052] In some implementations, after generating the time-domain representation of the dereverberated audio signal, the dereverberated audio signal may be played or presented (e.g., by one or more speaker devices of the user device). In some implementations, the dereverberated audio signal may be stored in local memory of the user device, etc. In some implementations, the dereverberated audio signal may be transmitted to another user device for presentation by the user device, to a server for storage, or the like.

[0053] In some implementations, a machine learning model for dereverberating an audio signal may be trained using a training set. The training set may include any suitable number of training samples (e.g., 100 training samples, 1000 training samples, 10,000 training samples, etc.), where each training sample includes a clean (e.g., unreverberated) audio signal and a corresponding reverberated audio signal. As described above with respect to FIGS. 2 and 3, the machine learning model may be trained using the training set to generate a predicted dereverberation mask that, when applied to a particular reverberant audio signal, generates a predicted dereverberated audio signal.

[0054] Training a machine learning model that can robustly generate predicted dereverberation masks for different reverberant audio signals can depend on the quality of the training set. For example, for a machine learning model to be robust, the training set may need to capture reverberation from a vast number of different room types (e.g., rooms with different sizes, layouts, furniture, etc.), a vast number of different speakers, etc. Collecting such a training set can be challenging. For example, a training set may be generated by applying various AIRs, each characterizing room reverberation, to a clean audio signal, thereby generating pairs of clean audio signals and corresponding reverberant audio signals generated by convolving the AIRs with the clean audio signal. However, the number of available real AIRs may be limited, and available real AIRs may not fully characterize potential reverberation effects (e.g., by not capturing enough rooms with different dimensions, layouts, etc.).

[0055] Disclosed herein are techniques for generating an augmented training set that can be used to train a robust machine learning model for dereverberating audio signals. In some implementations, real AIR is used to generate a set of synthesized AIR. The synthesized AIR may be generated by altering and / or modifying various characteristics of early and / or late reflections of measured AIR, as shown and described below in connection with FIGS. 4A, 4B, and 5A. In some implementations, real AIR may be measured AIR measured in a room environment (e.g., using one or more microphones placed in the room). Alternatively, in some implementations, real AIR may be modeled AIR generated using a room acoustics model that incorporates, for example, the shape of the room, materials in the room, the layout of the room, objects in the room (e.g., furniture), and / or any combination thereof. In contrast, synthesized AIR may be AIR generated based on real AIR (e.g., by modifying the components and / or characteristics of real AIR), regardless of whether the real AIR is measured or generated using a room acoustics model. In other words, real AIR may be considered a starting point for generating one or more synthesized AIRs. A technique for generating synthetic AIR is shown in and described below in connection with FIG. 5A. The real AIR and / or the synthesized AIR can then be used to generate a training set including training samples generated based on the real AIR and the synthesized AIR, as shown and described below in connection with FIG. 5B. For example, the training samples may include a clean audio signal and a corresponding reverberant audio signal generated by convolving the synthesized AIR with the clean audio signal.Because many synthetic AIRs can be generated from a single real AIR, and because multiple reverberant audio signals can be generated from a single clean audio signal and a single AIR (whether measured or synthesized), the augmented training set can contain a larger number of training samples that better capture the potential expansion of reverberation effects, thereby resulting in a more robust machine learning model when trained with the augmented training set.

[0056] FIG. 4A shows an example of AIR measured in a reverberant environment. As shown, early reflections 402 may arrive at the recipient simultaneously with or shortly after direct sound 406. In contrast, late reflections 404 may arrive at the recipient after the early reflections 402. The late reflections 404 are associated with a duration 408, which may be on the order of 100 milliseconds, 0.5 seconds, 1 second, 1.5 seconds, etc. The late reflections 404 are also associated with a decay 410, which characterizes how the amplitude of the late reflections 404 decays or decreases over time. In some examples, the decay 410 may be characterized as an exponential decay, a linear function, a portion of a polynomial function, etc. The boundary between early and late reflections may be in the range of approximately 50 milliseconds to 80 milliseconds.

[0057] FIG. 4B shows a schematic diagram of how the AIR shown in FIG. 4A may be modified to generate synthetic AIR. In some implementations, the time of the components of the early reflections 402 may be modified. For example, as shown in FIG. 4B, the time of the early reflection component 456 in the synthetic AIR may be modified, e.g., to be earlier or later than the time of the early reflection component in the measured AIR. In some implementations, the duration of the late reflections may be modified. For example, with reference to the synthetic AIR shown in FIG. 4B, the duration 458 is truncated relative to the corresponding duration 408 of the measured AIR. In some implementations, the shape of the decay of the late reflections may be modified in the synthesized AIR. For example, with reference to the synthetic AIR shown in FIG. 4B, the decay 458 is steeper than the corresponding decay 408 of the measured AIR, causing the late reflection component in the synthetic AIR to decay more than in the measured AIR.

[0058] FIG. 5A illustrates an example process 500 for generating one or more synthetic AIRs from a single real AIR. In some implementations, the blocks of process 500 may be implemented by a device, such as a server, desktop computer, or laptop computer, that generates an augmented training set for training a machine learning model for dereverberating an audio signal. In some implementations, two or more blocks of process 500 may be performed substantially in parallel. In some implementations, the blocks of process 500 may be performed in an order not shown in FIG. 5A. In some implementations, one or more blocks of process 500 may be omitted.

[0059] Process 500 can begin at 502 by obtaining AIR. The AIR can be real AIR. For example, the AIR can be measured using a set of microphones in a reverberant room environment. As another example, the AIR can be AIR generated using a room acoustic model. The AIR can be obtained from any suitable source, such as a database that stores measured AIR.

[0060] At 504, process 500 may identify a first portion of AIR corresponding to early reflections of the direct sound and a second portion of AIR corresponding to late reflections of the direct sound. In some implementations, process 500 may identify the first and second portions by identifying a separation boundary between early and late reflections in AIR. The separation boundary may correspond to a point in the AIR that divides the AIR into early and late reflections. In some implementations, the separation boundary may be identified by selecting a random value from within a predetermined range. Examples of the predetermined range include 15 ms to 85 ms, 20 ms to 80 ms, 30 ms to 70 ms, etc. In some implementations, the separation boundary may be a random value selected from any suitable distribution (e.g., uniform distribution, normal distribution, etc.) corresponding to the predetermined range.

[0061] At 506, process 500 may generate one or more synthetic AIRs by modifying portions of the early reflections and / or late reflections of AIR. In some implementations, early reflections and late reflections may be identified within AIR based on the separation boundary identified in block 504. In some implementations, process 500 may generate synthetic AIRs by modifying portions of the early reflections of AIR. For example, as shown in connection with FIG. 4B and described above, process 500 may modify the time point of one or more components of the early reflections. In some implementations, process 500 may modify the order of one or more components of the early reflections. For example, in some implementations, process 500 may modify the order of one or more components of the early reflections so that one or more components of the early reflections have different time points within the early reflection portion of AIR. In some implementations, the components of the early reflection portion of AIR may be randomized.

[0062] In some implementations, process 500 can generate synthetic AIR by modifying portions of late reflections in AIR. For example, as illustrated in connection with FIG. 4B and described above, process 500 may modify the duration of late reflections in synthesized AIR by randomly selecting a duration after which late reflections should be truncated from a predetermined range. In some implementations, the predetermined range may be determined based on a point in time (e.g., a separation boundary) separating the first portion of AIR and the second portion of AIR identified in block 502. For example, in some implementations, late reflections may be truncated at a randomly selected duration selected from a range such as from the separation boundary to 1 second, from the separation boundary to 1.5 seconds, etc.

[0063] As another example, in some implementations, process 500 may generate synthetic AIR by modifying the decay associated with late reflections. As a more specific example, in some implementations, process 500 may generate a decay function (e.g., an exponential decay function, a linear decay, etc.). Continuing with this more specific example, process 500 may then modify the amplitude of the components of the late reflections according to the generated decay function. In some implementations, this may result in the synthesized AIR having late reflection components that are attenuated relative to the corresponding late reflection components of the measured AIR. Conversely, in some implementations, this may result in the synthesized AIR having late reflection components that are amplified or boosted relative to the corresponding late reflection components of the measured AIR. Modifying the decay associated with late reflections may change the reverberation time (RT), such as the time for reverberation to decrease by 60 dB (e.g., RT).

[0064] It should be noted that in some implementations, the synthesized AIR may include modifications to both the early and late reflection components. Furthermore, in some implementations, the early and / or late reflection components may be modified in multiple ways in the synthesized AIR relative to the real AIR. For example, in some implementations, the synthesized AIR may include both truncated late reflections and late reflection components whose amplitudes have been modified based at least in part on the modified attenuation applied to the late reflections in the synthesized AIR.

[0065] Additionally, in some implementations, the synthesized AIR may be further modified, for example, in post-processing. For example, in some implementations, the direct-to-reverberant ratio (DRR) associated with the synthesized AIR may be modified. As a more specific example, in some implementations, the DRR associated with the synthesized AIR may be modified by applying a gain to a portion (e.g., an early reflection portion of the synthesized AIR) to increase or decrease the DRR. In some implementations, multiple modified synthetic AIRs may be generated from a single synthesized AIR. For example, in some implementations, multiple modified synthetic AIRs may be generated by applying different gains to a single synthesized AIR, each corresponding to a different modified synthetic AIR.

[0066] At 508, process 500 may determine whether additional synthesized AIR should be generated based on the AIR obtained in block 502. In some implementations, process 500 may determine whether additional synthesized AIR should be generated based on whether a target or threshold number of synthesized AIRs to be generated from the AIRs have been generated. For example, if N synthesized AIRs are to be generated from a particular AIR, process 500 may determine whether N synthesized AIRs have been generated from the AIR obtained in block 502. Note that N may be any suitable value, such as 1, 5, 10, 20, 50, 100, 500, 1000, 2000, etc.

[0067] If, at 508, process 500 determines that additional composite AIR should not be generated (“no” at block 508), process 500 may end at 510. Conversely, if, at block 508, process 500 determines that additional composite AIR should be generated (“yes” at block 508), process 500 may loop back to block 504 and identify different first portions of AIR and second portions of AIR obtained at block 502. By looping through blocks 504-508, process 500 may generate multiple composite AIRs from a single measured AIR.

[0068] FIG. 5B illustrates an example of a process 550 for generating an augmented training set using real AIR and / or synthesized AIR. The augmented training set may be used to train a machine learning model for dereverberating audio signals. In some implementations, the blocks of process 550 may be implemented by a device suitable for generating an augmented training set, such as a server, a desktop computer, or a laptop computer. In some implementations, the device may be the same device that implements the blocks of process 500 as shown in and described above with respect to FIG. 5A. In some implementations, two or more blocks of process 550 may be performed substantially in parallel. In some implementations, the blocks of process 550 may be performed in an order other than that shown in FIG. 5B. In some implementations, one or more blocks of process 550 may be omitted.

[0069] Process 550 may begin at 552 by obtaining a set of clean input audio signals (e.g., input audio signals without reverberation and / or noise). The clean input audio signals in the set of clean input audio signals may have been recorded by any suitable number of devices (or microphones associated with any suitable number of devices). For example, in some implementations, two or more of the clean input audio signals may have been recorded by the same device. As another example, in some implementations, each of the clean input audio signals may have been recorded by a different device. In some implementations, two or more of the clean input audio signals may have been recorded in the same room environment. In some implementations, each of the clean input audio signals may have been recorded in a different room environment. In some implementations, the clean input audio signals in the set of clean input audio signals may include any combination of audible sound types, such as speech, music, sound effects, etc. However, each clean input audio signal may be free of reverberation, echo, and / or noise.

[0070] At block 554, process 550 may obtain a set of AIRs including real AIRs and / or synthetic AIRs. The set of AIRs may include any suitable number of AIRs (e.g., 100 AIRs, 200 AIRs, 500 AIRs, etc.). The set of AIRs may include any suitable ratio of real AIRs to synthetic AIRs, such as 90% synthetic AIRs and 10% real AIRs, 80% synthetic AIRs and 20% real AIRs, etc. A more detailed technique for generating synthetic AIRs is shown in and described above in connection with FIG. 5A.

[0071] At block 556, process 550 may generate a reverberant audio signal based on the clean input audio signal and the AIR for each pairwise combination of a clean input audio signal in the set of clean input audio signals and an AIR in the set of AIRs. For example, in some implementations, process 550 may convolve the AIR with the clean input audio signal to generate a reverberant audio signal. In some implementations, given N clean input audio signals and M AIRs, process 550 may generate up to N x M reverberant audio signals.

[0072] In some implementations, at block 558, process 550 may add noise to one or more of the reverberant audio signals generated in block 556 to generate a noisy reverberant audio signal. Examples of noise that may be added include white noise, pink noise, Brownian noise, multi-talker speech babble, etc. Process 550 may add different types of noise to different reverberant audio signals. For example, in some implementations, process 550 may add white noise to a first reverberant audio signal to generate a first noisy reverberant audio signal. Continuing with this example, in some implementations, process 550 may add multi-talker speech babble-type noise to the first reverberant audio signal to generate a second noisy reverberant audio signal. Continuing with this example, in some implementations, process 550 may add Brownian noise to the second reverberant audio signal to generate a third noisy reverberant audio signal. In other words, in some implementations, different versions of the noisy reverberant audio signal may be generated by adding different types of noise to the reverberant audio signal. Note that in some implementations, block 558 may be omitted, and the training set may be generated without adding noise to any reverberant audio signal.

[0073] At the end of block 558, process 550 has generated a training set including multiple training samples. Each training sample may include a clean audio signal and a corresponding reverberant audio signal. The reverberant audio signal may or may not include added noise. Note that in some implementations, a single clean audio signal may be associated with multiple training samples. For example, a clean audio signal may be used to generate multiple reverberant audio signals by convolving the clean audio signal with multiple different AIRs. As another example, a single reverberant audio signal (e.g., generated by convolving a single clean audio signal with a single AIR) may be used to generate multiple noisy reverberant audio signals, each corresponding to a different type of noise added to the single reverberant audio signal. Thus, a single clean audio signal may be associated with 10, 20, 30, 100, etc. training samples, each containing a different corresponding reverberant audio signal (or noisy reverberant audio signal).

[0074] In some implementations, an augmented training set may be generated for a particular type of audio content. For example, the particular type of audio content may correspond to a type of audio content for which dereverberation may be particularly difficult. For example, it may be difficult to perform dereverberation on an audio signal that includes far-field noise, such as the noise of a dog barking or a baby crying in the background of an audio signal that includes near-field speech (e.g., from a video conference, an audio call, etc.). The difficulty of performing dereverberation on far-field noise may lead to poor noise management (e.g., denoising of the audio signal). Because dereverberation of far-field noise may depend on both the characteristics / acoustics of the room and / or the specific noise, it may be difficult to train a model to perform dereverberation on such far-field noise. For example, the training dataset used to train such a model may not have sufficient training samples of a particular type of far-field noise present in the augmented set of room acoustics, thereby making a model trained on such a limited training set less robust. Thus, generating an augmented training set for a particular type of audio content may allow a more robust model to be trained. In some implementations, the particular type of audio content may include a particular type of sound or event (e.g., a dog barking, a baby crying, an emergency siren passing, etc.) and / or a particular audio environment (e.g., an indoor environment, an outdoor environment, an indoor shared workspace, etc.). In some implementations, the augmented training set may be generated by first identifying a training set of audio signals that include a particular type of audio content. For example, a training set may be obtained that includes a dog barking in a background of near-field speech. As another example, a training set may be obtained that includes a far-field siren passing in a background of near-field speech.In some implementations, because reverberation is commonly present in indoor environments, a training set may be obtained that includes audio content captured in indoor environments (and does not include audio content generated in outdoor environments). Note that in some implementations, the training set may be obtained by applying audio signals from a corpus of audio signals that classifies each audio signal as being associated with a particular type of audio content. In some implementations, an augmented training set may be generated by applying synthesized AIR and / or a particular type of noise (e.g., speech noise, indoor room noise, etc.) to the identified training set to generate the augmented training set.

[0075] Note that in some implementations, the augmented training set may be used to train speech enhancement models other than dereverberation models. For example, in some implementations, such an augmented training set may be used to train a machine learning model for noise management (e.g., noise removal), a machine learning model that performs a combination of noise management and dereverberation, etc.

[0076] Machine learning models for dereverberating audio signals may have various types of architectures. The machine learning model may take as input a frequency-domain representation of a reverberant audio signal and generate as output a predicted dereverberation mask that, when applied to the frequency-domain representation of the reverberant audio signal, produces a frequency-domain representation of a dereverberated (e.g., clean) audio signal. Exemplary architecture types include CNN, LSTM, RNN, deep neural networks, etc. In some implementations, the machine learning model may combine two or more architecture types, such as a CNN and a recurrent element. In some such implementations, a CNN may be used to extract features of the input reverberant audio signal at different resolutions. In some implementations, the recurrent element may act as a memory gate that controls the amount of previously provided input data used by the CNN. Using a recurrent element in combination with a CNN may enable the machine learning model to produce smoother outputs. Also, using a recurrent element in combination with a CNN may enable the machine learning model to achieve higher accuracy and reduce training time. Thus, using recurrent elements in combination with CNNs may improve computational efficiency by reducing the time and / or computational resources used to train a robust and accurate machine learning model for dereverberating audio signals. Examples of types of recurrent elements that may be used include GRUs, LSTM networks, Elman RNNs, and / or any other suitable types of recurrent elements or architectures.

[0077] In some implementations, a recurrent element may be combined with a CNN such that the recurrent element and the CNN are parallel. For example, the output of the recurrent element may be provided to one or more layers of the CNN, such that the CNN generates an output based on the outputs of the layers of the CNN and based on the output of the recurrent element.

[0078] In some embodiments, a CNN utilized in a machine learning model may include multiple layers. Each layer may extract features of an input reverberant audio signal spectrum (e.g., a frequency-domain representation of a reverberant audio signal) at a different resolution. In some implementations, layers of a CNN may have different dilation factors. Using a dilation factor greater than 1 may effectively increase the receptive field of a convolution filter used for a particular layer with a dilation factor greater than 1 without increasing the number of parameters. Thus, using a dilation factor greater than 1 may allow a machine learning model to be trained more robustly (by increasing the receptive field size) without increasing complexity (e.g., by maintaining the number of parameters to be learned or optimized). In one example, a CNN may have a first layer group, each with an increasing dilation factor, and a second layer group, each with a decreasing dilation factor. In one particular example, the first layer group may include six layers with dilation factors of 1, 2, 4, 8, 12, and 20, respectively. Continuing with this example, the second layer group may include five layers with decreasing expansion factors (e.g., five layers with expansion factors of 12, 8, 4, 2, and 1, respectively). The size of the receptive field considered by the CNN is related to the expansion factor, convolutional filter size, stride size, and / or pad size (e.g., whether the model is causal or not). As an example, assuming six CNN layers with increasing expansion factors of 1, 2, 4, 8, 12, and 20, a convolutional filter size of 3 × 3, a stride of 0, and a causal model, the CNN may have a total receptive field of (2 × (1 + 2 + 4 + 8 + 12 + 20)) + 1 frames, or 95 frames. As another example, the same network with an expansion of 0 would have a receptive field size of (2 * (1 + 1 + 1 + 1 + 1)) + 1 = 13. In some implementations, the total receptive field may correspond to the delay line duration, which indicates the duration of the spectrum considered by the machine learning model. It should be noted that the above expansion factors are merely examples.In some implementations, a smaller expansion factor may be used to reduce the delay duration, for example, for real-time audio signal durations.

[0079] In some implementations, the machine learning model may be zero-latency. In other words, the machine learning model may not use look-ahead, or future, data points. This is sometimes referred to as the machine learning model being causal. Conversely, in some implementations, the machine learning model may implement layers that utilize look-ahead blocks.

[0080] 6 illustrates an example of a machine learning model 600 that combines a CNN 606 and a GRU 608 in parallel. As shown, the machine learning model 600 takes as input 602 a reverberant audio signal spectrum (e.g., a frequency domain representation of the reverberant audio signal) and generates an output 604 corresponding to a predicted dereverberation mask.

[0081] As shown, the CNN 606 includes a first set of layers 610 with increasing expansion factors. In particular, the first set of layers 610 includes six layers with expansion factors of 1, 2, 4, 8, 12, and 20, respectively. The first set of layers 610 is followed by a second set of layers 612 with decreasing expansion factors. In particular, the second set of layers 612 includes five layers with expansion factors of 12, 8, 4, 2, and 1. The second set of layers 612 is followed by a third set of layers 614, each of which has an expansion factor of 1. In some implementations, the first set of layers 610, the second set of layers 612, and the third set of layers 614 may each include a convolutional block. Each convolutional block may utilize a convolutional filter. Although CNN 606 utilizes convolutional filters of size 3x3, this is merely exemplary, and in some implementations, other filter sizes (e.g., 4x4, 5x5, etc.) may be used. As shown in FIG. 6, each layer of CNN 606 can feed forward to the next or subsequent layer of CNN 606. Furthermore, in some implementations, the output of a layer having a particular expansion factor may be provided as input to a second layer having the same expansion factor. For example, a layer of a first layer set 610 having an expansion factor of 2 may be provided via connection 614 to a layer of a second layer set 612 having an expansion factor of 2. Connections 616, 618, and 620 similarly provide connections between layers having the same expansion factor.

[0082] 6, the output of the GRU 608 may be provided to various layers of the CNN 606 such that the CNN 606 generates the output 604 based on the layers of the CNN 606 and the output of the GRU 608. For example, as shown in FIG. 6, the GRU 608 may provide output to layers with decreasing expansion factors (e.g., layers included in the second set of layers 612) via connections 622, 624, 626, 628, 630, and 632. The GRU 608 may have any suitable number of nodes (e.g., 48, 56, 64, etc.) and / or any suitable number of layers (e.g., 1, 2, 3, 4, 8, etc.). In some implementations, the GRU 608 may be preceded by a first reshaping block 634 that reshapes the dimensions of the input 602 to dimensions suitable for the GRU 608 and / or required by the GRU 310. A second reshaping block 636 may follow the GRU 608. The second reshaping block 636 may reshape the dimensionality of the output produced by the GRU 608 to be suitable for provision to each layer of the CNN 606 that receives the output of the GRU 608.

[0083] In some implementations, the machine learning model may be trained using a loss function that indicates the degree of reverberation associated with a predicted dereverberated audio signal generated using the predicted dereverberation mask generated by the machine learning model. By training the machine learning model to minimize the loss function that includes the indication of the degree of reverberation, the machine learning model may not only generate dereverberated audio signals that are similar in content to the corresponding reverberant audio signal (e.g., contain similar direct sound content as in the reverberant audio signal), but may also generate dereverberated audio signals that have less reverberation. In some implementations, the loss term for a particular training sample may be a combination of the difference between the predicted dereverberated audio signal and a ground truth clean audio signal and the degree of reverberation associated with the predicted dereverberated audio signal.

[0084] In some implementations, the degree of reverberation included in the loss function may be speech-to-reverberation modulation energy. In some implementations, the speech-to-reverberation modulation energy may be the ratio of modulation energy at relatively high modulation frequencies to modulation energy across all modulation frequencies. In some implementations, the speech-to-reverberation modulation energy may be the ratio of modulation energy at relatively high modulation frequencies to modulation energy across relatively low modulation frequencies. In some implementations, the relatively high modulation frequencies and the relatively low modulation frequencies may be identified based on a modulation filter. For example, if modulation energy is determined in M ​​modulation frequency bands, the highest N of the M (e.g., 3, 4, 5, etc.) modulation frequency bands may be considered to correspond to “high modulation frequencies,” and the remaining bands (e.g., MN) may be considered to correspond to “low modulation frequencies.”

[0085] FIG. 7 illustrates an example process 700 for training a machine learning model using a loss function that incorporates the degree of reverberation of a predicted dereverberated audio signal, according to some implementations. In some implementations, the blocks of process 700 may be implemented by a device such as a server, a desktop computer, a laptop computer, etc. If an augmented training set is constructed to train the machine learning model, the device that implements the blocks of process 700 may be the same device or a different device as that used to construct the augmented training set. In some implementations, two or more blocks of process 700 may be performed substantially in parallel. In some implementations, the blocks of process 700 may be performed in an order other than that shown in FIG. 7. In some implementations, one or more blocks of process 700 may be omitted.

[0086] Process 700 may begin at 702 by obtaining a training set including training samples that include pairs of reverberant and clean audio signals. In some implementations, the clean audio signals may be considered the "ground truth" signal that the machine learning model should be trained to predict or generate. In some implementations, the training set may be an augmented training set constructed using synthetic AIR, as described above with respect to FIGS. 4A, 4B, 5A, and 5B. In some implementations, process 700 may obtain the training set from a database, a remote server, etc.

[0087] At 704, for a given training sample (e.g., for a given pair of a reverberant audio signal and a clean audio signal), process 700 may provide the reverberant audio signal to a machine learning model to obtain a predicted dereverberation mask. In some implementations, process 700 may provide the reverberant audio signal by determining a frequency domain representation of the reverberant audio signal and providing the frequency domain representation of the reverberant audio signal. In some implementations, the frequency domain representation of the reverberant audio signal may have been filtered or otherwise transformed using a filter that approximates the filtering of the human cochlea, as shown and described above with respect to block 304 of FIG. 3 .

[0088] It should be noted that the machine learning model may have any suitable architecture. For example, the machine learning model may include a deep neural network, a CNN, a LSTM, an RNN, etc. In some implementations, the machine learning model may combine two or more architectures, such as a CNN and a recurrent element. In some implementations, the CNN may use inflation factors in different layers. Specific examples of machine learning models that may be used are shown in and described above in connection with FIG. 6.

[0089] At 706, process 700 may use the predicted dereverberation mask to obtain a predicted dereverberated audio signal. For example, in some implementations, process 700 may apply the predicted dereverberation mask to a frequency-domain representation of the dereverberated audio signal to obtain a frequency-domain representation of the dereverberated audio signal, as shown and described above with respect to block 310 of FIG. 3. Continuing with this example, in some implementations, process 700 may then generate a time-domain representation of the dereverberated audio signal, as shown and described above with respect to block 312 of FIG.

[0090] At 708, the process 700 can determine a value of a reverberation metric associated with the predicted dereverberated audio signal. The reverberation metric is a measure of the speech-to-reverberation modulation energy (generally referred to herein as f srmr (z), where z is the predicted dereverberated audio signal. An exemplary formula for determining the speech-to-reverberation modulation energy, which considers the ratio of energy at relatively high modulation frequencies to energy at relatively low modulation frequencies, is given by:

number

[0091] In the formula given above, z j,k represents the average modulation energy over frames of the jth critical band grouped by the kth modulation filter, where there are 23 critical bands and 8 modulation bands. srmr Higher values ​​of (z) indicate a higher degree of reverberation. Note that other numbers of critical bands and / or modulation bands may be used to determine the speech-to-reverberation modulation energy.

[0092] At 710, the process 700 may determine a loss term based on the clean audio signal, the predicted dereverberated audio signal, and the value of the reverberation metric. In some implementations, the loss term may be a combination of the difference between the clean audio signal and the predicted dereverberated audio signal and the value of the reverberation metric. In some implementations, the combination may be a weighted sum, where the value of the reverberation metric is weighted by the importance of minimizing reverberation in the output generated using the machine learning model. A particular predicted dereverberated audio signal (referred to herein as y pre ) and a specific clean audio signal (denoted herein as y ref An exemplary expression for the loss term for σ (denoted as σ) is given by: loss=(y pre -y ref ) 2 +w*f srmr (z)

[0093] As shown in the above equation, the loss term may be increased if there is a relatively high degree of reverberation in the predicted clean audio signal and / or if the predicted dereverberated audio signal differs substantially from the ground truth clean audio signal.

[0094] At 712, process 700 may update weights of the machine learning model based at least in part on the loss terms. For example, in some implementations, process 700 may use gradient descent and / or any other suitable technique to calculate updated weight values ​​associated with the machine learning model. The weights may also be updated based on other factors such as a learning rate, a dropout rate, etc. The weights may be associated with various nodes, layers, etc. of the machine learning model.

[0095] At block 714, process 700 may determine whether the machine learning model should continue to be trained. Process 700 may determine whether the machine learning model should continue to be trained based on a determination of whether a stopping criterion has been reached. The stopping criterion may include a determination that an error associated with the machine learning model has decreased below a predetermined error threshold, a determination that a change in weights associated with the machine learning model from one iteration to the next is less than a predetermined change threshold, and / or similar determinations.

[0096] If, at block 714, process 700 determines that the machine learning model should not continue to be trained (“no” at block 714), process 700 may end at 716. Conversely, if, at block 714, process 700 determines that the machine learning model should continue to be trained (“yes” at block 714), process 700 may loop back to 704 and loop through blocks 704-714 using different training samples.

[0097] In some implementations, an augmented training set (e.g., as described above with respect to FIGS. 4A, 4B, 5A, and 5B) may be used in conjunction with a machine learning model that utilizes a loss function that incorporates the degree of reverberation of a predicted clean audio signal, such as described above with respect to FIG. 7. In some implementations, the machine learning model may have an architecture that incorporates CNN and GRU in parallel, as shown and described above with respect to FIG. 6. By combining an augmented training set including training samples generated using synthesized AIR with a machine learning model that utilizes a reverberation metric in an optimized loss function and may optionally have an architecture that utilizes both CNN and GRU, the machine learning model may be able to be trained efficiently (e.g., in a manner that minimizes computational resources) while achieving both high accuracy in the predicted dereverberated audio signal and low degree of reverberation in the predicted dereverberated audio signal. Such a system may be particularly useful for real-time audio signal dereverberation, which may require training on an augmented training set and a low-latency machine learning model architecture. FIG. 8 shows a schematic diagram of an exemplary system 800 that utilizes an augmented training set in the context of a machine learning model that utilizes a loss function that incorporates a reverberation metric.

[0098] As shown, the system 800 includes a training set creation component 802. The training set creation component 802 may generate an augmented training set that can be used by a machine learning model to dereverberate an audio signal. In some implementations, the training set component 802 may be implemented, for example, on a device that generates and / or stores the augmented training set. The training set creation component 802 may retrieve measured AIR from an AIR database 806. The training set creation component 802 may then generate synthetic AIR based on the measured AIR retrieved from the AIR database 806. More detailed techniques for generating synthetic AIR are shown in and described above in connection with FIGS. 4A, 4B, and 5A. The training set creation component 802 may retrieve a clean audio signal from a clean audio signal database 804. The training set creation component 802 may then generate an augmented training set 808 based on the measured AIR, the synthesized AIR, and the clean audio signal. A more detailed technique for generating the augmented training set is shown in and described above in connection with Figure 5B. The augmented training set 808 may include multiple (e.g., 100, 1000, 10,000, etc.) training samples, each training sample being a pair of a clean audio signal (e.g., retrieved from the clean audio signal database 804) and a corresponding reverberant audio signal generated by the training set creation component 802 based on a single AIR (either measured AIR or synthesized AIR).

[0099] The augmented training set 808 can then be used to train a machine learning model 810a. In some implementations, the machine learning model 810a may have an architecture that includes a CNN and a recurrent element (e.g., a GRU, an LSTM network, an Elman RNN, etc.) in parallel. In particular, the CNN may generate an output based on the outputs of the layers of the CNN and the outputs of the recurrent element. An example of such an architecture is shown in FIG. 6 and described above in connection with FIG. 6. The machine learning model 810a may include a prediction component 812a and a reverberation determination component 814. The prediction component 812a may generate a predicted dereverberated audio signal for a reverberant audio signal obtained from the augmented training set 808. Examples for generating a predicted dereverberated audio signal are described in more detail above in connection with FIGS. 2, 3, and 7. The reverberation determination component 814 may determine the degree of reverberation in the predicted dereverberated audio signal. For example, the degree of reverberation may be based on speech versus reverberation modulation energy, as described above in connection with block 708 of Figure 7. The degree of reverberation may be used to update weights associated with prediction component 812a. For example, the degree of reverberation may be included in a loss function that is minimized or optimized to update weights associated with prediction component 812a, as shown in connection with blocks 710 and 712 of Figure 7 and described above.

[0100] After training, the trained machine learning model 810b may utilize the trained prediction component 812b (e.g., corresponding to the determined weights) to generate a dereverberated audio signal. For example, the trained machine learning model 810b may take the reverberated audio signal 814 as input and generate the dereverberated audio signal 816 as output. Note that the trained machine learning model 810b may have the same architecture as the machine learning model 810a, but may not determine the degree of reverberation at inference time.

[0101] 9 is a block diagram illustrating example components of a device capable of implementing various aspects of the present disclosure. As with other figures provided herein, the types and number of elements shown in FIG. 9 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 900 may be configured to perform at least some of the methods disclosed herein. In some implementations, device 900 may be or include a television, one or more components of an audio system, a mobile device (such as a cellular phone), a laptop computer, a tablet device, a smart speaker, or another type of device.

[0102] According to some alternative implementations, apparatus 900 may be or include a server. In some such examples, apparatus 900 may be or include an encoder. Thus, in some examples, apparatus 900 may be a device configured for use in an audio environment, such as a home audio environment, while in other examples, apparatus 900 may be a device configured for use in the "cloud," e.g., a server.

[0103] In this example, device 900 includes an interface system 905 and a control system 910. The interface system 905, in some implementations, may be configured to communicate with one or more other devices in an audio environment. The audio environment, in some examples, may be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 905, in some implementations, may be configured to exchange control information and associated data with audio devices in the audio environment. The control information and associated data, in some examples, may relate to one or more software applications that device 900 is executing.

[0104] The interface system 905, in some implementations, may be configured to receive or provide a content stream. The content stream may include audio data. The audio data may include, but is not limited to, an audio signal. In some examples, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.

[0105] The interface system 905 may include one or more network interfaces and / or one or more external device interfaces (e.g., one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 905 may include one or more wireless interfaces. The interface system 905 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, the interface system 905 may include one or more interfaces between the control system 910 and a memory system, such as the optional memory system 915 shown in FIG. 9 . However, the control system 910 may include a memory system in some examples. In some implementations, the interface system 905 may be configured to receive input from one or more microphones in the environment.

[0106] The control system 910 may include, for example, a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0107] In some implementations, the control system 910 may reside in more than one device. For example, in some implementations, a portion of the control system 910 may reside in a device within one of the environments described herein, and another portion of the control system 910 may reside in a device outside of the environment, such as a server, a mobile device (e.g., a smartphone or tablet computer), or the like. In other examples, a portion of the control system 910 may reside in a device within one environment, and another portion of the control system 910 may reside in one or more other devices of the environment. For example, a portion of the control system 910 may reside in a device implementing a cloud-based service, such as a server, and another portion of the control system 910 may reside in another device implementing the cloud-based service, such as another server, memory device, or the like. The interface system 905 may also reside in more than one device, in some examples.

[0108] In some implementations, the control system 910 may be configured to at least partially perform the methods disclosed herein. According to some examples, the control system 910 may be configured to implement a method for dereverberating an audio signal, a method for training a machine learning model for performing dereverberation of an audio signal, a method for generating a training set for a machine learning model for performing dereverberation of an audio signal, a method for generating synthetic AIR for inclusion in the training set, etc.

[0109] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices such as those described herein, including, but not limited to, random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may reside, for example, in optional memory system 915 and / or control system 910 shown in FIG. 9 . Accordingly, various inventive aspects of the subject matter described in this disclosure may be implemented in one or more non-transitory media having software stored thereon. The software may include, for example, instructions for dereverberating an audio signal using a trained machine learning model, instructions for training a machine learning model that performs dereverberation of the audio signal, instructions for generating one or more synthesized AIRs, instructions for generating a training set for training a machine learning model that performs dereverberation of the audio signal, etc. The software may be executable by one or more components of a control system, such as, for example, control system 910 of FIG. 9 .

[0110] In some examples, the device 900 may include an optional microphone system 920 shown in FIG. 9 . The optional microphone system 920 may include one or more microphones. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker of a speaker system, a smart audio device, or the like. In some examples, the device 900 may not include the microphone system 920. However, in some such implementations, the device 900 may still be configured to receive microphone data for one or more microphones in the audio environment via the interface system 910. In some such implementations, a cloud-based implementation of the device 900 may be configured to receive microphone data, or noise metrics corresponding at least in part to the microphone data, from one or more microphones in the audio environment via the interface system 910.

[0111] According to some implementations, device 900 may include an optional loudspeaker system 925 shown in FIG. 9. Optional loudspeaker system 925 may include one or more loudspeakers, which may also be referred to herein as a "speaker" or more generally as an "audio reproduction transducer." In some examples (e.g., cloud-based implementations), device 900 may not include loudspeaker system 925. In some implementations, device 900 may include headphones. The headphones may be connected or coupled to device 900 via a headphone jack or via a wireless connection (e.g., BLUETOOTH®).

[0112] Some aspects of the present disclosure include systems or devices configured (e.g., programmed) to perform one or more examples of the disclosed methods, and tangible computer-readable media (e.g., disks) storing code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems may be or include a programmable general-purpose processor, digital signal processor, or microprocessor programmed with software or firmware and / or configured to perform any of various operations on data, including embodiments of the disclosed methods or steps thereof. Such a general-purpose processor may be or include a computer system that includes an input device, a memory, and a processing subsystem programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to asserted data.

[0113] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) configured (e.g., programmed and otherwise configured) to perform necessary processing on audio signal(s), including performing one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed system (or elements thereof) may be implemented as a general-purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include input devices and memory) that is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations, including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general-purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general-purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or keyboard), memory, and a display device.

[0114] Another aspect of the present disclosure is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code (e.g., executable code) for performing one or more examples of the disclosed methods or steps thereof.

[0115] While particular embodiments of and applications of the present disclosure have been described herein, it will be apparent to those skilled in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the present disclosure as described and claimed herein. While certain forms of the present disclosure have been illustrated and described, it should be understood that the disclosure should not be limited to the specific embodiments described and illustrated, or in the specific manner described.

[0116] Several aspects will be described. [Aspect 1] 1. A method for dereverberating an audio signal, the method comprising: obtaining an actual acoustic impulse response (AIR) by a control system; identifying, by the control system, a first portion of the real AIR corresponding to early reflections of the direct sound and a second portion of the real AIR corresponding to late reflections of the direct sound; generating, by the control system, one or more synthesized AIRs by modifying the first portion of the real AIR and / or the second portion of the real AIR; generating, by the control system, a plurality of training samples using the real AIR and the one or more synthesized AIRs, each training sample including an input audio signal and a reverberant audio signal, the reverberant audio signal being generated based at least in part on the input audio signal and one of the real AIRs or one of the one or more synthesized AIRs, the plurality of training samples being used to train a machine learning model that takes a reverberant test audio signal as input and generates a dereverberated audio signal as output. method. [Aspect 2] The method of claim 1, wherein identifying the first portion of the actual AIR corresponding to an early reflection and the second portion of the actual AIR corresponding to a late reflection includes selecting a random time value within a predetermined range, wherein the first portion includes a portion of the actual AIR before the random time value and the second portion includes a portion of the actual AIR after the random time value. Aspect 3 3. The method of claim 2, wherein the predetermined range is from about 20 milliseconds to about 80 milliseconds. Aspect 4 4. The method of any one of aspects 1 to 3, wherein modifying the first portion of the real AIR comprises randomizing the time points of the responses included in the first portion of the real AIR. Aspect 5 4. The method of any one of aspects 1 to 3, wherein modifying the second portion of the actual AIR includes truncating the second portion of the actual AIR after a duration randomly selected from a predetermined range of late reflection durations. Aspect 6 A method according to any one of aspects 1 to 3, wherein modifying the second portion of the real AIR includes modifying the amplitude of one or more responses included in the second portion of the real AIR. Aspect 7 Modifying the amplitude of the one or more responses included in the second portion of the real AIR comprises: determining a target attenuation function associated with the second portion of the actual AIR; modifying the amplitude of the one or more responses included in the second portion of the real AIR according to the target attenuation function. The method of embodiment 6. Aspect 8 4. The method of any one of aspects 1 to 3, wherein the reverberant audio signal is generated by convolving the input audio signal with the one of the real AIRs or the one of the one or more synthesized AIRs. Aspect 9 4. The method of any one of aspects 1 to 3, further comprising adding noise to a convolution of the input audio signal with the one of the real AIRs or the one of the one or more synthesized AIRs to generate the reverberant audio signal. Aspect 10 identifying an updated first portion of the actual AIR and an updated second portion of the actual AIR; By modifying the updated first portion of the actual AIR and / or the updated second portion of the actual AIR. 4. The method of any one of embodiments 1 to 3, further comprising generating additional synthesized AIR. Aspect 11 4. The method of any one of aspects 1 to 3, further comprising providing the plurality of training samples to the machine learning model to generate a machine learning model that takes the test audio signal with reverberation as the input and generates the dereverberated audio signal as the output. Aspect 12 12. The method of claim 11, wherein the test audio signal is a live-captured audio signal. Aspect 13 4. The method of any one of aspects 1 to 3, wherein the actual AIR is a measured AIR measured in a physical room. Aspect 14 4. The method of any one of aspects 1 to 3, wherein the real AIR is generated using a room acoustic model. Aspect 15 4. The method of any one of aspects 1 to 3, wherein the input audio signal is associated with a particular audio content type. Aspect 16 16. The method of claim 15, wherein the specific audio content type includes far-field noise. Aspect 17 16. The method of claim 15, wherein the particular audio content type includes audio content captured in an indoor environment. Aspect 18 16. The method of aspect 15, further comprising, before generating the plurality of training samples, obtaining a training set of a plurality of input audio signals each associated with the particular audio content type. Aspect 19 4. An apparatus configured to perform the method of any one of aspects 1 to 3. Aspect 20 One or more non-transitory media storing software, the software including instructions for controlling one or more devices to perform the method of any one of aspects 1 to 3.

Claims

[Claim 1] 1. A method for dereverberating an audio signal, the method comprising: obtaining an actual acoustic impulse response (AIR) by a control system; identifying, by the control system, a first portion of the real AIR corresponding to early reflections of the direct sound and a second portion of the real AIR corresponding to late reflections of the direct sound; generating, by the control system, one or more synthesized AIRs by modifying the first portion of the real AIR and / or the second portion of the real AIR, wherein modifying the second portion of the real AIR includes modifying the amplitude of one or more responses included in the second portion of the real AIR; generating, by the control system, a plurality of training samples using the real AIR and the one or more synthesized AIRs, each training sample including an input audio signal and a reverberant audio signal, the reverberant audio signal being generated based at least in part on the input audio signal and one of the real AIRs or one of the one or more synthesized AIRs, the plurality of training samples being used to train a machine learning model that takes a reverberant test audio signal as input and generates a dereverberated audio signal as output. method.