Low latency noise suppression

Time-domain filtering with machine learning-derived coefficients addresses latency issues in noise suppression, providing high-quality noise suppression in wearable devices without significant delay.

JP2026514356APending Publication Date: 2026-05-11QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-03-21
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Existing noise suppression technologies in wearable devices introduce significant latency due to domain transformation operations, affecting real-time audio processing and user satisfaction.

Method used

Implementing time-domain filtering using time-domain filter coefficients derived from frequency-domain noise suppression by machine learning models, allowing for high-quality noise suppression without excessive latency.

Benefits of technology

Achieves low-latency noise suppression with improved quality compared to conventional methods, maintaining real-time audio processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514356000001_ABST
    Figure 2026514356000001_ABST
Patent Text Reader

Abstract

The device includes one or more processors configured to acquire audio data representing one or more audio signals. The audio data includes a first segment and a second segment following the first segment. The processors are configured to perform one or more transformation operations on the first segment to generate frequency-domain audio data. The processors are configured to provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output. The processors are configured to perform one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients. The processors are configured to perform time-domain filtering on the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims priority from U.S. Provisional Patent Application No. 63 / 493,158, filed on March 30, 2023, and U.S. Non - Provisional Patent Application No. 18 / 611,308, filed on March 20, 2024, both owned by the same applicant, and the entire contents of which are hereby expressly incorporated by reference herein.

[0002] This disclosure generally relates to low - latency noise suppression.

Background Art

[0003] Various types of hearing - related problems affect a significant number of people. For example, one common problem is that even people with relatively normal hearing may find it difficult to hear speech in noisy environments, and this problem can be significantly exacerbated for people with hearing loss. For some individuals, speech can be easily understood only when the signal - to - noise ratio (of the voice relative to ambient noise) exceeds a certain level.

[0004] Wearable devices (e.g., earbuds, headphones, hearing aids, etc.) can be used in many situations to improve hearing, situation awareness, speech intelligibility, etc. Generally, such devices apply a relatively simple noise suppression process to remove as much ambient noise as possible. Such noise suppression processes can improve the signal - to - noise ratio sufficiently for speech to be understandable, but these noise suppression processes can also reduce the user's situation awareness because they simply try to remove as much noise as possible, thereby potentially removing important environmental cues such as traffic sounds. The use of more complex noise suppression processes can introduce significant latency. Latency in processing real - time audio can lead to user dissatisfaction.

Summary of the Invention

[0005] According to one implementation of the present disclosure, the device includes one or more processors configured to acquire audio data representing one or more audio signals. The audio data includes a first segment and a second segment following the first segment. One or more processors are configured to perform one or more transformation operations on the first segment to generate frequency-domain audio data. One or more processors are configured to provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output. One or more processors are configured to perform one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients. One or more processors are configured to perform time-domain filtering on the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0006] According to another implementation of the present disclosure, the method includes obtaining audio data representing one or more audio signals. The audio data includes a first segment and a second segment following the first segment. The method includes performing one or more transformation operations on the first segment to generate frequency-domain audio data. The method includes providing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output. The method includes performing one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients. The method includes performing time-domain filtering on the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0007] According to another implementation of this disclosure, a non-temporary computer-readable medium stores instructions executable by one or more processors to cause one or more processors to acquire audio data representing one or more audio signals. The audio data includes a first segment and a second segment following the first segment. Instructions are executable to cause one or more processors to perform one or more transformation operations on the first segment to generate frequency-domain audio data. Instructions are executable to cause one or more processors to provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output. Instructions are executable to cause one or more processors to perform one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients. Instructions are executable to cause one or more processors to perform time-domain filtering on the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0008] According to another implementation of the present disclosure, the apparatus includes means for performing one or more transformation operations on a first segment of audio data to generate frequency-domain audio data, the audio data including a first segment and a second segment following the first segment. The apparatus also includes means for processing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output. The apparatus also includes means for performing one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients. The apparatus also includes means for performing time-domain filtering on the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0009] After reviewing the entire application, including the following sections, namely the brief description of the drawings, the modes for carrying out the invention, and the claims, other aspects, advantages, and features of the present disclosure will become apparent. [Brief explanation of the drawing]

[0010] [Figure 1] This is a block diagram of a particular embodiment of a device capable of performing low-latency noise suppression, as illustrated by some examples of the present disclosure. [Figure 2] Figure 1 is a block diagram of an exemplary embodiment of the device, which is operable to perform low latency noise suppression, as shown in some examples of the present disclosure. [Figure 3] Figure 1 is a block diagram of an exemplary embodiment of the device, which is operable to perform low latency noise suppression, as shown in some examples of the present disclosure. [Figure 4] Figure 1 is a block diagram of an exemplary embodiment of the device, which is operable to perform low latency noise suppression, as shown in some examples of the present disclosure. [Figure 5] Figure 1 is a block diagram of an exemplary embodiment of the device, which is operable to perform low latency noise suppression, as shown in some examples of the present disclosure. [Figure 6] Figure 1 is a block diagram of an exemplary embodiment of the device, which is operable to perform low latency noise suppression, as shown in some examples of the present disclosure. [Figure 7] Figure 1 is a block diagram of an exemplary embodiment of the device, which is operable to perform low latency noise suppression, as shown in some examples of the present disclosure. [Figure 8] Figure 1 is a block diagram of an exemplary embodiment of the device, which is operable to perform low latency noise suppression, as shown in some examples of the present disclosure. [Figure 9] This figure shows an example of an integrated circuit capable of performing low-latency noise suppression, with some examples from the present disclosure. [Figure 10] Figures of headsets capable of performing low-latency noise suppression, as illustrated by some examples of the present disclosure. [Figure 11]Figures of headsets, such as virtual reality, mixed reality, or augmented reality headsets, that are capable of performing low latency noise suppression, as illustrated by some examples of the present disclosure. [Figure 12] Figures of augmented reality glasses capable of performing low latency noise suppression, as illustrated by some examples of the present disclosure. [Figure 13] Figures of wearable devices capable of performing low-latency noise suppression, as illustrated by some examples of the present disclosure. [Figure 14] The following are diagrams of earbuds capable of performing low-latency noise suppression, as illustrated by some examples of the present disclosure. [Figure 15] Figure 1 illustrates some examples of specific implementations of a method for performing low-latency noise suppression, which can be performed by the device shown in Figure 1. [Figure 16] This is a block diagram of a specific exemplary example of a device that is operable to perform low latency noise suppression, as illustrated by several examples of the present disclosure. [Modes for carrying out the invention]

[0011] In contexts where latency constraints are sufficiently flexible, machine learning-based noise suppression processes can be used to reduce the magnitude of noise components in audio data, to increase the magnitude of target sound components in audio data, or both. However, machine learning processes and associated pre- and post-processing can introduce significant delays. For example, a machine learning noise suppression model generally operates in the frequency domain, which involves converting audio data from the time domain to the frequency domain to generate input data for the machine learning model. Furthermore, after the machine learning model processes the input data to generate noise-suppressed audio data, the noise-suppressed audio data is converted back to the time domain for output to the user. Each of these operations introduces a delay, which can result in unacceptable latency, especially for real-time audio processing of speech.

[0012] The embodiments disclosed herein enable audio processing in a manner that provides high-quality noise suppression without introducing excessive latency. According to certain embodiments, noise suppression is performed in the time domain by applying one or more time-domain filters to the received audio data. The coefficients of one or more time-domain filters are determined based on the noise suppression output generated in the frequency domain by one or more machine learning models. The time-domain filter coefficients determined based on the noise suppression output of one or more machine learning models provide significantly better noise suppression than conventional time-domain processes such as adaptive noise cancellation. Furthermore, since the time-domain filter coefficients are applied to receive the audio data in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data.

[0013] Specific aspects of this disclosure are described below with reference to the drawings. In this description, common features are indicated by common reference numerals. Where used herein, various terms are used solely for the purpose of describing specific implementations and are not intended to limit the implementations. For example, the singular forms "a," "an," and "the" are intended to include the plural form unless the context otherwise indicates. Furthermore, some features described herein are singular in some implementations and plural in others. For example, Figure 1 shows a device 100 containing one or more processors ("Processor(single or plural)" 190 in Figure 1), which indicates that in some implementations the device 100 contains a single processor 190, and in other implementations the device 100 contains multiple processors 190. For the sake of convenience of reference in this specification, such features are generally introduced as “one or more” features, and therefore, unless a description relates to a plurality of features, they are referred to in the singular or optional plural form (as indicated by “(singular or plural)”).

[0014] In some drawings, multiple instances of a particular type of feature are used. These features are physically and / or logically distinct, but the same reference number is used for each, and the different instances are distinguished by the addition of a letter to the reference number. When a feature is referred herein as a group or type (for example, when no particular feature is referred), the reference number is used without a distinguishing letter. However, when a particular feature of one of several features of the same type is referred herein, the reference number is used with a distinguishing letter. For example, referring to Figure 5, several microphones are illustrated and associated with reference numbers 102A and 102B. When referring to a particular one of these microphones, such as microphone 102A, the distinguishing letter "A" is used. However, when referring to any one of these microphones or these microphones as a group, the reference number 102 is used without a distinguishing letter.

[0015] As used herein, the terms "comprise", "comprises", and "comprising" may be used interchangeably with "include", "includes", or "including". Additionally, the term "wherein" may be used interchangeably with "where". As used herein, "exemplary" should not be construed as indicating an example, implementation, and / or aspect as limiting, or as indicating a preferred or favorable implementation. As used herein, terms used to indicate the order in which elements such as structures, components, operations, etc. are modified (e.g., "first", "second", "third", etc.) do not, by themselves, indicate the priority or order of the elements with respect to other elements, but merely distinguish the elements from other elements having the same name (if the use of the terms indicating the order is set aside). As used herein, the term "set" refers to one or more of a particular element, and the term "plurality" refers to multiple (e.g., two or more) of a particular element.

[0016] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and may also include (or alternatively) any combination thereof. Two devices (or components) may be directly or indirectly coupled via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or a combination thereof) (e.g., they may be communicatively coupled, electrically coupled, or physically coupled). Two electrically coupled devices (or components) may be contained in the same device or in different devices, and may be connected via electronic components, one or more connectors, or inductive coupling, as exemplary and non-limiting examples. Two communicatively coupled devices (or components), such as those communicating telecommunicates, may directly or indirectly transmit and receive signals (e.g., digital or analog signals) via one or more wires, buses, networks, etc. As used herein, “directly coupled” may include two devices that are coupled without any intermediary components (e.g., communicatively coupled, electrically coupled, or physically coupled).

[0017] In the present disclosure, terms such as "determining", "calculating", "estimating", "shifting", "adjusting", etc. may be used to represent how one or more operations are performed. It should be noted that such terms should not be construed as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, "generating", "calculating", "estimating", "using", "selecting", "accessing", and "determining" may be used interchangeably. For example, "generating", "calculating", "estimating", or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated by another component or device.

[0018] FIG. 1 is a block diagram of a particular aspect of a device 100 operable to perform low latency noise suppression according to some examples of the present disclosure. In FIG. 1, device 100 includes or is coupled to one or more microphones 102 and one or more speakers 108.

[0019] A microphone(s) 102 is configured to generate an audio signal(s) 104 based on sounds 170 detected in the surrounding environment. Sounds 170 may include target sounds such as speech, as well as non-target sounds (e.g., noise). Device 100 is configured to process audio data 106 representing the audio signal(s) 104 in order to generate a noise-suppressed output signal 114, which may be used to drive a speaker(s) 108 to generate an output sound 172. In the output sound 172, with respect to sound 170, components of the audio data 106 corresponding to target sounds(s) are emphasized, components of the audio data 106 corresponding to non-target sounds(s) are deemed, or both. In certain implementations, device 100 includes, is contained within, or corresponds to a wearable device. In such an implementation, the user may wear the device 100 on, near, or on one or both of their ears to improve the perception of target sounds, reduce the perception of non-target sounds, or both. For example, the user may wear the device 100 to improve the perception of speech in a noisy environment.

[0020] In Figure 1, one or more processors 190 are configured to perform operations related to two data paths, including a first data path 110 and a second data path 120. The first data path 110 is a low-latency data path. To reduce latency, the first data path 110 is configured to perform operations on audio data 106 in the time domain. Thus, latency associated with domain transformation operations is avoided in the first data path 110. In contrast, the second data path 120 is configured to provide high-quality noise suppression at the expense of higher latency than the first data path 110. For example, in some implementations, the first data path 110 is associated with latency of 1 millisecond or less, and the second data path 120 is associated with latency of more than 1 millisecond. As another example, in some implementations, the first data path 110 may be associated with latency of 2 milliseconds or less, and the second data path 120 may be associated with latency of more than 10 milliseconds.

[0021] The second data path 120 is configured to determine time-domain filter coefficients 132, which, in some implementations, are stored in a buffer 140 and then applied by one or more time-domain filters 112 in the first data path 110. In Figure 1, the second data path 120 includes an analysis filter bank 122, one or more machine learning models 126, and a time-domain filter designer 130. The buffer 140 may be included within the time-domain filter designer 130, within one or more time-domain filters 112, or separately from both the time-domain filter designer 130 and one or more time-domain filters 112.

[0022] The analysis filter bank 122 is configured to perform one or more transformation operations (e.g., a Fast Fourier Transform (FFT) operation) based on samples of the audio data 106 in order to generate frequency-domain audio data 124. According to some implementations, a set of samples of the audio data 106 is stored for processing by the analysis filter bank 122 (e.g., in one or more buffers of the analysis filter bank 122), and the frequency-domain audio data 124 for the set of samples contains information indicating the loudness of the sound in each frequency bin of a plurality of frequency bins. During or after the transformation of the set of samples to the frequency domain, subsequent sets of samples of the audio data 106 are stored so as to be transformed into subsequent sets of frequency-domain audio data 124.

[0023] A machine learning model(s) 126 includes one or more trained models, such as neural networks(s), configured to process input data 125 based on frequency-domain audio data 124 to generate output data 129. In certain implementations, the machine learning model(s) 126 is temporally dynamic such that the output data 129, based on a particular set of input data 125, is influenced by one or more previous sets of input data 125. For example, the machine learning model(s) 126 may include one or more recurrent neural networks, such as neural networks(s) containing one or more long-term memory layers, one or more gated recurrent units, or other recurrent structures.

[0024] The time-domain filter designer 130 is configured to process a noise suppression output 128 based on output data 129 in order to generate time-domain filter coefficients 132. As an example, the time-domain filter designer 130 may perform one or more inverse transform operations using the noise suppression output 128 to generate the time-domain filter coefficients 132. The time-domain filter designer 130 may be configured to generate the time-domain filter coefficients 132 as a real-valued mask or as a complex-valued mask. In some implementations, examples of real-valued masks that can be generated by the time-domain filter designer 130 include linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, infinite-impulse response (IIR) filter coefficients, and all-pole filter coefficients. Complex-valued masks may include FIR or IIR filters that indicate magnitude and phase. The technical advantage of using linear-phase FIR filter coefficients is that the delay associated with applying linear-phase FIR filter coefficients is predictable because it depends entirely on the length of the FIR filter. The technical advantage of using minimum-phase FIR filter coefficients or autoregressive filter coefficients is that the delay introduced by the application of such filter coefficients is frequency-dependent, while the delay is reduced because it is the minimum possible for a given input data.

[0025] The time-domain filter(s) 112 is updated periodically or from time to time (for example, when updated time-domain filter coefficients 132 become available). In this configuration, the time-domain filter(s) 112 applies the time-domain filter coefficients 132 to audio data 106 that is newer than the audio data 106 used to generate the time-domain filter coefficients 132. For example, the audio data 106 may include a first segment (for example, a portion of the audio data for a particular period) and a second segment (for example, a portion of the audio data for a later period) that follows the first segment (for example, immediately following it in time or following one or more other segments that immediately follow it). In this example, the time-domain filter coefficients 132 may be determined using the first segment and applied to the second segment. As one non-limiting example, the operation of the second data path 120 may take place over a period of about 16 milliseconds, while the operation of the first data path 110 may take place over a period of about 1 millisecond. Therefore, in this particular example, the time-domain filter coefficient 132 applied to a specific data sample in the first data path 110 is always at least 16 milliseconds older than the data sample. Except in unusual circumstances, ambient noise typically changes slowly enough that, even with such a delay, the time-domain filter coefficient 132 is sufficiently representative to provide significant noise suppression.

[0026] During operation, the microphone(s) 102 generates an audio signal(s) 104 based on sound(s) 170. Sound(s) 170 may include speech or other target sounds, as well as non-target sounds such as noise. Audio data(s) 106 representing the audio signal(s) 104 is provided to a first data path(s) 110 for low-latency processing and to a second data path(s) 120 to update the time-domain filter coefficient(s) 132.

[0027] In the first data path 110, audio data 106 is processed by a time-domain filter(s) 112 that applies a first set 132 of time-domain filter coefficients to the audio data 106. The first set 132 of time-domain filter coefficients is received from the second data path 120 after the processing of the previous set 106 of audio data. The time-domain filter(s) 112 generates a noise-suppressed output signal 114 by applying the first set 132 of time-domain filter coefficients to the audio data 106. The noise-suppressed output signal 114 is provided to the speaker 108 to generate an output sound 172 (for example, used to drive it). Optionally, the processor(s) 190 may perform other processing on the noise-suppressed output signal 114 before providing it to the speaker 108. For example, the processor(s) 190 may perform feedforward adaptive noise cancellation (ANC), feedback ANC, or hybrid ANC to further process the noise-suppressed output signal 114.

[0028] In the second data path 120, a set of samples of audio data 106 is accumulated to generate frequency-domain audio data 124 representing a set of samples, and undergoes one or more transformation operations by the analysis filter bank 122. In some implementations, the frequency-domain audio data 124 is provided as input to a machine learning model(s) 126 (e.g., as input data 125). In other implementations, the frequency-domain audio data 124 is processed to generate input data 125. For example, the frequency-domain audio data 124 may optionally undergo various frequency-domain noise suppression or signal augmentation operations to generate input data 125. For example, input data 125 may be generated by applying beamforming, blind-source separation, speech augmentation operations, or a combination thereof to the frequency-domain audio data 124. Performing such conventional frequency-domain operations to generate input data 125 can improve the operation of the machine learning model(s) 126 by providing cleaner input data 125. In the same or different implementations, generating input data 125 based on frequency-domain audio data 124 may include data aggregation, filtering, and resampling operations.

[0029] A machine learning model(s) 126 performs non-linear, temporally dynamic operations based on prior training of the machine learning model(s) 126 to generate output data 129. In some implementations, the machine learning model(s) 126 is trained to generate output data 129 that includes target audio components of frequency-domain audio data 124 and omits or suppresses non-target audio components of frequency-domain audio data 124. Such models are referred to herein as “inline” models. In contrast, in some implementations, the machine learning model(s) 126 is trained to generate output data 129 that includes non-target audio components of frequency-domain audio data 124 and omits or suppresses target audio components of frequency-domain audio data 124. Such models are referred to herein as “masking” models. The output data 129 from a masking model can be used directly as a noise suppression output 128. The output data 129 from the inline model can be further processed to generate a noise-suppressed output 128, as will be further explained with reference to Figure 3.

[0030] The noise suppression output 128 represents an estimate (by one or more machine learning models) 126 of a portion of the frequency-domain audio data 124 (e.g., audio components) corresponding to non-target audio (e.g., noise). The noise suppression output 128 is provided as input to the time-domain filter designer 130. The time-domain filter designer 130 performs inverse transform and parameterization operations to generate time-domain filter coefficients 132. The specific inverse transform and parameterization operations performed may differ for different implementations. For example, the inverse transform operation may include various inverse Fourier transform operations, such as the inverse fast Fourier transform (IFFT) operation. The parameterization operation may include, for example, windowing or shifting the time-domain data generated by the inverse transform operation to generate a specific number of time-domain filter coefficients 132 based on the number of filter coefficients applied by the time-domain filter(s) 112. Applying more filter coefficients may provide greater noise suppression at the expense of greater computational complexity.

[0031] While time-domain filter coefficients 132 are being generated based on a first set of samples of audio data 106, additional samples of audio data 106 may be received and aggregated to form a second set of samples of the audio data. After the second set of samples is collected, the second set of samples undergoes the same operation as described above to generate a second set of time-domain filter coefficients 132. In some implementations, the second set of samples is independent of the first set of samples. For example, there is no overlap between the first set of samples and the second set of samples. In other implementations, the second set of samples contains one or more samples from the first set of samples. That is, the first and second sets of samples have at least some overlap. The specific amount of overlap varies for different implementations and may be selected based on available processing resources, a specific sound environment, user settings, or other selection criteria.

[0032] In the example shown in Figure 1, the audio data 106 provided to the first data path 110 is the same as the audio data 106 provided to the second data path 120. In some implementations, the audio data 106 provided to the second data path 120 includes the audio data 106 provided to the first data path 110 as well as additional audio data 106. For example, device 100 may include two or more microphones 102, at least one of which is coupled to the second data path 120 and not to the first data path 110. For example, device 100 may include a self-talk microphone configured primarily to capture sounds corresponding to the utterances of the user of device 100. In this exemplary example, the audio data 106 generated by the self-talk microphone may be provided to the second data path 120 to generate time-domain filter coefficients 132, but not to the first data path 110 to generate a noise-suppressed output signal 114.

[0033] In addition, or alternatively, in some implementations, the audio data 106 provided to the first data path 110 includes the audio data 106 provided to the second data path 120, as well as additional audio data 106. For example, device 100 may include two or more microphones 102, at least one of which is coupled to the first data path 110 and not to the second data path 120. For example, device 100 may include a feedback microphone configured primarily to capture output sound 172 and generate a feedback signal. In this exemplary example, the audio data 106 generated by the feedback microphone may be processed in the time domain (for example, in a feedback ANC filter 816 shown in Figure 8) to generate or modify a noise-suppressed output signal 114.

[0034] In addition, or alternatively, in some implementations, the audio data 106 provided to the first data path 110 includes at least a portion of the audio data 106 provided to the second data path 120, and the audio data 106 provided to the second data path 120 includes at least a portion of the audio data 106 provided to the first data path 110, and each of the data paths 110, 110 is also provided with additional audio data 106 that is not provided to the other data paths 120, 120. For example, device 100 may include both a feedback microphone and a self-talk microphone, each of which operates as described above.

[0035] In some implementations, device 100 corresponds to or is included in one of various types of devices. In an exemplary example, processor 190 is integrated into a wearable device that includes or is coupled with a microphone(s) 102 and a speaker(s) 108. Examples of such wearable devices include, but are not limited to, a headset device as further described with reference to Figure 10, a virtual reality, mixed reality, or augmented reality headset as described with reference to Figure 11, augmented reality glasses as described with reference to Figure 12, a hearing aid device as described with reference to Figure 13, or earbuds as described with reference to Figure 14.

[0036] As described above, one technical advantage of implementing device 100 is that it can provide high-quality noise suppression without introducing excessive latency. For example, noise suppression of input audio data 106 is performed entirely in the time domain (e.g., by the first data path 110), thereby avoiding delays due to domain transformation and inverse transformation operations. Thus, the latency introduced by the noise suppression operation is very small (e.g., on the order of less than 2 milliseconds), comparable to the latency of conventional time-domain noise suppression operations such as ANC. However, the quality of noise suppression is higher than that which can be achieved using such conventional time-domain noise suppression operations.

[0037] Figure 2 is a block diagram of an exemplary embodiment of the device of Figure 1, which is operable to perform low-latency noise suppression, by some example of the present disclosure. The device 100 of Figure 2 represents one particular non-limiting example of the device 100 of Figure 1, and therefore the device 100 of Figure 2 includes many of the same components as those shown in Figure 1, each of which operates as described above. For example, the device 100 of Figure 2 includes a first data path 110 and a second data path 120. As described with reference to Figure 1, the first data path 110 includes one or more time-domain filters 112 configured to apply time-domain filter coefficients 132 (determined in the second data path 120) to audio data 106 in order to produce a noise-suppressed output signal 114.

[0038] In the example shown in Figure 2, the second data path 120 includes an analysis filter bank 122, a machine learning model(s) 126, and a time-domain filter designer 130. In some implementations, the processor(s) 190 in Figure 2 also includes a buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. In Figure 2, the machine learning model(s) 126 is configured to receive input data 125 containing or based on frequency-domain audio data 124 and to generate a frequency mask 204 as output data 129. The frequency mask 204 represents the estimated magnitude of noise in the frequency-domain audio data 124 for each frequency bin of a plurality of frequency bins. Since the frequency mask 204 is a frequency-domain estimate of noise in the audio data 106, it can be provided as input to the time-domain filter designer 130 to generate time-domain filter coefficients 132.

[0039] Figure 3 is a block diagram of an exemplary embodiment of the device of Figure 1, which is operable to perform low-latency noise suppression, according to several examples of the present disclosure. The device 100 of Figure 3 represents one particular non-limiting example of the device 100 of Figure 1, and therefore the device 100 of Figure 3 includes many of the same components as those shown in Figure 1, each of which operates as described above. For example, the device 100 of Figure 3 includes a first data path 110 and a second data path 120. As described with reference to Figure 1, the first data path 110 includes a time-domain filter (one or more) 112 configured to apply time-domain filter coefficients 132 (determined in the second data path 120) to audio data 106 in order to produce a noise-suppressed output signal 114.

[0040] In the example shown in Figure 3, the second data path 120 includes an analysis filter bank 122, a machine learning model(s) 126, and a time-domain filter designer 130. In some implementations, the processor(s) 190 in Figure 3 also includes a buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. In Figure 3, the machine learning model(s) 126 is configured to receive input data 125 containing or based on frequency-domain audio data 124 and to generate noise-suppressed audio data 304 as output data 129 in Figure 1. The noise-suppressed audio data 304 includes, for example, a frequency-domain estimate of the target audio in the audio data 106. In some embodiments, the noise-suppressed audio data 304 may include synthesized audio data. For example, the machine learning model(s) 126 may be configured to synthesize target audio (e.g., speech) based on the input data 125. The advantage of a machine learning model (one or more) 126 generating synthesized target audio (e.g., speech) is that the synthesized target audio may be free of non-target audio, and therefore the noise-suppressed audio data 304 in this example may contain only the target audio.

[0041] In Figure 3, the second data path 120 also includes a mask generator 306. The mask generator 306 is configured to generate a frequency mask 204 based on noise-suppressed audio data 304. For example, the mask generator 306 may determine the ratio of noise-suppressed audio data 304 to frequency-domain audio data 124 for each frequency bin in a set of frequency bins. Each frequency bin of the frequency-domain audio data 124 may be adjusted by the ratio associated with the frequency bin in order to generate the frequency mask 204. The frequency mask 204 may be contained in or correspond to the noise-suppressed output 128, which is provided as input to the time-domain filter designer 130 to generate time-domain filter coefficients 132.

[0042] Figure 4 is a block diagram of an exemplary embodiment of the device of Figure 1, which is operable to perform low-latency noise suppression, by some example of the present disclosure. The device 100 of Figure 4 represents one particular non-limiting example of the device 100 of Figure 1, and therefore the device 100 of Figure 4 includes many of the same components as those shown in Figure 1, each of which operates as described above. For example, the device 100 of Figure 4 includes a first data path 110 and a second data path 120. As described with reference to Figure 1, the first data path 110 includes a time-domain filter (one or more) 112 configured to apply time-domain filter coefficients 132 (determined in the second data path 120) to audio data 106 in order to produce a noise-suppressed output signal 114.

[0043] In the example shown in Figure 4, the second data path 120 includes an analysis filter bank 122, a frequency-domain signal expander 402, a machine learning model (one or more) 126, and a time-domain filter designer 130. In some implementations, the processor (one or more) 190 in Figure 4 also includes a buffer 140 between the time-domain filter designer 130 and the time-domain filter (one or more) 112. The frequency-domain signal expander 402 is configured to process the frequency-domain audio data 124 to generate expanded audio data 404. The expanded audio data 404 has an improved signal-to-noise ratio (SNR) with respect to the frequency-domain audio data 124. The frequency-domain signal expander 402 may use one or more of a variety of different operations to improve the SNR. Examples of such operations are described with reference to Figures 5 to 7.

[0044] In Figure 4, the input data 125 includes or is based on the augmented audio data 404. The machine learning model(s) 126 in Figure 4 generates output data 129, such as a frequency mask, noise-suppressed audio data, or both, based on the input data 125. The time-domain filter designer 130 generates time-domain filter coefficients 132 based on the noise-suppressed output 128, which includes or is based on the output data 129. For example, in some cases where the output data 129 of the machine learning model(s) 126 in Figure 4 is a frequency mask, the noise-suppressed output 128 includes or corresponds to the output data 129. In other cases where the output data 129 of the machine learning model(s) 126 in Figure 4 is noise-suppressed audio data, a second path 120 may further include the mask generator 306 in Figure 3 to generate a frequency mask that corresponds to or is contained in the noise-suppressed output 128.

[0045] Figure 5 is a block diagram of an exemplary embodiment of the device of Figure 1, which is operable to perform low-latency noise suppression, according to some examples of the present disclosure. The device 100 of Figure 5 represents one particular non-limiting example of the device 100 of Figure 1, and therefore the device 100 of Figure 5 includes many of the same components as those shown in Figure 1, each of which operates as described above. For example, the device 100 of Figure 5 includes a first data path 110 and a second data path 120. As described with reference to Figure 1, the first data path 110 includes a time-domain filter(s) 112 configured to apply time-domain filter coefficients 132 determined in the second data path 120 to audio data 106 in order to produce a noise-suppressed output signal 114.

[0046] In the example shown in Figure 5, the second data path 120 includes an analysis filter bank 122, a beamformer 502, a machine learning model (one or more) 126, and a time-domain filter designer 130. In some implementations, the processor (one or more) 190 in Figure 5 also includes a buffer 140 between the time-domain filter designer 130 and the time-domain filter (one or more) 112. The beamformer 502 is an example of the frequency-domain signal expander 402 in Figure 4. The beamformer 502 is configured to process the frequency-domain audio data 124 to generate beamformed audio data 504, thereby emphasizing or deemphasizing sounds received from a particular direction. For example, components of sound 170 originating from the direction of a target sound source may be emphasized in the beamformed audio data 504, or components of sound 170 originating from directions other than the direction of the target sound source may be deemed in the beamformed audio data 504, or both.

[0047] Audio signals 104 from two or more microphones 102 (e.g., a first microphone 102A and a second microphone 102B) are used to determine the directivity associated with the sound 170. Optionally, in some implementations, other information may be used to determine the direction toward the target sound source. For example, a camera may be used to generate image or video data that can be analyzed to determine the direction toward the person speaking from device 100 in Figure 5. In other examples, no other sensors are used. For example, a beamformer 502 can direct a beam toward a dominant sound source in a particular situation, under the assumption that the dominant sound source is more likely to be the target sound source than the background sound source.

[0048] In Figure 5, the input data 125 includes or is based on beamformed audio data 504. Since the beamformed audio data 504 represents an improvement in the SNR of the target sound, the computational complexity of the machine learning model(s) 126 in Figure 5 can be reduced compared to generating the input data 125 based on frequency domain audio data 124 without beamforming. Based on the input data 125, the machine learning model(s) 126 generates output data 129, such as a frequency mask, noise-suppressed audio data, or both. The time-domain filter designer 130 generates time-domain filter coefficients 132 based on a noise-suppressed output 128 that includes or is based on the output data 129.

[0049] Figure 6 is a block diagram of an exemplary embodiment of the device of Figure 1, which is operable to perform low-latency noise suppression, according to some examples of the present disclosure. The device 100 of Figure 6 represents one particular non-limiting example of the device 100 of Figure 1, and therefore the device 100 of Figure 6 includes many of the same components as those shown in Figure 1, each of which operates as described above. For example, the device 100 of Figure 6 includes a first data path 110 and a second data path 120. As described with reference to Figure 1, the first data path 110 includes a time-domain filter(s) 112 configured to apply time-domain filter coefficients 132 determined in the second data path 120 to audio data 106 in order to produce a noise-suppressed output signal 114.

[0050] In the example shown in Figure 6, the second data path 120 includes an analysis filter bank 122, a speech augmentation engine 602, a machine learning model(s) 126, and a time-domain filter designer 130. In some implementations, the processor(s) 190 in Figure 6 also includes a buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. The speech augmentation engine 602 is an example of the frequency-domain signal expander 402 in Figure 4. The speech augmentation engine 602 is configured to process the frequency-domain audio data 124 to enhance the components of the frequency-domain audio data 124 that represent the speech in order to generate speech augmented audio data 604. For example, the speech augmentation engine 602 may perform spectral modifications such as filtering, equalization, and spectral enhancement to enhance speech, deencode non-speech, or both.

[0051] In Figure 6, the input data 125 includes or is based on speech-augmented audio data 604. Since the speech-augmented audio data 604 represents the improvement in the SNR of the target sound when the target sound is speech, the computational complexity of the machine learning model(s) 126 in Figure 6 can be reduced compared to generating the input data 125 based on frequency-domain audio data 124 without speech augmentation. Based on the input data 125, the machine learning model(s) 126 generates output data 129, such as a frequency mask, noise-suppressed audio data, or both. The time-domain filter designer 130 generates time-domain filter coefficients 132 based on a noise-suppressed output 128 that includes or is based on the output data 129.

[0052] Figure 7 is a block diagram of an exemplary embodiment of the device of Figure 1, which is operable to perform low-latency noise suppression, according to some examples of the present disclosure. The device 100 of Figure 7 represents one particular non-limiting example of the device 100 of Figure 1, and therefore the device 100 of Figure 7 includes many of the same components as those shown in Figure 1, each of which operates as described above. For example, the device 100 of Figure 7 includes a first data path 110 and a second data path 120. As described with reference to Figure 1, the first data path 110 includes a time-domain filter (one or more) 112 configured to apply time-domain filter coefficients 132 determined in the second data path 120 to audio data 106 in order to produce a noise-suppressed output signal 114.

[0053] In the example shown in Figure 7, the second data path 120 includes an analysis filter bank 122, a source separation engine 702, a machine learning model(s) 126, and a time-domain filter designer 130. In some implementations, the processor(s) 190 in Figure 7 also includes a buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. The source separation engine 702 is an example of the frequency-domain signal expander 402 in Figure 4. The source separation engine 702 is configured to process frequency-domain audio data 124 to identify and highlight components of the frequency-domain audio data 124 that represent sound 170 from a target source. For example, the source separation engine 702 may use a blind source separation operation, such as independent component analysis, to perform source separation. In this example, a portion of the source-separated audio data corresponding to the target sound may be used to generate input data 125.

[0054] In Figure 7, the input data 125 includes or is based on source-separated audio data 704. Since the source-separated audio data 704 represents an improvement in the SNR of the target sound, the computational complexity of the machine learning model(s) 126 in Figure 7 can be reduced compared to generating the input data 125 based on frequency-domain audio data 124 without source separation. Based on the input data 125, the machine learning model(s) 126 generates output data 129, such as a frequency mask, noise-suppressed audio data, or both. The time-domain filter designer 130 generates time-domain filter coefficients 132 based on a noise-suppressed output 128 that includes or is based on the output data 129.

[0055] Figures 5 to 7 show implementations of device 100 in which specific examples of signal expansion are performed, but in some implementations, the frequency-domain signal expander 402 may perform two or more of these signal expansion operations or other frequency-domain signal expansion operations. For example, the frequency-domain signal expander 402 may include a beamformer 502 and a speech expansion engine 602. In another example, the frequency-domain signal expander 402 may include a beamformer 502 and a source separation engine 702. In yet another example, the frequency-domain signal expander 402 may include a speech expansion engine 602 and a source separation engine 702.

[0056] Figure 8 is a block diagram of an exemplary embodiment of the device of Figure 1, which is operable to perform low-latency noise suppression, according to some examples of the present disclosure. The device 100 of Figure 8 represents one particular non-limiting example of the device 100 of Figure 1, and therefore the device 100 of Figure 8 includes many of the same components as those shown in Figure 1, each of which operates as described above. For example, the device 100 of Figure 8 includes a first data path 110 and a second data path 120. As described with reference to Figure 1, the first data path 110 includes a time-domain filter(s) 112 configured to apply time-domain filter coefficients 132 determined in the second data path 120 to audio data 106 in order to produce a noise-suppressed output signal 114.

[0057] In the example shown in Figure 8, the second data path 120 includes an analysis filter bank 122, one or more machine learning models 126, and a time-domain filter designer 130. In some implementations, the processor 190 in Figure 8 also includes a buffer 140 between the time-domain filter designer 130 and one or more time-domain filters 112. In various implementations, the second data path 120 also optionally includes one or more input preprocessing engines 820, one or more output postprocessing engines 822, or both.

[0058] As described above, the analysis filter bank 122 is configured to receive audio data 106 representing an audio signal 104 from one or more microphones 102 (e.g., microphones 102A and 102B in Figure 8) and to generate frequency-domain audio data 124. In some implementations, the frequency-domain audio data 124 is used as input data 125 for a machine learning model(s) 126. In such implementations, the input preprocessing engine(s) 820 may be omitted. In other implementations, the frequency-domain audio data 124 is modified to generate input data 125 for a machine learning model(s) 126. In such implementations, the input preprocessing engine(s) 820 is configured to generate input data 125 based on the frequency-domain audio data 124. For example, the input preprocessing engine(s) 820 may include one or more frequency-domain signal expanders (e.g., frequency-domain signal expander 402 in Figure 4), such as the beamformer 502 in Figure 5, the speech augmentation engine 602 in Figure 6, the source separation engine 702 in Figure 7, or a combination thereof. The input preprocessing engine(s) 820 may also, or alternatively, perform other operations to generate input data 125 based on frequency-domain audio data 124. For example, the input preprocessing engine(s) 820 may perform data aggregation operations, filtering operations, resampling operations, other data manipulations, or a combination thereof to generate input data 125.

[0059] A machine learning model(s) 126 generates output data 129 based on input data 125. Depending on the configuration of the machine learning model(s) 126, the output data 129 may include noise-suppressed audio data, a frequency mask, or both. Optionally, the output data 129 may be modified by an output post-processing engine(s) 822 to generate a noise-suppressed output 128, which is provided to the time-domain filter designer(s) 130 to generate time-domain filter coefficients 132. For example, when the output data 129 includes noise-suppressed audio data, the output post-processing engine(s) 822 may perform operations to determine a frequency mask based on the noise-suppressed audio data, as described by the mask generator(s) 306 in Figure 3. The output post-processing engine(s) 822 may also, or alternatively, perform other operations to generate a noise-suppressed output 128 based on the output data 129. For example, the output post-processing engine(s) 822 may perform data aggregation, filtering, resampling, other data manipulation, or a combination thereof to generate the noise-suppressed output 128.

[0060] Figure 8 shows several optional components, some or all of which are included in some implementations and omitted in others. For example, in Figure 8, the processor(s) 190 includes a feedforward ANC filter 812 configured to perform feedforward adaptive noise cancellation based on audio data 106 from one or more microphones 102. As another example, in Figure 8, the device 100 includes or is coupled to at least one external microphone (e.g., microphones 102A and 102B) configured to generate audio data 106 representing sound 170 in the ambient environment around the device 100. In addition, in Figure 8, device 100 includes or is coupled to at least one feedback microphone (e.g., microphone 102C) configured to generate a feedback signal 814 based on sound present near speaker(s) 108, the feedback signal 814 may include the output sound 172 produced by speaker(s) 108 and components of the sound 170 transmitted directly to the user's ear canal 804 (subject to some transfer function P(z) 802). In this example, the feedback signal 814 may be provided to a feedback ANC filter 816 to generate a feedback noise signal which is subtracted from the noise-suppressed output signal 114.

[0061] Figure 9 shows an implementation form 900 of device 100 as an integrated circuit 902 including one or more processors 190. The integrated circuit 902 also includes an audio input 904, such as one or more bus interfaces, to enable audio data 106 to be received for processing. The integrated circuit 902 also includes a signal output 906, such as a bus interface, to enable the transmission of a noise-suppressed output signal 114. In Figure 9, the processor(s) 190 of the integrated circuit 902 includes one or more audio components 940, such as an analysis filter bank 122, a machine learning model(s) 126, a time-domain filter designer 130, and a time-domain filter(s) 112. Optionally, the audio component(s) 940 may include other components as described above with reference to Figures 1 to 8, such as the mask generator 306 in Figure 3, the frequency domain signal expander 402 in any of Figures 4 to 7, or any one or more of the input pre-processing engine(s) 820, output post-processing engine(s) 822, feedforward ANC filter 812, and feedback ANC filter 816 in Figure 8. The integrated circuit 902 enables a high-quality, low-latency noise suppression implementation as a component in a system such as a microphone-containing wearable device, such as a headset as shown in Figure 10, a virtual reality, mixed reality, or augmented reality headset as shown in Figure 11, augmented reality headset glasses as shown in Figure 12, a wearable device as shown in Figure 13, earbuds as shown in Figure 14, or another wearable device.

[0062] Figure 10 shows an implementation configuration 1000 in which device 100 includes a headset device 1002. The headset device 1002 includes one or more microphones 102 and one or more speakers 108. In the example shown in Figure 10, microphone 102A is primarily positioned to detect voice from the person wearing the headset device 1002, and microphone 102B is positioned to detect ambient sounds such as voice from another person or other sounds. Components of processor 190, including audio components 940, are integrated into the headset device 1002 and are shown using dashed lines to indicate components that are not generally visible to the user of the headset device 1002.

[0063] In a specific example of operation, the microphone 102B may detect sounds in the environment surrounding the headset device 1002 and generate audio data representing those sounds. The audio data may be provided to the audio component 940, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to the audio data subsequently received. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to the audio data received in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the headset device 1002 may provide high-quality low-latency noise suppression.

[0064] Figure 11 shows one implementation form 1100 in which device 100 includes a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset device 1102. The headset device 1102 includes one or more microphones 102 and one or more speakers 108. In addition, components of one or more processors 190, including an audio component 940, are integrated into the headset device 1102. In a particular example of operation, one or more microphones 102 may detect sounds in the environment around the headset device 1102 and generate audio data representing those sounds. The audio data may be provided to the audio component 940, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to subsequently received audio data. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to receive audio data in the time domain, using such time-domain filter coefficients to process the audio data results in little to no latency. Therefore, the headset device 1102 can provide high-quality, low-latency noise suppression.

[0065] Figure 12 shows one implementation form 1200 in which device 100 includes a portable electronic device corresponding to augmented reality or mixed reality glasses 1202. The glasses 1202 include a holographic projection unit 1204 configured to project visual data onto the surface of a lens 1206 or to reflect visual data from the surface of the lens 1206 onto the wearer's retina. The glasses 1202 also include a microphone(s) 102, a speaker(s) 108, and a processor(s) 190 including an audio component 940.

[0066] In a specific example of operation, a microphone(s) 102 may detect sounds in the environment around the glasses 1202 and generate audio data representing those sounds. The audio data may be provided to an audio component 940, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to the audio data subsequently received. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to receive the audio data in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the glasses 1202 may provide high-quality, low-latency noise suppression.

[0067] In some implementations, the holographic projection unit 1204 is configured to display information related to sound detected by the microphone(s) 102. For example, the holographic projection unit 1204 may display a notification indicating that sound has been detected. In another example, the holographic projection unit 1204 may display a notification indicating a detected audio event. For example, the notification may be superimposed on the user's field of view at a specific location that coincides with the location of the sound source associated with the audio event.

[0068] Figure 13 illustrates a wearable device capable of performing low-latency noise suppression, according to some examples of the present disclosure. In the example shown in Figure 13, the wearable device is a hearing aid device 1302. The hearing aid device 1302 includes a microphone(s) 102, a speaker(s) 108, and a processor(s) 190 including an audio component 940. In the example shown in Figure 13, the hearing aid device 1302 includes a portion 1304 configured to be worn behind the user's ear, a portion 1308 configured to extend over the ear, and a portion 1306 fitted in or near the user's ear canal. In other examples, the hearing aid device 1302 may have a different configuration or form factor. For example, the hearing aid device 1302 may be an in-ear device that does not include the portion 1304 configured to be worn behind the ear and the portion 1308 configured to extend over the ear.

[0069] In a specific example of the operation of the hearing aid device 1302, a microphone(s) 102 may detect sounds in the environment surrounding the hearing aid device 1302 and generate audio data representing those sounds. The audio data may be provided to an audio component 940, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to the audio data subsequently received. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to receive the audio data in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the hearing aid device 1302 may provide high-quality low-latency noise suppression.

[0070] Figure 14 shows one implementation configuration 1400, in which device 100 includes a portable electronic device that corresponds to one or more earbuds 1404 (e.g., a first earbud 1406, a second earbud 1402, or both). Although earbud 1406 is described, it should be understood that this technology may be applicable to other in-ear or over-ear audio devices.

[0071] In the example shown in Figure 14, the first earbud 1402 includes a first microphone 1410A, such as a high-signal-to-noise microphone, positioned to capture the voice of the wearer of the first earbud 1402; one or more other microphones, indicated as microphone(s) 1412A, configured to detect ambient sounds and spatially distributed to support beamforming; an "inner" microphone 1414A positioned close to the wearer's ear canal (for example, to assist with active noise cancellation); and a self-speaking microphone 1416A, such as a bone conduction microphone, configured to convert sound vibrations of the wearer's ossicles or skull into audio signals. In a particular implementation, microphone(s) 1412A corresponds to microphone(s) 102 in any of Figures 1-4, 6, or 7, or to microphone(s) 102A and / or 102B in Figure 5 or 8. In certain implementation configurations, microphone 1414A corresponds to microphone 102C in Figure 8.

[0072] The second earbud 1404 may be configured in substantially the same manner as the first earbud 1402. For example, the second earbud may include a microphone 1410B positioned to capture the voice of the wearer of the second earbud 1404, one or more other microphones 1412B configured to detect ambient sounds and spatially distributed to support beamforming, an "inner" microphone 1414B, and a self-speaking microphone 1416B.

[0073] In some implementations, the earbuds 1402 and 1404 are configured to automatically switch between various operating modes, such as a pass-through mode in which ambient sound is processed by the audio component 940 for output via speaker(s) 108, and a playback mode in which non-ambient sound (e.g., streaming audio corresponding to telephone conversations, media playback, video games, etc.) is played through speaker(s) 108. In other implementations, the earbuds 1402 and 1404 may support fewer modes, or may support one or more other modes in place of or in addition to the described modes.

[0074] In an exemplary example of operation in pass-through mode, one or more of the microphones 102 (e.g., microphones 1412A, 1412B) may detect sound in the environment around the earbuds 1402, 1404 and generate audio data representing that sound. The audio data may be provided to one or both of the audio components 940A, 940B, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to the received audio data. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to the received audio data in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the earbuds 1402, 1404 can provide high-quality low-latency noise suppression.

[0075] Figure 15 illustrates certain implementations of a method for performing low-latency noise suppression, which may be performed by the device of Figure 1, according to some examples of the present disclosure. In certain embodiments, one or more operations of Method 1500 are performed by at least one of the devices 100, processors (one or more) 190, or audio components 940, or a combination thereof, as described in various ways with reference to Figures 1-14.

[0076] Method 1500 includes, in block 1502, acquiring audio data representing one or more audio signals. The audio data includes a first segment and a second segment following the first segment. For example, referring to Figure 1, audio data 106 represents an audio signal(s) 104 received from a microphone(s) 102. In this example, the microphone(s) 102 generates the audio signal(s) 104 based on sounds 170 present in the environment around device 100. Each segment may include a time segment of the audio signal, such as a sample representing a few milliseconds of the audio signal.

[0077] Method 1500 includes performing one or more transformation operations on a first segment in block 1504 to generate frequency-domain audio data. For example, referring to Figure 1, the analysis filter bank 122 performs one or more transformation operations (e.g., time-domain to frequency-domain transformation operations such as FFT) to generate frequency-domain audio data 124.

[0078] Method 1500 includes, in block 1506, providing input data based on frequency-domain audio data as input to one or more machine learning models to generate noise-suppressed outputs. For example, in some implementations, the input data includes or corresponds to frequency-domain audio data. In other implementations, the frequency-domain audio data is modified or manipulated to generate the input data. For example, in some such implementations, Method 1500 includes performing a beamforming operation to determine beamformed audio data that distinguishes between portions of audio data from a target audio source and portions of audio data from a non-target audio source. In this example, the input data is based on (e.g., includes or corresponds to) beamformed audio data. As another example, in some such implementations, Method 1500 includes performing a speech augmentation operation to determine speech augmented audio data. In this example, the input data is based on (e.g., includes or corresponds to) speech augmented audio data. As yet another example, in some such implementations, Method 1500 includes performing a source separation operation to determine source-separated audio data. In this example, the input data is based on (e.g., including or corresponding to) source-separated audio data. In addition to, or instead of, performing signal enhancement operations (such as beamforming, speech enhancement, or source separation), method 1500 may include performing other data operations, such as data aggregation and filtering, to generate input data based on frequency-domain audio data.

[0079] In some implementations, one or more machine learning models directly generate a noise suppression output (for example, the noise suppression output is output data from one or more machine learning models). For example, a masking machine learning model outputs a frequency mask that can be used as or included in the noise suppression output.

[0080] In some implementations, one or more machine learning models generate output data that is modified to produce noise-suppressed output. For example, an inline machine learning model outputs noise-suppressed audio data. In this example, a mask generator (e.g., mask generator 306 in Figure 3) may generate a frequency mask based on the noise-suppressed audio data and frequency-domain audio data.

[0081] Method 1500 includes performing one or more inverse transform operations on the noise suppression output in block 1508 to generate time-domain filter coefficients. For example, the time-domain filter designer 130 may perform inverse transform and parameterization operations to generate time-domain filter coefficients 132 based on the noise suppression output 128. The time-domain filter coefficients may include, for example, linear phase finite impulse response (FIR) filter coefficients, minimum phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.

[0082] Method 1500 includes, in block 1510, performing time-domain filtering on a second segment using time-domain filter coefficients to generate a noise-suppressed output signal. For example, a time-domain filter(s) 112 applies time-domain filter coefficients 132 to a segment (e.g., a second segment) of audio data 106 following a segment (e.g., a first segment) used to generate the time-domain filter coefficients. One technical advantage of Method 1500 is that the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and because the time-domain filter coefficients are applied to audio data received in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data.

[0083] Optionally, in some implementations, Method 1500 also includes performing adaptive noise cancellation operation based on the noise-suppressed output signal. For example, referring to Figure 8, the noise-suppressed output signal 114 may be modified using a feedforward ANC filter 812, a feedback ANC filter 816, or a hybrid ANC filter.

[0084] The method 1500 in Figure 15 may be performed by a processing unit such as a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a central processing unit (CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. For example, the method 1500 in Figure 15 may be performed by one or more processors that execute instructions, as described, for example, with reference to Figure 16.

[0085] Referring to Figure 16, a block diagram of a specific exemplary implementation of the device is shown, denoted as 1600 overall. In various implementations, device 1600 may have more or fewer components than those shown in Figure 16. In the exemplary implementation, device 1600 may correspond to device 100. In the exemplary implementation, device 1600 may perform one or more operations described with reference to Figures 1 to 15.

[0086] In certain implementations, device 1600 includes a processor 1606 (e.g., a central processing unit (CPU)). Device 1600 may include one or more additional processors 1610 (e.g., one or more DSPs). In certain embodiments, the processor(s) 190 in Figure 1 corresponds to processor 1606, processor(s) 1610, or a combination thereof. Processor(s) 1610 may include a speech and music codec 1608, which includes a voice coder ("vocoder") encoder 1636 and a vocoder decoder 1638. Processor(s) 1610 and / or speech and music codec 1608 also include low-latency noise suppression components such as an analysis filter bank 122, a machine learning model(s) 126, a time-domain filter designer 130, and a time-domain filter(s) 112.

[0087] In Figure 16, device 1600 includes memory 1686 and codec 1634. Memory 1686 includes instructions 1656 that can be executed by one or more additional processors 1610 (or processor 1606) to perform functions described with reference to device 100 in Figure 1 (e.g., to store). In Figure 16, device 1600 also includes a modem 1670 coupled to antenna 1652 via transceiver 1650. The modem 1670, transceiver 1650, and antenna 1652 enable device 1600 to exchange data with one or more other devices via wireless communication. For example, in some implementations, device 1600 may generate audio output in speaker(s) 108 based on data received via wireless communication with another device.

[0088] Device 1600 may include a display 1628 coupled to a display controller 1626. Speakers 108 and microphones 102 may be coupled to codec 1634. In Figure 16, codec 1634 includes a digital-to-analog converter (DAC) 1602 and an analog-to-digital converter (ADC) 1604. In a particular implementation, codec 1634 may receive an analog signal (e.g., audio signal 104 in Figures 1-8) from microphones 102, convert the analog signal to a digital signal (e.g., audio data 106 in Figures 1-8) using analog-to-digital converter 1604, and provide the digital signal to speech and music codec 1608. Speech and music codec 1608 may process the digital signal. The digital signal may be further processed by the analysis filter bank 122, one or more machine learning models 126, a time-domain filter designer 130, and one or more time-domain filters 112. For example, audio data may be provided to one or more time-domain filters 112, which applies the current set of time-domain filter coefficients to the audio data to generate a noise-suppressed output signal. In addition, in this example, the audio data may be provided to the analysis filter bank 122 to generate frequency-domain audio data. Input data based on the frequency-domain audio data may be provided to one or more machine learning models 126 to generate output data, and the noise-suppressed output based on the output data may be provided to the time-domain filter designer 130 to generate updated time-domain filter coefficients. In this example, the updated time-domain filter coefficients may then be applied to the received audio data.

[0089] In certain implementations, the speech and music codec 1608 may provide the codec 1634 with a digital signal representing a noise-suppressed output signal and / or other audio content. The codec 1634 may use the digital-to-analog converter 1602 to convert the digital signal to an analog signal and provide the analog signal to one or more speakers 108.

[0090] In certain implementations, device 1600 may be contained within a system-in-package or system-on-chip device 1622. In certain implementations, memory 1686, processor 1606, processor 1610, display controller 1626, codec 1634, and modem 1670 are contained within a system-in-package or system-on-chip device 1622. In certain implementations, input device 1630 and power supply 1644 are coupled to the system-in-package or system-on-chip device 1622. Furthermore, in certain implementations, as shown in Figure 16, the display 1628, input device 1630, speaker(s) 108, microphone(s) 102, antenna 1652, and power supply 1644 are located outside the system-in-package or system-on-chip device 1622. In certain implementations, each of the display 1628, input device 1630, speaker(s) 108, microphone(s) 102, antenna 1652, and power supply 1644 may be coupled to components of the system-in-package or system-on-chip device 1622, such as interfaces or controllers.

[0091] Device 1600 may include wearable devices such as wearable mobile communication devices, wearable personal digital assistants, wearable display devices, wearable game systems, wearable music players, wearable radios, wearable cameras, wearable navigation devices, headsets, augmented reality headsets, mixed reality headsets, virtual reality headsets, voice-activated devices, portable electronic devices, wearable computing devices, wearable communication devices, virtual reality (VR) devices, one or more earbuds, hearing aid devices, or any combination thereof.

[0092] In relation to the described implementation, the apparatus includes means for performing one or more transformation operations on a first segment of audio data in order to generate frequency-domain audio data, the audio data including a first segment and a second segment following the first segment. For example, the means for performing one or more transformation operations may correspond to device 100, one or more processors 190, analysis filter bank 122, processor 1606, one or more processors 1610, one or more other circuits or components configured to perform one or more transformation operations in order to generate frequency-domain audio data, or any combination thereof.

[0093] The apparatus also includes means for processing input data based on frequency-domain audio data as input to one or more machine learning models in order to generate a noise-suppressed output. For example, the means for processing input data may correspond to device 100, one or more processors 190, one or more machine learning models 126 (optionally together with an input pre-processing engine 820, an output post-processing engine 822, or both), a mask generator 306, processor 1606, one or more processors 1610, one or more other circuits or components configured to process input data in order to generate a noise-suppressed output, or any combination thereof.

[0094] The apparatus also includes means for performing one or more inverse transform operations on the noise suppression output to generate time-domain filter coefficients. For example, the means for performing one or more inverse transform operations may correspond to device 100, one or more processors 190, time-domain filter designer 130, processor 1606, one or more processors 1610, one or more other circuits or components configured to perform one or more inverse transform operations, or any combination thereof.

[0095] The apparatus also includes means for performing time-domain filtering of a second segment using time-domain filter coefficients to generate a noise-suppressed output signal. For example, the means for performing time-domain filtering may correspond to device 100, one or more processors 190, one or more time-domain filters 112, processor 1606, one or more processors 1610, one or more other circuits or components configured to perform time-domain filtering, or any combination thereof.

[0096] In some implementations, a non-temporary computer-readable medium (e.g., a computer-readable storage device such as memory 1686) includes an instruction (e.g., instruction 1656) which, when executed by one or more processors (e.g., one or more processors 1610 or processor 1606), causes one or more processors to acquire audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment following the first segment; to perform one or more transformation operations on the first segment to generate frequency-domain audio data; to provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; to perform one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients; and to perform time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0097] Specific aspects of this disclosure are described below in a set of interrelated embodiments.

[0098] According to Embodiment 1, the device includes one or more processors configured to acquire audio data representing one or more audio signals, wherein the audio data comprises a first segment and a second segment following the first segment; perform one or more transformation operations on the first segment to generate frequency-domain audio data; provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; perform one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients; and perform time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0099] Example 2 includes the device of Example 1, wherein a noise-suppressed output signal is generated with a latency of less than 1 millisecond, and time-domain filter coefficients are generated with a latency of more than 1 millisecond.

[0100] Example 3 includes the device of Example 1 or Example 2, and the input data includes frequency domain audio data.

[0101] Example 4 includes any of the devices from Examples 1 to 3, wherein one or more machine learning models are configured to generate an output that includes a frequency mask representing the estimated magnitude of noise in the frequency domain audio data for each of the multiple frequency bins, and the noise suppression output includes the frequency mask.

[0102] Example 5 includes any of the devices from Examples 1 to 3, wherein one or more machine learning models are configured to generate an output containing noise-suppressed audio data, and one or more processors are configured to determine a frequency mask representing the estimated magnitude of noise in the frequency-domain audio data for each of a plurality of frequency bins based on the noise-suppressed audio data, and the noise-suppressed output contains the frequency mask.

[0103] Example 6 includes any of the devices from Examples 1 to 5, and one or more machine learning models include one or more recurrent neural networks.

[0104] Example 7 includes any of the devices from Examples 1 to 6 and is configured such that, in order to generate a noise suppression output, one or more processors perform a beamforming operation on frequency-domain audio data to determine beamformed audio data that distinguishes between portions of audio data from a target audio source and portions of audio data from a non-target audio source, and the input data includes or is based on beamformed audio data.

[0105] Example 8 includes any of the devices from Examples 1 to 7, and is configured such that one or more processors perform speech augmentation operations to determine speech augmented audio data in order to process frequency domain audio data to generate a noise suppression output, and the input data includes or is based on speech augmented audio data.

[0106] Example 9 includes any of the devices from Examples 1 to 8, and is configured such that one or more processors perform source isolation operations to determine source isolation audio data in order to process frequency domain audio data to generate a noise suppression output, and the input data includes or is based on source isolation audio data.

[0107] Example 10 includes any device from Examples 1 to 9, wherein the time-domain filter coefficients include linear phase finite impulse response (FIR) filter coefficients, minimum phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.

[0108] Example 11 includes any device from Examples 1 to 10, in which one or more processors are integrated into the wearable device.

[0109] Example 12 includes any device from Examples 1 to 11, in which one or more processors are integrated into one or more earbuds.

[0110] Example 13 includes any of the devices from Examples 1 to 12, further including one or more microphones, and one or more audio signals are received from one or more microphones.

[0111] Example 14 includes the device of Example 13 and further includes an adaptive noise cancellation circuit coupled to at least one of the one or more microphones.

[0112] Example 15 includes any of the devices from Examples 1 to 14, and further includes one or more speakers and one or more microphones coupled to one or more processors and integrated into a wearable device, wherein one or more microphones include at least one external microphone configured to generate audio data and at least one feedback microphone configured to generate a feedback signal based on the sound generated by one or more speakers in response to a noise-suppressed output signal.

[0113] According to Example 16, the method includes: acquiring audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment following the first segment; performing one or more transformation operations on the first segment to generate frequency-domain audio data; providing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; performing one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients; and performing time-domain filtering on the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0114] Example 17 includes the method of Example 16, wherein a noise-suppressed output signal is generated with a latency of less than 1 millisecond, and time-domain filter coefficients are generated with a latency of more than 1 millisecond.

[0115] Example 18 includes the method of Example 16 or Example 17, wherein the input data includes frequency domain audio data.

[0116] Example 19 comprises any of the methods of Examples 16 to 18, wherein one or more machine learning models are configured to generate an output that includes a frequency mask representing the estimated magnitude of noise in frequency-domain audio data for each frequency bin of a plurality of frequency bins, and the noise suppression output includes the frequency mask.

[0117] Example 20 comprises any of the methods of Examples 16 to 18, wherein one or more machine learning models are configured to generate an output containing noise-suppressed audio data, the method further comprises determining a frequency mask representing the estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, based on the noise-suppressed audio data, and the noise-suppressed output contains the frequency mask.

[0118] Example 21 includes any of the methods from Examples 16 to 20, wherein one or more machine learning models include one or more recurrent neural networks.

[0119] Example 22 comprises any method of Examples 16 to 21, further comprising performing a beamforming operation on frequency-domain audio data to determine beamformed audio data that distinguishes between portions of audio data from a target audio source and portions of audio data from a non-target audio source, wherein the input data is based on or includes beamformed audio data.

[0120] Example 23 comprises any method of Examples 16 to 22, further comprising performing a speech augmentation operation to determine speech augmented audio data, wherein the input data is based on or includes speech augmented audio data.

[0121] Example 24 comprises any method of Examples 16 to 23, further comprising performing a source separation operation to determine source-separated audio data, wherein the input data is based on or includes source-separated audio data.

[0122] Example 25 includes any method from Example 16 to Example 24, wherein the time-domain filter coefficients include linear phase finite impulse response (FIR) filter coefficients, minimum phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.

[0123] Example 26 includes any of the methods of Examples 16 to 25, wherein one or more audio signals are received from one or more microphones.

[0124] Example 27 includes any method from Examples 16 to 26, and further includes performing an adaptive noise cancellation operation based on a noise-suppressed output signal.

[0125] According to Example 28, the device includes a memory configured to store instructions and a processor configured to execute instructions in order to perform any of the methods of Examples 16 to 27.

[0126] According to Example 29, a non-temporary computer-readable medium stores instructions, and when the instructions are executed by the processor, the processor causes the processor to perform any of the methods in Examples 16 to 27.

[0127] According to Example 30, the apparatus includes means for carrying out any of the methods of Examples 16 to 27.

[0128] According to Example 31, a non-temporary computer-readable medium stores instructions, which can be executed by one or more processors to cause one or more processors to acquire audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment following the first segment; to cause one or more transformation operations to be performed on the first segment to generate frequency-domain audio data; to cause input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; to cause one or more inverse transformation operations to be performed on the noise-suppressed output to generate time-domain filter coefficients; and to cause time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0129] Example 32 includes the non-temporary computer-readable medium of Example 31, wherein the instructions cause one or more processors to generate a noise-suppressed output signal with a latency of less than 1 millisecond, and cause one or more processors to generate time-domain filter coefficients with a latency of more than 1 millisecond.

[0130] Example 33 includes the non-temporary computer-readable medium of Example 31, wherein the input data includes frequency-domain audio data.

[0131] Example 34 includes a non-temporary computer-readable medium from any of Examples 31 to 33, wherein one or more machine learning models are configured to generate an output that includes a frequency mask representing the estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and the noise suppression output includes the frequency mask.

[0132] Example 35 includes a non-temporary computer-readable medium from any of Examples 31 to 33, wherein one or more machine learning models are configured to generate an output containing noise-suppressed audio data, and instructions cause one or more processors to determine a frequency mask representing the estimated magnitude of noise in the frequency-domain audio data for each of a plurality of frequency bins, based on the noise-suppressed audio data, and the noise-suppressed output contains the frequency mask.

[0133] Example 36 includes a non-temporary computer-readable medium from any of Examples 31 to 35, and one or more machine learning models include one or more recurrent neural networks.

[0134] Example 37 includes a non-temporary computer-readable medium from any of Examples 31 to 36, and the instructions are sent to one or more processors to generate a noise-suppressed output. To determine beamformed audio data that distinguishes between portions of audio data from a target audio source and portions of audio data from a non-target audio source, a beamforming operation is performed on frequency-domain audio data, and the input data is based on or includes beamformed audio data.

[0135] Example 38 includes a non-temporary computer-readable medium of any of Examples 31 to 37, and in order to process frequency-domain audio data to produce a noise-suppressed output, the instruction causes one or more processors to perform speech augmentation operations to determine speech augmented audio data, and the input data is based on or includes speech augmented audio data.

[0136] Example 39 includes a non-temporary computer-readable medium of any of Examples 31 to 38, and in order to process frequency-domain audio data to produce a noise-suppressed output, the instruction causes one or more processors to perform a source-separation operation to determine source-separated audio data, and the input data is based on or includes source-separated audio data.

[0137] Example 40 includes a non-temporary computer-readable medium of any of Examples 31 to 39, wherein the time-domain filter coefficients include linear phase finite impulse response (FIR) filter coefficients, minimum phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.

[0138] Example 41 includes the non-temporary computer-readable medium of Example 40, wherein the instruction causes one or more processors to perform adaptive noise cancellation operations based on a noise-suppressed output signal.

[0139] According to Example 42, the apparatus includes, for generating frequency-domain audio data, performing one or more transformation operations on a first segment of audio data, wherein the audio data includes a first segment and a second segment following the first segment; means for processing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; means for performing one or more inverse transformation operations on the noise-suppressed output to generate time-domain filter coefficients; and means for performing time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.

[0140] Example 43 includes the apparatus of Example 42, wherein means for performing time-domain filtering are operable to produce a noise-suppressed output signal with a latency of 1 millisecond or less, and means for performing one or more transformation operations, means for processing input data, and means for determining time-domain filter coefficients are operable to produce time-domain filter coefficients with a latency of more than 1 millisecond.

[0141] Example 44 includes the apparatus of Example 42, and the input data includes frequency domain audio data.

[0142] Example 45 includes the apparatus of any of Examples 42 to 44, wherein one or more machine learning models are configured to generate an output that includes a frequency mask representing the estimated magnitude of noise in frequency-domain audio data for each frequency bin of a plurality of frequency bins, and the noise suppression output includes the frequency mask.

[0143] Example 46 comprises any apparatus from Examples 42 to 44, wherein one or more machine learning models are configured to generate an output containing noise-suppressed audio data, and further comprises means for determining a frequency mask representing the estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins based on the noise-suppressed audio data, wherein the noise-suppressed output includes the frequency mask.

[0144] Example 47 includes the apparatus of any of Examples 42 to 46, and one or more machine learning models include one or more recurrent neural networks.

[0145] Example 48 includes the apparatus of any of Examples 42 to 47 and further includes means for performing a beamforming operation on frequency-domain audio data to determine beamformed audio data that distinguishes between portions of audio data from a target audio source and portions of audio data from a non-target audio source, wherein the input data is based on or includes beamformed audio data.

[0146] Example 49 includes the apparatus of any of Examples 42 to 48, and further includes means for performing a speech augmentation operation to determine speech augmented audio data, wherein the input data is based on or includes speech augmented audio data.

[0147] Example 50 includes the apparatus of any of Examples 42 to 49, and further includes means for performing a source separation operation to determine source-separated audio data, wherein the input data is based on or includes source-separated audio data.

[0148] Example 51 includes the apparatus of any of Examples 42 to 50, wherein the time-domain filter coefficients include linear phase finite impulse response (FIR) filter coefficients, minimum phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.

[0149] Example 52 includes the apparatus of any of Examples 42 to 51, wherein means for performing time-domain filtering, means for performing one or more transformation operations, means for processing input data, and means for performing one or more inverse transformation operations are integrated into a wearable device.

[0150] Example 53 includes the apparatus of any of Examples 42 to 52, wherein means for performing time-domain filtering, means for performing one or more transformation operations, means for processing input data, and means for performing one or more inverse transformation operations are integrated into one or more earbuds.

[0151] Example 54 includes the apparatus of any of Examples 42 to 53 and further includes means for generating one or more audio signals based on ambient sound.

[0152] Example 55 includes the apparatus of any of Examples 42 to 54 and further includes means for performing adaptive noise cancellation based on a noise-suppressed output signal.

[0153] Example 56 includes the apparatus of any of Examples 42 to 55 and further includes means for generating an output sound based on a noise-suppressed output signal, means for generating a feedback signal based on the output sound, and means for performing adaptive noise cancellation based on the noise-suppressed output signal and the feedback signal.

[0154] Those skilled in the art will further understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithmic steps described herein with respect to the implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The functions of various exemplary components, blocks, configurations, modules, circuits, and steps have been described conceptually above. Whether such functions are implemented as hardware or as processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art will understand that the described functions can be implemented in various ways for each specific application, and such a decision on implementation should not be construed as a departure from the scope of this disclosure.

[0155] Steps of the methods or algorithms described in relation to the implementations disclosed herein may be embodied directly in hardware, in software modules executed by a processor, or in a combination of the two. The software modules may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compact disc read-only memory (CD-ROM), or any other form of non-temporary storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside within a computing device or user terminal. Alternatively, the processor and storage medium may reside as separate components within the computing device or user terminal.

[0156] The foregoing description of the disclosed embodiments is provided to enable those skilled in the art to create or use the disclosed embodiments. Various modifications of these embodiments will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other embodiments without departing from the scope of this disclosure. Accordingly, this disclosure is not intended to be limited to the embodiments shown herein, but should be given the broadest possible scope to coincide with the principles and novel features defined by the following claims.

Claims

1. It is a device, It comprises one or more processors, and the one or more processors Acquiring audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment following the first segment. To generate frequency domain audio data, one or more transformation operations are performed on the first segment, To generate a noise-suppressed output, input data based on the frequency-domain audio data is provided as input to one or more machine learning models. To generate time-domain filter coefficients, one or more inverse transform operations are performed on the noise suppression output, To generate a noise-suppressed output signal, time-domain filtering of the second segment is performed using the time-domain filter coefficients, A device configured to perform the following actions.

2. The device according to claim 1, wherein the input data includes the frequency domain audio data.

3. The device according to claim 1, wherein one or more machine learning models are configured to generate an output including a frequency mask representing the estimated magnitude of noise in the frequency domain audio data for each frequency bin of a plurality of frequency bins, and the noise suppression output includes the frequency mask.

4. The device according to claim 1, wherein one or more machine learning models are configured to generate an output including noise-suppressed audio data, and one or more processors are configured to determine a frequency mask representing the estimated magnitude of noise in the frequency domain audio data for each frequency bin of a plurality of frequency bins, based on the noise-suppressed audio data, and the noise-suppressed output includes the frequency mask.

5. The device according to claim 1, wherein, in order to generate the noise suppression output, one or more processors are configured to perform a beamforming operation on the frequency domain audio data to determine beamformed audio data that distinguishes between portions of the audio data from a target audio source and portions of the audio data from a non-target audio source, and the input data includes the beamformed audio data.

6. The device according to claim 1, wherein, in order to process the frequency domain audio data to generate the noise suppression output, one or more processors are configured to perform speech augmentation operations to determine speech augmented audio data, and the input data includes the speech augmented audio data.

7. The device according to claim 1, wherein, in order to process the frequency domain audio data to generate the noise suppression output, one or more processors are configured to perform a source isolation operation to determine source isolated audio data, and the input data includes the source isolated audio data.

8. The device according to claim 1, wherein the time-domain filter coefficients include linear phase finite impulse response (FIR) filter coefficients, minimum phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.

9. The device according to claim 1, wherein the one or more processors are integrated into a wearable device.

10. The device according to claim 1, further comprising one or more microphones, wherein one or more audio signals are received from the one or more microphones.

11. The device according to claim 10, further comprising an adaptive noise cancellation filter coupled to at least one of the one or more microphones.

12. The device according to claim 1, further comprising one or more speakers and one or more microphones coupled to the one or more processors and integrated into a wearable device, wherein the one or more microphones include at least one external microphone configured to generate the audio data and at least one feedback microphone configured to generate a feedback signal based on the sound generated by the one or more speakers in response to the noise-suppressed output signal.

13. Acquiring audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment following the first segment. To generate frequency domain audio data, one or more transformation operations are performed on the first segment, To generate a noise-suppressed output, input data based on the frequency-domain audio data is provided as input to one or more machine learning models. To generate time-domain filter coefficients, one or more inverse transform operations are performed on the noise suppression output, To generate a noise-suppressed output signal, time-domain filtering of the second segment is performed using the time-domain filter coefficients, Methods that include...

14. The method according to claim 13, wherein the input data includes the frequency domain audio data.

15. The method according to claim 13, wherein the one or more machine learning models are configured to generate an output that includes a frequency mask representing the estimated magnitude of noise in the frequency domain audio data for each frequency bin of a plurality of frequency bins, and the noise suppression output includes the frequency mask.

16. The method according to claim 13, wherein the one or more machine learning models are configured to generate an output including noise-suppressed audio data, and the method further comprises determining a frequency mask representing the estimated magnitude of noise in the frequency domain audio data for each frequency bin of a plurality of frequency bins, and the noise-suppressed output includes the frequency mask.

17. The method according to claim 13, further comprising performing a beamforming operation on the frequency domain audio data to determine beamformed audio data that distinguishes between a portion of the audio data from a target audio source and a portion of the audio data from a non-target audio source, wherein the input data includes the beamformed audio data.

18. The method according to claim 13, further comprising performing a speech augmentation operation to determine speech augmented audio data, wherein the input data includes the speech augmented audio data.

19. The method according to claim 13, further comprising performing a source separation operation to determine source-separated audio data, wherein the input data includes the source-separated audio data.

20. A non-temporary computer-readable medium for storing instructions, wherein when an instruction is executed by one or more processors, the one or more processors, Acquiring audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment following the first segment. To generate frequency domain audio data, one or more transformation operations are performed on the first segment, To generate a noise-suppressed output, input data based on the frequency-domain audio data is provided as input to one or more machine learning models. To generate time-domain filter coefficients, one or more inverse transform operations are performed on the noise suppression output, To generate a noise-suppressed output signal, time-domain filtering of the second segment is performed using the time-domain filter coefficients, A non-temporary computer-readable medium capable of performing the following actions.