DISPOSITIVO, MÉTODO, E, MEIO NÃO TRANSITÓRIO LEGÍVEL POR COMPUTADOR
Patent Information
- Application Number
- BR112025019967
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-20
- Filing Date
- 2024-03-21
- Publication Date
- 2026-08-04
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
1 / 71 “COMPUTER-READABLE NON-TRANSIENTIAL DEVICE, METHOD, AND MEANS I. Cross-Reference to Related Deposit Requests
[0001] The present application claims priority over provisional patent application No. US 63 / 493,158, filed March 30, 2023, and non-provisional patent application No. US 18 / 611,308, filed March 20, 2024, the contents of which are expressly incorporated herein by reference in their entirety. II. Field
[0002] This disclosure refers generally to low-latency noise suppression. III. Description of the Related Technique
[0003] Various types of hearing-related problems affect a significant number of people. For example, a common problem is that even people with relatively normal hearing may have difficulty hearing speech in noisy environments, and the problem can be considerably worse for those with hearing loss. For some individuals, speech is readily intelligible only when the signal-to-noise ratio (of speech relative to ambient noise) is above a certain level.
[0004] Wearable devices (e.g., in-ear headphones, earphones, hearing aids, etc.) can be used to improve hearing, situational awareness, speech intelligibility, etc. in many circumstances. Generally, such devices employ relatively simple noise suppression processes to remove as much ambient noise as possible. Although such Petition 870250084259, dated 09 / 18 / 2025, page 8 / 171 2 / 71 While noise suppression processes can improve the signal-to-noise ratio sufficiently for speech to be intelligible, these noise suppression processes can also reduce the user's situational awareness, since these processes simply attempt to remove as much noise as possible, thus possibly removing important environmental cues such as traffic sounds. The use of more complex noise suppression processes can introduce significant latency. Latency in real-time speech processing can lead to user dissatisfaction. IV. Summary
[0005] According to one implementation of this disclosure, a device includes one or more processors configured to obtain audio data representing one or more audio signals. The audio data includes a first segment and a second segment subsequent to the first segment. The one or more processors are configured to perform one or more transform operations on the first segment to generate frequency-domain audio data. The one or more processors are configured to provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressing output. The one or more processors are configured to perform one or more reverse transform operations on the noise-suppressing output to generate time-domain filter coefficients.One or more processors are configured to perform time-domain filtering of the second segment using time-domain filter coefficients to generate a signal. Petition 870250084259, dated 09 / 18 / 2025, page 9 / 171 3 / 71 output with suppressed noise.
[0006] According to another implementation of the present disclosure, a method includes obtaining audio data representing one or more audio signals. The audio data includes a first segment and a second segment subsequent to the first segment. The method includes performing one or more transform operations on the first segment to generate frequency-domain audio data. The method includes providing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output. The method includes performing one or more reverse transform operations on the noise-suppressed output to generate time-domain filter coefficients. The method includes performing time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.
[0007] According to another implementation of the present disclosure, a non-transient, computer-readable medium stores instructions that are executable by one or more processors to cause the one or more processors to obtain audio data representing one or more audio signals. The audio data includes a first segment and a second segment subsequent to the first segment. The instructions are executable to cause the one or more processors to perform one or more transform operations on the first segment to generate frequency-domain audio data. The instructions are executable to cause the one or more processors to provide input data based on the frequency-domain audio data. Petition 870250084259, dated 09 / 18 / 2025, page 10 / 171 4 / 71 as input to one or more machine learning models to generate a noise-suppressed output. The instructions are executable to cause one or more processors to perform one or more reverse transform operations on the noise-suppressed output to generate time-domain filter coefficients. The instructions are executable to cause one or more processors to perform second-segment time-domain filtering using the time-domain filter coefficients to generate a noise-suppressed output signal.
[0008] According to another implementation of the present disclosure, an apparatus includes means for performing one or more transform operations on a first segment of audio data to generate frequency-domain audio data, wherein the audio data includes a first segment and a second segment subsequent to the first segment. The apparatus also includes means for processing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output. The apparatus also includes means for performing one or more reverse transform operations on the noise-suppressed output to generate time-domain filter coefficients. The apparatus also includes means for performing time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.
[0009] Other aspects, advantages and attributes of this disclosure will become apparent after reviewing the entire application, including the following sections: Brief description Petition 870250084259, dated 09 / 18 / 2025, page 11 / 171 5 / 71 of the drawings, Detailed description and Claims. V. Brief Description of the Drawings
[0010] Figure 1 is a block diagram of particular aspects of an operable device for performing low-latency noise suppression, according to some examples in this disclosure.
[0011] Figure 2 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure.
[0012] Figure 3 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure.
[0013] Figure 4 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure.
[0014] Figure 5 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure.
[0015] Figure 6 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure.
[0016] Figure 7 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure. Petition 870250084259, dated 09 / 18 / 2025, page 12 / 171 6 / 71
[0017] Figure 8 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure.
[0018] Figure 9 illustrates an example of an operable integrated circuit for achieving low-latency noise suppression, according to some examples in this disclosure.
[0019] Figure 10 is a diagram of an operable headset for achieving low-latency noise suppression, according to some examples in this disclosure.
[0020] Figure 11 is a diagram of a headset, such as a virtual reality, mixed reality or augmented reality headset, operable to perform low-latency noise suppression, according to some examples in this disclosure.
[0021] Figure 12 is a diagram of operable augmented reality glasses for achieving low-latency noise suppression, according to some examples in this disclosure.
[0022] Figure 13 is a diagram of an operable wearable device for performing low-latency noise suppression, according to some examples in this disclosure.
[0023] Figure 14 is a diagram of operable in-ear headphones for achieving low-latency noise suppression, according to some examples in this disclosure.
[0024] Figure 15 is a diagram of a particular implementation of a realization method. Petition 870250084259, dated 09 / 18 / 2025, page 13 / 171 7 / 71 Low-latency noise suppression that can be achieved by the device in Figure 1, according to some examples in this disclosure.
[0025] Figure 16 is a block diagram of a particular illustrative example of a device that is operable to perform low-latency noise suppression, according to some examples in the present disclosure. VI. Detailed Description
[0026] In contexts where latency constraints are sufficiently flexible, machine learning-based noise suppression processes can be used to reduce the magnitude of noise components in audio data, to increase the magnitude of target sound components in audio data, or both. However, machine learning processes and related pre-processing and post-processing can introduce significant delay. To illustrate, machine learning noise suppression models generally operate in the frequency domain, which implies transforming audio data from a time domain to a frequency domain to generate input data for a machine learning model.Additionally, after the machine learning model processes the input data to generate de-noise audio data, the de-noise audio data is transformed back into the time domain for output to a user. Each of these operations introduces delay, which can lead to unacceptable latency, especially for real-time speech audio processing.
[0027] The aspects disclosed in the present invention enable audio processing in a way that Petition 870250084259, dated 09 / 18 / 2025, page 14 / 171 8 / 71 provides high-quality noise suppression without introducing undue latency. According to a particular aspect, noise suppression is achieved in the time domain by applying one or more time-domain filters to received audio data. The coefficients of the one or more time-domain filters are determined based on the noise suppression output generated in the frequency domain by one or more machine learning models. The time-domain filter coefficients determined based on the noise suppression output of the machine learning model(s) provide significantly better noise suppression than traditional time-domain processes such as adaptive noise cancellation.Additionally, once time-domain filter coefficients are applied to receive time-domain audio data, little or no latency is added when using such time-domain filter coefficients to process the audio data.
[0028] Particular aspects of the present disclosure are described below with reference to the drawings. In the description, common attributes are referred to as common reference numbers. As used in the present invention, various terminologies are used for the purpose of describing only particular implementations, and are not intended to be limiting of implementations. For example, the singular forms a, an, and are intended to include plural forms as well, unless the context clearly indicates otherwise. Furthermore, some attributes described in the present invention are singular in some implementations and plural in other implementations. To illustrate, Figure 1 depicts a device 100 that includes one or more processors. Petition 870250084259, dated 09 / 18 / 2025, p. 15 / 171 9 / 71 (processor(s) 190 of Figure 1), which indicates that, in some implementations, the device 100 includes a single processor 190 and, in other implementations, the device 100 includes multiple processors 190. For ease of reference in the present invention, such attributes are generally introduced as one or more attributes and are subsequently referred to in the singular or optional plural (as indicated by (s)), unless aspects relating to several of the attributes are being described.
[0029] In some drawings, multiple instances of a particular type of attribute are used. Although these attributes are physically and / or logically distinct, the same reference number is used for each, and the different instances are distinguished by adding a letter to the reference number. When attributes as a group or a type are referenced in the present invention, for example, when none of the particular attributes are being referenced, the reference number is used without a distinguishing letter. However, when reference is made in the present invention to a particular attribute of multiple attributes of the same type, the reference number is used with the distinguishing letter. For example, with reference to Figure 5, multiple microphones are illustrated and associated with the reference numbers 102A and 102B. When reference is made to one microphone among these microphones, such as a 102A microphone, the distinguishing letter A is used.However, when reference is made to any arbitrary microphone among these microphones, or to these microphones as a group, the reference number 102 is used without a distinguishing letter.
[0030] As used in the present invention, the Petition 870250084259, dated 09 / 18 / 2025, page 16 / 171 10 / 71 The terms comprise, includes, and comprising can be used interchangeably with include, includes, or including. Additionally, the term in which can be used interchangeably with where. As used in the present invention, exemplifier indicates an example, an implementation, and / or an aspect, and should not be interpreted as limiting or as indicating a preference or a preferred implementation. As used in the present invention, an ordinal term (e.g., first, second, third, etc.) used to modify an element, such as a structure, a component, an operation, etc., does not, by itself, indicate any priority or order of the element in relation to another element, but simply distinguishes the element from another element that has the same name (but for use of the ordinal term).As used in the present invention, the term "set" refers to one or more of a particular element, and the term "plurality" refers to multiples (e.g., two or more) of a particular element.
[0031] As used in the present invention, coupled may include communicatively coupled, electrically coupled or physically coupled, and may also (or alternatively) include any combination thereof. Two devices (or components) may be coupled (e.g., communicatively coupled, electrically coupled or physically coupled), directly or indirectly, via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network or a combination thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or Petition 870250084259, dated 09 / 18 / 2025, page 17 / 171 11 / 71 in different devices, and can be connected via electronic circuits, one or more connectors, or inductive coupling, as illustrative and non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, as in an electrical communication, can send and receive signals (e.g., digital signals or analog signals) directly or indirectly, via one or more wires, buses, networks, etc. As used in the present invention, directly coupled can include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without intervening components.
[0032] In the present disclosure, terms such as determine, calculate, estimate, move, adjust, etc. may be used to describe how one or more operations are performed. It should be noted that such terms should not be interpreted as limiting, and other techniques may be used to perform similar operations. Additionally, as referred to in the present invention, generate, calculate, estimate, use, select, access, and determine may be used interchangeably. For example, generate, calculate, estimate, or determine a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or it may refer to using, selecting, or accessing the parameter (or signal) that has already been generated, such as by another component or device.
[0033] Figure 1 is a block diagram of particular aspects of an operable device 100 for performing low-latency noise suppression, according to some examples in this disclosure. In Figure 1, the Petition 870250084259, dated 09 / 18 / 2025, p. 18 / 171 12 / 71 device 100 includes or is coupled to one or more microphones 102 and one or more speakers 108.
[0034] The microphone(s) 102 is / are configured to generate audio signal(s) 104 based on sound 170 detected in an environment. Sound 170 may include a target sound, such as speech, as well as non-target sounds (e.g., noise). Device 100 is configured to process audio data 106 representing the audio signal(s) 104 to generate a noise-suppressed output signal 114, which can be used to drive the speaker(s) 108 to generate output sound 172. In the output sound 172, the components of the audio data 106 corresponding to the target sound(s) are emphasized, the components of the audio data 106 corresponding to the non-target sound(s) are minimized, or both, relative to sound 170. In particular implementations, device 100 includes, is included in, or corresponds to a wearable device.In such implementations, a user can wear the 100 device in, on, or near one or both of the user's ears to enhance perception of target sounds, to decrease perception of non-target sounds, or both. For example, the user can wear the 100 device to enhance speech perception in a noisy environment.
[0035] In Figure 1, processor(s) 190 is / are configured to perform operations associated with two data paths, including a first data path 110 and a second data path 120. The first data path 110 is a low-latency data path. To reduce the delay, the first data path 110 is configured to perform operations on the audio data 106 in the time domain. Consequently, Petition 870250084259, dated 09 / 18 / 2025, page 19 / 171 13 / 71 The delays associated with domain transform operations are avoided in the first data path 110. In contrast, the second data path 120 is configured to provide high-quality noise suppression at the cost of higher latency than the first data path 110. For example, in some implementations, the first data path 110 is associated with a latency of 1 millisecond or less, and the second data path 120 is associated with a latency of more than 1 millisecond. As another example, in some implementations, the first data path 110 may be associated with a latency of 2 milliseconds or less, and the second data path 120 is associated with a latency of more than 10 milliseconds.
[0036] The second data path 120 is configured to determine time-domain filter coefficients 132 which, in some implementations, are stored in a buffer 140 to be subsequently applied by one or more time-domain filters 112 in the first data path 110. In Figure 1, the second data path 120 includes an analysis filter bank 122, one or more machine learning models 126, and a time-domain filter designer 130. The buffer 140 may be included in the time-domain filter designer 130, in the time-domain filter(s) 112, or be distinct from both the time-domain filter designer 130 and the time-domain filter(s) 112.
[0037] The analysis filter bank 122 is configured to perform one or more transform operations (for example, fast-Fourier transforms (FFT) operations), based on Petition 870250084259, dated 09 / 18 / 2025, page 20 / 171 14 / 71 samples of audio data 106, to generate frequency domain audio data 124. According to some implementations, a sample set of audio data 106 is accumulated for processing by the analysis filter bank 122 (e.g., in one or more buffers of the analysis filter bank 122), and the frequency domain audio data 124 for the sample set includes information indicating the magnitude of the sound within each frequency bin of a plurality of frequency bins. During or after the transformation of the sample set into the frequency domain, a subsequent sample set of audio data 106 is accumulated to be transformed into a subsequent set of frequency domain audio data 124.
[0038] The machine learning model(s) 126 include(s) one or more trained models, such as neural network(s), that is / are configured to process input data 125 based on frequency domain audio data 124 to generate output data 129. In particular implementations, the machine learning model(s) 126 is / are temporally dynamic, such that the output data 129 that is based on a particular set of input data 125 is affected by one or more previous sets of input data 125. For example, the machine learning model(s) 126 may include one or more recurrent neural networks, such as neural network(s) including one or more long-term memory layers, one or more synchronized recurrent units, or other recurrent structures.
[0039] Time-domain filter designer 130 is configured to process the suppression output of Petition 870250084259, dated 09 / 18 / 2025, p. 21 / 171 15 / 71 noise 128 that relies on output data 129 to generate time-domain filter coefficients 132. As an example, the time-domain filter designer 130 can perform one or more reverse transform operations using the noise suppression output 128 to generate time-domain filter coefficients 132. The time-domain filter designer 130 can be configured to generate time-domain filter coefficients 132 as real-value masks or as complex-value masks. Examples of real-value masks that can be generated by the time-domain filter designer 130 in some implementations include linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, all-pole filter coefficients, etc.Complex value masks may include FIR or IIR filters indicating magnitude and phase. A technical benefit of using linear-phase FIR filter coefficients is a predictable delay because the delay associated with applying linear-phase FIR filter coefficients is entirely dependent on the length of the FIR filter. A technical benefit of using minimum-phase FIR filter coefficients or autoregressive filter coefficients is the reduced delay because, although the delay introduced by applying such filter coefficients is frequency-dependent, the delay is minimized for the particular input data.
[0040] Time-domain filter(s) 112 Petition 870250084259, dated 09 / 18 / 2025, page 22 / 171 16 / 71 is / are updated periodically or occasionally (for example, when updated time-domain filter coefficients 132 become available). In this arrangement, the time-domain filter(s) 112 apply time-domain filter coefficients 132 to audio data 106 that are newer than the audio data 106 used to generate the time-domain filter coefficients 132. For example, the audio data 106 might include a first segment (for example, a portion of the audio data for a particular time period) and a second segment (for example, a portion of the audio data for a later time period) that is subsequent (for example, immediately after in time, or after one or more other segments that immediately follow) to the first segment. In this example, the time-domain filter coefficients 132 can be determined using the first segment and can be applied to the second segment.As a non-limiting example, the operations on the second data path 120 can be performed over a period of about 16 milliseconds; while the operations on the first data path 110 can be performed over a period of about 1 millisecond. Thus, in this specific example, the time-domain filter coefficients 132 applied to a particular data sample on the first data path 110 are always at least 16 milliseconds older than the data sample. Except in extraordinary circumstances, ambient noise typically changes slowly enough that, even with such a delay, the time-domain filter coefficients 132 are sufficiently representative to provide significant noise suppression. Petition 870250084259, dated 09 / 18 / 2025, page 23 / 171 17 / 71
[0041] During operation, microphone(s) 102 generate(s) audio signal(s) 104 based on sound 170. Sound 170 may include speech or other target sounds, as well as non-target sounds such as noise. Audio data 106 representing audio signal(s) 104 are provided to the first data path 110 for low-latency processing and to the second data path 120 to update the time-domain filter coefficients 132.
[0042] In the first data path 110, the audio data 106 is processed by the time-domain filter(s) 112 by applying a first set of time-domain filter coefficients 132 to the audio data 106. The first set of time-domain filter coefficients 132 is received from the second data path 120 after processing a previous set of audio data 106. The time-domain filter(s) 112 generate(s) the noise-suppressed output signal 114 by applying the first set of time-domain filter coefficients 132 to the audio data 106. The noise-suppressed output signal 114 is provided to (e.g., is used to drive) the loudspeaker 108 to generate the output sound 172. Optionally, the processor(s) 190 may perform further processing on the noisy output signal. suppressed 114 before providing the suppressed noise output signal 114 to the loudspeaker 108.For example, processor(s) 190 may perform feedforward adaptive noise cancellation (ANC), feedback ANC or hybrid ANC to further process the noise-suppressed output signal 114.
[0043] In the second data path 120, one Petition 870250084259, dated 09 / 18 / 2025, p. 24 / 171 18 / 71 The sample set of audio data 106 is accumulated and subjected to one or more transformation operations by the analysis filter bank 122 to generate the frequency domain audio data 124 that represent the sample set. In some implementations, the frequency domain audio data 124 is provided as input (e.g., as the input data 125) to the machine learning model(s) 126. In other implementations, the frequency domain audio data 124 is processed to generate the input data 125. For example, the frequency domain audio data 124 may optionally be subjected to a variety of frequency domain noise suppression or signal boosting operations to generate the input data 125.To illustrate, input data 125 can be generated by subjecting frequency-domain audio data 124 to beamforming operations, blind source separation operations, speech augmentation operations, or combinations thereof. Performing such conventional frequency-domain operations to generate input data 125 can improve the operation of the machine learning model(s) 126 by providing cleaner input data 125. In the same or different implementations, the generation of input data 125 based on frequency-domain audio data 124 may include data aggregation operations, filtering operations, resampling operations, etc.
[0044] The machine learning model(s) 126 perform(s) nonlinear, temporally dynamic operations, based on the prior training of the machine learning model(s) 126 to generate the output data 129. Petition 870250084259, dated 09 / 18 / 2025, p. 25 / 171 19 / 71 In some implementations, the machine learning model(s) 126 is / are trained to generate output data 129 that includes target audio components from the frequency-domain audio data 124 and omits or suppresses non-target audio components from the frequency-domain audio data 124. Such models are referred to in the present invention as inline models. Conversely, in some implementations, the machine learning model(s) 126 is / are trained to generate output data 129 that includes the non-target audio components from the frequency-domain audio data 124 and omits or suppresses the target audio components from the frequency-domain audio data 124. Such models are referred to in the present invention as masking models. The output data 129 of masking models can be used directly as the noise suppression output 128.The output data 129 from online models can be further processed to generate the noise suppression output 128, as further described with reference to Figure 3.
[0045] The noise suppression output 128 represents an estimate (by the machine learning model(s) 126) of a portion (e.g., audio components) of the frequency domain audio data 124 that corresponds to non-target audio (e.g., noise). The noise suppression output 128 is provided as input to the time domain filter designer 130. The time domain filter designer 130 performs inverse transform operations and parameterization operations to generate the time domain filter coefficients 132. The specific inverse transform operations and parameterization operations performed may differ. Petition 870250084259, dated 09 / 18 / 2025, page 26 / 171 20 / 71 for different implementations. For example, inverse transform operations may include multiple inverse Fourier transform operations, such as inverse fast Fourier transform (IFFT) operations. Parameterization operations may include, for example, windowing or shifting time-domain data generated by inverse transform operations to generate a specific number of time-domain filter coefficients 132 based, for example, on multiple filter coefficients applied by the time-domain filter(s) 112. Applying a larger number of filter coefficients may provide greater noise suppression at the cost of greater computational complexity.
[0046] Although time-domain filter coefficients 132 based on a first set of audio data samples 106 are being generated, additional samples of audio data 106 can be received and aggregated to form a second set of audio data samples. After the second set of samples is collected, the second set of samples is subjected to the same operations described above to generate a second set of time-domain filter coefficients 132. In some implementations, the second set of samples is independent of the first set of samples. For example, there is no overlap between the first and second sets of samples. In other implementations, the second set of samples includes one or more samples from the first set of samples. That is, the first and second sets of samples have at least some overlap. The specific amount of overlap is different for Petition 870250084259, dated 09 / 18 / 2025, page 27 / 171 21 / 71 different implementations and can be selected based on available processing capabilities, a specific sound environment, user settings, or other selection criteria.
[0047] In the example illustrated in Figure 1, the audio data 106 provided for the first data path 110 is the same as the audio data 106 provided for the second data path 120. In some implementations, the audio data 106 provided for the second data path 120 includes the audio data 106 provided for the first data path 110, as well as additional audio data 106. For example, the device 100 may include two or more microphones 102, and at least one of the microphones 102 is coupled to the second data path 120 and not to the first data path 110. To illustrate, the device 100 may include a self-speech microphone that is configured primarily to capture the sound corresponding to the speech of a user of the device 100.In this illustrative example, the audio data 106 generated by the self-speech microphone can be provided to the second data path 120 to generate time-domain filter coefficients 132 and not provided to the first data path 110 to generate the suppressed noise output signal 114.
[0048] Additionally or alternatively, in some implementations, the audio data 106 provided for the first data path 110 includes the audio data 106 provided for the second data path 120, as well as additional audio data 106. For example, the device 100 may include two or more microphones 102, and at least one of them Petition 870250084259, dated 09 / 18 / 2025, page 28 / 171 22 / 71 microphones 102 is coupled to the first data path 110 and not to the second data path 120. To illustrate, the device 100 may include a feedback microphone that is configured primarily to capture the output sound 172 and to generate a feedback signal. In this illustrative example, the audio data 106 generated by the feedback microphone can be processed in the time domain (e.g., in a feedback ANC filter 816, shown in Figure 8) to generate or modify the noise-suppressed output signal 114.
[0049] Additionally or alternatively, in some implementations, the audio data 106 provided for the first data path 110 includes at least some of the audio data 106 provided for the second data path 120, the audio data 106 provided for the second data path 120 includes at least some of the audio data 106 provided for the first data path 110, and each of the data paths 110, 120 is also provided with additional audio data 106 that is not provided for another data path 110, 120. For example, the device 100 may include both the feedback microphone and the self-speech microphone, each of which operates as described above.
[0050] In some implementations, device 100 corresponds to or is included in one of several types of devices. In an illustrative example, processor 190 is integrated into a wearable device that includes or is coupled to microphone(s) 102 and speaker(s) 108. Examples of such wearable devices include, without limitation, a headset device, as further described with reference to Figure 10; a Petition 870250084259, dated 09 / 18 / 2025, p. 29 / 171 23 / 71 virtual reality, mixed reality or augmented reality headset, as described with reference to Figure 11; augmented reality glasses, as described with reference to Figure 12; a hearing aid, as described with reference to Figure 13; or in-ear headphones, as described with reference to Figure 14.
[0051] A technical advantage of implementing device 100 as described above is that device 100 can provide high-quality noise suppression without introducing undue latency. For example, noise suppression for the input audio data 106 is performed entirely in the time domain (e.g., by the first data path 110), thus avoiding delays due to domain transform and inverse transform operations. Therefore, the latency introduced by noise suppression operations is very small (e.g., on the order of 2 milliseconds or less) and is comparable to the latency of conventional time-domain noise suppression operations such as ANC. However, the quality of the noise suppression is higher than can be achieved using such conventional time-domain noise suppression operations.
[0052] Figure 2 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure. Device 100 in Figure 2 represents a particular and non-limiting example of device 100 in Figure 1; thus, device 100 in Figure 2 includes many of the same components as illustrated in Figure 1, each of which operates as described above. For example, the Petition 870250084259, dated 09 / 18 / 2025, page 30 / 171 24 / 71 device 100 of Figure 2 includes the first data path 110 and the second data path 120. As described with reference to Figure 1, the first data path 110 includes the time-domain filter(s) 112 which is / are configured to apply time-domain filter coefficients 132 (which are determined in the second data path 120) to the audio data 106 to generate the suppressed noise output signal 114.
[0053] In the example illustrated in Figure 2, the second data path 120 includes the analysis filter bank 122, the machine learning model(s) 126, and the time-domain filter designer 130. In some implementations, the processor(s) 190 in Figure 2 also include(s) the buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. In Figure 2, the machine learning model(s) 126 is / are configured to receive input data 125 that includes or is based on the frequency-domain audio data 124 and to generate a frequency mask 204 as output data 129. The frequency mask 204 represents an estimated magnitude of noise in the frequency-domain audio data 124 for each frequency bin of a plurality of frequency bins.Since frequency mask 204 is an estimate of noise in the frequency domain in audio data 106, frequency mask 204 can be provided as input to the time domain filter designer 130 to generate the time domain filter coefficients 132.
[0054] Figure 3 is a block diagram of illustrative aspects of the device in Figure 1, which is Petition 870250084259, dated 09 / 18 / 2025, p. 31 / 171 25 / 71 operable to perform low-latency noise suppression, according to some examples in this disclosure. Device 100 of Figure 3 represents a particular and non-limiting example of device 100 of Figure 1; thus, device 100 of Figure 3 includes many of the same components as illustrated in Figure 1, each of which operates as described above. For example, device 100 of Figure 3 includes the first data path 110 and the second data path 120. As described with reference to Figure 1, the first data path 110 includes the time-domain filter(s) 112 which is / are configured to apply time-domain filter coefficients 132 (which are determined in the second data path 120) to the audio data 106 to generate the suppressed noise output signal 114.
[0055] In the example illustrated in Figure 3, the second data path 120 includes the analysis filter bank 122, the machine learning model(s) 126, and the time-domain filter designer 130. In some implementations, the processor(s) 190 in Figure 3 also include the buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. In Figure 3, the machine learning model(s) 126 is / are configured to receive input data 125 that includes or is based on the frequency-domain audio data 124 and to generate de-noise audio data 304 as the output data 129 in Figure 1. The de-noise audio data 304 includes, for example, a frequency-domain estimate of the target audio in the audio data. 106. In some modalities, the data of Petition 870250084259, dated 09 / 18 / 2025, page 32 / 171 26 / 71 audio with suppressed noise 304 may include synthesized audio data. For example, the machine learning model(s) 126 may be configured to synthesize target audio (e.g., speech) based on the input data 125. An advantage of the machine learning model(s) 126 that generate(s) synthesized target audio (e.g., speech) is that the synthesized target audio may be devoid of non-target audio; thus, the audio data with suppressed noise 304 in this example may include only target audio.
[0056] In Figure 3, the second data path 120 also includes a mask generator 306. The mask generator 306 is configured to generate the frequency mask 204 based on the noise-suppressed audio data 304. For example, the mask generator 306 can determine, for each frequency bin of a set of frequency bins, a ratio between the noise-suppressed audio data 304 and the frequency-domain audio data 124. Each frequency bin of the frequency-domain audio data 124 can be adjusted by a ratio associated with the frequency bin to generate the frequency mask 204. The frequency mask 204 can be included in or correspond to the noise suppression output 128 that is provided as input to the time-domain filter designer 130 to generate the time-domain filter coefficients 132.
[0057] Figure 4 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure. Device 100 in Figure 4 represents a particular example. Petition 870250084259, dated 09 / 18 / 2025, page 33 / 171 27 / 71 and not limiting device 100 of Figure 1, thus, device 100 of Figure 4 includes many of the same components as illustrated in Figure 1, each of which operates as described above. For example, device 100 of Figure 4 includes the first data path 110 and the second data path 120. As described with reference to Figure 1, the first data path 110 includes the time-domain filter(s) 112 which is / are configured to apply time-domain filter coefficients 132 (which are determined in the second data path 120) to the audio data 106 to generate the suppressed noise output signal 114.
[0058] In the example illustrated in Figure 4, the second data path 120 includes the analysis filter bank 122, a frequency-domain signal booster 402, the machine learning model(s) 126, and the time-domain filter designer 130. In some implementations, the processor(s) 190 in Figure 4 also include the buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. The frequency-domain signal booster 402 is configured to process the frequency-domain audio data 124 to generate augmented audio data 404. The augmented audio data 404 has an improved signal-to-noise ratio (SNR) compared to the frequency-domain audio data 124. The frequency-domain signal booster 402 can use one or more of several Different operations to improve SNR. Examples of such operations are described with reference to Figures 5 to 7.
[0059] In Figure 4, the input data 125 Petition 870250084259, dated 09 / 18 / 2025, p. 34 / 171 28 / 71 include or are based on the augmented audio data 404. The machine learning model(s) 126 in Figure 4 generate(s) output data 129, such as a frequency mask, noise-suppressed audio data, or both, based on the input data 125. The time-domain filter designer 130 generates the time-domain filter coefficients 132 based on a noise-suppressed output 128 that includes or is based on the output data 129. To illustrate, in some examples where the output data 129 of the machine learning model(s) 126 in Figure 4 is a frequency mask, the noise-suppressed output 128 includes or matches the output data 129.In other examples where the output data 129 of the machine learning model(s) 126 in Figure 4 are audio data with suppressed noise, the second path 120 may additionally include the mask generator 306 in Figure 3 to generate a frequency mask that matches or is included in the noise suppression output 128.
[0060] Figure 5 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure. Device 100 in Figure 5 represents a particular and non-limiting example of device 100 in Figure 1; thus, device 100 in Figure 5 includes many of the same components as illustrated in Figure 1, each of which operates as described above. For example, device 100 in Figure 5 includes the first data path 110 and the second data path 120. As described with reference to Figure 1, the first data path 110 includes the time-domain filter(s) 112 Petition 870250084259, dated 09 / 18 / 2025, page 35 / 171 29 / 71 which is / are configured to apply time-domain filter coefficients 132 that are determined in the second data path 120 to the audio data 106 to generate the suppressed noise output signal 114.
[0061] In the example illustrated in Figure 5, the second data path 120 includes the analysis filter bank 122, a beamformer 502, the machine learning model(s) 126, and the time-domain filter designer 130. In some implementations, the processor(s) 190 of Figure 5 also include the buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. The beamformer 502 is an example of the frequency-domain signal booster 402 of Figure 4. The beamformer 502 is configured to process the frequency-domain audio data 124 to emphasize or minimize the sound received from a particular direction to generate beamforming audio data 504.For example, sound components 170 originating from the direction of a target sound source may be emphasized in beamforming audio data 504, sound components 170 originating from a direction other than the direction of a target sound source may be minimized in beamforming audio data 504, or both.
[0062] Audio signals 104 from two or more microphones 102 (for example, a first microphone 102A and a second microphone 102B) are used to determine the directionality associated with the sound 170. Optionally, in some implementations, other information may be used to determine the direction to the target sound source. For example, a camera may be used to generate data of Petition 870250084259, dated 09 / 18 / 2025, page 36 / 171 30 / 71 image or video data that can be analyzed to determine a direction from device 100 in Figure 5 toward a person who is speaking. In other examples, no other sensor is used. To illustrate, beamformer 502 can direct a beam toward the dominant sound source in a particular environment, assuming that the dominant sound source is more likely to be the target sound source than a background sound source.
[0063] In Figure 5, the input data 125 includes or is based on the beamforming audio data 504. Since the beamforming audio data 504 represents an improvement in the SNR of the target sound, the computational complexity of the machine learning model(s) 126 in Figure 5 can be reduced compared to generating the input data 125 based on the non-beamforming frequency domain audio data 124. The machine learning model(s) 126 generate(s) output data 129, such as a frequency mask, noise-suppressed audio data, or both, based on the input data 125. The time-domain filter designer 130 generates the time-domain filter coefficients 132 based on a noise-suppressed output 128 that includes or is based on the output data 129.
[0064] Figure 6 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure. Device 100 in Figure 6 represents a particular and non-limiting example of device 100 in Figure 1; thus, device 100 in Figure 6 includes many of the same features. Petition 870250084259, dated 09 / 18 / 2025, page 37 / 171 31 / 71 components as illustrated in Figure 1, each of which operates as described above. For example, device 100 in Figure 6 includes the first data path 110 and the second data path 120. As described with reference to Figure 1, the first data path 110 includes the time-domain filter(s) 112 which is / are configured to apply time-domain filter coefficients 132 that are determined in the second data path 120 to the audio data 106 to generate the suppressed noise output signal 114.
[0065] In the example illustrated in Figure 6, the second data path 120 includes the analysis filter bank 122, a speech augmentation mechanism 602, the machine learning model(s) 126, and the time-domain filter designer 130. In some implementations, the processor(s) 190 of Figure 6 also include(s) the buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. The speech augmentation mechanism 602 is an example of the frequency-domain signal booster 402 of Figure 4. The speech augmentation mechanism 602 is configured to process the frequency-domain audio data 124 to emphasize components of the frequency-domain audio data 124 that represent speech to generate speech-augmented audio data 604.For example, the 602 speech augmentation mechanism can perform spectral modification, such as filtering, equalization, and spectral enhancement, to emphasize speech, minimize non-speech, or both.
[0066] In Figure 6, the input data 125 includes or is based on the augmented speech audio data 604. Since the augmented speech audio data 604 Petition 870250084259, dated 09 / 18 / 2025, page 38 / 171 32 / 71 represent an improvement in the SNR of the target sound when the target sound is speech; the computational complexity of the machine learning model(s) 12 6 of Figure 6 can be reduced compared to generating the input data 125 based on the frequency domain audio data 124 without speech augmentation. The machine learning model(s) 126 generate(s) output data 129, such as a frequency mask, noise-suppressed audio data, or both, based on the input data 125. The time domain filter designer 130 generates the time domain filter coefficients 132 based on a noise suppression output 128 that includes or is based on the output data 129.
[0067] Figure 7 is a block diagram of illustrative aspects of the device of Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure. Device 100 of Figure 7 represents a particular and non-limiting example of device 100 of Figure 1; thus, device 100 of Figure 7 includes many of the same components as illustrated in Figure 1, each of which operates as described above. For example, device 100 of Figure 7 includes the first data path 110 and the second data path 120. As described with reference to Figure 1, the first data path 110 includes the time-domain filter(s) 112 which is / are configured to apply time-domain filter coefficients 132 that are determined in the second data path 120 to the audio data 106 to generate the suppressed noise output signal 114.
[0068] In the example illustrated in Figure 7, the Petition 870250084259, dated 09 / 18 / 2025, page 39 / 171 33 / 71 The second data path 120 includes the analysis filter bank 122, a source separation mechanism 702, the machine learning model(s) 126, and the time-domain filter designer 130. In some implementations, the processor(s) 190 of Figure 7 also include the buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. The source separation mechanism 702 is an example of the frequency-domain signal booster 402 of Figure 4. The source separation mechanism 702 is configured to process the frequency-domain audio data 124 to identify and emphasize components of the frequency-domain audio data 124 that represent the sound 170 of a target source. For example, the 702 source separation mechanism can use blind source separation operations, such as independent component analysis, to perform source separation.In this example, a portion of the audio data with a separate source that corresponds to a target sound can be used to generate the input data 125.
[0069] In Figure 7, the input data 125 includes or is based on the source-separated audio data 704. Since the source-separated audio data 704 represents an improvement in the SNR of the target sound, the computational complexity of the machine learning model(s) 126 in Figure 7 can be reduced compared to generating the input data 125 based on the frequency-domain audio data 124 without source separation. The machine learning model(s) 126 generate(s) output data 129, such as a frequency mask, noise-suppressed audio data, or both, based on the input data 125. The Petition 870250084259, dated 09 / 18 / 2025, p. 40 / 171 34 / 71 time-domain filter designer 130 generates time-domain filter coefficients 132 based on a noise suppression output 128 that includes or is based on output data 129.
[0070] Although Figures 5 to 7 illustrate implementations of device 100 in which particular examples of signal boosting are performed, in some implementations, the frequency-domain signal booster 402 may perform more than one of these signal boosting operations or other frequency-domain signal boosting operations. For example, the frequency domain signal booster 402 may include the beamformer 502 and the speech augmentation mechanism 602. As another example, the frequency domain signal booster 402 may include the beamformer 502 and the source separation mechanism 702. As another example, the frequency domain signal booster 402 may include the speech augmentation mechanism 602 and the source separation mechanism 702. In yet another example, the frequency domain signal booster 402 may include the beamformer 502, the speech augmentation mechanism 602, and the source separation mechanism 702.
[0071] Figure 8 is a block diagram of illustrative aspects of the device in Figure 1, which is operable to perform low-latency noise suppression, according to some examples in this disclosure. Device 100 in Figure 8 represents a particular and non-limiting example of device 100 in Figure 1; thus, device 100 in Figure 8 includes many of the same components as illustrated in Figure 1, each of which operates as described above. For example, the Petition 870250084259, dated 09 / 18 / 2025, page 41 / 171 Device 100 of Figure 8 includes the first data path 110 and the second data path 120. As described with reference to Figure 1, the first data path 110 includes the time-domain filter(s) 112 which is / are configured to apply time-domain filter coefficients 132 that are determined in the second data path 120 to the audio data 106 to generate the suppressed noise output signal 114.
[0072] In the example illustrated in Figure 8, the second data path 120 includes the analysis filter bank 122, the machine learning model(s) 126, and the time-domain filter designer 130. In some implementations, the processor(s) 190 in Figure 8 also include the buffer 140 between the time-domain filter designer 130 and the time-domain filter(s) 112. In several implementations, the second data path 120 also optionally includes one or more input preprocessing mechanisms 820, one or more output postprocessing mechanisms 822, or both.
[0073] As described above, the analysis filter bank 122 is configured to receive audio data 106 representing audio signals 104 from one or more microphones 102 (e.g., microphones 102A and 102B in Figure 8) and to generate frequency domain audio data 124. In some implementations, the frequency domain audio data 124 is used as the input data 125 for the machine learning model(s) 126. In such implementations, the input preprocessing mechanism(s) 820 may be omitted. In other implementations, the frequency domain audio data Petition 870250084259, dated 09 / 18 / 2025, page 42 / 171 36 / 71 frequency 124 are modified to generate the input data 125 for the machine learning model(s) 126. In such implementations, the input preprocessing mechanism(s) 820 is / are configured to generate the input data 125 based on the frequency domain audio data 124. For example, the input preprocessing mechanism(s) 820 may include one or more frequency domain signal boosters (e.g., the frequency domain signal boosters 402 of Figure 4), such as the beamformer 502 of Figure 5, the speech boosting mechanism 602 of Figure 6, the source separation mechanism 702 of Figure 7, or combinations thereof. The input preprocessing mechanism(s) 820 may also, or alternatively, perform other operations to generate the input data 125 based on the frequency domain audio data 124.For example, the input preprocessing mechanism(s) 820 may perform data aggregation operations, filtering operations, resampling operations, other data manipulations, or combinations thereof to generate the input data 125.
[0074] The machine learning model(s) 126 generate(s) the output data 129 based on the input data 125. Depending on the configuration of the machine learning model(s) 126, the output data 129 may include audio data with suppressed noise, a frequency mask, or both. Optionally, the output data 129 may be modified by the output post-processing mechanism(s) 822 to generate the noise suppression output 128 which is provided to the time-domain filter designer 130 to generate the filter coefficients of Petition 870250084259, dated 09 / 18 / 2025, page 43 / 171 37 / 71 time domain 132. For example, when the output data 129 includes audio data with suppressed noise, the output post-processing mechanism(s) 822 may perform operations to determine a frequency mask based on the suppressed audio data, as described for the mask generator 306 in Figure 3. The output post-processing mechanism(s) 822 may also, or alternatively, perform other operations to generate the noise-suppressed output 128 based on the output data 129. For example, the output post-processing mechanism(s) 822 may perform data aggregation operations, filtering operations, resampling operations, other data manipulations, or combinations thereof to generate the noise-suppressed output 128.
[0075] Figure 8 illustrates several optional components, some or all of which are included in some implementations and omitted in other implementations. For example, in Figure 8, processor(s) 190 include a feedforward ANC filter 812 configured to perform adaptive feedforward noise cancellation based on audio data 106 from one or more microphones 102. As another example, in Figure 8, device 100 includes or is coupled to at least one external microphone (e.g., microphones 102A and 102B) configured to generate audio data 106 representing the sound 170 in the environment surrounding device 100. Additionally, in Figure 8, device 100 includes or is coupled to at least one feedback microphone (e.g., microphone 102C) configured to generate a feedback signal 814 based on the sound present near the speaker(s) 108, which may include the sound of Petition 870250084259, dated 09 / 18 / 2025, page 44 / 171 38 / 71 output 172 produced by the loudspeaker(s) 108 and sound components 170 transported directly to a user's ear canal 804 (subject to some transfer function P(z) 802). In this example, the feedback signal 814 can be provided to a feedback ANC filter 816 to generate a feedback noise signal that is subtracted from the suppressed noise output signal 114.
[0076] Figure 9 depicts an implementation 900 of the device 100 as an integrated circuit 902 that includes one or more processors 190. The integrated circuit 902 also includes an audio input 904, as one or more bus interfaces, to enable audio data 106 to be received for processing. The integrated circuit 902 also includes a signal output 906, as a bus interface, to enable the sending of the output signal with suppressed noise 114. In Figure 9, the processor(s) 190 of the integrated circuit 902 include(s) one or more audio components 940, such as the analysis filter bank 122, the machine learning model(s) 126, the time-domain filter designer 130, and the time-domain filter(s) 112.Optionally, the audio component(s) 940 may include other components as described above with reference to Figures 1 to 8, such as the mask generator 306 of Figure 3; the frequency domain signal booster 402 of any of Figures 4 to 7; or any one or more of the input pre-processing mechanism(s) 820, the output post-processing mechanism(s) 822, the feedforward ANC filter 812 and the feedback ANC filter 816 of Figure 8. The integrated circuit 902 enables the implementation of high-quality, low-noise suppression. Petition 870250084259, dated 09 / 18 / 2025, page 45 / 171 39 / 71 latency as a component in a system, such as a wearable device that includes microphones, such as the headset as shown in Figure 10, a virtual reality, mixed reality or augmented reality headset as shown in Figure 11, augmented reality headset glasses as shown in Figure 12, a wearable device as shown in Figure 13, in-ear headphones as shown in Figure 14 or other wearable device.
[0077] Figure 10 depicts an implementation 1000 in which the device 100 includes a headset device 1002. The headset device 1002 includes the microphone(s) 102 and the speaker(s) 108. In the example illustrated in Figure 10, microphone 102A is positioned primarily to detect the speech of a person using the headset device 1002, and microphone 102B is positioned to detect ambient sound, such as the speech of another person or other sounds. The components of the processor(s) 190, including the audio component 940, are integrated into the headset device 1002 and are depicted using dashed lines to indicate components that are generally not visible to a user of the headset device 1002.
[0078] In a particular example of operation, microphone 102B can detect sound in an environment around headset device 1002 and generate audio data representing the sound. The audio data can be provided to audio components 940, which can process the audio data in the time domain to generate a noise-suppressed output signal and can process the audio data in the frequency domain to generate (or update) time-domain filter coefficients that are applied to the data. Petition 870250084259, dated 09 / 18 / 2025, page 46 / 171 40 / 71 audio received subsequently. In this example, the time-domain filter coefficients determined through frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to the received audio data in the time domain, little or no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the 1002 headset device can provide high-quality, low-latency noise suppression.
[0079] Figure 11 depicts an implementation 1100 in which the device 100 includes a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset device 1102. The headset device 1102 includes the microphone(s) 102 and the speaker(s) 108. Additionally, the processor(s) components 190, including the audio component 940, are integrated into the headset device 1102. In a particular example of operation, the microphone(s) 102 can detect sound in an environment around the headset device 1102 and generate audio data representing the sound.Audio data can be provided to the 940 audio components, which can process the audio data in the time domain to generate a noise-suppressed output signal and can process the audio data in the frequency domain to generate (or update) time-domain filter coefficients that are subsequently applied to the received audio data. In this example, the time-domain filter coefficients determined through frequency-domain processing provide... Petition 870250084259, dated 09 / 18 / 2025, page 47 / 171 41 / 71 high-quality noise suppression, and since time-domain filter coefficients are applied to receive time-domain audio data, little or no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the 1102 headset device can provide high-quality, low-latency noise suppression.
[0080] Figure 12 depicts an implementation 1200 in which the device 100 includes a portable electronic device corresponding to augmented reality or mixed reality glasses 1202. The glasses 1202 include a holographic projection unit 1204 configured to project visual data onto a lens surface 1206 or to reflect visual data off a lens surface 1206 and onto the user's retina. The glasses 1202 also include microphone(s) 102, speaker(s) 108 and processor(s) 190, which include audio component 940.
[0081] In a particular example of operation, microphone(s) 102 can detect sound in an environment around glasses 1202 and generate audio data representing the sound. The audio data can be provided to audio components 940, which can process the audio data in the time domain to generate a noise-suppressed output signal and can process the audio data in the frequency domain to generate (or update) time-domain filter coefficients that are subsequently applied to the received audio data. In this example, the time-domain filter coefficients determined through frequency-domain processing provide high-quality noise suppression, and since the Petition 870250084259, dated 09 / 18 / 2025, page 48 / 171 42 / 71 time-domain filter coefficients are applied to receive audio data in the time domain; little or no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the 1202 glasses can provide high-quality, low-latency noise suppression.
[0082] In some implementations, the holographic projection unit 1204 is configured to display information related to the sound detected by the microphone(s) 102. For example, the holographic projection unit 1204 may display a notification indicating that speech has been detected. In another example, the holographic projection unit 1204 may display a notification indicating a detected audio event. For example, the notification may be superimposed on the user's field of view at a particular position that coincides with the location of the sound source associated with the audio event.
[0083] Figure 13 is a diagram of an operable wearable device for performing low-latency noise suppression, according to some examples in this disclosure. In the example illustrated in Figure 13, the wearable device is a hearing aid device 1302. The hearing aid device 1302 includes microphone(s) 102, speaker(s) 108, and processor(s) 190, which include audio component 940. In the example illustrated in Figure 13, the hearing aid device 1302 includes a portion 1304 configured to be worn behind a user's ear, a portion 1308 configured to extend over the ear, and a portion 1306 to be worn in or near a user's ear canal. In other examples, the hearing aid device Petition 870250084259, dated 09 / 18 / 2025, page 49 / 171 43 / 71 hearing aid 1302 has a different configuration or form factor. For example, hearing aid device 1302 may be an in-the-ear device that does not include portion 1304 configured to be worn behind an ear and portion 1308 configured to extend over the ear.
[0084] In a particular example of operation of the hearing aid device 1302, the microphone(s) 102 can detect sound in an environment around the hearing aid device 1302 and generate audio data representing the sound. The audio data can be provided to the audio components 940, which can process the audio data in the time domain to generate a noise-suppressed output signal and can process the audio data in the frequency domain to generate (or update) time-domain filter coefficients that are subsequently applied to the received audio data.In this example, the time-domain filter coefficients determined through frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to receive time-domain audio data, little or no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the 1302 hearing aid device can provide high-quality, low-latency noise suppression.
[0085] Figure 14 depicts an implementation 1400 in which the device 100 includes a portable electronic device corresponding to one or more in-ear headphones 1406 (for example, a first in-ear headphone 1402, a second in-ear headphone 1404 or Petition 870250084259, dated 09 / 18 / 2025, page 50 / 171 44 / 71 both). Although the 1406 in-ear headphones are described, it should be understood that the present technology can be applied to other in-ear or over-the-ear audio devices.
[0086] In the example illustrated in Figure 14, the first in-ear headphone 1402 includes a first microphone 1410A, as a high signal-to-noise ratio microphone positioned to capture the voice of a user of the first in-ear headphone 1402, one or more other microphones configured to detect ambient sounds and spatially distributed to support beamforming, illustrated as microphone(s) 1412A, an internal microphone 1414A adjacent to the user's ear canal (e.g., to assist in active noise cancellation), and a self-speech microphone 1416A, as a bone conduction microphone configured to convert sound vibrations from the ear bone or skull of a user into an audio signal. In a particular implementation, microphone(s) 1412a correspond(s) to microphone(s) 102 in any of Figures 1 to 4, 6 or 7, or correspond(s) to microphone(s) 102A and / or 102B in Figures 5 or 8.In one particular implementation, microphone 1414A corresponds to microphone 102C in Figure 8.
[0087] The second in-ear headset 1404 can be configured in a substantially similar manner to the first in-ear headset 1402. For example, the second in-ear headset may include a microphone 1410B positioned to capture the voice of a user of the second in-ear headset 1404, one or more other microphones 1412B configured to detect ambient sounds and spatially distributed to Petition 870250084259, dated 09 / 18 / 2025, page 51 / 171 45 / 71 supports beamforming, an internal 1414B microphone and a self-talking 1416B microphone.
[0088] In some implementations, the in-ear headphones 1402, 1404 are configured to automatically switch between various modes of operation, such as a pass-through mode in which ambient sound is processed by the audio components 940 for output through a speaker(s) 108 and a playback mode in which non-ambient sound (e.g., streaming audio corresponding to a telephone conversation, media playback, video game, etc.) is played through the speaker(s) 108. In other implementations, the in-ear headphones 1402, 1404 may support a smaller number of modes or may support one or more other modes instead of or in addition to the modes described.
[0089] In an illustrative example of pass-through operation, one or more of the microphones 102 (for example, microphones 1412A, 1412B) can detect sound in an environment around the in-ear headphones 1402, 1404 and generate audio data representing the sound. The audio data can be provided to one or both of the audio components 940A, 940B, which can process the audio data in the time domain to generate a noise-suppressed output signal and can process the audio data in the frequency domain to generate (or update) time-domain filter coefficients that are applied to subsequently received audio data. In this example, the time-domain filter coefficients determined through frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to Petition 870250084259, dated 09 / 18 / 2025, page 52 / 171 46 / 71 audio data received in the time domain, little or no latency is added with the use of such time-domain filter coefficients to process the audio data. Thus, the 1402, 1404 in-ear headphones can provide high-quality, low-latency noise suppression.
[0090] Figure 15 is a diagram of a particular implementation of a method for achieving low-latency noise suppression that can be performed by the device of Figure 1, according to some examples in this disclosure. In a particular aspect, one or more operations of method 1500 are performed by at least one of the device 100, the processor(s) 190 or the audio components 940, described variously with reference to Figures 1 to 14, or a combination thereof.
[0091] Method 1500 includes, in block 1502, obtaining audio data representing one or more audio signals. The audio data includes a first segment and a second segment subsequent to the first segment. For example, with reference to Figure 1, audio data 106 represents the audio signal(s) 104 received from microphone(s) 102. In this example, microphone(s) 102 generate(s) audio signal(s) 104 based on the sound 170 present in an environment around device 100. Each of the segments may include a time segment of the audio signal, such as a sample representing a few milliseconds of the audio signal.
[0092] Method 1500 includes, in block 1504, performing one or more transform operations on the first segment to generate frequency-domain audio data. For example, with reference to Figure 1, analysis filter bank 122 performs one or more transform operations. Petition 870250084259, dated 09 / 18 / 2025, page 53 / 171 47 / 71 (for example, time-domain to frequency-domain transform operations, such as FFT) to generate the frequency-domain audio data 124.
[0093] Method 1500 includes, in block 1506, providing input data based on frequency-domain audio data as input to one or more machine learning models to generate a noise suppression output. For example, in some implementations, the input data includes or matches the frequency-domain audio data. In other implementations, the frequency-domain audio data is modified or manipulated to generate the input data. For example, in some of these implementations, method 1500 includes performing beamforming operations to determine beamforming audio data that distinguishes a portion of the audio data from a target audio source and a portion of the audio data from a non-target audio source. In this example, the input data is based on (e.g., includes or matches) the beamforming audio data.As another example, in some of these implementations, method 1500 includes performing speech augmentation operations to determine audio data with augmented speech. In this example, the input data is based on (e.g., includes or matches) the audio data with augmented speech. As another example, in some of these implementations, method 1500 includes performing source separation operations to determine audio data with separated sources. In this example, the input data is based on (e.g., includes or matches) the audio data with separated sources. In addition to, or instead of, performing signal augmentation operations. Petition 870250084259, dated 09 / 18 / 2025, page 54 / 171 48 / 71 (such as beamforming, speech augmentation, or source separation), the 1500 method may include performing other data manipulations to generate the input data based on the frequency domain audio data, such as data aggregation, filtering, etc.
[0094] In some implementations, one or more machine learning models directly generate the noise suppression output (for example, the noise suppression output consists of the output data from one or more machine learning models). For example, masking machine learning models output a frequency mask that can be used as or included in the noise suppression output.
[0095] In some implementations, one or more machine learning models generate the output data that is modified to generate the noise-suppressed output. For example, inline machine learning models output noise-suppressed audio data. In this example, a mask generator (e.g., mask generator 306 in Figure 3) can generate a frequency mask based on the noise-suppressed audio data and the frequency-domain audio data.
[0096] Method 1500 includes, in block 1508, performing one or more reverse transform operations on the noise suppression output to generate time-domain filter coefficients. For example, time-domain filter designer 130 can perform inverse transform operations and parameterization operations to generate time-domain filter coefficients 132 based on noise suppression output 128. The coefficients of Petition 870250084259, dated 09 / 18 / 2025, page 55 / 171 49 / 71 Time-domain filters may include, for example, linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.
[0097] Method 1500 includes, in block 1510, performing time-domain filtering of the second segment using time-domain filter coefficients to generate a noise-suppressed output signal. For example, time-domain filter(s) 112 apply time-domain filter coefficients 132 to a segment (e.g., the second segment) of audio data 106 that is subsequent to the segment (e.g., the first segment) used to generate the time-domain filter coefficients. A technical advantage of method 1500 is that the time-domain filter coefficients determined through frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to the received audio data in the time domain, little or no latency is added by using such time-domain filter coefficients to process the audio data.
[0098] Optionally, in some implementations, method 1500 also includes performing adaptive noise cancellation operations based on the suppressed noise output signal. For example, with reference to Figure 8, the suppressed noise output signal 114 can be modified using a feedforward ANC filter 812, a feedback ANC filter 816, or a Petition 870250084259, dated 09 / 18 / 2025, p. 56 / 171 50 / 71 hybrid ANC filter.
[0099] Method 1500 of Figure 15 can be implemented by a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Method 1500 of Figure 15 can be implemented by one or more processors executing instructions, as described with reference to Figure 16.
[00100] With reference to Figure 16, a block diagram of a particular illustrative implementation of a device, generally designated 1600, is depicted. In various implementations, the 1600 device may have a greater or lesser number of components than illustrated in Figure 16. In one illustrative implementation, the 1600 device may correspond to the 100 device. In another illustrative implementation, the 1600 device may perform one or more of the operations described with reference to Figures 1 to 15.
[00101] In a particular implementation, the 1600 device includes a 1606 processor (e.g., a central processing unit (CPU)). The 1600 device may include one or more additional 1610 processors (e.g., one or more DSPs). In a particular aspect, the 190 processor(s) in Figure 1 correspond(s) to the 1606 processor, the 1610 processor(s), or a combination thereof. The 1610 processor(s) may include an encoder-decoder (CODEC). Petition 870250084259, dated 09 / 18 / 2025, p. 57 / 171 51 / 71 decoder) 1608 for speech and music which includes a voice encoder (vocoder) 1636 and a vocoder decoder 1638. The processor(s) 1610 and / or the speech and music CODEC 1608 also include low-latency noise suppression components such as the analysis filter bank 122, the machine learning model(s) 126, the time-domain filter designer 130 and the time-domain filter(s) 112.
[00102] In Figure 16, device 1600 includes a memory 1686 and a CODEC 1634. The memory 1686 includes (for example, stores) instructions 1656, which are executable by one or more additional processors 1610 (or by processor 1606) to implement the functionality described with reference to device 100 in Figure 1. In Figure 16, device 1600 also includes a modem 1670 coupled, via a transceiver 1650, to an antenna 1652. The modem 1670, the transceiver 1650, and the antenna 1652 enable device 1600 to exchange data with one or more other devices via wireless communications. For example, in some implementations, device 1600 may generate audio output on speaker(s) 108 based on data received via wireless communication with another device.
[00103] Device 1600 may include a display 1628 coupled to a display controller 1626. Speaker(s) 108 and microphone(s) 102 may be coupled to CODEC 1634. In Figure 16, CODEC 1634 includes a digital-to-analog converter (DAC) 1602 and an analog-to-digital converter (ADC) 1604. In a particular implementation, CODEC 1634 may receive analog signals (by Petition 870250084259, dated 09 / 18 / 2025, page 58 / 171 52 / 71 For example, the audio signal(s) 104 (from Figures 1 to 8) from the microphone(s) 102, convert the analog signals into digital signals (e.g., audio data 106 from Figures 1 to 8) using the analog-to-digital converter 1604, and provide the digital signals to the speech and music codec 1608. The speech and music codec 1608 can process the digital signals. Digital signals can be further processed by analysis filter bank 122, machine learning model(s) 126, time-domain filter designer 130, and time-domain filter(s) 112. For example, audio data can be provided to time-domain filter(s) 112 which apply a current set of time-domain filter coefficients to the audio data to generate a noise-suppressed output signal. Additionally, in this example, audio data can be provided to analysis filter bank 122 to generate frequency-domain audio data.Input data based on frequency domain audio data can be provided to machine learning model(s) 12 6 to generate output data, and a noise suppression output based on the output data can be provided to time domain filter designer 130 to generate updated time domain filter coefficients. In this example, the updated time domain filter coefficients can be applied to subsequently received audio data.
[00104] In a particular implementation, the 1608 speech and music codec can provide digital signals representing the suppressed noise output signal and / or other audio content to the 1634 codec. The 1634 codec can Petition 870250084259, dated 09 / 18 / 2025, page 59 / 171 53 / 71 convert the digital signals into analog signals using the digital / analog converter 1602 and can provide the analog signals to the loudspeaker(s) 108.
[00105] In a particular implementation, device 1600 may be included in a packaged system or system-on-chip device 1622. In a particular implementation, memory 1686, processor 1606, processors 1610, display controller 1626, CODEC 1634, and modem 1670 are included in the packaged system or system-on-chip device 1622. In a particular implementation, an input device 1630 and a power supply 1644 are coupled to the packaged system or system-on-chip device 1622. Furthermore, in a particular implementation, as illustrated in Figure 16, the display 1628, the input device 1630, the speaker(s) 108, the microphone(s) 102, the antenna 1652, and the power supply 1644 are external to the packaged system device. or system-on-a-chip 1622.In a particular implementation, each of the display 1628, the input device 1630, the speaker(s) 108, the microphone(s) 102, the antenna 1652 and the power supply 1644 can be coupled to a component of the packaged system or system-on-chip device 1622, such as an interface or a controller.
[00106] The 1600 device may include a wearable device, such as a wearable mobile communication device, a wearable personal digital assistant, a wearable display device, a wearable gaming system, a wearable music player, a wearable radio, a wearable camera, a wearable navigation device, a headset, an augmented reality headset, a reality headset Petition 870250084259, dated 09 / 18 / 2025, pp. 60 / 171 54 / 71 mixed, a virtual reality headset, a voice-activated device, a portable electronic device, a wearable computing device, a wearable communication device, a virtual reality (VR) device, one or more in-ear headphones, a hearing aid, or any combination thereof.
[00107] In conjunction with the implementations described, an apparatus includes means for performing one or more transform operations on a first segment of audio data to generate frequency-domain audio data, wherein the audio data includes a first segment and a second segment subsequent to the first segment. For example, the means for performing the one or more transform operations may correspond to device 100, processor(s) 190, analysis filter bank 122, processor 1606, processor(s) 1610, one or more other circuits or components configured to perform one or more transform operations to generate frequency-domain audio data, or any combination thereof.
[00108] The device also includes means for processing input data based on frequency-domain audio data as input to one or more machine learning models to generate a noise-suppression output. For example, the means for processing the input data may correspond to device 100, processor(s) 190, machine learning model(s) 126 (optionally, in conjunction with input pre-processing mechanism 820, output post-processing mechanism 822, or both), and mask generator. Petition 870250084259, dated 09 / 18 / 2025, pp. 61 / 171 55 / 71 306, to processor 1606, to processor(s) 1610, to one or more other circuits or components configured to process input data to generate a noise-suppressed output, or any combination thereof.
[00109] The apparatus also includes means for performing one or more reverse transform operations on the noise suppression output to generate time-domain filter coefficients. For example, the means for performing the one or more reverse transform operations may correspond to device 100, processor(s) 190, time-domain filter designer 130, processor 1606, processor(s) 1610, one or more other circuits or components configured to perform one or more reverse transform operations, or any combination thereof.
[00110] The apparatus also includes means for performing second-segment time-domain filtering using time-domain filter coefficients to generate a noise-suppressed output signal. For example, the means for performing time-domain filtering may correspond to device 100, processor(s) 190, time-domain filter(s) 112, processor 1606, processor(s) 1610, one or more other circuits or components configured to perform time-domain filtering, or any combination thereof.
[00111] In some implementations, a non-transient computer-readable medium (for example, a computer-readable storage device such as 1686 memory) includes instructions (for example, 1656 instructions) that, when executed by one or more Petition 870250084259, dated 09 / 18 / 2025, page 62 / 171 56 / 71 processors (for example, one or more 1610 processors or the 1606 processor) cause the one or more processors to obtain audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment subsequent to the first segment; perform one or more transform operations on the first segment to generate frequency domain audio data; provide input data based on the frequency domain audio data as input to one or more machine learning models to generate a noise-suppressed output; perform one or more reverse transform operations on the noise-suppressed output to generate time domain filter coefficients; and perform time domain filtering of the second segment using the time domain filter coefficients to generate a noise-suppressed output signal.
[00112] The particular aspects of disclosure are described below in sets of interrelated examples:
[00113] According to Example 1, a device includes one or more processors configured to obtain audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment subsequent to the first segment; perform one or more transform operations on the first segment to generate frequency-domain audio data; provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressing output; perform one or more reverse transform operations on the noise-suppressing output. Petition 870250084259, dated 09 / 18 / 2025, page 63 / 171 57 / 71 of noise to generate time-domain filter coefficients; and perform time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.
[00114] Example 2 includes the device from Example 1, wherein the suppressed noise output signal is generated with a latency of 1 millisecond or less and the time-domain filter coefficients are generated with a latency of more than 1 millisecond.
[00115] Example 3 includes the device from Example 1 or Example 2, where the input data includes frequency domain audio data.
[00116] Example 4 includes the device from any of Examples 1 through 3, wherein one or more machine learning models are configured to generate output including a frequency mask representing an estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask.
[00117] Example 5 includes the device from any of Examples 1 through 3, wherein one or more machine learning models are configured to generate output including suppressed audio data, wherein one or more processors are configured to determine, based on the suppressed audio data, a frequency mask representing an estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask. Petition 870250084259, dated 09 / 18 / 2025, pp. 64 / 171 58 / 71
[00118] Example 6 includes the device from any of Examples 1 through 5, where one or more machine learning models include one or more recurrent neural networks.
[00119] Example 7 includes the device from any of Examples 1 to 6, wherein, to generate the noise suppression output, one or more processors are configured to perform beamforming operations on the frequency-domain audio data to determine beamformed audio data by distinguishing a portion of the audio data from a target audio source and a portion of the audio data from a non-target audio source, wherein the input data include or are based on the beamformed audio data.
[00120] Example 8 includes the device from any of Examples 1 to 7, where, to process the frequency domain audio data to generate the noise suppression output, one or more processors are configured to perform speech augmentation operations to determine augmented speech audio data, where the input data includes or is based on the augmented speech audio data.
[00121] Example 9 includes the device from any of Examples 1 through 8, wherein, to process the frequency domain audio data to generate the noise suppression output, one or more processors are configured to perform source separation operations to determine source-separated audio data, wherein the input data includes or is based on the source-separated audio data. Petition 870250084259, dated 09 / 18 / 2025, page 65 / 171 59 / 71
[00122] Example 10 includes the device from any of Examples 1 through 9, wherein the time-domain filter coefficients include linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.
[00123] Example 11 includes the device from any of Examples 1 to 10, in which one or more processors are integrated into a wearable device.
[00124] Example 12 includes the device of any of Examples 1 to 11, in which one or more processors are integrated into one or more in-ear headphones.
[00125] Example 13 includes the device of any of Examples 1 to 12 and additionally includes one or more microphones, where one or more audio signals are received from one or more microphones.
[00126] Example 14 includes the device from Example 13 and additionally includes a set of adaptive noise-canceling circuits coupled to at least one or more microphones.
[00127] Example 15 includes the device of any of Examples 1 to 14 and additionally includes one or more loudspeakers and one or more microphones coupled to one or more processors and integrated into a wearable device, wherein the one or more microphones include at least one external microphone configured to generate the audio data and at least one feedback microphone configured to generate a signal of Petition 870250084259, dated 09 / 18 / 2025, pp. 66 / 171 60 / 71 feedback based on the sound produced by one or more speakers responsive to the suppressed noise output signal.
[00128] According to Example 16, a method includes obtaining audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment subsequent to the first segment; performing one or more transform operations on the first segment to generate frequency-domain audio data; providing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; performing one or more reverse transform operations on the noise-suppressed output to generate time-domain filter coefficients; and performing time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.
[00129] Example 17 includes the method from Example 16, where the suppressed noise output signal is generated with a latency of 1 millisecond or less and the time-domain filter coefficients are generated with a latency of more than 1 millisecond.
[00130] Example 18 includes the method from Example 16 or Example 17, where the input data includes frequency domain audio data.
[00131] Example 19 includes the method from any of Examples 16 through 18, where one or more machine learning models are configured to generate output including a frequency mask representing an estimated magnitude of noise in the domain audio data. Petition 870250084259, dated 09 / 18 / 2025, pp. 67 / 171 61 / 71 of the frequency for each frequency bin of a plurality of frequency bins, and where the noise suppression output includes the frequency mask.
[00132] Example 20 includes the method of any of Examples 16 to 18, wherein one or more machine learning models are configured to generate output including audio data with suppressed noise, and which further comprises determining, based on the suppressed audio data, a frequency mask representing an estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask.
[00133] Example 21 includes the method from any of Examples 16 to 20, where one or more machine learning models include one or more recurrent neural networks.
[00134] Example 22 includes the method of any of Examples 16 to 21 and additionally includes performing beamforming operations on frequency-domain audio data to determine beamformed audio data by distinguishing a portion of audio data from a target audio source and a portion of audio data from a non-target audio source, where the input data are based on or include beamformed audio data.
[00135] Example 23 includes the method of any of Examples 16 to 22 and additionally includes performing speech augmentation operations to determine speech augmented audio data, where the input data are based on or include the speech augmented audio data. Petition 870250084259, dated 09 / 18 / 2025, pp. 68 / 171 62 / 71
[00136] Example 24 includes the method of any of Examples 16 through 23 and additionally includes performing source separation operations to determine source-separated audio data, where the input data is based on or includes source-separated audio data.
[00137] Example 25 includes the method of any of Examples 16 to 24, wherein the time-domain filter coefficients include linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.
[00138] Example 26 includes the method of any of Examples 16 to 25, in which one or more audio signals are received from one or more microphones.
[00139] Example 27 includes the method of any of Examples 16 to 26 and additionally includes performing adaptive noise cancellation operations based on the suppressed noise output signal.
[00140] According to Example 28, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method of any of Examples 16 to 27.
[00141] According to Example 29, a non-transient, computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform the method of any of Examples 16 to 27.
[00142] According to Example 30, a device Petition 870250084259, dated 09 / 18 / 2025, pp. 69 / 171 63 / 71 includes means to carry out the method of any of Examples 16 to 27.
[00143] According to Example 31, a non-transient, computer-readable medium stores instructions that are executable by one or more processors to cause the one or more processors to obtain audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment subsequent to the first segment; perform one or more transform operations on the first segment to generate frequency-domain audio data; provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; perform one or more reverse transform operations on the noise-suppressed output to generate time-domain filter coefficients; and perform time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.
[00144] Example 32 includes the non-transient, computer-readable medium of Example 31, wherein the instructions cause one or more processors to generate the suppressed noise output signal with a latency of 1 millisecond or less and cause one or more processors to generate the time-domain filter coefficients with a latency of more than 1 millisecond.
[00145] Example 33 includes the non-transient, computer-readable medium of Example 31, where the input data includes frequency-domain audio data. Petition 870250084259, dated 09 / 18 / 2025, pp. 70 / 171 64 / 71
[00146] Example 34 includes the non-transient, computer-readable medium of any of Examples 31 to 33, wherein one or more machine learning models are configured to generate output including a frequency mask representing an estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask.
[00147] Example 35 includes the non-transient, computer-readable medium of any of Examples 31 to 33, wherein one or more machine learning models are configured to generate output including suppressed-noise audio data, wherein the instructions cause one or more processors to determine, based on the suppressed-noise audio data, a frequency mask representing an estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise-suppressed output includes the frequency mask.
[00148] Example 36 includes the non-transient, computer-readable medium of any of Examples 31 to 35, in which one or more machine learning models include one or more recurrent neural networks.
[00149] Example 37 includes the computer-readable nontransient medium of any of Examples 31 through 36, where, to generate the noise suppression output, the instructions cause one or more processors to perform beamforming operations on the frequency-domain audio data to determine the Petition 870250084259, dated 09 / 18 / 2025, pp. 71 / 171 65 / 71 beamforming audio data distinguishing a portion of the audio data from a target audio source and a portion of the audio data from a non-target audio source, where the input data are based on or include the beamforming audio data.
[00150] Example 38 includes the non-transient, computer-readable means of any of Examples 31 to 37, wherein, to process the frequency-domain audio data to generate the noise-suppressed output, the instructions cause one or more processors to perform speech augmentation operations to determine augmented speech audio data, wherein the input data are based on or include the augmented speech audio data.
[00151] Example 39 includes the non-transient, computer-readable medium of any of Examples 31 through 38, in which, to process the frequency-domain audio data to generate the noise-suppressed output, the instructions cause one or more processors to perform source separation operations to determine source-separated audio data, wherein the input data is based on or includes source-separated audio data.
[00152] Example 40 includes the computer-readable nontransient medium of any of Examples 31 to 39, wherein the time-domain filter coefficients include linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, response filter coefficients of Petition 870250084259, dated 09 / 18 / 2025, pp. 72 / 171 66 / 71 infinite impulse (IIR) or all-pole filter coefficients.
[00153] Example 41 includes the non-transient, computer-readable medium of Example 40, in which the instructions cause one or more processors to perform adaptive noise cancellation operations based on the suppressed noise output signal.
[00154] According to Example 42, an apparatus includes means for performing one or more transform operations on a first segment of audio data to generate frequency-domain audio data, wherein the audio data includes a first segment and a second segment subsequent to the first segment; means for processing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; means for performing one or more reverse transform operations on the noise-suppressed output to generate time-domain filter coefficients; and means for performing time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.
[00155] Example 43 includes the apparatus of Example 42, wherein the means for performing time-domain filtering are operable to generate the output signal with suppressed noise with a latency of 1 millisecond or less, and wherein the means for performing one or more transform operations, the means for processing the input data, and the means for determining the time-domain filter coefficients are operable to generate the filter coefficients. Petition 870250084259, dated 09 / 18 / 2025, pp. 73 / 171 67 / 71 time domain with a latency of more than 1 millisecond.
[00156] Example 44 includes the apparatus of Example 42, where the input data includes frequency domain audio data.
[00157] Example 45 includes the apparatus of any of Examples 42 to 44, wherein one or more machine learning models are configured to generate output including a frequency mask representing an estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask.
[00158] Example 46 includes the apparatus of any of Examples 42 to 44, wherein one or more machine learning models are configured to generate output including suppressed audio data, further comprising means for determining, based on the suppressed audio data, a frequency mask representing an estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask.
[00159] Example 47 includes the apparatus of any of Examples 42 to 46, in which one or more machine learning models include one or more recurrent neural networks.
[00160] Example 48 includes the apparatus of any of Examples 42 to 47 and additionally includes means for performing beamforming operations on the data of Petition 870250084259, dated 09 / 18 / 2025, pp. 74 / 171 68 / 71 frequency domain audio to determine beamforming audio data by distinguishing a portion of the audio data from a target audio source and a portion of the audio data from a non-target audio source, where the input data are based on or include beamforming audio data.
[00161] Example 49 includes the apparatus of any of Examples 42 to 48 and additionally includes means for performing speech augmentation operations to determine augmented speech audio data, where the input data are based on or include augmented speech audio data.
[00162] Example 50 includes the apparatus of any of Examples 42 to 49 and additionally includes means for performing source separation operations to determine source-separated audio data where the input data are based on or include source-separated audio data.
[00163] Example 51 includes the apparatus of any of Examples 42 to 50, wherein the time-domain filter coefficients include linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.
[00164] Example 52 includes the apparatus of any of Examples 42 to 51, in which the means for performing time-domain filtering, the means for performing one or more transform operations, the means for processing the input data, and the means for performing Petition 870250084259, dated 09 / 18 / 2025, pp. 75 / 171 69 / 71 as one or more reverse transform operations are integrated into a wearable device.
[00165] Example 53 includes the apparatus of any of Examples 42 to 52, in which the means for performing time-domain filtering, the means for performing one or more transform operations, the means for processing the input data, and the means for performing one or more reverse transform operations are integrated into one or more in-ear headphones.
[00166] Example 54 includes the apparatus of any of Examples 42 to 53 and additionally includes means for generating one or more audio signals based on ambient sound.
[00167] Example 55 includes the apparatus of any of Examples 42 to 54 and additionally includes means for performing adaptive noise cancellation based on the suppressed noise output signal.
[00168] Example 56 includes the apparatus of any of Examples 42 to 55 and additionally includes means for generating output sound based on the suppressed noise output signal; means for generating a feedback signal based on the output sound; and means for performing adaptive noise cancellation based on the suppressed noise output signal and the feedback signal.
[00169] Those skilled in the art will further understand that the various logic blocks, configurations, modules, circuits, and illustrative algorithm steps described in connection with the implementations disclosed in the present invention can be implemented as electronic hardware, computer software running on a Petition 870250084259, dated 09 / 18 / 2025, pp. 76 / 171 70 / 71 processor, or combinations of both. Several illustrative components, blocks, configurations, modules, circuits, and stages have been described above in terms of their functionalities. Whether such functionality is implemented in the form of hardware or processor-executable instructions depends on the particular application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in various ways for each particular application; such implementation decisions should not be interpreted as causing a departure from the scope of this disclosure.
[00170] The steps of a method or algorithm described in connection with the implementations disclosed in the present invention may be incorporated directly into hardware, in a software module executed by a processor, or in a combination of both. A software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transient storage medium known in the art.An exemplary storage medium is coupled to the processor so that the processor can read information from it and write to it. Petition 870250084259, dated 09 / 18 / 2025, pp. 77 / 171 71 / 71 information on the storage medium. Alternatively, the storage medium may be an integral part of the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and storage medium may reside as distinct components in a computing device or user terminal.
[00171] The above description of the disclosed aspects is provided to enable a person skilled in the art to perform or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined in the present invention can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown in the present invention, but should be applied to the widest possible scope consistent with the innovative principles and attributes as defined by the following claims. Petition 870250084259, dated 09 / 18 / 2025, pp. 78 / 171
Claims
1 / 6 CLAIMS 1. Device characterized by comprising: one or more processors configured to: obtain audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment subsequent to the first segment; perform one or more transform operations on the first segment to generate frequency domain audio data; provide input data based on the frequency domain audio data as input to one or more machine learning models to generate a noise suppression output; perform one or more reverse transform operations on the noise suppression output to generate time domain filter coefficients; and perform time domain filtering of the second segment using the time domain filter coefficients to generate a noise-suppressed output signal.
2. Device according to claim 1, characterized in that the input data includes frequency domain audio data.
3. Device, according to claim 1, characterized in that one or more machine learning models are configured to generate output including a frequency mask representing an estimated magnitude of noise in the frequency domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask.
4. A device according to claim 1, characterized in that one or more machine learning models are configured to generate output including suppressed audio data, wherein one or more processors are configured to determine, based on the suppressed audio data, a frequency mask representing an estimated magnitude of noise in the frequency domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask.
5. Device according to claim 1, characterized in that, to generate the noise suppression output, one or more processors are configured to perform beamforming operations on the frequency domain audio data to determine beamforming audio data by distinguishing a portion of the audio data from a target audio source and a portion of the audio data from a non-target audio source, wherein the input data include the beamforming audio data.
6. Device according to claim 1, characterized in that, to process frequency domain audio data to generate noise suppression output, one or more processors are configured to perform speech augmentation operations to determine augmented speech audio data, wherein the input data include augmented speech audio data.
7. Device according to claim 1, characterized in that, to process frequency domain audio data to generate noise suppression output, one or more processors are configured to perform source separation operations to determine audio data with a separate source, wherein the input data includes audio data with a separate source.
8. Device according to claim 1, characterized in that the time-domain filter coefficients include linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, or all-pole filter coefficients.
9. Device according to claim 1, characterized in that one or more processors are integrated into a wearable device.
10. Device according to claim 1, characterized by further comprising one or more microphones, wherein one or more audio signals are received from one or more microphones.
11. Device according to claim 10, characterized by further comprising an adaptive noise-canceling filter coupled to at least one or more microphones.
12. Device according to claim 1, characterized by further comprising one or more loudspeakers and one or more microphones coupled to one or more processors and integrated into a wearable device, wherein the one or more microphones include at least one external microphone configured to generate the audio data and at least one feedback microphone configured to generate a feedback signal based on the sound produced by the one or more loudspeakers responsive to the noise-suppressed output signal.
13. A method characterized by comprising: obtaining audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment subsequent to the first segment; performing one or more transform operations on the first segment to generate frequency-domain audio data; providing input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; performing one or more reverse transform operations on the noise-suppressed output to generate time-domain filter coefficients; and performing time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal.
14. Method according to claim 13, characterized in that the input data includes frequency domain audio data.
15. A method according to claim 13, characterized in that one or more machine learning models are configured to generate output including a frequency mask representing an estimated magnitude of noise in the frequency-domain audio data for each frequency bin of a plurality of frequency bins, and wherein the noise suppression output includes the frequency mask. Petition 870250084259, dated 09 / 18 / 2025, pp. 154 / 171 5 / 6 16. A method according to claim 13, characterized in that one or more machine learning models are configured to generate output including suppressed audio data, and further comprising determining, based on the suppressed audio data, a frequency mask representing an estimated magnitude of noise in the frequency domain audio data for each frequency bin of a plurality of frequency bins, wherein the noise suppression output includes the frequency mask.
17. A method according to claim 13, characterized by further comprising performing beamforming operations on frequency-domain audio data to determine beamformed audio data by distinguishing a portion of the audio data from a target audio source and a portion of the audio data from a non-target audio source, wherein the input data includes the beamformed audio data.
18. A method according to claim 13, characterized by further comprising performing speech augmentation operations to determine audio data with augmented speech, wherein the input data include audio data with augmented speech.
19. A method according to claim 13, characterized by further comprising performing source separation operations to determine source-separated audio data, wherein the input data includes source-separated audio data.
20. Non-transitory, computer-readable medium characterized by storing instructions that are executable by Petition 870250084259, dated 09 / 18 / 2025, page.155 / 171 6 / 6 one or more processors to perform the following: obtain audio data representing one or more audio signals, wherein the audio data includes a first segment and a second segment subsequent to the first segment; perform one or more transform operations on the first segment to generate frequency-domain audio data; provide input data based on the frequency-domain audio data as input to one or more machine learning models to generate a noise-suppressed output; perform one or more reverse transform operations on the noise-suppressed output to generate time-domain filter coefficients; and perform time-domain filtering of the second segment using the time-domain filter coefficients to generate a noise-suppressed output signal. Petition 870250084259, dated 09 / 18 / 2025, pp. 156 / 171.