Target signal-to-noise ratio based audio processing

By combining time-domain and frequency-domain processing techniques, and utilizing low-latency recurrent neural networks and frequency-domain filter coefficients to separate target audio from non-target audio, the problem of excessive delay in improving signal-to-noise ratio and loss of context awareness in wearable devices in noisy environments is solved, achieving low-latency, high-quality target signal-to-noise ratio audio processing.

CN120883276APending Publication Date: 2025-10-31QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480018969.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-24
Filing Date
2024-03-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing wearable devices struggle to effectively separate target and non-target audio in noisy environments, resulting in excessive delays in improving signal-to-noise ratio or loss of context awareness.

Method used

By combining time-domain and frequency-domain processing techniques, low-latency recurrent neural networks and frequency-domain processing are used to generate time-domain filter coefficients, achieving high-quality separation of target audio and non-target audio, and adjusting the gain according to the target signal-to-noise ratio to generate the output audio signal.

Benefits of technology

It enables the generation of output audio signals with a target signal-to-noise ratio under low latency, improving the user's perception of the target sound without losing environmental cues and providing a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120883276A_ABST
    Figure CN120883276A_ABST
Patent Text Reader

Abstract

An apparatus includes one or more processors configured to obtain data specifying a target signal-to-noise ratio based on a hearing condition of a person, and configured to obtain audio data representing one or more audio signals. The one or more processors are configured to determine a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on the target signal-to-noise ratio. The one or more processors are configured to apply a first gain to a first component of the audio data to generate a target signal, and to apply a second gain to a second component of the audio data to generate a noise signal. The one or more processors are further configured to combine the target signal and the noise signal to generate an output audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] I. Cross-references to related applications

[0002] This application claims the priority of jointly owned U.S. Nonprovisional Patent Application No. 18 / 323,176, filed May 24, 2023, and U.S. Provisional Patent Application No. 63 / 493,158, filed March 30, 2023, the contents of each of which are expressly incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure generally relates to audio processing based on a target signal-to-noise ratio.

[0004] III. Related Technical Descriptions

[0005] Various types of hearing-related problems affect a large number of people. For example, a common problem is that even people with relatively normal hearing may find it difficult to hear speech in noisy environments, and the problem can be much worse for people with hearing loss. For some individuals, speech is only easily understood when the signal-to-noise ratio (of speech relative to ambient noise) is above a certain level.

[0006] In many cases, wearable devices (e.g., earbuds, headphones, hearing aids, etc.) can be used to improve hearing, contextual awareness, and speech intelligibility. Typically, such devices employ relatively simple noise suppression procedures to remove as much ambient noise as possible. While these noise suppression procedures can adequately improve the signal-to-noise ratio to make speech intelligible, they can also reduce the user's contextual awareness because they attempt to simply remove as much noise as possible, potentially missing important environmental cues such as traffic sounds. Using more complex noise suppression procedures can introduce significant latency. Latency when processing real-time speech can lead to user dissatisfaction. Summary of the Invention

[0007] According to one embodiment of this disclosure, an apparatus includes one or more processors configured to acquire data with a specified target signal-to-noise ratio based on a person's hearing condition, and configured to acquire audio data representing one or more audio signals. The one or more processors are configured to determine, based on the target signal-to-noise ratio, a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data. The one or more processors are configured to apply the first gain to the first component of the audio data to generate a target signal, and are configured to apply the second gain to the second component of the audio data to generate a noise signal. The one or more processors are further configured to combine the target signal and the noise signal to generate an output audio signal.

[0008] According to another embodiment of this disclosure, a method includes obtaining data with a specified target signal-to-noise ratio based on a person's hearing condition at one or more processors, and obtaining audio data representing one or more audio signals at the one or more processors. The method further includes determining, at the one or more processors, a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on the target signal-to-noise ratio. The method further includes applying the first gain to the first component of the audio data to generate a target signal, and applying the second gain to the second component of the audio data to generate a noise signal. The method further includes combining the target signal and the noise signal to generate an output audio signal.

[0009] According to another embodiment of this disclosure, a non-transitory computer-readable medium stores instructions executable by one or more processors to cause the processors to obtain data representing a specified target signal-to-noise ratio based on human hearing conditions, and to obtain audio data representing one or more audio signals. The instructions further cause the processors to determine, based on the target signal-to-noise ratio, a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data. The instructions further cause the processors to apply the first gain to the first component of the audio data to generate a target signal, and to apply the second gain to the second component of the audio data to generate a noise signal. The instructions further cause the processors to combine the target signal and the noise signal to generate an output audio signal.

[0010] According to another embodiment of this disclosure, an apparatus includes components for obtaining data with a specified target signal-to-noise ratio based on a person's hearing condition. The apparatus also includes components for obtaining audio data representing one or more audio signals. The apparatus further includes components for determining a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on the target signal-to-noise ratio. The apparatus also includes components for applying the first gain to the first component of the audio data to generate a target signal. The apparatus further includes components for applying the second gain to the second component of the audio data to generate a noise signal. The apparatus also includes components for combining the target signal and the noise signal to generate an output audio signal.

[0011] Other aspects, advantages, and features of this disclosure will become apparent upon examination of the entire application, which includes the description of the drawings, detailed description, and claims. Attached Figure Description

[0012] Figure 1This is a block diagram of a specific aspect of a device operable to perform audio processing based on a target signal-to-noise ratio, according to some examples of this disclosure.

[0013] Figure 2 Based on some examples of this disclosure Figure 1 A block diagram illustrating an exemplary aspect of a device operable to perform audio processing based on a target signal-to-noise ratio.

[0014] Figure 3 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target signal-to-noise ratio.

[0015] Figure 4 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target signal-to-noise ratio.

[0016] Figure 5 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target signal-to-noise ratio.

[0017] Figure 6 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target signal-to-noise ratio.

[0018] Figure 7 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target signal-to-noise ratio.

[0019] Figure 8 Examples of integrated circuits operable to perform audio processing based on a target signal-to-noise ratio, according to some examples of this disclosure, are illustrated.

[0020] Figure 9 This is a diagram illustrating a headset operable to perform audio processing based on a target signal-to-noise ratio, according to some examples of this disclosure.

[0021] Figure 10 This is an illustration of a headset (such as a virtual reality, mixed reality, or augmented reality headset) operable to perform audio processing based on a target signal-to-noise ratio, according to some examples of this disclosure.

[0022] Figure 11 This is an illustration of augmented reality glasses operable to perform audio processing based on a target signal-to-noise ratio, according to some examples of this disclosure.

[0023] Figure 12 This is an illustration of a wearable device operable to perform audio processing based on a target signal-to-noise ratio, according to some examples of this disclosure.

[0024] Figure 13 This is a diagram illustrating an earbud operable to perform audio processing based on a target signal-to-noise ratio, according to some examples of this disclosure.

[0025] Figure 14 Based on some examples of this disclosure, it is possible to... Figure 1 The diagram illustrates a specific implementation of an audio processing method based on a target signal-to-noise ratio, performed by the device.

[0026] Figure 15 This is a block diagram of a particular exemplary example of a device operable to perform audio processing based on a target signal-to-noise ratio, according to some examples of this disclosure. Detailed Implementation

[0027] The aspects disclosed herein achieve audio processing in a hearing-aiding manner by separately adjusting the levels of speech (or other target sounds) and noise (e.g., non-target sounds) and mixing the resulting signals to meet a target signal-to-noise ratio (SNR). As an example, if the target SNR is 15 dB, the gain applied to the target sound signal and the non-target sound signal is selected to provide a 15 dB SNR in the output audio signal. In this example, 15 dB is merely illustrative. In some specific implementations, the target SNR can be user-configurable (e.g., based on settings specified via an application and uploaded to a wearable device, such as earbuds, headphones, or hearing aids). As an example, the target SNR can be configured via a smartphone application used to control or manipulate settings on a wearable device.

[0028] One challenge in providing a consistent SNR for the output audio signal is the latency associated with reliably separating the input audio signal into target and non-target audio components. Higher quality source separation can be performed in the frequency domain compared to the time domain; however, frequency domain processing introduces significant latency, which can lead to a poor user experience. In contrast, time domain processing offers low latency but struggles to reliably separate target and non-target sources, especially in dynamic environments.

[0029] In some implementations, these challenges are addressed by using low-latency time-domain processing for source separation. For example, low-latency recurrent neural networks can be used to distinguish between target and non-target audio components. In other implementations, these challenges are addressed by combining time-domain and frequency-domain processes. For illustration, time-domain processing can be used for source separation, where the time-domain processing is guided by the output of frequency-domain processing (e.g., controlled or adjusted based on that output). For example, frequency-domain processing can be used to determine filter coefficients used by a time-domain filter to separate audio data into target and non-target components. In this example, the time-domain filter generates a signal representing the target audio component and a signal representing the non-target audio component. After the target and non-target audio components are separated, the resulting signals are individually gain-adjusted and mixed to generate an output audio signal with a target SNR.

[0030] Time-domain filters generated based on frequency-domain processing provide high-quality source separation and achieve low latency by separating the time-domain processing path from the frequency-domain processing path, ensuring that the time-domain processing path is updated whenever new time-domain filter coefficients are available from the frequency-domain processing path. In this specific implementation, time-domain filter coefficients determined based on frequency-domain processing provide significantly better noise suppression than traditional time-domain-only methods (such as adaptive noise cancellation). Furthermore, since the time-domain filter coefficients are applied to the received audio data in the time domain, processing the audio data using such time-domain filter coefficients adds little or no latency.

[0031] Specific aspects of this disclosure are described below with reference to the accompanying drawings. In this description, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular embodiments and are not intended to limit the scope of the embodiments. For example, the singular forms “a,” “an,” and “the” are intended to also include the plural forms, unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some embodiments and plural in others. For illustrative purposes, Figure 1 It describes a system that includes one or more processors ( Figure 1 The device 100 is referred to as "processor 190", which indicates that in some embodiments, device 100 includes a single processor 190, while in other embodiments, device 100 includes multiple processors 190. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular or optional plural form (as indicated by "(multiple)"), unless the aspect described relates to multiples of features.

[0032] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference numerals are used for each feature, and these different instances are distinguished by adding letters to the reference numerals. Reference numerals are used without distinguishing letters when a feature is referenced herein as a group or a type of feature (e.g., when a specific feature among these features is not referenced). However, reference numerals are used with distinguishing letters when a specific feature among multiple features of the same type is mentioned herein. For example, see reference... Figure 2 The figure illustrates multiple microphones, which are associated with reference numerals 102A and 102B. When referring to a specific microphone among these microphones (such as microphone 102A), the distinguishing letter "A" is used. However, when referring to any one of these microphones or to these microphones as a group, reference numeral 102 is used without the distinguishing letter.

[0033] As used herein, the term "comprise" is used interchangeably with "include". Similarly, the term "wherein" is used interchangeably with "where". As used herein, "exemplary" indicates an example, specific implementation, and / or aspect, and should not be construed as restrictive or indicating a preference or preferred implementation. As used herein, ordinal terms used to modify elements (such as structures, components, operations, etc.) (e.g., "first", "second", "third", etc.) do not themselves indicate any priority or order of that element relative to another element, but merely distinguish that element from another element with the same name (but using ordinal terms). As used herein, the term "set" refers to one or more specific elements among specific elements, while the term "multiple" refers to multiple (e.g., two or more) specific elements.

[0034] As used herein, “coupling” can include “communicationally coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof). As an illustrative, non-limiting example, two electrically coupled devices (or components) may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling. In some specific implementations, two communicationally coupled (e.g., electrically connected) devices (or components) may transmit and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, “direct coupling” can include two devices coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) without intermediate components.

[0035] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” and “adjust” can be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as restrictive, and similar operations can be performed using other techniques. Furthermore, as mentioned herein, “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “generating,” “calculating,” “estimate,” or “determining” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining the parameter (or signal), or it can refer to using, selecting, or accessing a parameter (or signal) that has already been generated (e.g., by another component or device).

[0036] Figure 1 This is a block diagram of a specific aspect of a device 100 operable to perform target SNR-based audio processing, according to some examples of this disclosure. Figure 1In this embodiment, device 100 includes one or more processors 190 coupled to one or more microphones 102 and one or more speakers 130. As an example, device 100 includes a wearable device (e.g., earbuds, headsets, hearing aids, or any similar device) or a corresponding wearable device configured to process audio data 104 representing one or more audio signals from microphones 102. For example, microphones 102 are configured to capture sounds 170 (e.g., ambient sounds) that may include speech (or another target sound) and noise (e.g., non-target sounds) to generate audio data 104. Processor 190 is configured to process audio data 104 to generate an output audio signal 132, which is adjusted to provide an output sound 172 with a target SNR (e.g., as indicated by setting 116) to improve the user's perception of the target sound without significantly sacrificing contextual awareness.

[0037] exist Figure 1 In this processor 190, a source separator 106, a gain determiner 114, a first gain module 120, a second gain module 124, and a mixer 128 are included. In a particular aspect, the source separator 106 is configured to perform a time-domain source separation operation to generate a first signal 108 representing a first component of audio data 104 and a second signal 110 representing a second component of audio data 104. The first component includes a first portion of the audio data 104 representing speech (or other target sound), and the second component includes a second portion of the audio data 104 representing non-speech (or other non-target) sound. For example, the source separator 106 may include one or more machine learning models (e.g., recurrent neural networks) configured and trained to perform low-latency time-domain sound source separation. As another example, the source separator 106 includes two or more time-domain filters. In this example, the source separator 106 is configured to apply first time-domain filter coefficients to generate the first signal 108 representing the first component and apply second time-domain filter coefficients to generate the second signal 110 representing the second component.

[0038] Gain determiner 114 is configured to determine a first gain to be applied to a first component of audio data 104 and a second gain to be applied to a second component of audio data 104 based on a target SNR indicated by setting 116. In some embodiments, processor 190 retrieves setting 116 from memory 162 accessible to processor 190. In some embodiments, setting 116 may be stored in memory 162 based on user input received via a user interface of device 100. In some embodiments, setting 116 may be received at device 100 from a second device 150. For example, device 100 may include a modem 144 coupled to processor 190. In such embodiments, device 100 may communicate with second device 150 via modem 144. In some such embodiments, second device 150 may include a computing device or a mobile communication device (e.g., a smartphone). For example, second device 150 may include an application configured to present a user interface (e.g., graphical user interface (GUI) 152) to allow a user to specify one or more parameters of setting 116. For illustration, in Figure 1 In the GUI 152, there are user-selectable elements 154 for specifying SNR settings (e.g., target SNR), user-selectable elements 156 for specifying volume settings, user-selectable elements 158 for specifying spectrum tilt settings, and user-selectable elements 160 for specifying one or more other settings.

[0039] Optionally, the gain determiner 114 receives a signal 112 from the source separator 106 indicating the relative amplitudes of the first signal 108 and the second signal 110, and determines the first gain and the second gain based at least in part on the signal 112. The gain determiner 114 provides a signal 140 to the first gain module 120 so that the first gain module 120 applies the first gain to a first component of the audio data 104 to generate a first gain adjustment signal 122 based on the first component. The gain determiner 114 also provides the signal 142 to the second gain module 124 so that the second gain module 124 applies the second gain to a second component of the audio data 104 to generate a second gain adjustment signal 126 based on the second component.

[0040] In a specific aspect, the gain determiner 114 selects a first gain and a second gain to preserve the target sound. For example, if signals 108 and 110 have the same power and the target SNR is 10dB, the gain determiner 114 may set the first gain to 0dB to preserve the target sound and set the second gain to -10dB to achieve the target SNR.

[0041] Mixer 128 is configured to combine a first gain adjustment signal 122 and a second gain adjustment signal 126 to generate an output audio signal 132. In a particular implementation, gain determiner 114 determines the first gain and the second gain such that the output audio signal 132 has an SNR based on a target SNR. By using time-domain source separation, processor 190 is able to generate the output audio signal 132 with very low latency. For example, the delay between obtaining specific audio data and generating the corresponding output audio signal representing the specific audio data is less than one millisecond.

[0042] In some specific implementations, device 100 corresponds to or is included in one of a variety of types of devices. In an exemplary example, processor 190 is integrated into a wearable device that includes or is coupled to microphone 102 and speaker 130. Examples of such wearable devices include, but are not limited to, those referenced in [reference missing]. Figure 9 Further description of the headset device; as referenced Figure 10 The virtual reality, mixed reality, or augmented reality headset described; as referenced Figure 11 The augmented reality glasses described; as referenced Figure 12 The hearing aid device described; or as referenced Figure 13 The earplugs described.

[0043] One technical advantage of device 100 is its ability to generate output sound 172 with a target SNR, thereby emphasizing the target audio (e.g., speech) without excessively losing environmental cues associated with non-target audio. Another technical advantage of device 100 is the generation of output sound 172 with low latency (e.g., less than a few milliseconds, such as less than 2 milliseconds, less than 1.5 milliseconds, or less than 1 millisecond), which provides a better user experience compared to longer latency. In some aspects, as further described below, source separator 106 operates in the time domain by applying filter coefficients based on frequency domain processing. This implementation offers the additional technical advantage of high-quality source separation without increasing latency.

[0044] Figure 2 This is a block diagram illustrating an exemplary aspect of a device 200 operable to perform target SNR-based audio processing according to some examples of this disclosure. In a particular aspect, device 200 is Figure 1 An example of a specific implementation of device 100. For example, device 200 includes components coupled to a reference. Figure 1 The processor 190 of the described microphone 102 and speaker 130. Figure 2 The processor 190 includes reference Figure 1The described components include a source splitter 106, a gain determiner 114, a first gain module 120, a second gain module 124, and a mixer 128. Device 200 includes a wearable device (e.g., an earbud, headset, hearing aid, or any similar device) or a corresponding wearable device configured to process audio data 104 representing one or more audio signals from microphone 102 to generate an output audio signal 132, which is adjusted to provide an output sound 172 with a target SNR (e.g., as indicated by setting 116) to improve the user's perception of the target sound without significantly sacrificing contextual awareness.

[0045] exist Figure 2 In the example illustrated, processor 190 is configured to perform operations associated with two data paths, including a first data path 250 and a second data path 260. The first data path 250 is a low-latency data path and includes a source splitter 106, a first gain module 120, a second gain module 124, a mixer 128, and optionally other components or modules. In a particular aspect, the operation of the first data path 250 is performed in the time domain. Therefore, the latency associated with domain transformation operations is avoided in the first data path 250. In contrast, the second data path 260 is configured to provide high-quality source separation at the cost of greater latency than the first data path 250. For example, in some implementations, the first data path 250 is associated with a latency of 1 millisecond or less, and the second data path 260 is associated with a latency greater than 1 millisecond. As another example, in some implementations, the first data path 250 may be associated with a latency of 2 milliseconds or less, and the second data path 260 with a latency greater than 10 milliseconds.

[0046] The second data path 260 is configured to determine filter coefficients (e.g., time-domain filter coefficients) for use by the filters of the source separator 106 (e.g., the first extraction filter 220 and the second extraction filter 222). As an example, the filter coefficients determined by the operation of the second data path 260 are stored in one or more buffers 280 and subsequently applied by the filters of the source separator 106 (in the time domain).

[0047] exist Figure 2 In this second data path 260, a filter bank 202, one or more frequency domain source splitters 204, and two or more filter designers 210, 212 are included. In some embodiments, a buffer 280 is included in filter designers 210, 212 or in another component or module of the second data path 260. In other embodiments, a buffer 280 is included in extraction filters 220, 222 or in another component or module of the first data path 250.

[0048] Filter bank 202 is configured to perform one or more transformation operations (e.g., Fast Fourier Transform (FFT) operations) based on samples of audio data 104 to generate frequency-domain audio data. According to some specific implementations, a sample set of audio data 104 is accumulated for processing by filter bank 202 (e.g., at one or more buffers of filter bank 202), and the frequency-domain audio data of the sample set includes information indicating the amplitude of the sound within each of a plurality of frequency intervals. During or after transforming the sample set to the frequency domain, subsequent sets of samples of audio data 104 are accumulated to be transformed into subsequent sets of frequency-domain audio data.

[0049] Frequency domain source separator 204 is configured to process frequency domain audio data from filter bank 202 to distinguish sounds from various sources, various types of sounds, or both. For example, the frequency domain audio data generated by filter bank 202 represents each set of samples of audio data 104 as a set of amplitudes associated with frequency intervals, and frequency domain source separator 204 generates frequency domain target audio data 206 indicating the amplitude of frequency intervals associated with target sounds, and frequency domain non-target audio data 208 indicating the amplitude of frequency intervals associated with non-target sounds. In some embodiments, target sounds include speech sounds, and non-target sounds include non-speech sounds. In other embodiments, target sounds include music, speech sounds from a specific person, sounds from a specific direction, etc. In such embodiments, non-target sounds include noise, non-speech ambient sounds, speech from people other than the target speaker, sounds from directions other than the target direction, etc. Although Figure 2 The example illustrates that the frequency domain source separator 204 generates two outputs corresponding to the frequency domain target audio data 206 and the frequency domain non-target audio data 208. However, in other specific implementations, the frequency domain source separator 204 generates more than two outputs, such as outputs representing multiple different target sounds (e.g., speech from two different speakers), outputs representing multiple different non-target sounds (e.g., crowd noise, traffic noise, and animal sounds), or both.

[0050] In some implementations, the frequency domain source separator 204 includes one or more machine learning models trained to distinguish various sound sources or sound types. For example, the frequency domain source separator 204 may include one or more recurrent neural networks (such as neural networks including one or more long short-term memory layers, one or more gated recurrent units, or other recurrent structures) trained to distinguish target sounds from non-target sounds. In the same or different implementations, the frequency domain source separator 204 may include one or more beamformers to distinguish target sounds from one or more directions from non-target sounds from other directions. In the same or different implementations, the frequency domain source separator 204 may perform operations to distinguish target sounds from non-target sounds based on the statistical properties of the sounds (e.g., using blind source separation techniques and / or speech enhancement techniques).

[0051] In some implementations, the frequency domain source splitter 204 is configured to use two or more techniques to determine the frequency domain target audio data 206 and the frequency domain non-target audio data 208. For example, the frequency domain source splitter 204 may include one or more beamformers configured to process the frequency domain audio data from the filter bank 202 to distinguish sounds from different directions. In this example, the beamformer generates directional data (e.g., directional frequency domain audio data or directional data associated with the frequency domain audio data from the filter bank 202) and provides the directional data as input to one or more machine learning models or other source splitters to generate the frequency domain target audio data 206 and the frequency domain non-target audio data 208. In this example, the frequency domain target audio data 206 and the frequency domain non-target audio data 208 may be distinguished based on both sound type and source direction. As another example, the frequency domain source splitter 204 may include one or more speech enhancement engines configured to process the frequency domain audio data from the filter bank 202 to enhance speech sounds. In this example, speech-enhanced audio data from the speech enhancement engine can be fed as input to one or more machine learning models or other source separators to generate frequency domain target audio data 206 and frequency domain non-target audio data 208.

[0052] In the examples above, beamforming and speech enhancement are described as examples of preprocessing operations that can be performed before using another process or technique (e.g., blind source separation, one or more machine learning models, etc.) to generate frequency-domain target audio data 206 and frequency-domain non-target audio data 208. In addition to or instead of such preprocessing operations, frequency-domain source separator 204 can be configured to perform post-processing operations to generate frequency-domain target audio data 206 and frequency-domain non-target audio data 208. For example, frequency-domain source separator 204 may include a machine learning model (sometimes referred to as an "inline" model) trained to generate output audio data representing the target sound. In this example, the output of the machine learning model includes frequency-domain target audio data 206, and post-processing operations can be performed to remove the frequency-domain target audio data 206 from the frequency-domain audio data from filter bank 202 to generate frequency-domain non-target audio data 208. In an alternative example, frequency-domain source separator 204 may include a machine learning model (sometimes referred to as a "mask" model) trained to generate output audio data representing non-target sounds. In this example, the output of the machine learning model includes frequency domain non-target audio data 208, and post-processing operations can be performed to remove the frequency domain non-target audio data 208 from the frequency domain audio data from the filter bank 202 to generate frequency domain target audio data 206.

[0053] Each of filter designers 210 and 212 is configured to generate time-domain filter coefficients based on frequency-domain audio data. For example, filter designer 210 is configured to generate a first coefficient 214 based on frequency-domain target audio data 206, and filter designer 212 is configured to generate a second coefficient 216 based on frequency-domain non-target audio data 208. Although filter designers 210 and 212 are configured to generate time-domain filter coefficients based on frequency-domain non-target audio data 208, the time-domain filter coefficients are generated based on frequency-domain non-target audio data 206. Figure 2While illustrated as a separate component, in some implementations, device 200 may include a single filter designer that processes the frequency-domain target audio data 206 to generate first coefficients 214 and processes the frequency-domain non-target audio data 208 to generate second coefficients 216. For example, the single filter designer may generate the first coefficients 214 based on the frequency-domain target audio data 206, and subsequently generate the second coefficients 216 for the frequency-domain non-target audio data 208 associated with the frequency-domain target audio data 206. For illustration, when the frequency-domain source separator 204 includes a post-processing operation to generate the frequency-domain non-target audio data 208 based on the frequency-domain target audio data 206, the frequency-domain target audio data 206 may be provided to the single filter designer in parallel with providing the frequency-domain target audio data 206 to the post-processing operation. In this example, the post-processing operation generates frequency domain non-target audio data 208, while a single filter designer generates a first coefficient 214, and when the post-processing operation is complete, the frequency domain non-target audio data 208 generated by the post-processing operation is provided to the single filter designer to generate a second coefficient 216.

[0054] In certain respects, each of the filter designers 210, 212 is configured to use frequency domain audio data (e.g., in...). Figure 2 In the example illustrated, one or more inverse transform operations are performed on the target audio data 206 and the non-target audio data 208 in the frequency domain, respectively, to generate time-domain filter coefficients. Filter designers 210 and 212 can be configured to generate coefficients 214 and 216 as real-valued masks or as complex-valued masks. Examples of real-valued masks that can be generated by filter designers 210 and 212 in some specific implementations include linear-phase finite impulse response (FIR) filter coefficients, minimum-phase FIR filter coefficients, autoregressive filter coefficients, infinite impulse response (IIR) filter coefficients, all-pole filter coefficients, etc. Complex-valued masks may include FIR filters or IIR filters that indicate amplitude and phase. The technical advantage of using linear-phase FIR filter coefficients is predictable delay, because the delay associated with applying linear-phase FIR filter coefficients depends entirely on the length of the FIR filter. The technical advantage of using minimum-phase FIR filter coefficients or autoregressive filter coefficients is reduced delay, because although the delay introduced by applying such filter coefficients is frequency-dependent, the delay is the minimum possible for a given input data.

[0055] Extraction filters 220 and 222 are updated periodically or occasionally (e.g., when updated coefficients 214 and 216 become available). In this arrangement, extraction filters 220 and 222 apply time-domain filter coefficients (e.g., coefficients 214 and 216) to audio data 104 that is newer than the audio data 104 used to generate coefficients 214 and 216. As a non-limiting example, the operation of the second data path 260 can be performed within a period of approximately 16 milliseconds; while the operation of the first data path 250 can be performed within a period of approximately 1 millisecond. Therefore, in this particular example, the time-domain filter coefficients applied to a particular data sample in the first data path 250 are always at least 16 milliseconds earlier than that data sample. Except in unusual circumstances, ambient noise typically changes slowly enough that even with this delay, the time-domain filter coefficients are sufficiently representative to provide reliable sound separation.

[0056] During operation, microphone 102 generates a signal based on sound 170. The signal may include (e.g., at a decoder / decoder (codec)) a digital signal (e.g., audio data 104) or an analog signal that is processed to generate audio data 104. Sound 170 may include speech or other target sounds as well as non-target sounds, such as noise. Audio data 104 representing sound 170 is provided to a first data path 250 for low-latency processing and to a second data path 260 for frequency domain processing to generate or update time-domain filter coefficients for use by extraction filters 220, 222.

[0057] In the first data path 250, audio data 104 is processed by a first extraction filter 220 to generate a first signal 108 representing a first component of the audio data 104, where the first component represents speech (or other target sound). For example, the first extraction filter 220 applies a first set of time-domain filter coefficients to the audio data 104. The first set of time-domain filter coefficients corresponds to a first coefficient 214 previously generated by the second data path 260 after processing a previous set of audio data 104. The first signal 108 is provided to a first gain module 120, which applies a first gain to the first signal 108 to generate a first gain adjustment signal 122. The first gain applied by the first gain module 120 is based on signal 140 from the gain determiner 114.

[0058] Additionally, in the first data path 250, audio data 104 is processed by a second extraction filter 222 to generate a second signal 110 representing a second component of the audio data 104, where the second component represents noise (or other non-target sound). For example, the second extraction filter 222 applies a second set of time-domain filter coefficients to the audio data 104. This second set of time-domain filter coefficients corresponds to second coefficients 216 previously generated by the second data path 260 after processing the previous set of audio data 104. The second signal 110 is provided to a second gain module 124, which applies a second gain to the second signal 110 to generate a second gain adjustment signal 126. The second gain applied by the second gain module 124 is based on signal 142 from the gain determiner 114.

[0059] Gain determiner 114 determines signals 140 and 142 based on setting 116. For example, setting 116 may specify a target SNR or be used to determine a target SNR, and gain determiner 114 may determine signals 140 and 142 such that the SNR of the output sound 172 substantially (e.g., within the operating tolerance of device 100) matches the target SNR. Optionally, in some implementations, gain determiner 114 causes signals 140 and 142 to be at least partially based on signal 112 (… Figure 1 (As shown). For example, signal 112 may indicate an estimate of the SNR of the audio data, which gain determiner 114 may use to determine the gain to be applied to the target sound and non-target sounds to generate an output sound 172 with a target SNR. In the same or different embodiments, gain determiner 114 may optionally receive feedback signal 278 from a microphone 102C positioned near speaker 130. In such embodiments, gain determiner 114 may use feedback signal 278 to estimate the SNR of output sound 172 and determine the gain to be applied to the target sound and non-target sounds to adjust the SNR of output sound 172 toward the target SNR.

[0060] Mixer 128 mixes a first gain adjustment signal 122 and a second gain adjustment signal 126. In some embodiments, the output of mixer 128 is used to drive speaker 130 to generate output sound 172. Optionally, in some embodiments, device 200 includes a feedback adaptive noise cancellation (ANC) filter 240. In such embodiments, feedback ANC filter 240 is configured to receive a feedback signal 278, which may include audio data representing the output sound 172 and portions of sound 170 as modified by transfer function (P(z)) 272. For example, transfer function 272 may be due to a portion of device 200 partially or completely blocking the user's ear canal 276. For example, device 200 may include earmuffs, ear covers, or another structure that at least partially blocks or closes the ear canal 276. In this case, a portion of sound 170 reaching the ear canal 276 may be modified due to passing through earmuffs, ear covers, or other structures, resulting in an unnatural sound (e.g., due to some frequencies attenuating more than others). Feedback signal 278 represents user-perceptible sounds, which may include output sound 172 and sound 170 as modified by transfer function 272. Feedback ANC filter 240 generates feedback ANC signal 242, which is used to modify the output of mixer 128 to account for the portion of sound 170 that is perceptible to the user. For example, the portion of sound 170 represented in feedback signal 278 may be subtracted from the output of mixer 128.

[0061] Optionally, in some embodiments, device 200 further includes, or alternatively includes, a feedforward ANC filter 244. In such embodiments, the feedforward ANC filter 244 is configured to receive audio data 104 and generate a feedforward ANC signal 246, which is used to modify the output of mixer 128 to account for noise in sound 170. For example, in some cases, the target audio includes speech, and the non-target audio includes non-speech sounds. In such cases, device 200 attempts to adjust the gain applied to non-speech sounds in a manner that gives the output sound 172 a target SNR, so that important audio information is not completely eliminated in order to improve the user's perception of speech. However, some sounds (such as hissing due to wind noise) are unlikely to carry audio information that provides any important environmental cues. Completely removing such sounds (e.g., pure noise) can produce a better user experience; therefore, such sounds can be completely or substantially removed from the output sound. Although Figure 2 The example illustrates the use of feedforward ANC signal 246 to modify the output of mixer 128, but in other implementations, feedforward ANC signal 246 may be used to modify audio data 104 provided to a first data path 250, a second data path 260, or both.

[0062] In the second data path 260, the sample set of audio data 104 is accumulated and undergoes one or more transformation operations performed by filter bank 202 to generate frequency domain audio data representing the sample set. In some specific embodiments, the frequency domain audio data is processed by frequency domain source separator 204 to generate frequency domain target audio data 206 (which includes the target audio components of the frequency domain audio data and omits or suppresses the non-target audio components of the frequency domain audio data) and frequency domain non-target audio data 208 (which includes the non-target audio components of the frequency domain audio data and omits or suppresses the target audio components of the frequency domain audio data).

[0063] Frequency-domain target audio data 206 is provided as input to filter designer 210, which performs inverse transform and parameterization operations to generate first coefficients 214. The specific inverse transform and parameterization operations performed may differ depending on the specific implementation. As an example, the inverse transform operation may include various inverse Fourier transform operations, such as the inverse fast Fourier transform (IFFT) operation. The parameterization operation may include, for example, windowing or shifting the time-domain data generated by the inverse transform operation to generate a specific number of time-domain filter coefficients based on, for example, the number of filter coefficients applied by the first extraction filter 220. Applying a larger number of filter coefficients can provide greater noise suppression at the cost of greater computational complexity.

[0064] Similarly, frequency-domain non-target audio data 208 is provided as input to filter designer 212, which performs inverse transform and parameterization operations to generate second coefficients 216. The specific inverse transform and parameterization operations performed may differ for different implementations. Furthermore, the inverse transform and parameterization operations performed by filter designer 210 may differ from those performed by filter designer 212. For example, the first extraction filter 220 may be more complex than the second extraction filter 222 (and therefore may use more filter coefficients) to improve the extraction of the target audio data.

[0065] When generating the first coefficients 214 and the second coefficients 216 of the first set of samples based on audio data 104, additional samples of audio data 104 may be received and aggregated to form a second set of samples of audio data. After collecting the second set of samples, the second set of samples undergoes the same operations described above to update the first coefficients 214 and the second coefficients 216.

[0066] Figure 3 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target SNR. Figure 3 Equipment 300 indicates Figure 2 Equipment 200 (equipment 200 itself is) Figure 1 A specific non-limiting example of the device 100, therefore, Figure 3 The device 300 includes reference Figure 1 and Figure 2 Many of the same components are illustrated and described, each of which operates as described above. For example, device 300 includes components coupled to, as referenced... Figure 1 The processor 190 of the described microphone 102 and speaker 130. Additionally, Figure 3 The processor 190 includes reference Figure 1 The described components include a source splitter 106, a gain determiner 114, a first gain module 120, a second gain module 124, and a mixer 128. Furthermore, in... Figure 3 In this context, source splitter 106, first gain module 120, second gain module 124, and mixer 128 are as shown in the reference. Figure 2 The first data path 250 described is associated with (e.g., a low-latency time-domain data path). Figure 3 The device 300 also includes, as referenced Figure 2 The described second data path 260 (e.g., frequency domain data path) includes filter bank 202, frequency domain source splitter 204, and filter designers 210, 212. Although Figure 3 Not shown, but device 300 may also include Figure 1 or Figure 2 Optional features, such as Figure 1 Signal 112 Figure 2 Feedforward ANC filter 244 Figure 2 Feedback ANC filters 240 or combinations thereof.

[0067] exist Figure 3 In the example illustrated, the target audio includes speech and the non-target audio includes noise. Therefore, in Figure 3 In this context, the frequency domain source separator 204 includes one or more speech / noise separators 302. For example, the speech / noise separator 302 may include one or more machine learning models trained to distinguish between speech and noise (or speech and non-speech) sounds in frequency domain audio data. As another example, the speech / noise separator 302 may include one or more machine learning models trained to generate (e.g., reconstruct) speech data based on input frequency domain audio data, and one or more spectral subtraction components for generating non-speech components based on the reconstructed speech data and the input frequency domain audio data. Other frequency domain speech / noise separation techniques may also be implemented by the speech / noise separator 302.

[0068] exist Figure 3 In the example illustrated, the frequency-domain target audio data 206 includes a frequency-domain representation of speech in sound 170, and the frequency-domain non-target audio data 208 includes a frequency-domain representation of the non-speech (e.g., noise) components of sound 170. Filter designer 210 generates first coefficients 214 such that first extraction filter 220 is able to extract the speech components of audio data 104 in the time domain to generate first signal 108. Therefore, Figure 3 The first extraction filter 220 is a speech extraction filter 304. Similarly, the filter designer 212 generates second coefficients 216 such that the second extraction filter 222 is able to extract the non-speech (e.g., noise) components of the audio data 104 in the time domain to generate the second signal 110. Therefore, Figure 3 The second extraction filter 222 in the filter is the noise extraction filter 306.

[0069] exist Figure 3 In the above, a first signal 108 representing the speech component of audio data 104 is provided to a first gain module 120, which in turn... Figure 3 The speech gain module 310 is located in the middle. The speech gain module 310 applies a first gain (e.g., speech gain) to the first signal 108 to generate a first gain adjustment signal 122 (e.g., gain-adjusted speech signal). The speech gain applied by the speech gain module 310 is based on the signal 140 from the gain determiner 114.

[0070] Similarly, in Figure 3 In this process, a second signal 110 representing the non-verbal (e.g., noise) components of audio data 104 is provided to a second gain module 120, which in turn... Figure 3 The noise gain module 312 is located in the middle. The noise gain module 312 applies a second gain (e.g., noise gain) to the second signal 110 to generate a second gain adjustment signal 126 (e.g., gain-adjusted noise signal). The noise gain applied by the noise gain module 312 is based on the signal 142 from the gain determiner 114. Figure 3 In this process, mixer 128 mixes a gain-adjusted speech signal and a gain-adjusted noise signal to generate an output audio signal 132, which is adjusted to provide an output sound 172 with a target SNR (e.g., as indicated by setting 116) to improve the user's perception of speech without significantly losing contextual awareness.

[0071] Figure 4 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target SNR. Figure 4 Equipment 400 indicates Figure 2Equipment 200 (Equipment 200 is) Figure 1 A specific non-limiting example of the device 100, therefore, Figure 4 The device 400 includes reference Figure 1 and Figure 2 Many of the same components are illustrated and described, each of which operates as described above. For example, device 400 includes components coupled to, as referenced... Figure 1 The processor 190 of the described microphone 102 and speaker 130. Additionally, Figure 4 The processor 190 includes reference Figure 1 The described components include a source splitter 106, a gain determiner 114, a first gain module 120, a second gain module 124, and a mixer 128. Furthermore, in... Figure 4 In this context, source splitter 106, first gain module 120, second gain module 124, and mixer 128 are as shown in the reference. Figure 2 The first data path 250 described is associated with (e.g., a low-latency time-domain data path). Figure 4 The device 400 also includes, as referenced Figure 2 The described second data path 260 (e.g., frequency domain data path) includes filter bank 202, frequency domain source splitter 204, and filter designers 210, 212. Although Figure 4 Not shown, but device 400 may also include Figure 1 or Figure 2 Optional features, such as Figure 1 Signal 112 Figure 2 Feedforward ANC filter 244 Figure 2 Feedback ANC filters 240 or combinations thereof.

[0072] exist Figure 4 In the example illustrated, the frequency domain source splitter 204 includes one or more beamformers 402 (e.g., one or more minimum variance distortionless response (MVDR) beamformers). In this example, the target sound corresponds to portions of sound 170 originating from a specific direction (e.g., the target direction) relative to the user of device 400. For illustration, as explained above, device 400 may be a wearable device (such as a hearing aid, earbuds, headset, etc.), and the target sound may include sound originating from the direction the user is facing.

[0073] exist Figure 4In the frequency domain, target audio data 206 includes a frequency domain representation of sound from the target direction, and non-target audio data 208 includes a frequency domain representation of components of sound 170 from one or more other directions. Filter designer 210 generates first coefficients 214 such that first extraction filter 220 can extract components of audio data 104 originating from the target direction in the time domain to generate first signal 108. Similarly, filter designer 212 generates second coefficients 216 such that second extraction filter 222 can extract components of audio data 104 originating from directions other than the target direction in the time domain. As described above, first gain module 120 and second gain module 124, mixer 128, and gain determiner 114 operate with respect to first signal 108 and second signal 110 to generate output audio signal 132, which is adjusted to provide output sound 172 with a target SNR (e.g., as indicated by setting 116) to improve the user's perception of sound from the target direction without significantly sacrificing contextual awareness.

[0074] Figure 5 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target SNR. Figure 5 Equipment 500 indicates Figure 2 Equipment 200 (Equipment 200 is) Figure 1 A specific non-limiting example of the device 100, therefore, Figure 5 The device 500 includes reference Figure 1 and Figure 2 Many of the same components are illustrated and described, each of which operates as described above. For example, device 500 includes components coupled to, as referenced... Figure 1 The processor 190 of the described microphone 102 and speaker 130. Additionally, Figure 5 The processor 190 includes reference Figure 1 The described components include a source splitter 106, a gain determiner 114, a first gain module 120, a second gain module 124, and a mixer 128. Furthermore, in... Figure 5 In this context, source splitter 106, first gain module 120, second gain module 124, and mixer 128 are as shown in the reference. Figure 2 The first data path 250 described is associated with (e.g., a low-latency time-domain data path). Figure 5 The device 500 also includes, as referenced Figure 2 The described second data path 260 (e.g., frequency domain data path) includes filter bank 202, frequency domain source splitter 204, and filter designers 210, 212. Although Figure 5Not shown, but device 500 may also include Figure 1 or Figure 2 Optional features, such as Figure 1 Signal 112 Figure 2 Feedforward ANC filter 244 Figure 2 Feedback ANC filters 240 or combinations thereof.

[0075] exist Figure 5 In the example illustrated, the frequency domain source separator 204 includes a blind source separation module 502. The blind source separation module 502 is operable to distinguish portions of sound 170 from different sources based on the statistical characteristics of the sounds from different sources. In this example, the target sound corresponds to a portion of sound 170 having specific characteristics or a portion of that sound from the dominant sound source. For illustration, the blind source separation module 502 can distinguish portions of sound 170 from different sources based on the statistical characteristics of sound 170 and designate the sound from the sound source having a specific direction of arrival at device 500 as the target sound.

[0076] exist Figure 5 In the frequency domain, target audio data 206 includes a frequency domain representation of the target sound, and non-target audio data 208 includes frequency domain representations of other components of sound 170. Filter designer 210 generates first coefficients 214 to enable first extraction filter 220 to extract the target sound in the time domain, and filter designer 212 generates second coefficients 216 to enable second extraction filter 222 to extract non-target sounds in the time domain. As described above, first gain module 120 and second gain module 124, mixer 128, and gain determiner 114 operate with respect to first signal 108 and second signal 110 to generate an output audio signal 132, which is adjusted to provide an output sound 172 with a target SNR (e.g., as indicated by setting 116) to improve the user's perception of the target sound without significantly sacrificing contextual awareness.

[0077] Figure 6 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target SNR. Figure 6 Equipment 600 indicates Figure 2 Equipment 200 (Equipment 200 is) Figure 1 A specific non-limiting example of the device 100, therefore, Figure 6 The device 600 includes reference Figure 1 and Figure 2 Many of the same components are illustrated and described, each of which operates as described above. For example, device 600 includes components coupled to, as referenced... Figure 1 The processor 190 of the described microphone 102 and speaker 130. Additionally, Figure 6 The processor 190 includes reference Figure 1 The described components include a source splitter 106, a gain determiner 114, a first gain module 120, a second gain module 124, and a mixer 128. Furthermore, in... Figure 6 In this context, source splitter 106, first gain module 120, second gain module 124, and mixer 128 are as shown in the reference. Figure 2 The first data path 250 described is associated with (e.g., a low-latency time-domain data path). Figure 6 The device 600 also includes, as referenced Figure 2 The described second data path 260 (e.g., frequency domain data path) includes filter bank 202, frequency domain source splitter 204, and filter designers 210, 212. Although Figure 6 Not shown, but device 600 may also include Figure 1 or Figure 2 Optional features, such as Figure 1 Signal 112 Figure 2 Feedforward ANC filter 244 Figure 2 Feedback ANC filters 240 or combinations thereof.

[0078] exist Figure 6 In the example illustrated, the frequency domain source separator 204 includes one or more machine learning models 602. Depending on the specific type of machine learning model 602 used and how it is trained, the machine learning model 602 may distinguish the components of sound 170 based on direction, source, sound type (e.g., speech, music, crowd noise, traffic sounds, etc.), or a combination thereof. Therefore, when the frequency domain source separator 204 includes a machine learning model 602, the properties of the target sound are implementation-specific. Thus, using a machine learning model 602 as the frequency domain source separator 204 enables the selection of a wider variety of target sounds than many other source separation techniques.

[0079] exist Figure 6In the frequency domain, target audio data 206 includes a frequency domain representation of the target sound, and non-target audio data 208 includes frequency domain representations of other components of sound 170. Filter designer 210 generates first coefficients 214 to enable first extraction filter 220 to extract the target sound in the time domain, and filter designer 212 generates second coefficients 216 to enable second extraction filter 222 to extract non-target sounds in the time domain. As described above, first gain module 120 and second gain module 124, mixer 128, and gain determiner 114 operate with respect to first signal 108 and second signal 110 to generate an output audio signal 132, which is adjusted to provide an output sound 172 with a target SNR (e.g., as indicated by setting 116) to improve the user's perception of the target sound without significantly sacrificing contextual awareness.

[0080] Figure 7 Based on some examples of this disclosure Figure 1 The diagram illustrates an exemplary aspect of a device operable to perform audio processing based on a target SNR. Figure 7 Equipment 700 indicates Figure 1 This is a specific, non-limiting example of device 100, therefore, Figure 7 The device 700 includes reference Figure 1 Many of the same components are illustrated and described, each of which operates as described above. For example, device 700 includes components coupled to, as referenced... Figure 1 The processor 190 of the described microphone 102 and speaker 130. Additionally, Figure 7 The processor 190 includes reference Figure 1 The described source splitter 106, gain determiner 114, first gain module 120, second gain module 124, and mixer 128.

[0081] exist Figure 7 In the source separator 106, speech extraction neural network 702 and noise extraction neural network 704 are included. Figure 7 The speech extraction neural network 702 and the noise extraction neural network 704 operate in the time domain. For example, the speech extraction neural network 702 and the noise extraction neural network 704 each process audio data 104 in the time domain. In order to limit the time delay introduced by the source separator 106, the speech extraction neural network 702 and the noise extraction neural network 704 are low-latency neural networks.

[0082] exist Figure 7 In the example illustrated, the target audio includes speech and the non-target audio includes noise. Therefore, the first signal 108 represents the speech component of the audio data 104, and the second signal 110 represents the noise component of the audio data 104.

[0083] The first signal 108 is provided to the first gain module 120, which in turn... Figure 7 The speech gain module 720 is located in the middle. The speech gain module 720 applies a first gain (e.g., speech gain) to the first signal 108 to generate a first gain adjustment signal 122 (e.g., gain-adjusted speech signal). The speech gain applied by the speech gain module 720 is based on the signal 140 from the gain determiner 114.

[0084] The second signal 110 is provided to the second gain module 120, which in turn... Figure 7 The noise gain module 724 is located in the middle. The noise gain module 724 applies a second gain (e.g., noise gain) to the second signal 110 to generate a second gain adjustment signal 126 (e.g., gain-adjusted noise signal). The noise gain applied by the noise gain module 724 is based on the signal 142 from the gain determiner 114. Figure 7 In this process, mixer 128 mixes a gain-adjusted speech signal and a gain-adjusted noise signal to generate an output audio signal 132, which is adjusted to provide an output sound 172 with a target SNR (e.g., as indicated by setting 116) to improve the user's perception of speech without significantly losing contextual awareness.

[0085] Figure 8 A specific embodiment 800 of a device 100 is depicted as an integrated circuit 802 operable to perform audio processing based on a target signal-to-noise ratio. The integrated circuit 802 includes one or more processors 190 and audio inputs 804 (such as one or more bus interfaces) to enable audio data 104 to be received for processing by the processors 190. The integrated circuit 802 also includes signal outputs 806 (such as bus interfaces) to output an audio signal 132. Figure 8 In the integrated circuit 802, the processor 190 includes one or more audio processing components 840, such as a source splitter 106, a gain determiner 114, a first gain module 120, a second gain module 124, etc. Optionally, the audio processing component 840 may include those referenced above. Figures 1 to 7 Other components described. Integrated circuit 802 enables audio processing based on a target SNR as a component in a system, such as a wearable device including a microphone, etc. Figure 9 The headset described in the article, such as Figure 10 The virtual reality, mixed reality, or augmented reality headsets described in the text, such as Figure 11 The augmented reality headset glasses depicted in the article, such as Figure 12 The wearable devices described in the article, such as Figure 13 The earplugs or other wearable device depicted in the image.

[0086] Figure 9 A specific implementation 900 of a headset device 902, in which device 100 includes operable to perform audio processing based on a target signal-to-noise ratio, is depicted. The headset device 902 includes a microphone 102 and a speaker 130. Figure 9 In the example illustrated, microphone 102A is positioned to primarily detect speech from a person wearing the headset 902, and microphone 102B is positioned to detect ambient sounds, such as speech or other sounds from another person. Components of processor 190 (including audio processing component 840) are integrated into the headset 902 and are depicted using dashed lines to indicate components that are generally not visible to the user of the headset 902.

[0087] In a specific example of operation, microphone 102B can detect ambient sounds around headset device 902 and generate audio data representing the sounds. The audio data can be provided to audio processing component 840, which can process the audio data. For example, source separator 106 in audio processing component 840 can (e.g., in the time domain) process the audio data to identify a first component of the audio data corresponding to a target sound and a second component of the audio data corresponding to a non-target sound. In this example, gain determiner 114 in the audio component can determine a first gain to be applied to the first component of the audio data and a second gain to be applied to the second component of the audio data based on a specified target SNR. First gain module 120 in audio processing component 840 can apply the first gain to the first audio component to generate gain-adjusted target audio data, and second gain module 124 in audio processing component 840 can apply the second gain to the second audio component to generate gain-adjusted non-target audio data. Audio processing component 840 can mix the gain-adjusted target audio data and the gain-adjusted non-target audio data to generate an output audio signal 132 such that the output audio signal 132 has a target SNR. In some implementations, the source separator 106 can identify the first and second components of the audio data based on frequency domain processing of the audio data.

[0088] Figure 10A specific implementation 1000 is depicted in which device 100 includes a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset 1002. The headset 1002 includes a microphone 102 and a speaker 130. Additionally, components of processor 190 (including an audio processing component 840) are integrated into the headset 1002. In a particular example of operation, microphone 102 may detect sounds in the environment surrounding headset 1002 and generate audio data representing the sounds. The audio data may be provided to audio processing component 840, which may process the audio data. For example, source separator 106 in audio processing component 840 may (e.g., in the time domain) process the audio data to identify a first component of the audio data corresponding to a target sound and a second component of the audio data corresponding to a non-target sound. In this example, gain determiner 114 in the audio component may determine a first gain to be applied to the first component of the audio data and a second gain to be applied to the second component of the audio data based on a specified target SNR. The first gain module 120 in the audio processing component 840 can apply a first gain to a first audio component to generate gain-adjusted target audio data, and the second gain module 124 in the audio processing component 840 can apply a second gain to a second audio component to generate gain-adjusted non-target audio data. The audio processing component 840 can mix the gain-adjusted target audio data and the gain-adjusted non-target audio data to generate an output audio signal 132, such that the output audio signal 132 has a target SNR. In some specific embodiments, the source separator 106 can identify the first and second components of the audio data based on frequency domain processing of the audio data.

[0089] Figure 11 A specific embodiment 1100 of the device 100 is depicted, comprising a portable electronic device corresponding to augmented reality glasses or mixed reality glasses 1102. Glasses 1102 include a holographic projection unit 1104 configured to project visual data onto the surface of a lens 1106, or to reflect the visual data from the surface of the lens 1106 onto the wearer's retina. Glasses 1102 also include a microphone 102, a speaker 130, and a processor 190, which includes an audio processing component 840.

[0090] In a specific example of operation, microphone 102 can detect sounds in the environment surrounding glasses 1102 and generate audio data representing the sounds. The audio data can be provided to audio processing component 840, which can process the audio data. For example, source separator 106 in audio processing component 840 can (e.g., in the time domain) process the audio data to identify a first component of the audio data corresponding to a target sound and a second component of the audio data corresponding to a non-target sound. In this example, gain determiner 114 in the audio component can determine a first gain to be applied to the first component of the audio data and a second gain to be applied to the second component of the audio data based on a specified target SNR. First gain module 120 in audio processing component 840 can apply the first gain to the first audio component to generate gain-adjusted target audio data, and second gain module 124 in audio processing component 840 can apply the second gain to the second audio component to generate gain-adjusted non-target audio data. Audio processing component 840 can mix the gain-adjusted target audio data and the gain-adjusted non-target audio data to generate an output audio signal 132 such that the output audio signal 132 has a target SNR. In some implementations, the source separator 106 can identify the first and second components of the audio data based on frequency domain processing of the audio data.

[0091] In some implementations, the holographic projection unit 1104 is configured to display information related to the sound detected by the microphone 102. For example, the holographic projection unit 1104 may display a notification indicating that speech has been detected. In another example, the holographic projection unit 1104 may display a notification indicating a detected audio event. For example, the notification may be superimposed on the user's field of view at a specific location that coincides with the location of the source of the sound associated with the audio event.

[0092] Figure 12 The present disclosure describes several examples of a wearable device 1200 operable to perform low-latency noise suppression. In Figure 12 In the example illustrated, the wearable device is a hearing aid device 1202. Hearing aid device 1202 includes a microphone 102, a speaker 130, and a processor 190, which includes an audio processing component 840. Figure 12In the example illustrated, hearing aid device 1202 includes a portion 1204 configured to be worn behind a user's ear, a portion 1208 configured to extend above the ear, and a portion 1206 worn in or near the user's ear canal. In other examples, hearing aid device 1202 has different configurations or form factors. For illustration, hearing aid device 1202 may be an in-ear device that does not include the portion 1204 configured to be worn behind the ear and the portion 1208 configured to extend above the ear.

[0093] In a specific example of the operation of hearing aid device 1202, microphone 102 can detect sounds in the environment surrounding hearing aid device 1202 and generate audio data representing the sounds. The audio data can be provided to audio processing component 840, which can process the audio data. For example, source separator 106 in audio processing component 840 can (e.g., in the time domain) process the audio data to identify a first component of the audio data corresponding to a target sound and a second component of the audio data corresponding to a non-target sound. In this example, gain determiner 114 in the audio component can determine a first gain to be applied to the first component of the audio data and a second gain to be applied to the second component of the audio data based on a specified target SNR. First gain module 120 in audio processing component 840 can apply the first gain to the first audio component to generate gain-adjusted target audio data, and second gain module 124 in audio processing component 840 can apply the second gain to the second audio component to generate gain-adjusted non-target audio data. The audio processing component 840 can mix the gain-adjusted target audio data and the gain-adjusted non-target audio data to generate an output audio signal 132, such that the output audio signal 132 has a target SNR. In some specific implementations, the source separator 106 can identify the first and second components of the audio data based on frequency domain processing of the audio data.

[0094] Figure 13 A specific embodiment 1300 of a portable electronic device is depicted in which device 100 includes one or more earbuds 1306 (e.g., a first earbud 1302, a second earbud 1304, or both). Although earbud 1306 is described, it should be understood that the technology disclosed herein can be applied to other in-ear or over-ear audio devices.

[0095] exist Figure 13In the example illustrated, the first earpiece 1302 includes: a first microphone 1310A, such as a high signal-to-noise ratio microphone positioned to capture the speech of the wearer of the first earpiece 1302; one or more other microphones, illustrated as microphone 1312A, configured to detect ambient sound and spatially distributed to support beamforming; an "inner" microphone 1314A located near the wearer's ear canal (e.g., to assist active noise cancellation); and a self-speech microphone 1316A, such as a bone conduction microphone configured to convert sound vibrations from the wearer's ear bones or skull into audio signals. In a particular embodiment, microphone 1312A and Figures 1 to 7 The microphone 102 of any of them corresponds to the microphone.

[0096] The second earpiece 1304 may be configured in a manner substantially similar to that of the first earpiece 1302. For example, the second earpiece may include: a microphone 1310B positioned to capture the voice of the wearer of the second earpiece 1304; one or more other microphones 1312B configured to detect ambient sound and spatially distributed to support beamforming; an "internal" microphone 1314B; and its own speech microphone 1316B.

[0097] In some embodiments, earbuds 1302 and 1304 are configured to automatically switch between various operating modes, such as a passthrough mode in which ambient sounds are processed by audio processing component 840 for output via speaker 130, and a playback mode in which non-ambient sounds (e.g., streaming audio corresponding to telephone conversations, media playback, video games, etc.) are played back via speaker 130. In other embodiments, earbuds 1302 and 1304 may support fewer modes, or may support one or more other modes in place of the described modes, or may support one or more other modes in addition to the described modes.

[0098] In an exemplary example of operation in transparency mode, one or more of the microphones 102 (e.g., microphones 1312A, 1312B) can detect sounds in the environment surrounding earpieces 1302, 1304 and generate audio data representing the sounds. The audio data can be provided to one or both of the audio processing components 840A, 840B capable of processing the audio data. For example, the source separator 106 in audio processing component 840 can (e.g., in the time domain) process the audio data to identify a first component of the audio data corresponding to a target sound and a second component of the audio data corresponding to a non-target sound. In this example, the gain determiner 114 in the audio component can determine a first gain to be applied to the first component of the audio data and a second gain to be applied to the second component of the audio data based on a specified target SNR. The first gain module 120 in audio processing component 840 can apply the first gain to the first audio component to generate gain-adjusted target audio data, and the second gain module 124 in audio processing component 840 can apply the second gain to the second audio component to generate gain-adjusted non-target audio data. The audio processing component 840 can mix the gain-adjusted target audio data and the gain-adjusted non-target audio data to generate an output audio signal 132, such that the output audio signal 132 has a target SNR. In some specific implementations, the source separator 106 can identify the first and second components of the audio data based on frequency domain processing of the audio data.

[0099] Figure 14 Based on some examples of this disclosure, it is possible to... Figure 1 A diagram illustrating a specific implementation of method 1400, an audio processing method based on a target signal-to-noise ratio, performed by a device. In a particular aspect, one or more operations of method 1400 are defined by a reference... Figures 1 to 13 At least one of the device 100, processor 190, or audio processing component 840, described in various ways, or a combination thereof, shall be executed.

[0100] Method 1400 includes: at box 1402, obtaining data for a specified target signal-to-noise ratio (SNR) based on a person's hearing condition. For example, refer to... Figure 1 The target SNR can be indicated by setting 116. Setting 116 can be specified by a user via the user interface of the wearable device (e.g., device 100) or via the user interface of another device (e.g., second device 150) that transmits setting 116 to the wearable device (e.g., in user-specified settings).

[0101] Method 1400 includes: at block 1404, obtaining audio data representing one or more audio signals. For example, refer to... Figure 1Audio data 104 represents the audio signal received from microphone 102. In this example, microphone 102 generates the audio signal based on sound 170 present in the environment surrounding device 100.

[0102] Method 1400 includes: at block 1406, determining a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on a target signal-to-noise ratio. For example, Figure 1 The gain determiner 114 can be based on settings 116 and optionally other information (such as signal 112 from source separator 106, Figure 2 The feedback signal 278 (or both) is used to determine the first gain and the second gain. In some specific implementations, the first component includes first audio data representing speech, and the second component includes second audio data representing non-speech sounds.

[0103] The method may also include performing time-domain source separation to generate a first signal representing the first component and a second signal representing the second component. In some specific implementations, as referenced... Figure 7 As described, method 1400 may include providing input data based on at least one of one or more audio signals to one or more neural networks (e.g., Figure 7 The speech extraction neural network 702 and noise extraction neural network 704 are trained to generate a first signal representing a first component and a second signal representing a second component based on input data. In some such embodiments, the neural networks are low-latency networks configured to operate on time-domain audio data (e.g., audio data 104).

[0104] In other specific implementations, performing time-domain source separation can use frequency-domain processing to determine the time-domain filter coefficients to be applied to the audio data 104 to determine a first signal representing a first component and a second signal representing a second component. For example, the first signal can be generated by applying the first time-domain filter coefficients to at least one of one or more audio signals, and the second signal can be generated by applying the second time-domain filter coefficients to at least one of one or more audio signals. In this example, based on at least one of the one or more audio signals, one or more frequency-domain operations can be performed to determine the first and second time-domain filter coefficients.

[0105] For illustration, performing one or more frequency domain operations involves performing one or more transform operations based on one or more samples of audio data to generate frequency domain audio data. In different implementations, frequency domain audio data can be processed in various ways. In some implementations, it can be (e.g., by...) Figure 4The beamformer 402 performs beamforming operations based on frequency domain audio data to distinguish a first frequency range associated with target audio in one or more samples of the audio data and a second frequency range associated with non-target audio in one or more samples of the audio data. In other specific implementations, it may be (e.g., by...) Figure 5 The blind source separation module 502 performs blind source separation based on frequency domain audio data to distinguish a first frequency range associated with target audio in one or more samples of the audio data and a second frequency range associated with non-target audio in one or more samples of the audio data. In other specific implementations, the frequency domain audio data may be input into one or more machine learning models (e.g., Figure 6 Machine learning model 602), which is trained to distinguish a first frequency range associated with target audio in one or more samples of audio data and a second frequency range associated with non-target audio in one or more samples of audio data. In other specific embodiments, other frequency domain source separators (e.g., Figure 2 The frequency domain source separator 204 can perform frequency domain source separation operation based on frequency domain audio data to distinguish a first frequency range associated with target audio in one or more samples of the audio data and a second frequency range associated with non-target audio in one or more samples of the audio data.

[0106] In a specific implementation of processing frequency-domain audio data to distinguish or generate a first frequency range associated with target audio and a second frequency range associated with non-target audio, the method may include determining first time-domain filter coefficients based on the first frequency range and determining second time-domain filter coefficients based on the second frequency range. For example, Figure 2 The filter designer 210 can use one or more inverse domain transform operations and parameterization operations to generate first time-domain filter coefficients based on a first frequency range associated with the target audio, and Figure 2 The filter designer 212 can use the same or different inverse domain transform operations and parameterization operations to generate second time-domain filter coefficients based on a second frequency range associated with the non-target audio.

[0107] exist Figure 14In this method 1400, at block 1408, a first gain is applied to a first component of audio data to generate a target signal; and at block 1410, a second gain is applied to a second component of the audio data to generate a noise signal. For example, a first gain module 120 applies a first gain to a first signal 108, which includes a first component of audio data 104 (e.g., a component corresponding to the target audio). In this example, a second gain module 124 applies a second gain to a second signal 110, which includes a second component of audio data 104 (e.g., a component corresponding to a non-target audio).

[0108] Method 1400 includes: at block 1412, combining the target signal and the noise signal to generate an output audio signal. For example, Figure 1 The mixer 128 combines the target signal (e.g., the first gain adjustment signal 122) and the non-target signal (e.g., the second gain adjustment signal 126) to generate the output audio signal 132.

[0109] Optionally, in some implementations, method 1400 may also include other audio processing operations. For example, in some implementations, method 1400 may include performing feedback adaptive noise cancellation based on a feedback signal, performing feedforward adaptive noise cancellation based on audio data, or both, to generate an output audio signal.

[0110] Therefore, method 1400 enables the generation of an output audio signal based on a target SNR that can be specified by the user. In a particular implementation, the delay between obtaining specific audio data and generating a corresponding output audio signal representing the specific audio data is less than one millisecond. This low-latency processing is desirable for audio processing in wearable devices, as high latency can lead to a poor user experience. In some implementations, method 1400 uses certain relatively high-latency processes (such as frequency domain source separation) to improve audio source separation; however, such high-latency processes are applied indirectly. For example, frequency domain audio processing is used to generate time-domain filter coefficients in a manner that does not delay the processing of audio data for output to the user. To avoid delaying the processing of audio data for output to the user, frequency domain audio processing is performed in parallel with time-domain processing to generate output audio. In this arrangement, the time-domain filter coefficients generated via frequency domain processing are delayed relative to the audio data processed in the time domain.

[0111] Figure 14 Method 1400 can be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 14Method 1400 can be executed by one or more processors (such as reference processors) that execute instructions. Figure 15 (As described) Execution.

[0112] refer to Figure 15 The diagram depicts a block diagram of a specific, exemplary embodiment of the device, and generally designates the device as 1500. In various embodiments, device 1500 may have... Figure 15 The illustrated component may include more or fewer components. In an illustrative embodiment, device 1500 may correspond to device 100. In an illustrative embodiment, device 1500 may perform the reference... Figures 1 to 14 One or more operations as described.

[0113] In a particular implementation, device 1500 includes a processor 1506 (e.g., a central processing unit (CPU)). Device 1500 may include one or more additional processors 1510 (e.g., one or more DSPs). In a particular aspect, Figure 1 Processor 190 corresponds to processor 1506, processor 1510, or a combination thereof. Processor 1510 may include a speech and music decoder-decoder (codec) 1508, which includes a speech decoder (“vocoder”) encoder 1536 and a vocoder decoder 1538. Processor 1510 and / or speech and music codec 1508 include audio processing components 840, such as source splitter 106, gain determiner 114, first gain module 120, and second gain module 124. Optionally, audio processing component 840 may include components as referenced above. Figures 1 to 7 Other components as described.

[0114] exist Figure 15 In this device 1500, memory 1586 and codec 1534 are included. Memory 1586 includes (e.g., storage) functions that can be executed by one or more additional processors 1510 (or processor 1506) to implement a reference. Figure 1 The device 100 contains instructions 1556 describing its functions. The memory 1586 may also store settings 116 indicating the target SNR of the output audio signal to be generated based on the received audio data. Figure 15 In this device 1500, the device also includes a modem 144 coupled to an antenna 1552 via a transceiver 1550. The modem 144, transceiver 1550, and antenna 1552 enable the device 1500 to exchange data with one or more other devices wirelessly. For example, in some implementations, the device 1500 may be based on communication with another device (e.g., Figure 1The second device 150 receives data wirelessly (such as setting 116 or another indication of the target SNR), and generates an audio output at the speaker 130.

[0115] Device 1500 may include a display 1528 coupled to a display controller 1526. A speaker 130 and a microphone 102 may be coupled to a codec 1534. Figure 15 In this implementation, codec 1534 includes a digital-to-analog converter (DAC) 1502 and an analog-to-digital converter (ADC) 1504. In a particular embodiment, codec 1534 may receive analog signals from microphone 102 and use ADC 1504 to convert the analog signals into digital signals (e.g., ...). Figures 1 to 7 The audio data 104 is processed and the digital signal is provided to the speech and music codec 1508. The speech and music codec 1508 can process the digital signal. The digital signal can be further processed by the audio processing component 840. For example, the source separator 106 in the audio processing component 840 can (e.g., in the time domain) process the audio data to identify a first component of the audio data corresponding to a target sound and a second component of the audio data corresponding to a non-target sound. In this example, the gain determiner 114 in the audio component can determine a first gain to be applied to the first component of the audio data and a second gain to be applied to the second component of the audio data based on a specified target SNR. The first gain module 120 in the audio processing component 840 can apply the first gain to the first audio component to generate gain-adjusted target audio data, and the second gain module 124 in the audio processing component 840 can apply the second gain to the second audio component to generate gain-adjusted non-target audio data. The audio processing component 840 can mix the gain-adjusted target audio data and the gain-adjusted non-target audio data to generate an output audio signal 132 such that the output audio signal 132 has a target SNR. In some implementations, the source separator 106 can identify the first and second components of the audio data based on frequency domain processing of the audio data.

[0116] In a particular implementation, speech and music codec 1508 may provide a digital signal representing the output audio signal and / or other audio content generated by audio processing component 840 to codec 1534. Codec 1534 may use digital-to-analog converter 1502 to convert the digital signal into an analog signal and may provide the analog signal to speaker 130.

[0117] In a particular embodiment, device 1500 may be included in a system-in-package (SoC) or system-on-a-chip (SoC) 1522. In a particular embodiment, memory 1586, processor 1506, processor 1510, display controller 1526, codec 1534, and modem 144 are included in the SoC or SoC 1522. In a particular embodiment, input device 1530 and power source 1544 are coupled to the SoC or SoC 1522. Furthermore, in a particular embodiment, such as... Figure 15 As illustrated, display 1528, input device 1530, speaker 130, microphone 102, antenna 1552, and power source 1544 are external to system-in-package or system-on-chip device 1522. In a particular implementation, each of display 1528, input device 1530, speaker 130, microphone 102, antenna 1552, and power source 1544 may be coupled to components of system-in-package or system-on-chip device 1522, such as interfaces or controllers.

[0118] Device 1500 may include wearable devices such as wearable mobile communication devices, wearable personal digital assistants, wearable display devices, wearable gaming systems, wearable music players, wearable radio components, wearable cameras, wearable navigation devices, headsets, augmented reality headsets, mixed reality headsets, virtual reality headsets, voice-activated devices, portable electronic devices, wearable computing devices, wearable communication devices, virtual reality (VR) devices, one or more earbuds, hearing aid devices, or any combination thereof.

[0119] In conjunction with the described specific embodiments, an apparatus includes components for obtaining data for a specified target signal-to-noise ratio based on a person's hearing condition. For example, the components for obtaining data for the specified target signal-to-noise ratio may correspond to device 100, second device 150, processor 190, gain determiner 114, modem 144, memory 162, input device 1530, transceiver 1550, processor 1506, processor 1510, one or more other circuits or components configured to obtain data for a specified target signal-to-noise ratio based on a person's hearing condition, or any combination thereof.

[0120] The device also includes components for acquiring audio data representing one or more audio signals. For example, the components for acquiring audio data representing one or more audio signals may correspond to device 100, microphone 102, processor 190, source splitter 106, audio input 804, codec 1534, processor 1506, processor 1510, one or more other circuits or components configured to acquire audio data representing one or more audio signals, or any combination thereof.

[0121] The apparatus also includes components for determining a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on a target signal-to-noise ratio. For example, the components for determining the first and second gains may correspond to device 100, gain determiner 114, processor 190, processor 1506, processor 1510, one or more other circuits or components, or any combination thereof, configured to determine the first gain to be applied to the first component of the audio data and the second gain to be applied to the second component of the audio data based on a target signal-to-noise ratio.

[0122] The apparatus also includes components for applying a first gain to a first component of the audio data to generate a target signal. For example, the components for applying a first gain to a first component of the audio data to generate a target signal may correspond to device 100, first gain module 120, processor 190, speech gain module 310, speech gain module 720, processor 1506, processor 1510, one or more other circuits or components configured to apply a first gain to a first component of the audio data to generate a target signal, or any combination thereof.

[0123] The device also includes components for applying a second gain to a second component of the audio data to generate a noise signal. For example, the components for applying a second gain to a second component of the audio data to generate a noise signal may correspond to device 100, second gain module 120, processor 190, noise gain module 312, noise gain module 724, processor 1506, processor 1510, one or more other circuits or components configured to apply a second gain to a second component of the audio data to generate a noise signal, or any combination thereof.

[0124] The apparatus also includes components for combining the target signal and the noise signal to generate an output audio signal. For example, the components for combining the target signal and the noise signal to generate an output audio signal may correspond to device 100, mixer 128, processor 190, processor 1506, processor 1510, one or more other circuits or components configured to combine the target signal and the noise signal to generate an output audio signal, or any combination thereof.

[0125] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 1586) includes instructions (e.g., instructions 1556) that, when executed by one or more processors (e.g., one or more processors 1510 or 1506), cause the one or more processors to: obtain data with a specified target signal-to-noise ratio based on a person's hearing condition; obtain audio data representing one or more audio signals; determine a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on the target signal-to-noise ratio; apply the first gain to the first component of the audio data to generate a target signal; apply the second gain to the second component of the audio data to generate a noise signal; and combine the target signal and the noise signal to generate an output audio signal.

[0126] Specific aspects of this disclosure are described below in a collection of related embodiments:

[0127] According to Embodiment 1, a device includes one or more processors configured to: acquire data with a specified target signal-to-noise ratio based on a person's hearing condition; acquire audio data representing one or more audio signals; determine a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on the target signal-to-noise ratio; apply the first gain to the first component of the audio data to generate a target signal; apply the second gain to the second component of the audio data to generate a noise signal; and combine the target signal and the noise signal to generate an output audio signal.

[0128] Example 2 includes the device according to Example 1, wherein the data specifying the target signal-to-noise ratio is obtained from user-specified settings.

[0129] Example 3 includes the device according to Example 1 or Example 2, wherein the first component includes first audio data representing speech, and the second component includes second audio data representing non-speech sounds.

[0130] Example 4 includes the device according to any one of Examples 1 to 3, wherein the delay between obtaining specific audio data and generating a corresponding output audio signal representing the specific audio data is less than one millisecond.

[0131] Example 5 includes the device according to any one of Examples 1 to 4, and further includes one or more microphones coupled to the one or more processors, wherein the audio data represents sound captured by the one or more microphones.

[0132] Example 6 includes a device according to any one of Examples 1 to 5, wherein one or more processors are integrated into the wearable device.

[0133] Example 7 includes the device according to any one of Examples 1 to 6, and further includes a modem coupled to the one or more processors and configured to receive the data with a specified target signal-to-noise ratio from the second device.

[0134] Example 8 includes the device according to any one of Examples 1 to 7, and further includes one or more speakers and one or more microphones coupled to the one or more processors, wherein the one or more microphones include at least one microphone configured to generate the audio data and at least one microphone configured to generate a feedback signal based on sound generated by the one or more microphones in response to the output audio signal.

[0135] Example 9 includes a device according to any one of Examples 1 to 8, wherein one or more processors are configured to determine the first gain and the second gain based at least in part on a feedback signal.

[0136] Example 10 includes a device according to any one of Examples 1 to 9, wherein one or more processors are configured to perform feedback adaptive noise cancellation based on a feedback signal to generate the output audio signal.

[0137] Example 11 includes a device according to any one of Examples 1 to 10, wherein one or more processors are configured to perform feedforward adaptive noise cancellation based on the audio data to generate the output audio signal.

[0138] Example 12 includes a device according to any one of Examples 1 to 11, wherein one or more processors are configured to perform a time-domain source separation operation to generate a first signal representing the first component and a second signal representing the second component, wherein a first gain is applied to the first signal and a second gain is applied to the second signal.

[0139] Example 13 includes the device according to Example 12, wherein performing the time-domain source separation operation includes applying first time-domain filter coefficients to at least one of the one or more audio signals to generate a first signal representing the first component, and applying second time-domain filter coefficients to at least one of the one or more audio signals to generate a second signal representing the second component.

[0140] Example 14 includes the device according to Example 13, wherein the one or more processors are configured to perform one or more frequency domain operations based on at least one of the one or more audio signals to determine the first time domain filter coefficients and the second time domain filter coefficients.

[0141] Example 15 includes the device according to Example 13 or Example 14, wherein the one or more processors are configured to: perform one or more transformation operations based on one or more samples of the audio data to generate frequency domain audio data; perform beamforming operations based on the frequency domain audio data to distinguish a first frequency range associated with target audio in the one or more samples of the audio data and a second frequency range associated with non-target audio in the one or more samples of the audio data; determine first time domain filter coefficients based on the first frequency range; and determine second time domain filter coefficients based on the second frequency range.

[0142] Example 16 includes the device according to Example 13 or Example 14, wherein the one or more processors are configured to: perform one or more transformation operations based on one or more samples of the audio data to generate frequency domain audio data; perform blind source separation operations based on the frequency domain audio data to distinguish a first frequency range associated with target audio in the one or more samples of the audio data and a second frequency range associated with non-target audio in the one or more samples of the audio data; determine first time domain filter coefficients based on the first frequency range; and determine second time domain filter coefficients based on the second frequency range.

[0143] Example 17 includes the device according to Example 12, wherein performing the time-domain source separation operation includes providing input data based on at least one of the one or more audio signals to one or more neural networks, the one or more neural networks being trained to generate a first signal representing the first component and a second signal representing the second component based on the input data.

[0144] According to embodiment 18, a method includes: obtaining data with a specified target signal-to-noise ratio based on a person's hearing condition at one or more processors; obtaining audio data representing one or more audio signals at the one or more processors; determining, at the one or more processors, a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on the target signal-to-noise ratio; applying the first gain to the first component of the audio data to generate a target signal; applying the second gain to the second component of the audio data to generate a noise signal; and combining the target signal and the noise signal to generate an output audio signal.

[0145] Example 19 includes the method according to Example 18, wherein the data specifying the target signal-to-noise ratio is obtained from user-specified settings.

[0146] Example 20 includes the method according to Example 18 or Example 19, wherein the first component includes first audio data representing speech, and the second component includes second audio data representing non-speech sounds.

[0147] Example 21 includes the method according to any one of Examples 18 to 20, wherein the delay between obtaining specific audio data and generating a corresponding output audio signal representing the specific audio data is less than one millisecond.

[0148] Example 22 includes the method according to any one of Examples 18 to 21, wherein the audio data represents sound captured by one or more microphones coupled to the one or more processors.

[0149] Example 23 includes the method according to any one of Examples 18 to 22, wherein one or more processors are integrated into a wearable device.

[0150] Example 24 includes the method according to any one of Examples 18 to 23, and further includes performing feedforward adaptive noise cancellation based on the audio data.

[0151] Example 25 includes the method according to any one of Examples 18 to 24, and further includes performing feedback adaptive noise cancellation based on the feedback signal to generate the output audio signal.

[0152] Example 26 includes the method according to any one of Examples 18 to 25, and further includes determining the first gain and the second gain based at least in part on the feedback signal.

[0153] Example 27 includes the method according to any one of Examples 18 to 26, and further includes performing time-domain source separation to generate a first signal representing the first component and a second signal representing the second component, wherein a first gain is applied to the first signal and a second gain is applied to the second signal.

[0154] Example 28 includes the method according to Example 27, wherein performing time-domain source separation includes: applying first time-domain filter coefficients to at least one of the one or more audio signals to generate a first signal representing the first component; and applying second time-domain filter coefficients to at least one of the one or more audio signals to generate a second signal representing the second component.

[0155] Example 29 includes the method according to Example 28, and further includes performing one or more frequency domain operations based on at least one of the one or more audio signals to determine the first time domain filter coefficients and the second time domain filter coefficients.

[0156] Example 30 includes the method according to Example 28 or Example 29, and further includes: performing one or more transformation operations based on one or more samples of the audio data to generate frequency domain audio data; performing beamforming operations based on the frequency domain audio data to distinguish a first frequency range associated with target audio in the one or more samples of the audio data and a second frequency range associated with non-target audio in the one or more samples of the audio data; determining first time domain filter coefficients based on the first frequency range; and determining second time domain filter coefficients based on the second frequency range.

[0157] Example 31 includes the method according to Example 28 or Example 29, and further includes: performing one or more transformation operations based on one or more samples of the audio data to generate frequency domain audio data; performing blind source separation operations based on the frequency domain audio data to distinguish a first frequency range associated with target audio in the one or more samples of the audio data and a second frequency range associated with non-target audio in the one or more samples of the audio data; determining first time domain filter coefficients based on the first frequency range; and determining second time domain filter coefficients based on the second frequency range.

[0158] Example 32 includes the method according to Example 27, wherein performing the temporal source separation includes providing input data based on at least one of the one or more audio signals to one or more neural networks, the one or more neural networks being trained to generate a first signal representing the first component and a second signal representing the second component based on the input data.

[0159] According to embodiment 33, an apparatus includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform a method according to any one of embodiments 18 to 32.

[0160] According to Embodiment 34, a non-transitory computer-readable medium storage instruction, when executed by a processor, causes the processor to perform the method according to any one of Embodiments 18 to 32.

[0161] According to embodiment 35, an apparatus includes components for performing the method according to any one of embodiments 18 to 32.

[0162] According to Embodiment 36, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to: obtain data with a specified target signal-to-noise ratio based on a person's hearing condition; obtain audio data representing one or more audio signals; determine a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on the target signal-to-noise ratio; apply the first gain to the first component of the audio data to generate a target signal; apply the second gain to the second component of the audio data to generate a noise signal; and combine the target signal and the noise signal to generate an output audio signal.

[0163] Example 37 includes a non-transitory computer-readable medium according to Example 36, wherein the data specifying the target signal-to-noise ratio is obtained from user-specified settings.

[0164] Example 38 includes a non-transitory computer-readable medium according to Example 36 or Example 37, wherein the first component includes first audio data representing speech, and the second component includes second audio data representing non-speech sounds.

[0165] Example 39 includes a non-transitory computer-readable medium according to any one of Examples 36 to 38, wherein the delay between obtaining specific audio data and generating a corresponding output audio signal representing the specific audio data is less than one millisecond.

[0166] Example 40 includes a non-transitory computer-readable medium according to any one of Examples 36 to 39, wherein the audio data represents sound captured by one or more microphones coupled to the one or more processors.

[0167] Example 41 includes a non-transitory computer-readable medium according to any one of Examples 36 to 40, wherein the instructions cause the one or more processors to perform feedforward adaptive noise cancellation based on the audio data to generate the output audio signal.

[0168] Example 42 includes a non-transitory computer-readable medium according to any one of Examples 36 to 41, wherein the instructions cause the one or more processors to perform feedback adaptive noise cancellation based on a feedback signal to generate the output audio signal.

[0169] Example 43 includes a non-transitory computer-readable medium according to any one of Examples 36 to 42, wherein the instructions cause the one or more processors to determine the first gain and the second gain at least in part based on feedback signals.

[0170] Example 44 includes a non-transitory computer-readable medium according to any one of Examples 36 to 43, wherein the instructions cause the one or more processors to perform a time-domain source separation operation to generate a first signal representing the first component and a second signal representing the second component, wherein a first gain is applied to the first signal and a second gain is applied to the second signal.

[0171] Example 45 includes a non-transitory computer-readable medium according to Example 44, wherein performing the time-domain source separation operation includes applying first time-domain filter coefficients to at least one of the one or more audio signals to generate a first signal representing the first component, and applying second time-domain filter coefficients to at least one of the one or more audio signals to generate a second signal representing the second component.

[0172] Example 46 includes a non-transitory computer-readable medium according to Example 45, wherein the instructions cause the one or more processors to perform one or more frequency domain operations based on at least one of the one or more audio signals to determine the first time-domain filter coefficients and the second time-domain filter coefficients.

[0173] Example 47 includes a non-transitory computer-readable medium according to Example 45 or Example 46, wherein the instructions cause the one or more processors to: perform one or more transformation operations based on one or more samples of the audio data to generate frequency domain audio data; perform beamforming operations based on the frequency domain audio data to distinguish a first frequency range associated with target audio in the one or more samples of the audio data and a second frequency range associated with non-target audio in the one or more samples of the audio data; determine first time domain filter coefficients based on the first frequency range; and determine second time domain filter coefficients based on the second frequency range.

[0174] Example 48 includes a non-transitory computer-readable medium according to Example 45 or Example 46, wherein the instructions cause the one or more processors to: perform one or more transformation operations based on one or more samples of the audio data to generate frequency domain audio data; perform blind source separation operations based on the frequency domain audio data to distinguish a first frequency range associated with target audio in the one or more samples of the audio data and a second frequency range associated with non-target audio in the one or more samples of the audio data; determine first time-domain filter coefficients based on the first frequency range; and determine second time-domain filter coefficients based on the second frequency range.

[0175] Example 49 includes a non-transitory computer-readable medium according to Example 44, wherein performing the time-domain source separation operation includes providing input data based on at least one of the one or more audio signals to one or more neural networks, the one or more neural networks being trained to generate a first signal representing the first component and a second signal representing the second component based on the input data.

[0176] According to embodiment 50, an apparatus includes: components for obtaining data with a specified target signal-to-noise ratio based on a person's hearing condition; components for obtaining audio data representing one or more audio signals; components for determining a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data based on the target signal-to-noise ratio; components for applying the first gain to the first component of the audio data to generate a target signal; components for applying the second gain to the second component of the audio data to generate a noise signal; and components for combining the target signal and the noise signal to generate an output audio signal.

[0177] Example 51 includes the apparatus according to Example 50, wherein the data specifying the target signal-to-noise ratio is obtained from user-specified settings.

[0178] Example 52 includes the apparatus according to Example 50 or Example 51, wherein the first component includes first audio data representing speech, and the second component includes second audio data representing non-speech sounds.

[0179] Example 53 includes the apparatus according to any one of Examples 50 to 52, wherein the delay between obtaining specific audio data and generating a corresponding output audio signal representing the specific audio data is less than one millisecond.

[0180] Example 54 includes the apparatus according to any one of Examples 50 to 53, and further includes components for generating the audio data based on ambient sound.

[0181] Example 55 includes an apparatus according to any one of Examples 50 to 54, wherein the components for obtaining data specifying the target signal-to-noise ratio, the components for obtaining the audio data, the components for determining the first gain and the second gain, the components for applying the first gain to the first component of the audio data, the components for applying the second gain to the second component of the audio data, and the components for combining are integrated into a wearable device.

[0182] Example 56 includes the apparatus according to any one of Examples 50 to 55, and further includes: a component for generating sound based on the output audio signal; and a component for generating a feedback signal based on the sound.

[0183] Example 57 includes the apparatus according to any one of Examples 50 to 56, and further includes components for performing feedback adaptive noise cancellation based on the feedback signal to generate the output audio signal.

[0184] Example 58 includes the apparatus according to any one of Examples 50 to 57, wherein the first gain and the second gain are at least partially based on the feedback signal.

[0185] Example 59 includes the apparatus according to any one of Examples 50 to 58, and further includes components for performing feedforward adaptive noise cancellation based on the audio data.

[0186] Example 60 includes the apparatus according to any one of Examples 50 to 59, and further includes components for performing time-domain source separation to generate a first signal representing the first component and a second signal representing the second component, wherein the first gain is applied to the first signal and the second gain is applied to the second signal.

[0187] Example 61 includes the apparatus according to Example 60, wherein performing the time-domain source separation includes applying first time-domain filter coefficients to at least one of the one or more audio signals to generate a first signal representing the first component, and applying second time-domain filter coefficients to at least one of the one or more audio signals to generate a second signal representing the second component.

[0188] Example 62 includes the apparatus according to Example 61, and further includes components for performing one or more frequency domain operations based on at least one of the one or more audio signals to determine the first time domain filter coefficients and the second time domain filter coefficients.

[0189] Example 63 includes the apparatus according to Example 61 or Example 62, and further includes: means for performing one or more transformation operations based on one or more samples of the audio data to generate frequency domain audio data; means for performing beamforming operations based on the frequency domain audio data to distinguish a first frequency range associated with target audio in the one or more samples of the audio data and a second frequency range associated with non-target audio in the one or more samples of the audio data; means for determining first time domain filter coefficients based on the first frequency range; and means for determining second time domain filter coefficients based on the second frequency range.

[0190] Example 64 includes the apparatus according to Example 61 or Example 62, and further includes: means for performing one or more transformation operations based on one or more samples of the audio data to generate frequency domain audio data; means for performing a blind source separation operation based on the frequency domain audio data to distinguish a first frequency range associated with target audio in the one or more samples of the audio data and a second frequency range associated with non-target audio in the one or more samples of the audio data; means for determining first time domain filter coefficients based on the first frequency range; and means for determining second time domain filter coefficients based on the second frequency range.

[0191] Example 65 includes the apparatus according to Example 60, wherein performing the temporal source separation includes providing input data based on at least one of the one or more audio signals to one or more neural networks, the one or more neural networks being trained to generate a first signal representing the first component and a second signal representing the second component based on the input data.

[0192] Those skilled in the art will also understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithm steps described in connection with the specific embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various exemplary components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or as processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, and such implementation decisions shall not be construed as departing from the scope of this disclosure.

[0193] The steps of the methods or algorithms described in conjunction with the specific embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compressed optical disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and storage medium may reside as discrete components in a computing device or a user terminal.

[0194] The prior description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but should be granted the broadest scope that may be consistent with the principles and novel features as defined by the following claims.

Claims

1. An apparatus, the apparatus comprising: One or more processors, said one or more processors being configured to: Data for a specified target signal-to-noise ratio is obtained based on a person's hearing condition; To obtain audio data representing one or more audio signals; Based on the target signal-to-noise ratio, determine the first gain to be applied to the first component of the audio data and the second gain to be applied to the second component of the audio data; The first gain is applied to the first component of the audio data to generate the target signal; The second gain is applied to the second component of the audio data to generate a noise signal; as well as The target signal and the noise signal are combined to generate an output audio signal.

2. The device of claim 1, wherein the data specifying the target signal-to-noise ratio is obtained from user-specified settings.

3. The device of claim 1, wherein the first component comprises first audio data representing speech, and the second component comprises second audio data representing non-speech sounds.

4. The device of claim 1, wherein the delay between obtaining specific audio data and generating a corresponding output audio signal representing the specific audio data is less than one millisecond.

5. The device of claim 1, further comprising one or more microphones coupled to the one or more processors, wherein the audio data represents sound captured by the one or more microphones.

6. The device of claim 1, wherein the one or more processors are integrated into the wearable device.

7. The device of claim 1, further comprising a modem coupled to the one or more processors and configured to receive the data specifying the target signal-to-noise ratio from the second device.

8. The device of claim 1, further comprising one or more speakers and one or more microphones coupled to the one or more processors, wherein the one or more microphones include at least one microphone configured to generate the audio data and at least one microphone configured to generate a feedback signal based on sound generated by the one or more microphones in response to the output audio signal.

9. The device of claim 1, wherein the one or more processors are configured to determine the first gain and the second gain based at least in part on a feedback signal.

10. The device of claim 1, wherein the one or more processors are configured to perform feedback adaptive noise cancellation based on a feedback signal, perform feedforward adaptive noise cancellation based on the audio data, or both, to generate the output audio signal.

11. The apparatus of claim 1, wherein the one or more processors are configured to perform a time-domain source separation operation to generate a first signal representing the first component and a second signal representing the second component, wherein a first gain is applied to the first signal and a second gain is applied to the second signal.

12. The apparatus of claim 11, wherein performing the time-domain source separation operation comprises applying first time-domain filter coefficients to at least one of the one or more audio signals to generate the first signal representing the first component, and applying second time-domain filter coefficients to at least one of the one or more audio signals to generate the second signal representing the second component.

13. The device of claim 12, wherein the one or more processors are configured to perform one or more frequency domain operations based on at least one of the one or more audio signals to determine the first time-domain filter coefficients and the second time-domain filter coefficients.

14. The apparatus of claim 11, wherein performing the time-domain source separation operation comprises providing input data based on at least one of the one or more audio signals to one or more neural networks, the one or more neural networks being trained to generate a first signal representing the first component and a second signal representing the second component based on the input data.

15. A method comprising: Data with a specified target signal-to-noise ratio is obtained at one or more processors based on a person's hearing condition; Audio data representing one or more audio signals is obtained at the one or more processors; At one or more processors, a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data are determined based on the target signal-to-noise ratio. The first gain is applied to the first component of the audio data to generate the target signal; The second gain is applied to the second component of the audio data to generate a noise signal; as well as The target signal and the noise signal are combined to generate an output audio signal.

16. The method of claim 15, wherein the data specifying the target signal-to-noise ratio is obtained from user-specified settings.

17. The method of claim 15, wherein the first component comprises first audio data representing speech, and the second component comprises second audio data representing non-speech sounds.

18. The method of claim 15, wherein the delay between obtaining specific audio data and generating a corresponding output audio signal representing the specific audio data is less than one millisecond.

19. The method of claim 15, wherein the audio data represents sound captured by one or more microphones coupled to the one or more processors.

20. The method of claim 15, further comprising performing feedback adaptive noise cancellation based on the feedback signal, performing feedforward adaptive noise cancellation based on the audio data, or both, to generate the output audio signal.

21. The method of claim 15, further comprising performing time-domain source separation to generate a first signal representing the first component and a second signal representing the second component, wherein the first gain is applied to the first signal and the second gain is applied to the second signal.

22. The method of claim 21, wherein performing temporal source separation comprises: The first time-domain filter coefficients are applied to at least one of the one or more audio signals to generate the first signal representing the first component; as well as The second time-domain filter coefficients are applied to at least one of the one or more audio signals to generate the second signal representing the second component.

23. The method of claim 22, further comprising performing one or more frequency domain operations based on at least one of the one or more audio signals to determine the first time-domain filter coefficients and the second time-domain filter coefficients.

24. The method according to claim 22, further comprising: One or more transformation operations are performed based on one or more samples of the audio data to generate frequency domain audio data; Beamforming is performed based on the frequency domain audio data to distinguish a first frequency range associated with target audio in one or more samples of the audio data and a second frequency range associated with non-target audio in one or more samples of the audio data. The coefficients of the first time-domain filter are determined based on the first frequency range; as well as The coefficients of the second time-domain filter are determined based on the second frequency range.

25. The method according to claim 22, further comprising: One or more transformation operations are performed based on one or more samples of the audio data to generate frequency domain audio data; Blind source separation is performed based on the frequency domain audio data to distinguish a first frequency range associated with target audio in one or more samples of the audio data and a second frequency range associated with non-target audio in one or more samples of the audio data; The coefficients of the first time-domain filter are determined based on the first frequency range; as well as The coefficients of the second time-domain filter are determined based on the second frequency range.

26. The method of claim 21, wherein performing the temporal source separation comprises providing input data based on at least one of the one or more audio signals to one or more neural networks, the one or more neural networks being trained to generate a first signal representing the first component and a second signal representing the second component based on the input data.

27. A non-transitory computer-readable medium storing instructions, said instructions causing said one or more processors, when executed, to: Data for a specified target signal-to-noise ratio is obtained based on a person's hearing condition; To obtain audio data representing one or more audio signals; Based on the target signal-to-noise ratio, determine the first gain to be applied to the first component of the audio data and the second gain to be applied to the second component of the audio data; The first gain is applied to the first component of the audio data to generate the target signal; The second gain is applied to the second component of the audio data to generate a noise signal; as well as The target signal and the noise signal are combined to generate an output audio signal.

28. The non-transitory computer-readable medium of claim 27, wherein the first component comprises first audio data representing speech, and the second component comprises second audio data representing non-speech sounds.

29. The non-transitory computer-readable medium of claim 27, wherein the delay between obtaining specific audio data and generating a corresponding output audio signal representing the specific audio data is less than one millisecond.

30. An apparatus comprising: A component used to obtain data for a specified target signal-to-noise ratio based on a person's hearing condition; A component used to obtain audio data representing one or more audio signals; A component for determining, based on the target signal-to-noise ratio, a first gain to be applied to a first component of the audio data and a second gain to be applied to a second component of the audio data; Components for applying the first gain to the first component of the audio data to generate a target signal; Components for applying the second gain to the second component of the audio data to generate a noise signal; and A component for combining the target signal and the noise signal to generate an output audio signal.