Hearable devices and methods for playback of sound originating from sources within a threshold distance

A neural network-based headset system enhances audio perception by creating a 'sound bubble' around the user, selectively amplifying nearby speakers and suppressing distant noise using multiple microphones and advanced audio processing techniques, addressing the limitations of noise-cancelling headphones in discerning speaker distances.

WO2025160411A1PCT designated stage Publication Date: 2025-07-31UNIV OF WASHINGTON
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/012970
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-26
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Noise-cancelling headphones struggle to selectively perceive and program acoustic scenes based on speaker distances, and the human auditory system has limited ability to discern distance in noisy environments, making it difficult to focus on nearby speakers amidst interference and noise.

Method used

A neural network-based headset system uses multiple microphones to process acoustic signals, extracting audio signals from sources within a threshold distance through features like interchannel phase difference, interchannel level difference, head-related transfer functions, and direct-to-reverberant ratio, creating a 'sound bubble' where desired sounds are amplified while suppressing noise outside this threshold.

Benefits of technology

The system effectively enhances the ability to hear nearby speakers by suppressing distant and ambient noise, allowing users to focus on targeted audio sources within a programmable distance, even in challenging environments with varying reverberations and user-specific head-related transfer functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025012970_31072025_PF_FP_ABST
    Figure US2025012970_31072025_PF_FP_ABST
Patent Text Reader

Abstract

An example system includes a plurality of microphones placed proximate a head of a user. The microphones receive acoustic signals from the environment and generate corresponding input audio signals. The example system includes a trained machine learning model. This model receives input audio signals and extracts target audio signals corresponding to sources within a specified threshold distance from the user. These target audio signals are then provided to the speaker for playback to the user. This may create a sound bubble for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No. 0077145-07001 (UW 49996.02WO2) HEARABLE DEVICES AND METHODS FOR PLAYBACK OF SOUND ORIGINATING FROM SOURCES WITHIN A THRESHOLD DISTANCE CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit under 35 U.S.C. § 119 of the earlier filing date of U.S. Provisional Application Serial No.63 / 625,732 filed January 26, 2024, the entire contents of which are hereby incorporated by reference in their entirety for any purpose. TECHNICAL FIELD

[0002] Examples described herein relate generally to hearable systems. Examples of systems for playing sounds originating within a target region are described. BACKGROUND

[0003] Noise-cancelling headphones can suppress sounds around a wearer, but generally cannot perceive distance or selectively program acoustic scenes based on speaker distances. The distance perception of the human auditory system is also limited, and although people may be able to determine an angular direction of a sound source, estimating distance is more challenging. Distance perception becomes even more challenging with unfamiliar sounds in unknown environments. As interference and noise become louder, our ability to focus on nearby speakers becomes more difficult. SUMMARY

[0004] Examples of methods of training a neural network to extract audio signals originating from audio sources within a threshold distance are described herein. An example method may include generating a set of training data, the training data including audio signals generated by multiple microphones responsive to acoustic signals from audio sources in an environment and distance information to the audio sources. The example method may include generating augmented training data by performing operations including shifting the audio signals, adjusting an amplitude of the audio signals, attenuating frequencies of the audio signals, and / or varying a speed of the audio signals. The example method may include providing the set of training data including the augmented training data to a machine learning engine. The example method may include generating a set of parameters representing a trained neural network to infer distance to an audio source.

[0005] In some examples, the set of training data includes data obtained from varying positions of the multiple microphones in the environment to represent multiple HRTFs.Docket No. 0077145-07001 (UW 49996.02WO2)

[0006] In some examples, methods may further include refining the set of parameters based on further training data gathered from human wearers of multiple microphones.

[0007] In some examples, the set of training data includes audio signals generated in multiple environments including different reverberation characteristics.

[0008] In some examples, a system includes a plurality of microphones configured for placement proximate a head of a user, the plurality of microphones configured to receive acoustic signals from an environment and generate corresponding input audio signals, at least one speaker, at least one processor, at least one computer readable media encoded with instructions which, when executed by the at least one processor, cause the system to implement a trained machine learning model to receive the input audio signals and extract target audio signals corresponding to sources within the environment within a threshold distance from the head of the user and provide the target audio signals to the at least one speaker for playback to the user.

[0009] The system may also include a user interface configured to receive an indication of the threshold distance from the user.

[0010] The system may also include where the machine learning model is trained on interchannel phase difference (IPD) and interchannel level difference (ILD) features.

[0011] In some examples, the machine learning model utilizes a head-related transfer function (HRTF). The HRTF may be associated with the user. In some examples, the trained machine learning model performs thresholding on the input audio signals to classify input audio signals as closer or further than the threshold distance. In some examples, the trained machine learning model uses frequency-dependent variations where a phase of a particular audio signal changes as a function of distance and the trained machine learning model uses the phase difference across frequencies to compute a distance to an audio source associated with the particular audio signal. In some examples, the trained machine learning model utilizes a direct to reverberate ratio calculated based on multipath instances of the input audio signals to extract the target audio signals.

[0012] In some examples, the trained machine learning model utilizes a combination of features including IPD, ILD, head-related transfer function, thresholding, and a direct to reverberate ratio to extract the target audio signals.

[0013] In some examples, the plurality of microphones are included in ear buds. In some examples, the plurality of microphones are positioned along a headband.

[0014] Examples of methods are described herein. An example method may include wearing a plurality of microphones and at least one speaker; selecting a threshold distance; receivingDocket No. 0077145-07001 (UW 49996.02WO2) acoustic signals from an environment at the plurality of microphones to generate input audio signals; extracting, using a trained machine learning model implemented by at least one processor, target audio signals from the input audio signals, the target audio signals corresponding to audio signals generated by sources within the threshold distance; and listening to the target audio signals played back through the at least one speaker.

[0015] In some examples, selecting a threshold distance includes manually adjusting a user interface element to indicate the threshold distance.

[0016] In some examples, the trained machine learning model utilizes a combination of features including IPD, ILD, head-related transfer function, thresholding, and a direct to reverberate ratio to extract the target audio signals.

[0017] In some examples, the plurality of microphones are worn proximate a head of a user.

[0018] Some example methods include transmitting transmission signals, based on the input audio signals, to a computing system configured to implement the trained machine learning model.

[0019] Some example methods include receiving, from a plurality of microphones positioned proximate a head of a user, input audio signals associated with a plurality of audio sources in an environment; receiving an embedding indicative of a threshold distance from the head of the user, operating a trained machine learning model configured to extract target audio signals from the input audio signals, the target audio signals associated with one or more target audio sources of the plurality of audio sources in the environment which are within the threshold distance, and outputting the target audio signals to at least one speaker positioned proximate the head of the user.

[0020] In some examples, the trained machine learning model utilizes a combination of features including IPD, ILD, head-related transfer function, thresholding, and a direct to reverberate ratio to extract the target audio signals.

[0021] In some examples, the trained machine learning model utilizes a head-related transfer function to increase an effective aperture size of the plurality of microphones.

[0022] In some examples, one or more of the plurality of audio sources outside the threshold distance are louder than a loudest audio source within the threshold distance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIG. 1 is a schematic illustration of a system arranged in accordance with embodiments described herein.Docket No. 0077145-07001 (UW 49996.02WO2)

[0024] FIG. 2 is a schematic illustration of a computing system which may be utilized to train machine learning models described herein, arranged in accordance with embodiments described herein.

[0025] FIG. 3 is a schematic illustration of an environment in which example systems described herein may be utilized.

[0026] FIG. 4 is a schematic illustration of a neural network architecture arranged in accordance with examples described herein.

[0027] FIG. 5 is a schematic illustration of a separation block arranged in accordance with examples described herein.

[0028] FIG. 6 is a schematic illustration of a headset arranged in accordance with examples described herein.

[0029] FIG.7 is a schematic illustration of eyeglasses arranged in accordance with examples described herein.

[0030] FIG. 8 is a sheet of equations referred to herein. DETAILED DESCRIPTION

[0031] Certain details are set forth herein to provide an understanding of described embodiments of technology. However, other examples may be practiced without various of these particular details. In some instances, well-known circuits, control signals, timing protocols, and / or software operations have not been shown in detail in order to avoid unnecessarily obscuring the described embodiments. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter and / or claims presented here.

[0032] An example system includes a plurality of microphones placed proximate a head of a user. The microphones receive acoustic signals from the environment and generate corresponding input audio signals. The example system includes a trained machine learning model. This model receives input audio signals and extracts target audio signals corresponding to sources within a specified threshold distance from the user. These target audio signals are then provided to the speaker for playback to the user. This may create a sound bubble for the user.

[0033] The human auditory system generally has a limited ability to perceive distance and distinguish speakers in crowded settings. A headset technology that can create a sound bubble in which all speakers within the bubble are audible but speakers and noise outside the bubble are suppressed could augment human hearing. However, developing such technology has beenDocket No. 0077145-07001 (UW 49996.02WO2) challenging. Examples described herein include intelligent headset systems capable of creating sound bubbles. Example systems utilize real-time neural networks that use acoustic data from multiple microphones integrated into headsets (e.g., noise-cancelling headsets) and may be run on the device. In one example headset, the headset may process 8 ms audio chunks in 6.36 ms on an embedded central processing unit.

[0034] Example neural networks can generate sound bubbles with programmable radii. Example radii include between 1 m and 2 m. Output signals may reduce the intensity of sounds outside the bubble, such as by 49 dB in some examples. With previously unseen environments and wearers, examples of systems described herein can focus on multiple audio sources (e.g., human speakers) within the bubble, with multiple interfering human speakers and / or noise outside the bubble.

[0035] Examples of headset technology described herein can create a sound bubble in which audio sources (e.g., human speakers) within the bubble are audible, but speakers and noise outside the bubble are suppressed. Such a system may find use in a variety of scenarios. For example, it could be used by someone who wishes to eliminate noise in a restaurant and tune- in to a conversation at their table. The threshold distance for the bubble may be specified at a distance that may encompass the table but not other tables in the restaurant. In other examples, headset systems described herein may be used in a conference room with simultaneous conversations, where a person can exclusively hear the discussion within their bubble. In some examples, headset technology described herein may be used in noisy environments such as airplanes, tuning out distant speakers and ambient noises, but retaining the ability to hear a passing flight attendant who enters the wearer’s bubble. Accordingly, noise-cancelling headsets may be used herein that can suppress all sounds and play back sounds originating within a threshold distance (e.g., within the bubble).

[0036] Examples of systems described herein may identify audio sources as within a target region (e.g., a bubble) based on their distance from the wearer. Moreover, the systems may separate the audio sources from within the target region from those outside. This may occur even when the audio source(s) outside the bubble are louder. Accordingly, systems described herein may rely on more than just amplitude information to extract target audio signals from within a threshold distance. Moreover, the sound output of example headset systems described herein may be synchronized with a user’s visual senses. Accordingly, neural networks described herein may run on the device (e.g., on the headset) and in real time, using only limited computing capabilities and satisfying latency requirements of generally less than 20– 30 ms. Further, examples of neural networks described herein have the ability to create different bubble sizes and support configurable distance thresholds. Further, examples ofDocket No. 0077145-07001 (UW 49996.02WO2) neural networks described herein may generalize to real world settings, unseen users and / or reverberant environments. This is challenging as reverberations in real-world environments can vary across rooms. Additionally, the signals received at the microphones experience reflections from the human head (head related transfer functions (HRTFs)). These reflections can vary across users and, in turn, affect the distance estimates. Accordingly, neural networks described herein may be trained to account for these kinds of variations.

[0037] FIG. 1 is a schematic illustration of a system arranged in accordance with embodiments described herein. The system may be a hearable system. The system of FIG. 1 includes microphone 102, microphone 144, microphone 146, microphone 148, microphone 150, speaker 104, microphone 106, speaker 108, processor 110, and trained neural network 152. The trained neural network 152 is depicted as implemented using processor 110 and / or computer readable media 112. The system of FIG. 1 includes processor 110, and computer readable media 112. The computer readable media 112 may include executable instructions for extraction of audio signals originating within a threshold distance 116, embedding(s) 118, and / or executable instructions for noise cancellation 114. The system may also include input device 124.

[0038] The components shown in FIG. 1 are exemplary only. Additional, fewer, and / or different components may be used in other examples.

[0039] Examples of systems described herein may include one or more microphones and one or more speakers, such as microphone 102, microphone 106, microphone 144, microphone 146, microphone 148, microphone 150, speaker 104, and speaker 108 of FIG. 1. The microphones may receive acoustic signals from an environment and generate corresponding audio signals (e.g., electrical signals). The microphones and / or speakers may be provided in any of a variety of form factors. For example, microphones and / or speakers may be provided in one or more ear buds. For example, the microphone 102 and speaker 104 may be provided in an enclosure formed as an ear bud. The microphone 106 and speaker 108 may be provided in an enclosure formed as an ear bud. Any number of microphones and speakers may be used. In some examples, two microphones and two speakers may be present, e.g., microphone 102, microphone 106, speaker 104, and speaker 108. The two microphones and two speakers may be implemented in two earbuds in some examples. In some examples, two speakers may be present (e.g., one associated with each ear) and a greater number of microphones may be present - e.g., 3, 4, 5, 6, 7, 8, 9, or 10 microphones. For example, microphones microphone 102, microphone 106, microphone 144, microphone 146, microphone 148, and / or microphone 150 may be used. In some examples, multiple microphones may be positioned along a headband of a headset including the speakers. In some examples, two microphones may beDocket No. 0077145-07001 (UW 49996.02WO2) used and the two microphones may be referred to as binaural microphones. The multiple microphones may be referred to as an array of microphones.

[0040] Microphones of systems described herein may generally be placed proximate a head of a user. For example, one or more microphones may be supported by, on, or in the head of the user. In some examples, microphones may be positioned on or around one or more ears of the user (e.g., in one or more ear buds or headphone earcups). In some examples, microphones may be positioned on or around the head of the user (e.g., in one or more headbands of a headset). In some examples, microphones may be positioned on or around eyeglasses worn by a user. In some examples, microphones may be positioned on or around earrings, nose rings, necklaces, scarves, earmuffs, or other head or neck worn devices or objects. A combination of these positions may be used.

[0041] Each of the microphones in a system - such as microphone 102, microphone 106, microphone 144, microphone 146, microphone 148, and microphone 150 of FIG. 1 - may be referred to as a channel. Acoustic signals in an environment may be incident on some or all of the microphones. Accordingly, one audio source in the environment may produce an acoustic signal that is incident on each of the multiple microphones. Because the microphones are located at different positions, the incident acoustic signal will arrive at a particular time and direction relative of the audio source for each of the multiple microphones - this may be referred to as multiple channels.

[0042] Trained neural networks described herein, such as trained neural network 152 of FIG. 1 may be used to extract target audio signals from the input audio signals provided by the various microphones in the system. The target audio signals may be those originating from audio sources within a target region (e.g., within a threshold distance) from the microphones (e.g., from the head of the user). While described as a neural network, it is to be understood that other types of machine learning models may be trained and used to implement target audio signal extraction described herein.

[0043] During operation, speakers described herein may play back (e.g., produce) target audio signals. The target audio signals may be audio signals that originated from one or more audio sources within a threshold distance of a user (e.g., within a threshold distance of the collection of microphones used to capture the audio signals). In this manner, a user may create a sound bubble - e.g., a region of audio sources that may be heard, while audio sources outside the region (e.g., outside the bubble) may not be heard, and / or may not be heard as well.Docket No. 0077145-07001 (UW 49996.02WO2)

[0044] During operation, speakers described herein may produce optional noise-cancelling signals. While each of speakers 104 and 108 are shown in FIG. 1 as generating both noise- cancelling signals and target audio signals, in other examples one speaker may generate noise- cancelling signals and another speaker may generate target audio signals. In this manner, systems described herein may playback sounds from audio sources within a threshold distance of a user while suppressing environmental noise and sounds originating from outside the region. In some examples, the noise-cancelling signals may be generated to suppress all sounds in an environment, both within and outside the threshold distance. Accordingly, all sounds may be suppressed to generate a clean or silent environment, into which the target audio signals may be played back for increased clarity.

[0045] Examples of systems described herein may include one or more processors and computer readable media, such as processor 110 and computer readable media 112 of FIG. 1. Processors described herein may generally be implemented using any processor circuitry, including one or more field programmable gate arrays (FPGAs), central processing units (CPUs), graphic processing units (GPUs), application specific integrated circuits (ASICs), microcontrollers, and / or embedded processors. Computer readable media may generally be implemented using memory, random access memory (RAM), solid state drives (SSDs), read only memory (ROM), SD cards, and / or disk drives. While a single processor 110 and computer readable media 112 are shown in FIG. 1, any number may be used. In some examples, the processor 110 and computer readable media 112 may be incorporated into a headset (e.g., a noise cancelling headset). In some examples, the processor 110 and computer readable media 112 may be wholly or partially implemented in another computing device (e.g., a smartphone, tablet, desktop, wearable device, appliance, and / or vehicle) and may be in communication with a headset including the speaker(s) and microphone(s).

[0046] Software may be used to implement all or portions of operations described herein. For example, the computer readable media 112 of FIG. 1 may be encoded with instructions which, when executed, cause the processor 110 to perform operations described herein. For example, the computer readable media 112 may include executable instructions for noise cancellation 114, and / or executable instructions for extraction of audio signals originating within a threshold distance 116. The computer readable media 112 may store data received and / or used in operations described herein, such as embedding(s) 118. While a single computer readable media 112 is shown as including executable instructions for noise cancellation 114, executable instructions for extraction of audio signals originating within a threshold distance 116, and embedding(s) 118, it is to be understood that multiple computerDocket No. 0077145-07001 (UW 49996.02WO2) readable media may be used and the components shown as stored on computer readable media 112 may be stored on separate media in some examples.

[0047] Moreover, techniques described herein may be implemented in hardware, software, or combinations thereof. The trained neural network 152 may be implemented using a variety of hardware and / or software, including circuitry. In some examples, the processor 110 may implement techniques to reduce and / or minimize latency and / or power consumption. For example, integer quantization, may be used to reduce the network size and improve the inference time. In some examples an ASIC may be used to implement processor 110 to reduce power consumption.

[0048] The processor may be coupled to the microphones and speakers described herein using wired and / or wireless connections. For example, audio cables may be used. In some examples, Bluetooth, Zigbee, Wi-Fi, and / or other wireless protocols may be used.

[0049] Examples described herein may include one or more input devices, such as input device 124 of FIG. 1. The input device may be implemented, for example, as a button, a touchscreen, a speaker for receipt of audio input commands (which may be implemented using another speaker described herein). Other input devices may also be used. The input device may also be referred to as a user interface. The input device may be used to receive an indication of a target region from a user. For example, the input device 124 may receive an indication of a threshold distance from the user. In some examples, the input device 124 may include a display which may present a slider bar to a user to select the threshold distance. Accordingly, a user may manually adjust a user interface element (e.g., such as a slider bar or button) to indicate the threshold distance. In some examples, the input device may be located on a headband and / or earbud and / or it may be located on another device used by the user and in communication with the system of FIG. 1.

[0050] During operation, acoustic signals that are received by microphones described herein may be processed to extract signals originating within a predetermined region (e.g., within a threshold distance from the hearer in some examples). In this manner, users of systems described herein may hear speakers (and / or other audio sources) with improved clarity, even in the presence of background noise, audio sources outside of the predetermined region (e.g., beyond the threshold distance) and even as the user moves freely around the environment. The region of interest may generally be of any size or shape (e.g., spherical, cylindrical, etc.). In some examples, the region may be a region within a threshold distance of the user. The region may be relative to the hearer (e.g., relative to the hearable system), such as within a particular distance or distances away from the hearable system, such as aDocket No. 0077145-07001 (UW 49996.02WO2) distance from one or more of the microphones. The region may be fixed in some examples (e.g., a specific geographic location), which may not change as the user moves. The region may be predetermined in some examples and may be programmed, for example, into the executable instructions for extraction of audio signals originating within a threshold distance 116. The region may be changed in some examples responsive to an input signal. For example, the input device 124 may include one or more buttons which may allow a user to request an increase and / or decrease to a size or shape of the target region. Responsive to input signals from the input device 124, the executable instructions for extraction of audio signals originating within a threshold distance 116 may change a shape, size, or location of the target region.

[0051] During operation, acoustic signals may be received by microphones, such as the microphones 102 and 106 of FIG.1. The acoustic signals may include sounds made by audio sources within an environment. The audio signals may be processed by the processor 110 in accordance with the executable instructions for noise cancellation 114. In some examples the executable instructions for noise cancellation 114 may cause the processor 110 to generate control signals to have speakers generate noise-cancelling signals to cancel some or all sounds in the environment. For example, the noise-cancelling signals may be generated in accordance with the executable instructions for noise cancellation to provide a silent, or near- silent, sound condition in the ear of a user of the system.

[0052] In some examples, the noise-cancelling system used may nonetheless allow some residual sounds. In some examples, an inner microphone may be provided (e.g., within an ear bud and / or headset earcup) to monitor and wholly or partially mask this noise.

[0053] The acoustic signals may be converted by the microphones into electronic audio signals. The audio signals may be provided to and processed by the processor 110 in accordance with techniques described herein. For example, the trained neural network 152 may be used to extract target audio signals from the audio signals. The target audio signals may be those which were generated by an audio source within a threshold distance of the system. The executable instructions for extraction of audio signals originating within a threshold distance 116 may include instructions for a deep learning or other machine learning model or artificial intelligence model which may separate sounds from within the target region from sounds outside the target region (e.g., separate sounds originating from within the threshold distance from those originating outside the threshold distance).

[0054] In some examples, trained neural networks described herein (e.g., executable instructions for extraction of audio signals originating within the threshold distance), such asDocket No. 0077145-07001 (UW 49996.02WO2) trained neural network 152 of FIG. 1, may determine a distance to an audio signal source based on differences in receipt of the audio signal at multiple microphones (e.g., localization of the audio signal). Sounds originating at particular distances or locations may be determined to be within the target region.

[0055] Examples of trained neural network 152 may be a real-time neural network. A real- time neural network generally refers to a neural network which may process incoming signals within a threshold latency to present output signals to a user for use in an environment. Generally, the latency taken by the neural network to process a chunk of audio signals may be less than 30 ms in some examples, less than 20 ms in some examples, or less than 10 ms in some examples. In some examples, the amount of time taken to process a chunk of audio signals may be less than an amount of time taken to playback the audio signals. In one example, the trained neural network 152 may process 8 ms audio chunks in 6.36 ms.

[0056] Training of the machine learning model may be performed by the system used to implement the machine learning model in some examples. However, in some examples, the machine learning model may be trained by another computing system (e.g., a cloud computing system, server, desktop, or other computing system). The trained machine learning model may then be communicated to the processor 110 and / or computer readable media 112 of FIG. 1. The machine learning model may be trained to utilize a variety of features which may allow the trained machine learning model to extract target audio signals originating from sources within a threshold distance.

[0057] In some examples, interchannel phase difference (IPD) and / or interchannel level difference (ILD) are features which may be used by trained machine learning models described herein. IPD may refer to a phase difference between audio signals received at different channels (e.g., at different microphones described herein). ILD may refer to an amplitude difference between audio signals received at difference channels (e.g., at different microphones described herein). In some examples, the executable instructions for extraction of audio signals originating within a threshold distance 116 may include instructions for extraction based on observation or calculation of IPD and / or ILD.

[0058] In some examples, head related transfer functions (HRTFs) may be used by trained machine learning models described herein. For example, the executable instructions for extraction of audio signals originating within a threshold distance 116 may include instructions for analyzing and / or calculating head related transfer functions (HRTFs) associated with incident acoustic signals. The HRTFs may be used to determine which sounds are originating from within the target region. As a user's head moves relative to the audioDocket No. 0077145-07001 (UW 49996.02WO2) sources, the HRTFs may be used to determine a location of the audio source. Accordingly, training data used to train machine learning models described herein may include data from a variety of potential user head positions such that a resulting machine learning model may adjust to different HRTFs during operation (e.g., during extraction and / or inference).

[0059] Note that the use of HRTFs by neural networks described herein may effectively increase the aperture size of microphone arrays. Generally, systems described herein may dispose multiple microphones proximate a head of a user. This may mean that, relative to distances between the user and audio sources in the environment, the spacing between the microphones may be small. That is, the distance between microphones may be smaller than a distance between a user and a target audio source. The relatively close placement of the microphones may increase a difficulty associated with determining distance to an audio source, at least because the variations in the received signals between closely-positioned microphones may also be relatively small. Having a neural network trained over various HRTFs may allow for these smaller variations to be more accurately mapped to distances in some examples, allowing the microphone array positioned proximate a head of a user to effectively operate as though it were a larger microphone array (e.g., a larger aperture).

[0060] In some examples, machine learning models may be trained to utilize thresholding on input audio signals. Thresholding may be used to classify the input audio signals as either closer or further than the threshold distance. By classifying into two categories (e.g., closer or further), the trained machine learning model may be able to be more accurate and have less complexity than if a specific distance to each audio source needed to be calculated every time.

[0061] In some examples, machine learning models may be trained to use frequency- dependent variations of audio signals. For example, a phase of a particular audio signal generated by a microphone may change as a function of distance. Trained machine learning models described herein may use the phase difference across frequencies to compute a distance to an audio source associated with the particular audio signal. For example, consider an audio source emitting acoustic signals at a variety of frequencies. The different frequencies may arrive at a particular microphone with different phases. These phase differences across frequencies may be used to determine a distance to the audio source.

[0062] The model may be trained on signals arising in a variety of different environments. For example, training may occur in environments having different reverberation characteristics. This may allow for the model to be robust in operation in different reverberation environments. In some examples, trained machine learning models described herein may utilize a direct to reverberate ratio to extract target audio signals. ForDocket No. 0077145-07001 (UW 49996.02WO2) example, a ratio between an acoustic signal directly incident on a microphone and a reverberation of that signal incident on the microphone may be used. The signal directly incident may refer, for example, to an acoustic signal originating from an audio source and incident on a microphone. The reverberation may refer to the acoustic signal originating from the audio source, scattering or reflecting off a surface in the environment, and the scattered or reflected signal incident on the microphone. Accordingly, the reverberation signal typically arrives at the microphone at a later time than the direct signal.

[0063] Combinations of these instructions and techniques may be used in some examples. Trained machine learning models described herein may be trained to utilize a combination of features including IPD, ILD, HRTF, thresholding, frequency-dependent phase variation, and / or a direct to reverberate ratio to extract target audio signals originating from audio sources within a target region (e.g., within a threshold distance).

[0064] In this manner, systems described herein may extract target audio signals originating from audio sources within a target region. These target audio signals may be played back for a user, allowing the user to have improved hearing of sound sources within the threshold distance in some examples. It is to be understood that the extraction of the target audio signals from the input audio signals may not be perfect and / or complete in all examples. Rather, systems described herein may have attenuated the audio signals from sound sources outside the threshold distance (e.g., outside the target region or bubble). In some examples, the energy decay of signals from sound sources outside the threshold distance may be greater than 50 dB, or greater than 60 dB in some examples.

[0065] During operation, users of systems described herein may specify a threshold distance to be used in extracting audio signals. For example, an indication of the threshold distance may be provided using input device 124 of FIG. 1. The threshold distance may be used to generate one or more embeddings, such as embedding(s) 118 of FIG.1. The embeddings may be provided as an input to trained neural network 152 in some examples. The threshold distance may be programmable. The threshold distance may be expressed as a radius (e.g., a radius of a bubble). In some examples, the programmable radius may range from 1 m to 2 m, although other radii may be used.

[0066] In some examples, responsive to the indication of the threshold distance, processors described herein, such as processor 110, may enroll audio source(s) from within the threshold distance and / or may facilitate enrollment of one or more target audio sources. Enrollment generally refers to the process of generating an enrollment signal (e.g., an embedding vector) that may be indicative of characteristics of the target audio source and may subsequently beDocket No. 0077145-07001 (UW 49996.02WO2) used to extract target audio signals from received audio signals. In some examples, the processor 110 may perform the enrollment (e.g., the computer readable media 112 may include executable instructions for generating enrollment signals). In some examples, another computing system may perform the enrollment. For example, the processor 110 may provide audio signals received after receipt of the indication of the target region to another computing system (e.g., a cloud computing system). The processor 110 may transmit the audio signals through a wired or wireless connection. The other computing system may provide the enrollment signals back to the system of FIG. 1, and the processor 110 may store the enrollment signals in the computer readable media 112, such as embedding(s) 118.

[0067] In this manner, speakers described herein may provide noise-cancelling signals and signals from a target region. The noise-cancelling signals generally may be intended to create a silent environment in the ear of a hearer, and then the signals from the target region may be intended to provide the audio generated from within a particular region. In this manner, users of systems described herein may hear sounds from a particular region with increased clarity.

[0068] In some examples, because enrollment signals may be used to characterize audio sources from within the particular region, the extraction may continue to be robust as to enrolled audio sources as the user and / or the audio sources move through the environment in the presence of other interfering signals.

[0069] Examples of systems described herein, such as the example system of FIG. 1, may be implemented using a variety of form factors including one or more noise-cancelling headsets, smartphones, ear buds, earcups, headbands, hearing aids and / or other wearable devices.

[0070] Examples of systems described herein may be used in a variety of use cases. For example, systems described herein may be used to more clearly hear sounds from a target region (e.g., from within a threshold distance) in a crowded or noisy area. Examples include waiting rooms, classrooms, performances, outdoor environments, and / or indoor environments having mechanical or other noise. Examples of systems described herein may be used to provide signals from a target region in hearing aids. Note that noise-cancelling may be of less benefit and may not be used in hearing aids. This may be due to the poor hearing of noise sounds by the user in any event.

[0071] Users may accordingly wear one or more microphones described herein - such as in earbuds, on headphones, or otherwise placing the microphones proximate a head of a user. The user may select a threshold distance and / or other indication of a target region, such as by utilizing a user interface. For example, the input device 124 may be used to provide anDocket No. 0077145-07001 (UW 49996.02WO2) indication of a threshold distance. Systems described herein may extract target audio signals originating within the target region (e.g., within the threshold distance) and play back the target audio signals to the user. Note that other audio signals (e.g., those originating outside of the threshold distance) may not be played back to the user. Those signals may be cancelled in some examples using noise-cancelling signals, and / or simply may not be played back by the speakers described herein. In this manner, users may preferentially hear sounds originating from within a threshold distance or other target region (e.g., within a bubble). Note that the techniques described herein may not rely on an overall volume of the competing audio sources. In some examples, audio sources outside the target region (e.g., further than the threshold distance) may be louder than audio sources within the target region. Nonetheless, systems described herein may advantageously play back only the audio signals originating from audio sources within the target region, even if those audio signals are quieter than other audio signals from the environment.

[0072] Accordingly, the system of FIG. 1 may receive a multichannel noisy input. For example, each microphone of the system, including microphone 102, microphone 106, microphone 144, microphone 146, microphone 148, and microphone 150 may generate a channel of audio input signals. These audio input signals may include noise from the environment and audio from audio sources outside of a target region (e.g., outside a bubble). The processor 110 may be an embedded CPU which may implement a trained neural network 152 which may process the noisy multichannel input in real-time to generate a single-channel output. The single-channel output may include extracted audio signals originating from within a target region (e.g., from within a threshold distance).

[0073] FIG. 2 is a schematic illustration of a computing system which may be utilized to train machine learning models (e.g., neural networks) described herein, arranged in accordance with embodiments described herein. The computing system 202 of FIG. 2 may include processor 204 and computer readable media 212. The computer readable media 212 may include executable instructions for neural network training 206 and trained neural network data 210. During operation, the computing system 202 may generate, receive and / or access training data 208. The training data 208 may be used to train a neural network to develop the trained neural network data 210.

[0074] The components of FIG. 2 are exemplary only. Additional, fewer, and / or different components may be used in other examples. The trained neural network data 210 of FIG. 2 may be used to implement the trained neural network 152 of FIG. 1 in some examples. Accordingly, the computing system 202 of FIG. 2 may be used to develop the trained neural network 152 of FIG. 1 in some examples. In some examples, the processor 110 of FIG. 1Docket No. 0077145-07001 (UW 49996.02WO2) may be used to implement the processor 204 of FIG. 2, although in some examples the processors may be different.

[0075] The processor 204 may generally be implemented using any processor circuitry, including one or more FPGAs, CPUs, GPUs, ASICs, microcontrollers, and / or embedded processors. Computer readable media 212 may generally be implemented using memory, RAM, SSDs, ROM, SD cards, and / or disk drives. While a single processor 204 and computer readable media 212 are shown in FIG. 2, any number may be used.

[0076] The executable instructions for neural network training 206 may provide instructions for training a neural network to extract audio signals originating within a threshold distance. In some examples, the executable instructions of neural network training 206 may include instructions for generating training data, such as all or portions of training data 208. The training will utilize training data 208. Generally, the training data 208 may be provided such that the data may cause the machine learning model to utilize certain features in order to calculate a distance to audio source(s) and / or to extract target audio signals originating within a threshold distance. Accordingly, the training data may be selected to develop a machine learning model trained to utilize a combination of features including IPD, ILD, HRTF, thresholding, frequency-dependent phase variation, and / or a direct to reverberate ratio to extract target audio signals originating from audio sources within a target region (e.g., within a threshold distance). The trained neural network may be used, for example, to infer a distance to audio sources in an environment. Those inferred distances may be utilized to extract sounds originating from within the target region.

[0077] The executable instructions for neural network training 206 may include instructions to generate and / or augment training data in some examples. The training data 208 may be cleaned, normalized, and / or augmented in some examples. The training methodology may allow the machine learning model to achieve real-world generalization to previously unseen acoustic environments and wearers. Various data augmentation techniques may be used to account for diverse human anthropometric measurements, environmental noise and / or small changes in the microphone positions on the headset.

[0078] The executable instructions for neural network training 206 may include instructions to provide the training data to a machine learning engine (e.g., classification, regression, and / or clustering) to cause the neural network to adjust its internal parameters. The machine learning engine may be executed by one or more processors, such as processor 204 of FIG.2. The executable instructions for neural network training 206 may include instructions for evaluating the model's performance using the training data 208 and metrics such as accuracy.Docket No. 0077145-07001 (UW 49996.02WO2) An error function may be used to evaluate a difference between an inference made by the model and a known value associated with the training data (e.g., distance). The executable instructions for neural network training 206 may include instructions for adjusting parameters and / or hyperparameters of the machine learning model to improve performance (e.g., to minimize and / or reduce an error calculated by the error function).

[0079] Once trained, a set of parameters and / or hyperparameters may be available to represent all or a portion of the trained machine learning model. For example, the trained neural network data 210 may be used to implement the trained neural network 152 of FIG. 1.

[0080] Examples described herein may advantageously generate training data 208 in a manner which results in a robust data set to allow the trained model to utilize a combination of features including IPD, ILD, HRTF, thresholding, frequency-dependent phase variation, and / or a direct to reverberate ratio to extract target audio signals originating from audio sources within a target region (e.g., within a threshold distance). To generate data, a number of audio sources may be positioned in an environment. The environment may be real or simulated in some examples. A variety of different environments may be used, which may have different reverberation characteristics. A microphone array as described herein is used to record received audio signals. The microphone array may be moved robotically and / or worn by a human wearer. Examples of headsets described herein, such as the headset of FIG. 1, may be used to record the audio signals. Accordingly, data may be recorded from multiple known audio sources at known distances by a human wearer. The data may be recorded during movement of the head and / or body of the wearer as well as different rooms with different acoustic properties.

[0081] Generalization to practical wearers may advantageously utilize a dataset for training data 208 that captures how distance and HRTFs interact with real-world acoustic environments. Audio signals will generally reflect off the wearer’s head, causing the signal to undergo various transformations, known as HRTFs. Moreover, since the headset or other device supporting microphones, may not be perfectly rigid, the microphone positions may change slightly with different wearers’ head sizes. To facilitate training of a neural network that may generalize to unseen real-world wearers and speakers, a large dataset collected with real-world headsets is advantageous. The training data set (e.g., training data 208) should advantageously include distance information (e.g., distances between the headset and each audio source) to allow the neural network to learn to calculate distance using features including HRTF.Docket No. 0077145-07001 (UW 49996.02WO2)

[0082] In some examples, a robotic platform was used to automate data collection for training data 208 using headsets. The system automatically collected audio recordings using a mannequin head mounted on the robotic platform. The mannequin head was mounted on a rotating robotic platform and audio was recorded with the head rotated at different angles. A loudspeaker was placed on an upright linear actuator to change its height. One example robotic system can collect combinations of 20 angles and 3 heights within 25 minutes.

[0083] In addition to the mannequin dataset, human wearer data may be collected for inclusion in training data 208 in various environments to capture the acoustic properties of real humans and environments, allowing the neural network to be trained using both synthetic (e.g., mannequin) and human data for real-world evaluation. In an example of human data collection, wearers could move their head and / or body as they wear a headset including multiple microphones. Audio sources are placed in the environment at known distances from the wearer. Audio from the multiple sources is received at the headset. The training data may include the received audio signals from the multiple microphones worn by the user and the distances to the audio sources. Training data may be collected from multiple environments (e.g., multiple rooms) to allow for training across multiple different reverberation characteristics.

[0084] Accordingly, training data collected from human wearers may include audio signals generated by multiple microphones proximate the head of a human wearer. The audio signals may be augmented to allow the data collected to represent additional collection scenarios without actually having to collect additional data in some examples.

[0085] For example, an audio channel (e.g., audio signals generated by a microphone) may be shifted by a number of samples (e.g., four samples in one example) to simulate changes in the microphone position. Accordingly, additional training data may be generated by shifting audio signals generated by a microphone in time. The shifted data may be used to represent collection of acoustic signals at another microphone, although that microphone may not have been present in the testing environment. The shifted data may be used to represent collection of acoustic signals at the microphone in a different position. In some examples, shifting the audio signals may represent a microphone placed millimeters or centimeters away from the microphone used to generate the unshifted audio signals. In one example, each audio channel generated by a microphone worn by a human wearer may be shifted by a respective number of samples (e.g., between one and five samples in some examples). In one example, a shift of four samples represents an audio signal received by the same or another microphone 5.7 cm from the position of the microphone used to generate the unshifted audio signal. In this manner, training data may include audio signals generated by microphones responsive toDocket No. 0077145-07001 (UW 49996.02WO2) acoustic input and also shifted versions of those audio signals to represent additional channels (e.g., additional microphones) and / or collection positions. Accordingly, the computing system 202 may generate the additional training data. The computing system 202 may include executable instructions for shifting the audio signals. In some examples, the computing system 202 may include circuitry for shifting the audio signals.

[0086] As another example, the amplitude of one or more audio signals received at one or more microphones (e.g., one or more channels) may be adjusted. This may generate additional and / or replacement channels of training data having an adjusted amplitude. In some examples, the amplitude may be adjusted by up to 2 dB, up to 3 dB, or up to 5 dB in some examples. In this manner, training data may include audio signals generated by microphones responsive to acoustic input and also amplitude-adjusted versions of those audio signals. Accordingly, the computing system 202 may generate the additional training data. The computing system 202 may include executable instructions for adjusting an amplitude of the audio signals. In some examples, the computing system 202 may include circuitry for adjusting an amplitude of the audio signals.

[0087] As another example, certain frequencies of audio signals received at one or more microphones may be attenuated and / or removed to create training data and / or additional training data. Accordingly, the computing system 202 may implement a frequency filter and / or executable instructions for frequency attenuation. The training data may include the full frequency audio signals received at a microphone in addition to a version of the audio signals with certain frequencies attenuated in some examples. In some examples, the audio signals with certain frequencies attenuated may replace the full frequency audio signals in the audio data. By attenuating frequencies in the training data, the neural network may be trained to increase robustness to frequency distortions in an environment. Training data may be augmented in some examples by removing one or more frequency bins from the training data, and / or from selected channels of the training data. In some examples, one or more frequency bins may be zero-ed out.

[0088] As another example, an audio speed of recorded audio data received at a microphone array may be varied. This may generate additional training data which increases audio source diversity. For example, training data may be generated based on acoustic signals generated by an audio source in an environment and incident on one or more microphones. Additional training data may be generated by playing that data back at different speeds. In some examples, the speed may be varied by 80 – 120%, 75-125% in some examples, or 70-130% in other examples. For example, training data may be generated by receiving audio signals responsive to sounds produced at an audio source. Additional training data may be generatedDocket No. 0077145-07001 (UW 49996.02WO2) by changing the audio speed of the collected data, such as by changing to 80% speed. Additional training data may be generated by changing to 120% speed. This may allow the trained neural network to account for speaker diversity.

[0089] In this manner, training data may generally be generated by placing one or more audio sources (e.g., speakers) in an environment. The audio to be played by those sources may be known. The distance between receiving microphones and each audio source may be known. Training data may include audio signals generated by the microphones responsive to the audio sources. The training data may include the known audio signals played by the sources. The training data may include the distance to the sources. Additional training data may be generated by shifting the audio signals generated by the microphones. Additional training data may be generated by adjusting channel amplitudes. Additional training data may be generated by attenuating frequencies in the audio signals. Additional training data may be generated by varying an audio speed of the recorded signals. Combinations of these additional training data may be used.

[0090] The training data may be used to train a neural network. For example, the neural network may infer a distance to a particular audio source based on the training data. The inferred distance may be compared with the known distance in the training data and an error generated using an error function. Parameters of the neural network may be adjusted to minimize and / or remove error in inferences. Once the error function meets a predetermined criteria (e.g., is minimized), the parameters may be stored and may represent a trained neural network. In this manner, systems described herein may train a machine learning model on augmented data. The network may be trained to determine a distance to each of the known audio sources using an error function based on a comparison of a predicted distance with the known distances.

[0091] In all examples of the generation of training data, the microphones may be varied in position. For example, the microphones generating the audio data included in training data may be moved in an automated fashion – e.g., using one or more motor(s) and / or platforms. In some examples, the microphones may be worn by a human wearer and the wearer may change the position of their body and / or head. In this manner, training data may be generated at multiple positions of the microphones which may be positions in which the microphones may be expected to be during wear by a human user (e.g., head down, head turned, head tilted). This may allow the neural network to be trained across various microphone positions and / or HRTFs.Docket No. 0077145-07001 (UW 49996.02WO2)

[0092] In some examples, four data augmentation techniques were used to generate additional training data: (1) shifting each audio channel by up to four samples (e.g., 5.7 cm) to simulate changes in the microphone position, (2) adjusting the channel amplitude independently by up to 2 dB, (3) randomly zeroing out up to 512 frequency bins to enhance robustness to frequency distortions and (4) varying the audio speed by 80–120% to increase the speaker diversity. Selected ones, sub-combinations, or other augmentation operations may be used in other examples.

[0093] This provided a dataset of end-to-end audio recordings that captured both head- related reflections and reverberations as a function of distance in real-world environments, with mannequins as well as humans, spanning a total of 15.85 hours of recording in 22 different acoustic indoor environments in an implemented example (including offices, living spaces, conference rooms, classrooms and laboratories).

[0094] Accordingly, training data used herein may include audio signals generated responsive to acoustic signals received at multiple microphones, together with distance information to the audio sources in the environment. A neural network may be trained to accurately infer the distance to the audio sources. In this manner, the neural network may extract the audio signals associated with audio sources within a threshold distance.

[0095] While examples of the collection of training data in indoor environments like offices, conference rooms and noisy restaurants has been discussed, outdoor data collection may additionally or instead be used. Outdoor data collection may use precise acoustic simulators and / or be collected in quieter outdoor settings or semi-enclosed areas, such as stadiums and sports fields, with controlled noise for improved ability to obtain clean data.

[0096] Accordingly, training data used to train neural networks described herein may include data collected in one or more automated processes including automated movement of the microphones to represent a variety of HRTFs. In some examples, neural networks may be trained using the data collected during automation followed by fine-tuning using data collected by human wearers.

[0097] FIG. 3 is a schematic illustration of an environment in which example systems described herein may be utilized. FIG. 3 illustrates an indoor environment. Present in the environment are person 304, person 306, person 310, and person 312. The person 304 is wearing a headset 308 arranged in accordance with examples described herein. The person 310 is wearing a headset 314 arranged in accordance with examples described herein. The person 304 has configured the headset 308 to extract target audio signals originating within a threshold distance shown by bubble 316. The person 310 has configured the headset 314 toDocket No. 0077145-07001 (UW 49996.02WO2) extract target audio signals originating from within a threshold distance shown by bubble 318. Person 306 is within the bubble 316 with person 304, but person 310 and person 312 are outside of bubble 316. Person 312 and person 310 are within bubble 318, but person 304 and person 306 are outside of bubble 318.

[0098] In this manner, in accordance with techniques described herein, the headset 314 may playback for the person 310 the sounds generated by person 312, but not the sounds generated by person 304 and / or person 306. The headset 308 may playback for the person 304 the sounds generated by person 306, but not the sounds generated by person 310 and / or person 312. In this manner, the person 304 may better be able to hear and / or focus on a conversation with person 306 without distractions from the conversation being held by person 312 and person 310. Similarly, the person 310 may better be able to hear and / or focus on a conversation with person 312 without distractions from the conversation being held by person 304 and person 306.

[0099] The environment and arrangement of people, bubbles, and headsets in FIG. 3 is exemplary only. Additional, fewer, or different people, headsets, and / or bubbles may be used in other examples. Other environments may be used in other examples - including outdoor environments, lecture halls, conference halls, restaurants, automobiles, vehicles, airplanes, and / or medical facilities. The headset 308 and / or headset 314 may additionally cancel noise in some examples such as birds, insects, automobile noise, HVAC system noise, background music, or airplane noise.

[0100] Real-time on-device neural networks may advantageously be used that can run on hearable devices to extract audio signals from audio sources within a threshold distance (e.g., to create sound bubbles). Accordingly, examples of neural networks described herein, such as trained neural network 152 of FIG. 1, may process incoming audio in real time and on the device, such as on an embedded CPU, and achieve an end-to-end latency of less than 20–30 ms between the audio and visual scenes; adapt to diverse reverberant conditions and wearers; support configurable bubble sizes; and / or work even when farther audio sources are louder than those within the bubble, as well as when there are multiple or no speakers inside or outside the bubble.

[0101] Accordingly, neural networks described herein may evaluate distance to an audio source, not just direction. Neural networks described herein may achieve latency usable in a real-time application. Neural networks described herein may also support different HRTFs found in hearable devices.Docket No. 0077145-07001 (UW 49996.02WO2)

[0102] FIG. 4 is a schematic illustration of a neural network architecture arranged in accordance with examples described herein. The example architecture of FIG. 4 includes microphone 402, microphone 404, microphone 406, short-time Fourier transform (STFT) 408, features 410, concatenation 412, convolution 414, embedding 416, separation blocks 418, deconvolution 420, inverse STFT (iSTFT) 422, and speaker 424.

[0103] The components of FIG. 4 are exemplary only. Additional, fewer, and / or different components may be used in other examples.

[0104] The neural network architecture shown in FIG. 4 may be used to implement the trained neural network 152 of FIG.1 or other neural networks described herein. For example, the microphone 402 may be used to implement microphone 102, the microphone 404 may be used to implement microphone 406, the microphone 406 may be used to implement microphone 144, the speaker 424 may be used to implement speaker 104. The processor 110 may be used to implement the computational blocks shown in FIG. 4 - e.g., the circuitry may be configured to perform the operations shown and / or the executable instructions for extraction of audio signals originating within a threshold distance 116 may include instructions for performing the operations.

[0105] The real-time neural network architecture of FIG.4 generally includes four modules: a feature encoder, a distance-embedding network, a separation module and a feature decoder. To achieve low latency, the multichannel input audio may be processed in small chunks of LC (ms) each, with lookahead LF (ms).

[0106] As shown in FIG. 4, multiple channels of input audio signals are provided, one for each microphone. The audio input signals may be processed in chunks - a chunk LC is shown in FIG. 4 together with a lookahead amount LF.

[0107] The feature encoder includes a transform which is applied to the audio chuck. In the example of FIG.4, a short-time Fourier transform (STFT 408) is used. Other transforms may be used in other examples. The transform generally converts the audio chunk from a time domain representation to a time-frequency (TF) representation. The STFT 408 is applied to LCwith LFof lookahead on audio from multichannel microphones. Any number of microphones may be used, including six microphones. In one example, LC may be 8 ms and LFis 4 ms, although other values may be used.

[0108] IPD and ILD features are extracted from the frequency domain representation Xl output from the transform. These features are the phase and level differences, respectively, between the input channels. Accordingly, features 410 are obtained. Other features may be obtained from the frequency-domain representations in other examples. The features 410 mayDocket No. 0077145-07001 (UW 49996.02WO2) be stored, for example in computer readable media 112 of FIG. 1. The features may be combined with TF representations of the raw mixtures, such as in concatenation 412.

[0109] The features (e.g., IPD and ILD) and / or the combined features and raw chunks may be input to a set of separation blocks 418. Note that optionally, the features may be concatenated and / or convoluted. Accordingly, concatenation 412 and convolution 414, which may be a two-dimensional convolution, are shown in FIG. 4. The features may be applied to the set of separation blocks 418.

[0110] The set of separation blocks 418 may be referred to as a separation module. Generally, the separation module may be the most computationally intensive component of the neural network architecture of FIG. 4. The set of separation blocks 418 may extracts the TF representations of the audio sources within the threshold distance (e.g., within the bubble). TF-GridNet, a state-of-the-art speech separation network, may be used to implement the set of separation blocks 418 in some examples. However, in some examples TF-GridNet itself may not run as real time on an embedded CPU (e.g., Raspberry Pi) as would be advantageous. The frequency-dimension long short-term memory (LSTM) and attention blocks may have the highest computational complexity. Accordingly, in some examples, the set of separation blocks 418 may be implemented using TF-GridNet with the attention block(s) removed.

[0111] FIG. 5 is a schematic illustration of a separation block arranged in accordance with examples described herein. The separation block shown in FIG. 5 may be used to implement one or more of the set of separation blocks 418 of FIG. 4. A normalization may be provided followed by a strided convolution layer. The strided convolution layer may be used to compress the frequency dimension before applying an intraframe spectral LSTM. This may reduce the sequence length used by the LSTM and therefore reduce its execution time. Then, a unidirectional LSTM (e.g., causal LSTM) may be applied along the time dimension.

[0112] The set of separation blocks 418 may be conditioned by the threshold distance (e.g., the programmable bubble size). A threshold distance, d, also referred to as bubble size in FIG. 4, may be provided as an embedding, embedding 416, M. The embedding 416 may be applied to the set of separation blocks 418 to configure the set of separation blocks 418 for that threshold distance. The threshold distance (e.g., size of the bubble) may be encoded as a distance-embedding mask, generated from a one-hot vector representing discrete threshold distances and / or bubble sizes. This mask may then be applied to an output of the frequency- dimension LSTM using a FiLM layer.

[0113] In some examples, the threshold distance (e.g., bubble size) may not be selected from a fixed set. In some examples, the threshold distance (e.g., bubble size) may be selected fromDocket No. 0077145-07001 (UW 49996.02WO2) a continuous range of distances or sizes. Accordingly, instead of using one-hot vectors to generate a distance-embedding mask, examples may use a single continuous value to represent the threshold distance. A linear layer and a rectified linear unit layer may be used to convert the single value into a vector.

[0114] An output of the set of separation blocks 418, Z, may be optionally deconvolved using deconvolution 420, and then applied to an inverse transform block, such as iSTFT 422. For example, the TF representations of the extracted audio sources may be passed to an STFT decoder (e.g.,, an inverse STFT (iSTFT) operation) to recover the time-domain signal containing only the audio signals originating within a threshold distance (e.g., within the bubble). The inverse transform block may be used to convert the spectrogram back to time- domain audio. The resulting output may be extracted signals originating within a threshold distance of the microphone array. The output audio signals may be stored in a playback buffer. The playback buffer may be implemented, for example using computer readable media 112 of FIG. 1. The microphone 402 may play the audio output signals from the playback buffer.

[0115] For each input chunk, systems described herein may further reduce the inference runtime per chunk by caching computations made as prior chunks are processed. For example, a set of state buffers may be maintained. The buffers may be maintained, for example, by processor 110 of FIG. 1 in computer readable media 112 of FIG. 1. The state buffers may contain intermediate computation results, which may be passed to the neural network during inference. For example, when computing the final output of the model, the STFT decoder may utilize a chunk (e.g., 4 ms) of past context from the previous inference, which can be reused from the appropriate state buffer.

[0116] Several such state buffers may be maintained, including one for the initial convolution layer, a hidden-state buffer and a cell-state buffer for the intraframe LSTM in each separation block, one for the final deconvolution layer and one for the iSTFT layer. In some examples, the model may be converted into Open Neural Network Exchange (ONNX) for inference and run with ONNX Runtime. ONNX-specific optimizations may be performed to fine-tune the performance for real-time operation on an embedded CPU.

[0117] FIG. 6 is a schematic illustration of a headset arranged in accordance with examples described herein. The headset includes headband 614, earcup 616, and earcup 618. The position of six microphones is illustrated along the headband 614. Microphone 602 may be located on or about earcup 616. Microphone 604, microphone 606, microphone 608,and microphone 610 are positioned along the headband 614. Microphone 612 is positioned on orDocket No. 0077145-07001 (UW 49996.02WO2) about earcup 618. The microphones may be approximately evenly-spaced along headband 614 in some examples, although they may be unevenly spaced in some examples. The headset of FIG. 6 may be used to implement the system of FIG. 1 in some examples. For example, the microphones shown in FIG. 6 may be used to implement the microphones shown in FIG. 1. Although not shown in FIG. 6, the additional components shown in FIG.1 may be present on, integrated with, connected to, or in communication with the headset of FIG. 6 in some examples. While six microphones are shown in FIG. 6, any number may be used. In some examples, two microphones may be used, such as only microphone 602 and microphone 612. In some examples, four microphones may be used, such as only microphone 602, microphone 604, microphone 610, and microphone 612. Other arrangements may be used in other examples.

[0118] FIG.7 is a schematic illustration of eyeglasses arranged in accordance with examples described herein. Microphones may be worn proximate a head of a user by being included in eyeglass frames for example. In the example of FIG. 7, a six microphone configuration is shown. microphone 702 and microphone 704 are positioned at the respective outer corners of the rims of frames. Microphone 706 is positioned on the rim beneath a right lens, and microphone 708 is positioned on the rim beneath a left lens. Microphone 710 is positioned on a right temple arm. Microphone 712 is positioned on a left temple arm. Other arrangements are possible in other examples. The eyeglasses of FIG. 7 may be used to implement the system of FIG. 1 in some examples. For example, the microphones shown in FIG. 7 may be used to implement the microphones shown in FIG. 1. Although not shown in FIG. 7, the additional components shown in FIG. 1 may be present on, integrated with, connected to, or in communication with the eyeglasses of FIG. 7 in some examples. For example, the eyeglasses may be in communication with and / or include one or more speakers. In some examples, the eyeglasses may include a transmitter to transmit target audio signals to one or more earbuds or other headset worn by a user. In some examples, the eyeglasses may include a speaker, including one or more bone conduction speakers. While six microphones are shown in FIG. 7, any number may be used.

[0119] Accordingly, examples described herein include hearable systems capable of preserving the sounds originating inside a user-programmable target region (e.g., threshold distance and / or bubble) around the wearer and suppressing other sounds. Examples of systems described herein use a trained neural network to identify and extract sounds (e.g., speech) from audio sources inside the target region. The trained neural network may be trained on data collected in a variety of real-world conditions, allowing it to generalize to unseen users and environments.Docket No. 0077145-07001 (UW 49996.02WO2)

[0120] From the present specification it will be appreciated that, although specific embodiments have been described herein for purposes of illustration, various modifications may be made while remaining within the scope of the claimed technology.

[0121] Examples described herein may refer to various components as “coupled” or signals as being “provided to” or “received from” certain components. It is to be understood that in some examples the components are directly coupled one to another, while in other examples the components are coupled with intervening components disposed between them. Similarly, signal may be provided directly to and / or received directly from the recited components without intervening components, but also may be provided to and / or received from the certain components through intervening components.

[0122] IMPLEMENTED EXAMPLE

[0123] In an implemented example, an Orange Pi 5B was used to implement a neural network model to extract audio signals originating within a threshold distance. The model could process 8 ms audio chunks at 24 kHz in an average inference time of 7.30 ms with a standard deviation of 0.11 ms. Using a Raspberry Pi 4B, the audio chunks could be processed in an average inference time of 6.36 ms with a standard deviation of 0.18 ms.

[0124] A six-channel microphone array was used similar in size to a noise-cancelling headset or smart glasses. 30,000 rooms were generated with random sizes ranging from 5 × 4 × 2 m3to 8 × 8 × 4 m3. The reverberation of each room was randomly selected, with the wall absorption coefficient uniformly sampled from 0.1 to 0.9, and the maximum reflection order uniformly sampled between 10 and 72. In each room, 0–2 speakers were placed inside and 1– 2 outside the sound bubble, with a 50% chance of adding diffuse noise. The speech dataset was split into training, validation and testing sets. The signal-to-noise ratio (SNR) of speech inside the bubble ranged from –5 dB to 5 dB, with a median of 0 dB, and in 50% of the samples, the outside interference was louder than the speech inside. Across different reverberant conditions, the network could suppress the outside speakers with an average energy decay of 69 dB, 66 dB and 65 dB for the 1 m, 1.5 m and 2 m bubbles, respectively.

[0125] The bubble size was uniformly sampled from 0.8 m to 2 m. Zero, one, or two speakers were placed inside the bubble with 1–2 interfering speakers outside as well as background noise for training. During evaluation, six different bubble sizes were selected of 1 m, 1.2 m, 1.4 m, 1.6 m, 1.8 m and 2 m, and the decay–distance curves for each bubble were computed.

[0126] The neural network system was integrated with off-the-shelf noise-cancelling headsets, using a ReSpeaker six-channel microphone, Sony WH-1000XM4 headphones and a Raspberry Pi 4B. The noise-cancelling headphones suppress all sounds, allowing the systemDocket No. 0077145-07001 (UW 49996.02WO2) to reintroduce target sounds from inside the bubble. When activated, the system captures audio in 8 ms chunks, processes it on the Raspberry Pi using a trained neural network as described herein and plays it back through the headphones. Both Raspberry Pi and Orange Pi can process 8 ms chunks of audio using the corresponding networks in real time. The headphones’ noise-cancelling capability can attenuate signals across all frequencies, providing a clean slate on which only the desired sounds can be played back.

[0127] The neural network system included an architecture as shown in FIG. 4. Consider an acoustic environment with N speakers at two-dimensional locations p0, p1,…, pN–1. The headset is centered at location q, with M microphones. The speech signal for speaker j received by microphone i may be denoted as sij(t). sij(t) contains the channel response from source j to microphone i, including the multipath effect and HRTF. The background noise at microphone i is denoted by ni(t). Then, the mixture xi(t) at each microphone can be written as shown in Equation 1 in FIG. 8.

[0128] Describing the reference microphone index as c and the bubble size as d, the model ℱ should advantageously extract audio sources within the bubble at the reference microphone c as given by Equation 2 in FIG. 8, where is an indicator function that outputs 1 if the input expression is true, or 0 otherwise. ∥⋅∥2 is the Euclidean norm operation.

[0129] All the signals in this example implementation are sampled at 24 kHz for two reasons: LibriTTS, our largest speech dataset, uses 24 kHz, and higher rates like 48 kHz increase runtime latency. Although a lower rate of 16 kHz could reduce latency, 24 kHz was selected to preserve more high-frequency components, aligning with modern hearing aids, which use 20–30 kHz . This balances latency and audio quality for this example speech bubble system.

[0130] An STFT encoder was used to provide a stable and large-enough encoder window size to stabilize the performance under reverberation. The input audio frame is LC+ LF= 12 ms long, where LC = 8 ms is the current chunk size and LF = 4 ms is the lookahead. Accordingly, the example STFT encoder uses a discrete Fourier transform length of 288 samples, a step size of 192 samples and a Hann window. This allows us to compute the TF representation Xi(ω) ∈ ℂF×Kfor the signal at microphone i: xi(t) ∈ ℝT, where F is the number of frequency bins, K is number of time frames and T is the time duration of the input signal.

[0131] An iSTFT decoder was used with the same discrete Fourier transform length, step size and window as the encoder. This produces a 12 ms output, but the last 4 ms was removed, which will be affected by future chunks, leaving an 8 ms audio output. The 4 ms lookahead chunk is stored for overlap-and-add operation with the next chunk. Both STFT and iSTFT were implemented with convolution layers using Asteroid35 and Hann windows, offeringDocket No. 0077145-07001 (UW 49996.02WO2) flexibility in choosing discrete Fourier transform points, unlike fast Fourier transform, which utilizes powers of 2.

[0132] Interchannel features IPD and ILD were extracted as described in Equations 3 and 4 of FIG. 8, where ∣⋅∣ denotes the complex magnitude. The IPD and ILD were computed between the reference channel c and all other microphone channels and these features were appended to Xfeature∈ ℝ(3M−3)×F×Kas shown in Equation 5 of FIG. 8. The domain-specific features were stacked with the real and imaginary parts of the TF representations {Real(Xi), Imag(Xi)}, ∀i ∈ [0, M), to obtain Z ∈ ℝ(5M−3)×F×K. Finally, a two-dimensional convolution layer was used to associate the information across different microphone channels and features and to output a new TF representation, Z̃ ∈ ℝE×F×K, where E is the feature channel number.

[0133] The separation network included B repeated blocks. In each individual block, a strided convolution layer ^^ ∶ ℝE×F×K→ ℝE×F′×Kis applied on the frequency dimension to compress the frequency number from F to F′ to reduce the computational complexity. Next, a frequency-domain bidirectional LSTM is applied to associate the information across different frequency bands followed by a transposed convolution layer to uncompress to the original frequency dimension, Z̄ ∈ ℝE×F×K. Then, a unidirectional LSTM is applied to the frame dimension of the TF representations, Z̄. The hidden size for LSTM is H, which was set to 128.

[0134] A distance-embedded mask M ∈ ℝG×F, conditioned by the bubble size, was used to extract the TF representations of the speakers within the bubble from Z̄ using a FiLM layer. Here, G is the intermediate feature number, which is set to 4 in this example implementation. The distance-embedded mask was generated from one-hot vector d ∈ ℝDusing a linear layer, ℒ ∶ ℝD→ ℝG×F, where D is the number of discrete bubble sizes. The FiLM layer is composed of two convolution layers, ^^ω ∶ ℝG×F→ ℝE×Fand ^^b ∶ ℝG×F→ ℝE×F. Z̃ is updated to extract the TF representations of the speakers within the bubble as given in Equations 6-8 of FIG. 8. This updated Z̃ is input to the next separation block, and this process is repeated B times.

[0135] Example models were written and trained in PyTorch but ran inference on target hardware using ONNX Runtime. The PyTorch model was converted to ONNX (opset 19) and ONNX Runtime v. 16.1.0 was used for inference. To reduce latency, the PyTorch model was optimized by (1) replacing the convolution layers with a kernel size of 1 with faster PyTorch linear layers, (2) rewriting layer normalization to fuse into a single ONNX LayerNormalization node and (3) fusing shape-related operations to minimize the ONNX Transpose nodes.Docket No. 0077145-07001 (UW 49996.02WO2)

[0136] Using an example robotic platform, audio samples from LibriTTS36 and VCTK37 were played through a loudspeaker mounted on a linear actuator at various heights. The speaker was positioned at fixed distances from a mannequin head, and audio was recorded over 60 head orientations and three speaker heights. After testing all the angle and height combinations, the speaker was moved to a new distance. This process was repeated in four acoustic environments, with 8–11 distances per environment. Data were collected from two mannequin heads, totaling 10.8 h, all exclusively for training with minimal human intervention.

[0137] An additional dataset was collected from ten human wearers in 18 acoustic environments. The dataset was split for training, validation and testing, with each environment and wearer in only one split. Training data came from two users in ten environments, validation from three users in three environments and testing from five users in five environments.

[0138] To collect data, users sat on a rotating chair wearing our headset, with a loudspeaker on a linear actuator placed at various distances. The loudspeaker’s height was adjusted, and users rotated to random angles before audio playback. In addition to speech samples, the loudspeaker played WHAM! noise samples when outside the 1.5 m bubble. Participants were asked to naturally move their heads or bodies during playback for better generalization.

[0139] The synthetic and real-world mixtures were generated using three open-source datasets. For speech utterances of targeted speakers and interfering speakers, the VCTK37 and LibriTTS36 datasets were used, whereas the WHAM! dataset was used for the background noise samples. All of these datasets were split into three non-overlapping subsets for training, validation and testing. The synthetic dataset was generated using the PyRoomAcoustics simulator. Random three-dimensional rooms were created with dimensions sampled from 5 m to 8 m (length), 4 m to 8 m (width) and 2 m to 4 m (height). Wall absorption coefficients ranged from 0.1 to 0.9, with the maximum order between 10 and 72 for reverberation. A six-microphone array was randomly placed in the room at a height of 1.2–1.8 m. For each bubble size, 0–2 speakers were placed inside and 1–2 speakers were placed outside the bubble, with diffuse background noise added 50% of the time. Speaker distances and angles were randomly sampled, and heights varied by ±0.4 m from the wearer’s height. The input SNR uniformly ranged from –10 dB to 5 dB for training and –5 dB to 5 dB for testing and validation. The mixture of the reverberated signals of the speakers inside the bubble at the reference microphone was used as the ground-truth signal. Each sample was 5 seconds long, with 30k training samples, 3k validation samples and 3k testing samples. Real- world audio mixtures were created from recordings with mannequins and human wearers byDocket No. 0077145-07001 (UW 49996.02WO2) combining 5 s clips from multiple speakers and noise sources in the same room. Each mixture included 0–2 speakers within 1.5 m, 1–2 speakers beyond 1.5 m and 0–1 noise sources beyond 1.5 m. To avoid stacking ambient noise, the recordings were denoised using a 500 ms noise profile for each, mixing one original noisy clip with the denoised clips. To vary speaker amplitude outside the bubble, their signals were randomly scaled up to two times. The ground- truth signal was the sum of denoised speakers within 1.5 m, and interfering signals were rescaled to achieve an SNR between –10 dB and 5 dB for training and –5 dB to 5 dB for testing. Two datasets were generated: 60k mixtures from mannequin recordings (training only) and 20k training, 2k validation and 2k testing mixtures from human wearer recordings. The mannequin dataset included real-world ambient sounds like heating, ventilation and air conditioning noise.

[0140] The models were trained with two different loss functions. The first was a loss function based on the negative SNR, defined as given in Equation 9 of FIG. 8 where s is the target signal, ŝ is the network output signal, ∥⋅∥1 is the L1-norm (equivalently, the sum of element-wise absolute differences) and λ = 50 is a weighting factor. A combination of L1- loss and a variant of the multiresolution spectrogram loss was also used, defined as shown in Equation 10 of FIG. 8 where R is a set of STFT resolutions, k1= 20 is an L1-weighting factor and STFTRi is the STFT operation with resolution Ri. Three resolutions were used, where the fast Fourier transform length was varied over {1,024, 2,048, 512}, the step size over {120, 240, 50} and the window length over {600, 1,200, 240}. LSNR was used in the pre-training stage and LMultiResin the fine-tuning stage.

[0141] To train the model for the synthetic dataset, a single model was trained on the 30k dataset, containing 1 m, 1.5 m and 2 m bubbles, as well as 0–2 speakers inside the bubbles. The model was first trained using SNR-based loss LSNRfor 150 epochs, and then fine-tuned on the original 30k dataset using multiresolution loss LMultiRes for 100 epochs. The initial learning rate was set to 2 × 10−3. For all the models and training runs, the Adam optimizer was used, a batch size of 8 and a gradient clipping of 1. The real-world model was trained in two stages. First, the networks were pretrained on the mannequin dataset (150 epochs for Raspberry Pi and 161 epochs for Orange Pi) using data augmentations. The learning rate was increased from 2 × 10−4to 2 × 10−3over 10 epochs, held constant for 20 epochs and then reduced by 0.95 every 2 epochs. Next, the models were fine-tuned on the human wearer dataset for 200 epochs, halving the learning rate if validation loss did not decrease for 8 epochs. Pre-trained weights were used for fine-tuning with the LMultiRes loss function. Some models were trained only on the human wearer dataset using LSNRas a loss function. When applying multiple augmentations together, each was applied with 30% probability. TheDocket No. 0077145-07001 (UW 49996.02WO2) Raspberry Pi model was trained in three ways: (1) using only mannequin data, (2) using only human wearer data and (3) pre-training on mannequin data followed by fine-tuning with human data. All the strategies used LSNRas the loss function.

Claims

Docket No. 0077145-07001 (UW 49996.02WO2) CLAIMS What is claimed is:

1. A method of training a neural network to extract audio signals originating from audio sources within a threshold distance, the method comprising: generate a set of training data, the training data including audio signals generated by multiple microphones responsive to acoustic signals from audio sources in an environment and distance information to the audio sources; generate augmented training data by performing operations including: shifting the audio signals; adjusting an amplitude of the audio signals; attenuating frequencies of the audio signals; and varying a speed of the audio signals; provide the set of training data including the augmented training data to a machine learning engine; and generate a set of parameters representing a trained neural network to infer distance to an audio source.

2. The method of claim 1, wherein the set of training data includes data obtained from varying positions of the multiple microphones in the environment to represent multiple HRTFs.

3. The method of claim 1, further comprising refining the set of parameters based on further training data gathered from human wearers of multiple microphones.

4. The method of claim 1, wherein the set of training data includes audio signals generated in multiple environments including different reverberation characteristics.

5. A system comprising: a plurality of microphones configured for placement proximate a head of a user, the plurality of microphones configured to receive acoustic signals from an environment and generate corresponding input audio signals; at least one speaker; at least one processor; and at least one computer readable media encoded with instructions which, when executed by the at least one processor, cause the system to implement a trained machine learning model to receive the input audio signals and extract target audio signals corresponding to sources within the environment within a threshold distance from the headDocket No. 0077145-07001 (UW 49996.02WO2) of the user and provide the target audio signals to the at least one speaker for playback to the user.

6. The system of claim 5, further comprising a user interface configured to receive an indication of the threshold distance from the user.

7. The system of claim 6, wherein the threshold distance is provided as an embedding to the trained machine learning model.

8. The system of claim 5, wherein the machine learning model is trained on interchannel phase difference (IPD) and interchannel level difference (ILD) features.

9. The system of claim 5, wherein the machine learning model utilizes a head-related transfer function associated with the user.

10. The system of claim 5, wherein the trained machine learning model performs thresholding on the input audio signals to classify the input audio signals as closer or further than the threshold distance.

11. The system of claim 5, wherein the trained machine learning model uses frequency- dependent variations where a phase of a particular audio signal changes as a function of distance and the trained machine learning model uses a phase difference across frequencies to compute a distance to an audio source associated with the particular audio signal.

12. The system of claim 5, wherein the trained machine learning model utilizes a direct to reverberate ratio calculated based on multipath instances of the input audio signals to extract the target audio signals.

13. The system of claim 5, wherein the trained machine learning model utilizes a combination of features including IPD, ILD, head-related transfer function, thresholding, and a direct to reverberate ratio to extract the target audio signals.

14. The system of claim 5, wherein the plurality of microphones are included in ear buds.

15. The system of claim 5, wherein the plurality of microphones are positioned along a headband.

16. A method comprising: wearing a plurality of microphones and at least one speaker; selecting a threshold distance;Docket No. 0077145-07001 (UW 49996.02WO2) receiving acoustic signals from an environment at the plurality of microphones to generate input audio signals; extracting, using a trained machine learning model implemented by at least one processor, target audio signals from the input audio signals, the target audio signals corresponding to audio signals generated by sources within the threshold distance; and listening to the target audio signals played back through the at least one speaker.

17. The method of claim 16, wherein said selecting the threshold distance comprises manually adjusting a user interface element to indicate the threshold distance.

18. The method of claim 16, wherein the trained machine learning model utilizes a combination of features including IPD, ILD, head-related transfer function, thresholding, and a direct to reverberate ratio to extract the target audio signals.

19. The method of claim 16, wherein the plurality of microphones are worn proximate a head of a user.

20. The method of claim 16, further comprising transmitting transmission signals, based on the input audio signals, to a computing system configured to implement the trained machine learning model.

21. The method of claim 16, further comprising receiving, at the at least one speaker, from the computing system, the target audio signals.

22. A method comprising: receiving, from a plurality of microphones positioned proximate a head of a user, input audio signals associated with a plurality of audio sources in an environment; receiving an embedding indicative of a threshold distance from the head of the user; operating a trained machine learning model configured to extract target audio signals from the input audio signals, the target audio signals associated with one or more target audio sources of the plurality of audio sources in the environment which are within the threshold distance; and outputting the target audio signals to at least one speaker positioned proximate the head of the user.

23. The method of claim 22, wherein the trained machine learning model utilizes a combination of features including IPD, ILD, head-related transfer function, thresholding, and a direct to reverberate ratio to extract the target audio signals.Docket No. 0077145-07001 (UW 49996.02WO2) 24. The method of claim 22, wherein the trained machine learning model utilizes a head- related transfer function to increase an effective aperture size of the plurality of microphones.

25. The method of claim 22, wherein one or more of the plurality of audio sources outside the threshold distance are louder than a loudest audio source within the threshold distance.

Citation Information

Patent Citations

  • Techniques combining plural head-related transfer function (HRTF) spheres to place audio objects

    US20200382871A1

  • Method, device and software for applying an audio effect, in particular pitch shifting

    WO2021175460A1

  • Distance based sound separation using machine learning models

    WO2024006514A1