Artificial intelligence awareness modes for adjusting output of an audio device
Pre-trained machine learning models in wearable devices address the challenge of managing ambient noise by selectively adjusting audio output, providing real-time, personalized sound management with ultra-low latency and improved user awareness.
Patent Information
- Application Number
- PCT/US2025/013659
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-15
- Filing Date
- 2025-01-29
- Publication Date
- 2025-08-21
AI Technical Summary
Existing wearable audio devices struggle to manage ambient noise effectively, often isolating users from their surroundings and failing to differentiate between important and unimportant sounds, leading to disruptive adjustments in audio output.
Implementing pre-trained machine learning models in wearable devices to detect and separate spatialized sounds based on content, allowing users to selectively adjust audio output with ultra-low latency and personalized modes of awareness.
Enables real-time, personalized audio output that filters or enhances specific sounds based on user preferences, improving user experience by maintaining awareness of important sounds without manual intervention.
Smart Images

Figure US2025013659_21082025_PF_FP_ABST
Abstract
Description
ARTIFICIAL INTELLIGENCE AWARENESS MODES FOR ADJUSTING OUTPUT OFAN AUDIO DEVICEFIELD
[0001] Aspects of the disclosure generally relate to wearable devices, and, more particularly, to pre-trained artificial intelligence (Al) aware modes to enable a wearable device to manage ambient noise with ultra-low latency based on user input.BACKGROUND
[0002] Users rely on wearable audio output devices regularly for a multitude of purposes. For example, when users wake up, they may slip in earbuds to immerse themselves music, podcasts, or audiobooks during morning routines or commutes. At work, school, or the gym, in-ear audio devices offer a seamless integration with various activities, providing entertainment and / or a means of communication through hands-free calling features. Additionally, they serve as tools for relaxation, aiding in meditation sessions or providing white noise for increased concentration. The widespread use of in-ear audio output devices in various environments underscores their prevalence in modern lifestyles. It is desirable for such devices to seamlessly transition to allow a user hear various conversation and / or external noises based on the user’s preference. Accordingly, methods for facilitating user selections to filter or let through spatialized sounds using wearable audio output devices, as well as apparatuses and systems configured to implement these methods, are desired.SUMMARY
[0003] All examples and features mentioned herein can be combined in any technically possible manner.
[0004] Aspects of the present disclosure provide a method for managing ambient noise in a wearable device. The method includes detecting, by one or more feedforward microphones, one or more spatialized sounds proximate a user; separating, using at least one pre-trained machine learned (ML) model, at least a portion of the spatialized sounds, wherein the separation is basedon content of the sound; processing at least a portion of the spatialized sounds; and outputting, by the audio output device, an adaptive mix of the processed spatialized sounds.
[0005] In aspects, adjusting the output of the audio device includes detecting, by one or more feedback microphones, self-voice from the user, wherein separating, by the at least one pre-trained ML model, comprises identifying and separating at least a portion of non-user speech from the one or more spatialized sounds.
[0006] In aspects, at least one pre-trained ML model comprises a first pre-trained ML model that converts user selections into a cue, wherein the cue is passed to a second pre-trained ML model and the second pre-trained ML model identifies and separates a target sound selected by the user from the one or more spatialized sounds.
[0007] In aspects, identifying, by the second pre-trained ML model for sound separation, occurs without engaging the user in an enrolment phase.
[0008] In aspects the content comprises one or more of: wind noise, non-user speech, and user- selected sounds.
[0009] In aspects, the method further includes receiving user input regarding selection of one or more modes of awareness, wherein the processing and the outputting is further based on the selected one or more modes of awareness.
[0010] In aspects, the processing comprises adjusting mixing of any combination of self-voice, non-user speech, environmental noise, and streaming audio based on the selected more one or more modes of awareness.
[0011] In aspects, a first mode of the one or more modes uses a front facing beamformer targeting face-to-face conversation and a second mode of the one or more modes applies an omnidirectional beamformer.
[0012] In aspects, the one or more spatialized sounds comprise non-user speech, and the processing comprises muting or decreasing a gain applied to the non-user speech.
[0013] In aspects, the processing further comprises adjusting ambient noise in accordance with a user-selection.
[0014] In aspects, the adaptive mix output by the audio output device includes unprocessed ambient noise from the spatialized sounds.
[0015] In aspects, the one or more spatialized sounds comprise self-voice from the user, and the processing comprises ducking ambient sounds of the spatialized sounds.
[0016] In aspects, the at least one pre-trained ML model identifies and separates non-user speech by a target speaker from the ambient sounds.
[0017] In aspects, the ducking gradually occurs when energy levels of the self-voice exceeds an energy threshold range for a first period of time.
[0018] In aspects, the processing further comprises gradually increasing a gain applied to the ambient sounds when the energy level of the self-voice drops below the threshold range for a second period of time.
[0019] Aspects of the present disclosure provide a wearable audio output device. The wearable audio output device includes at least one memory for storing a plurality of modes of awareness, one or more feedforward microphones to detect one or more spatialized sounds proximate a user, one or more pre-trained machine learned (ML) models for separating at least a portion of the spatialized sounds, wherein the separation is based on content of the sounds, one or more circuits for processing at least a portion of the spatialized sounds, one or more speakers for outputting an adaptive mix of the processed spatialized sounds.
[0020] In aspects, the wearable audio output device further includes one or more feedback microphones for detecting self-voice from the user, wherein the separating, by the one or more pre-trained ML models, comprises identifying and separating at least a portion of non-user speech from the one or more spatialized sounds.
[0021] In aspects, the a first pre-trained ML model of the pre-trained ML models converts user selections into a cue, wherein the cue is passed to a second pre-trained ML model and the secondpre-trained ML model of the ML models identifies and separates a target sound selected by the user from the one or more spatialized sounds.
[0022] In aspects, the identifying, by the second pre-trained ML model for sound separation, occurs without engaging the user in an enrolment phase.
[0023] In aspects, the processing and outputting is based on a received user input regarding selection of one or more modes of awareness.
[0024] Two or more features described in this disclosure, including those described in this summary section, may be combined to form implementations not specifically described herein.
[0025] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG. 1 illustrates an example system, in which aspects of the present disclosure may be implemented.
[0027] FIG. 2 illustrates an exemplary wireless audio device, in which aspects of the present disclosure may be implemented.
[0028] FIG. 3 illustrates example operations performed by a wearable device worn by a user for adjusting audio output, according to certain aspects of the present disclosure.
[0029] FIG. 4 illustrates exemplary components of a target sound removal pipeline implementation performed by a wearable device for adjusting audio output, according to certain aspects of the present disclosure.
[0030] FIG. 5 illustrates exemplary components of a conversation boost pipeline implementation performed by a wearable device for adjusting audio output, according to certain aspects of the present disclosure.
[0031] FIG. 6 illustrates exemplary components of a multi-model pipeline implementation performed by a wearable device for adjusting audio output, according to certain aspects of the present disclosure.DETAILED DESCRIPTION
[0032] Certain aspects of the present disclosure provide techniques, and a device, and system implementing the techniques, for providing spatialized sound management by a wearable device. The spatialized sound management may involve using one or more pre-trained Al aware modes to output a real-time (e.g., with ultra-low latency), personalized mix of processed spatialized sounds by the wearable audio device.
[0033] Audio output devices, especially those utilizing noise cancellation, tend to isolate the user from the surrounding world, making it difficult for the user to be aware of sounds around them. Given the multitude of uses of audio output devices, at various times, a user may prefer to have differing levels of transparency of surrounding ambient noise or noise cancellation. In certain scenarios, a user may want to filter out sound based on the content of the noise. For example, a user may want to hear at least a portion of ambient sounds in a cafe that helps them study while removing noisy talkers in their vicinity. In other cases, a user may want to have the ability to cancel out ambient noises while ensuring they can hear specified noises such as a baby crying or fire alarm. In yet another case, a user may want to be able to boost speech of a target speaker talking to the user without amplifying the user’s own voice. In some cases, the user may want the wearable device’s audio level or noise cancellation to be quickly adjusted in real-time (e.g., with ultra-low latency) to respond to an important event, such as another person speaking to them, and enable a conversation with that nearby person. However, it is often cumbersome for users to manually or vocally control or adjust a setting on their headphone or take headphones off to respond to the event.
[0034] One possible solution to managing the proximate sounds and user voice is to embed sound event detection algorithms in the wearable device, so that device turns off noise cancellation or pauses audio content when an important event is detected (e.g., self-voice or a nearby sound event). However, this may be a jarring experience for the user wherein there is suddenly little to no ambient noise or transparency. Moreover, a user may want to filter out sound based on itscontent rather adjust an audio output based on, for example, a loud sound event. Further still, it may be difficult for a device to differentiate between different sounds with similar characteristics, such as differentiating between an event when someone is merely chatting nearby and when someone is attempting to talk to the user. Similarly, it may be difficult for a device to determine if a sound event comes from nearby entertainment (e.g., television, music, a podcast, etc.), which may not be important to the user, or from someone talking to the user (e.g., a family member), which may be important to the user. As a result of not being able to distinguish between when an event that is important to the user has been detected and when an event that is not important to the user has been detected, the device may not take appropriate actions in response to the detected event. For example, the device may greatly decrease the audio volume output, or even pause the audio output in response to a detected event that is not important to the user (e.g., co-workers conversing with each other), disrupting the user’s audio experience.
[0035] According to aspects of the present disclosure, a device may use at least one pre-trained ML model to identify sounds based on content of the sound and selectively augment a user’s hearing in near real time (e.g., with ultra-low latency). As described in more detail herein, a user’s hearing may be selectively augmented by, for example, boosting conversations in noisy environments, adjusting the output of nearby conversations the user does not wish to hear, and / or removing wind turbulence. As described in more detail below, a user may select one or more modes of awareness thereby enabling the user to have personalized, selective awareness.An Example System
[0036] FIG. 1 illustrates an example system 100, in which aspects of the present disclosure may be practiced. As shown, system 100 includes a wearable device 110 communicatively coupled with a computing device 120. The wearable device 110 may be configured to be worn by a user, and may be a headset that includes two or more speakers and two or more microphones, as illustrated in FIG. 1. The computing device 120 is illustrated as a smartphone or a tablet computer wirelessly paired with the wearable device 110. At a high level, the wearable device 110 may play audio content transmitted from the computing device 120. The user may use the graphical user interface (GUI) on the computing device 120 to select the audio content and / or adjust settings of the wearable device 110. In an example, the user may select desired modes of awareness as well as desired parameters associated with the modes of awareness. The wearable device 110 providessoundproofing, active noise cancellation, and / or other audio enhancement features to play the audio content transmitted from the computing device 120. According to aspects of the present disclosure, upon the determining or identifying a content of the sound, the wearable device 110 may take action in accordance with selected mode(s) of awareness of the user. The actions may include, for example, decreasing an audio volume of the wearable device 110, decreasing a noise cancellation of the wearable device 110, increasing transparency of the wearable device 110, pausing an audio output of the wearable device 110, enhancing a voice of a target speaker talking to the user, decreasing the voice from conversations occurring around the user, and / or decreasing wind turbulence.
[0037] In certain aspects, the wearable device 110 includes at least two microphones 111 and 112 to capture ambient sound. The captured sound may be used for active noise cancellation and / or event detection. For example, the microphones 111 and 112 may be positioned on opposite sides of the wearable device 110, as illustrated.
[0038] The wearable device 110 further includes hardware and circuitry including processor(s) / processing system and memory configured to implement one or more sound management capabilities or other capabilities including, but not limited to, noise canceling circuitry (not shown) and / or noise masking circuitry (not shown), body movement detecting device s / sensors and circuitry (e g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, etc.), geolocation circuitry and other sound processing circuitry. The noise cancelling circuitry is configured to reduce unwanted ambient sounds external to the wearable device 110 by using active noise cancelling (also known as active noise reduction). The sound masking circuitry is configured to reduce distractions by playing masking sounds via the speakers of the wearable device 110. The movement detecting circuitry is configured to use devices / sensors such as an accelerometer, gyroscope, magnetometer, or the like to detect whether the user wearing the wearable device 110 is moving (e.g., walking, running, in a moving mode of transport, etc.) or is at rest and / or the direction the user is looking or facing. The movement detecting circuitry may also be configured to detect a head position of the user for use in determining an event, as will be described herein, as well as in augmented reality (AR) applications where an AR sound is played back based on a direction of gaze of the user.
[0039] In an aspect, the wearable device 110 is wirelessly connected to the computing device 120 using one or more wireless communication methods including, but not limited to, Bluetooth, Wi-Fi, Bluetooth Low Energy (BLE), other radio frequency (RF) based techniques, or the like. In certain aspects, the wearable device 110 includes a transceiver that transmits and receives data via one or more antennae in order to exchange audio data and other information with the computing device 120.
[0040] The wearable device 110 is illustrated as over-the-head headphones; however, the techniques described herein apply to other wearable devices, such as wearable audio devices, including any audio output device that fits around, on, in, or near an ear (including open-ear audio devices worn on the head or shoulders of a user) or other body parts of a user, such as head or neck. The wearable device 110 may take any form, wearable or otherwise, including standalone devices (including automobile speaker system), stationary devices (including portable devices, such as battery powered portable speakers), headphones (including over-ear headphones, on-ear headphones, in-ear headphones), earphones, earpieces, headsets (including virtual reality (VR) headsets and AR headsets), goggles, headbands, earbuds, armbands, sport headphones, neckbands, or eyeglasses.
[0041] In certain aspects, the wearable device 110 is connected to the computing device 120 using a wired connection, with or without a corresponding wireless connection. The computing device 120 may be a smartphone, a tablet computer, a laptop computer, a digital camera, or other computing device that connects with the wearable device 110. As shown, the computing device 120 can be connected to a network 130 (e.g., the Internet) and may access one or more services over the network. As shown, these services can include one or more cloud services 140.
[0042] In certain aspects, the computing device 120 can access a cloud server in the cloud 140 over the network 130 using a mobile web browser or a local software application or “app” executed on the computing device 120. In certain aspects, the software application or “app” is a local application that is installed and runs locally on the computing device 120. In certain aspects, a cloud server accessible on the cloud 140 includes one or more cloud applications that are run on the cloud server. The cloud application may be accessed and run by the computing device 120. For example, the cloud application can generate web pages that are rendered by the mobile webbrowser on the computing device 120. In certain aspects, a mobile software application installed on the computing device 120 or a cloud application installed on a cloud server, individually or in combination, may be used to implement the techniques for low latency Bluetooth communication between the computing device 120 and the wearable device 110 in accordance with aspects of the present disclosure. In certain aspects, examples of the local software application and the cloud application include a gaming application, an audio AR or VR application, and / or a gaming application with audio AR or VR capabilities. The computing device 120 may receive signals (e.g., data and controls) from the wearable device 110 and send signals to the wearable device 110.
[0043] FIG. 2 illustrates an exemplary wearable device 110 and some of its components. Other components may be inherent in the wearable device 110 and not shown in FIG. 2. For example, the wearable device 110 may include an enclosure that houses an optional graphical interface (e.g., an OLED display) which can provide the user with information regarding currently playing (“Now Playing”) music.
[0044] The wearable device 110 includes one or more electro-acoustic transducers (or speakers) 214 for outputting audio. The wearable device 110 also includes a user input interface 217. The user input interface 217 may include a plurality of preset indicators, which may be hardware buttons. The preset indicators may provide the user with easy, one press access to entities assigned to those buttons. The assigned entities may be associated with different ones of the digital audio sources such that a single wearable device 110 may provide for single press access to various different digital audio sources. The preset indicators may allow the user to toggle between and select modes of awareness. In other aspects, the modes of awareness may be selected by haptic touch, such as tapping on the enclosure of the device 110.
[0045] The wearable device 110 may include a feedback sensor 111 and feedforward sensors 112. The feedback sensor 111 and feedforward sensors 112 may include two or more microphones (e.g., microphones 111, 112 as illustrated in FIG. 1) for capturing ambient sound and provide audio signals for determining location attributes of events. For example, the feedback sensor 111 may provide a mechanism for determining transmission delays between the computing device 120 and the wearable device 110. The transmission delays may be used to reduce errors in subsequent computation. The feedback sensor 111 may provide two or more channels of audio signals. Theaudio signals are captured by microphones that are spaced apart and may have different directional responses. The two or more channels of audio signals may be used for calculating directional attributes of an event of interest.
[0046] As shown in FIG. 2, the wearable device 110 includes an acoustic driver or speaker 214 to transduce audio signals to acoustic energy through audio hardware 223. The wearable device 110 also includes a network interface 219, at least one processor 221, the audio hardware 223, power supplies 225 for powering the various components of the wearable device 110, and memory 227. In certain aspects, the processor 221, the network interface 219, the audio hardware 223, the power supplies 225, and the memory 227 are interconnected using various buses 235, and several of the components can be mounted on a common motherboard or in other manners as appropriate.
[0047] The network interface 219 provides for communication between the wearable device 110 and other electronic computing devices via one or more communications protocols. The network interface 219 provides either or both of a wireless network interface 229 and a wired interface 231. The wireless interface 229 allows the wearable device 110 to communicate wirelessly with other devices in accordance with a wireless communication protocol such as IEEE 802.11. The wired interface 231 provides network interface functions via a wired (e.g., Ethernet) connection for reliability and fast transfer rate, for example, used when the wearable device 110 is not worn by a user. Although illustrated, the wired interface 231 is optional.
[0048] In certain aspects, the network interface 219 includes a network media processor 233 for supporting Apple AirPlay® and / or Apple Airplay® 2. For example, if a user connects an AirPlay® or Apple Airplay® 2 enabled device, such as an iPhone or iPad device, to the network, the user can then stream music to the network connected audio playback devices via Apple AirPlay® or Apple Airplay® 2. Notably, the audio playback device can support audio-streaming via AirPlay®, Apple Airplay® 2 and / or Digital Living Network Alliance’ s (DLNA) Universal Plug and Play (UPnP) protocols, all integrated within one device.
[0049] All other digital audio received as part of network packets may pass straight from the network media processor 233 through a USB bridge (not shown) to the processor 221 and runsinto the decoders, DSP, and eventually is played back (rendered) via the electro-acoustic transducer(s) 214.
[0050] The network interface 219 can further include Bluetooth circuitry 237 for Bluetooth applications (e.g., for wireless communication with a Bluetooth enabled audio source such as a smartphone or tablet) or other Bluetooth enabled speaker packages
[0051] Streamed data may pass from the network interface 219 to the processor 221. The processor 221 may execute instructions (e.g., for performing, among other things, digital signal processing, decoding, and equalization functions), including instructions stored in the memory 227. The processor 221 may be implemented as a chipset of chips that includes separate and multiple analog and digital processors. The processor 221 may provide, for example, for coordination of other components of the audio wearable device 110, such as control of user interfaces and the pre-processing and post-processing illustrated in FIGs. 4 - 6.
[0052] The processor 221 provides a processed digital audio signal to the audio hardware 223 which includes one or more digital-to-analog (D / A) converters for converting the digital audio signal to an analog audio signal. The audio hardware 223 also includes one or more amplifiers which provide amplified analog audio signals to the electro-acoustic transducer(s) 214 for sound output. In addition, the audio hardware 223 may include circuitry for processing analog input signals to provide digital audio signals for sharing with other devices, for example, other speaker packages for synchronized output of the digital audio.
[0053] The memory 227 can include, for example, flash memory and / or non-volatile random access memory (NVRAM). In some aspects, instructions (e.g., software) are stored in an information carrier. The instructions, when executed by one or more processing devices (e.g., the processor 221), perform one or more processes, such as those described elsewhere herein. The instructions can also be stored by one or more storage devices, such as one or more computer or machine-readable mediums (for example, the memory 227, or memory on the processor). The instructions can include instructions for performing decoding (i.e., the software modules include the audio codecs for decoding the digital audio streams), as well as digital signal processing and equalization. In certain aspects, the memory 227 and the processor 221 may collaborate in dataacquisition and real time processing with the feedback microphone 111 and feedforward microphones 112.Example Operations for Adjustins Sound Output
[0054] Aspects of the present disclosure provide techniques, including devices and system implementing the techniques, for providing spatialized sound management in a wearable device. The present disclosure may enable a user’s wearable device to provide adaptively mixed real-time sounds for selective user awareness.
[0055] In certain aspects, a wearable device may implement a pre-trained Al model to select sounds to remove or enhance based on its content and a user selection. In some cases, the sound selected by a user may be any combination of their self-voice, non-user speech, user selected sounds, environmental noises, and streaming audio. In these cases, the wearable device with the pre-trained Al model may determine when the user-specified sound content is present. The audio device may identify and separate the sounds based on the content, and output an adjusted mix of the any combination of the sounds. The wearable device may have options to select one or more user awareness modes by toggling through the selected modes on an app or by cycling through and selecting modes using buttons on the device or haptic touch on an enclosure of the device itself. In some cases, the Al model may support multiple modes of awareness, such as a target sound removal mode (where nearby conversation occurring in the vicinity of the user is adjusted such that the user may hear less of the conversation and, consequently, be less bothered by the conversation) and conversation boost mode (enhancing the voice of a target speaker without necessarily enhancing the user’s voice to allow the user to better hear a target speaker talking to the user). In some aspects, the target sound removal may have a subset hush mode, wherein a nearby noisy talker is removed. In some aspects, the Al model may also remove other environmental noises such as the sound of wind turbulence. As described herein, the Al model may be a multi-mode model, supporting any combination of multiple modes of awareness, such as, for example, the hush mode, conversation boost mode, and adjusting for wind turbulence.
[0056] In one example, the pre-trained Al model may have a target sound removal mode implemented based on user selection(s). In some cases, a user may want to sit in an environment with ambient noise, such as in a cafe. For example, a user may find the ambient noise helps a userfocus or study and the user may want the ability to hush or quiet a nearby noisy talker while maintaining the user’s ability to hear the other ambient noise. In some cases, the wearable device can use the Al model’ s hush aware mode to “hush” a noisy speaker. When the target sound removal mode with the hush subset mode is selected and the wearable device determines a nearby person is talking (e.g., far-field voice), the wearable device may separate out the speech from the nearby person, selectively suppress the non-user speech, and output a personalized, real-time adjusted audio output, wherein voice of the nearby talker is removed or the volume of the nearby talker’s speech is reduced. This approach may utilize an asymmetric window and short frames to enable low latency. This approach may help to mitigate delay and provide a real-time and personalized adjusted audio output for the user.
[0057] Another example of an Al model implementing an aware mode is boosting speech from a target speaker talking to the user as used in the conversation boost mode. In an example, a user may want to amplify or focus on speech from a speaker. In this case, the aware mode for conversation boost may be activated to enhance other people’s speech while suppressing a user’s self-speech. As such, the wearable device may determine the user’s self-voice using the feedback microphone rather than needing to engage a user in an enrollment phase. When the conversation boost mode is activated, the wearable device may determine the user’s own voice and reduce the volume of the user’s speech while also determining the target speaker’s speech and amplifying the speech from the target speaker. This approach may help mitigate delays and provide a personalized experience, as described above
[0058] FIG. 3 illustrates example operations 300 performed by the processor 221, audio hardware 223, and other components of the wearable device (e.g., the wearable device 110 of FIGs. 1-2), according to certain aspects of the present disclosure.
[0059] The operations may generally include, at block 302, detecting, by one or more feedforward microphones, one or more spatialized sounds proximate a user. In certain aspects, detecting one or more spatialized sounds proximate the user may involve at least one of measuring a sound using one or more microphones (e.g., microphones 111, 112) on the wearable device.
[0060] According to certain aspects, the operations 300 may further include, at block 304, separating, using at least one pre-trained ML model, at least a portion of the spatialized sounds,wherein the separation is based on content of the sound. The sound content may comprise of one or more of: wind noise, non-user speech, and user-selected sounds measured by the one or more microphones. In an example, the user-selected sounds may be an emergency alarm, a baby crying, a doorbell, or any sounds the user may wish to be made aware of in a mode of awareness. In some cases, the pre-processing and detection of spatialized sounds may have additional layers depending on the user selection of the aware mode and sound content to be mixed. For example, based on the selected modes of awareness, certain audio pre-processing may be selectively activated or deactivated, as illustrated in FIG. 5.
[0061] The separation of the spatialized sounds may include separating sound content based on one or more of a self-voice, speech from a target speaker, or environmental noise. In some cases, two Al models may be utilized, where one may identify / recognize the audio based on content and another may be used to separate the identified sounds. The settings for what sound content to separate may be configurable by the user through the user input interface (e.g., user input interface 216, 217) or may be configurable through the wearable device itself.
[0062] According to certain aspects, the operations 300 may further include, at block 306, processing at least a portion of the spatialized sounds. Processing a portion of the spatialized sounds may be implemented through any combination of the audio post-processing circuity as illustrated in FIG. 4-6, wherein the user has selected the sound content to be removed, enhanced, or mixed. In some cases, processing may be most effective when automatic noise cancellation (ANC) has been enabled.
[0063] According to certain aspects, the operations 300 may further include, at block 308, outputting, by the audio output device, an adaptive mix of the processed spatialized sounds. In some cases, the adaptive mix of the processed audio is output in real-time or with ultra-low latency.
[0064] According to certain aspects, the operations 300 may further include receiving user input regarding selection of one or more modes of awareness, wherein the processing and the outputting is further based on the selected one or more modes of awareness. As described above, the wearable device may receive input from a user to select a mode of awareness such as target sound removal, conversation boost, or multi-task mode. In some cases, a user may select by way of a toggle, slider, or other method on an app or directly on the wearable device. In certain aspects,a user could elect to remove other noisy talkers while maintaining ambient noise, reduce their own voice during a conversation, or remove environmental noises such as wind. Additionally, in some cases, a user may be able to select sound content to still hear such as fire alarms or babies crying while filtering out other sound content. Additionally, the user may determine what percentage of any specific sound content to filter out using for example, a slider on an app or through voice comments or haptic touch on the audio output device.
[0065] According to certain aspects, the processing at 306 may further include adjusting mixing of any combination of self-voice, non-user speech, environmental noise, streaming audio, and any other user selected audio based on the selected more one or more modes of awareness. In some cases, the audio output mix will be adjusted using the personalized selections discussed above.
[0066] According to certain aspects, the audio output device comprises a first pre-trained ML model that converts user selections into a cue, wherein the cue is passed to a second pre-trained ML model and the second pre-trained ML model identifies and separates a target sound selected by the user from the one or more spatialized sounds. In some aspects, there are multiple Al aware modes available to the user that may be able to be layered on to each other. The user may have an option to select conversation boost, wherein their own voice will be recognized and separated out by the second ML model. Furthermore, in some cases, the second ML model(s) for sound separation may perform the identification and / or separation without engaging the user in an enrollment phase, wherein the user engages in process to set up the device to recognize their voice (e.g., a user will not need to record speech samples).
[0067] In some cases, different beamformers may be used depending on the selected modes of awareness. According to certain aspects, the audio output device includes a beamformer targeting face-to-face conversation and an omnidirectional beamformer targeting larger conversations where people are distributed around the user. For example, it may be most effective for the device to have a front-facing beamformer when a user is looking at and engaging in a face-to-face conversation. In another example, it may be most effective for the device to utilize an omnidirectional beamformer for situations in which the user is engaged in a roundtable conversation format. In some cases, omnidirectional beamformers may have the advantage ofproviding an immersive audio in headphones that automatically duck when speaking. In an example implementation, beamformers that produce binaural output are used as they also provide the wearer with binaural cues that are enhance the conversation experience.
[0068] According to certain aspects, the device may implement a target sound removal mode of awareness, where the one or more spatialized sounds comprise non-user speech. In response to detecting the non-user speech, the device may mute or decrease a gain applied to the non-user speech. Furthermore, according to certain aspects, the processing 306 further comprises adjusting ambient noise in accordance with a user-selection. According to certain aspects, the outputting 308 may further include outputting unprocessed ambient noise from the spatialized sounds. As discussed above, for some users it may be jarring to adjust from transparency to noise cancellation. Thus, in this example, the processing 308 may be adjusted based on user selection wherein ambient noise may remain unprocessed while speech from other talkers are removed or ducked.
[0069] According to certain aspects, the operations 300 may further include detecting, by one or more feedback microphones, self-voice from the user. At 304, the separating may include separating, by the at least one pre-trained ML model, the self-voice from the one or more spatialized sounds. For example, where a user selects the conversation boost mode, the device will enable a user to remove or reduce their own voice through feedback microphone detection while maintaining external voices.
[0070] According to certain aspects, the one or more spatialized sounds may include self-voice from the user and a nearby voice. The processing 306 may include ducking ambient sounds of the spatialized sounds, adjusting a noise cancellation of the wearable device, adjusting a transparency of the wearable device, pausing an audio output of the wearable device, or outputting a notification sound from the wearable device. For example, when the wearable device measures a nearby voice (e g., determining the occurrence of a first event), the wearable device may decrease the audio volume (e.g., to 32 decibels (dB)) and decrease noise cancellation of the wearable device, to facilitate user awareness and interaction with the user’s environment.
[0071] According to certain aspects, in the conversation boost mode, the audio device gradually ducks at least some ambient noise when energy levels of the self-voice or the non-user speech exceeds an energy threshold range for a first period of time. In aspects, the device graduallyincreases a gain applied to the ambient sounds when the energy level of the self-voice and the energy level of the non-user speech drops below the threshold range for a second period of time. For example, where the energy level of a user’s voice or target voice exceeds a threshold for a given period of time, the conversation boost awareness mode may be triggered because the user is likely to be engaged in conversation with the target speaker. Likewise, in this example once either or both of the voices drops below the threshold for a certain amount of time, then the awareness mode may gradually increase the gain of the ambient noise back up to normal for user experience purposes because the conversation may have stopped.
[0072] FIG. 4 illustrates example components 400 of a wearable device 110, according to aspects of the present disclosure. The components 400 provide more details of the components illustrated in FIG. 2. For example the components 400 may be part of the processor 221, audio hardware 223, feedback sensor 111 and / or feedforward sensor 112 of the wearable device 110. The processing pipelines illustrated using components 400 is one exemplary embodiment for implementing a user selected, pre-trained Al aware mode.
[0073] According to certain aspects, the processing pipeline using components 400 are representative of the target sound removal mode of awareness. The outside microphones 403 receive audio input from the outside environment. As described above, audio input may comprise of self-voice, non-user speech, user selected sounds, and environmental noise.
[0074] Processing pipeline 404 may implement pre-processing steps. The pre-processing pipeline for the target sound removal implementation may include a beamformer and wind mixer 405. In an example, the wind mixer may mitigate the impact of wind noise through adaptively mixing or switching the beamformer output and raw mic signal(s). In some aspects, the wind mixer prevents a failure mode (e.g. amplification of wind noise rather than the desired target source) of the beamformer when wind is present. In some examples, the necessity and benefit of the wind mixing approach may depend on the acoustic architecture and design of the beamformer. As discussed above, beamformers may be focused in different directions depending on the mode of awareness. For example, in the target sound removal mode with the subset hush mode of awareness, the beamformer 405 may be omnidirectional. As previously discussed, omnidirectionalbeamformers may be most effective for situations in which the user is engaged in a roundtable conversation format or to pick up non-user speech generally occurring around the user.
[0075] Processing pipeline 404 may also include a left Asymmetric Window (“Asymwindow”) and Fast Fourier Transform (“FFT”) 407 and a right Asymmetric Window (“Asymwindow”) and Fast Fourier Transform (“FFT”) 409. In an example, the combination of the asymwindow and FFT and short frames serves as a benefit to enable ultra-low latency of the awareness mode. The ultra-low latency may provide the user with a real-time mixed audio output. Processing pipeline 404 may also include a function to average 411 the left and right Asymwindow and FFTs.
[0076] After determining an average 411 wherein the two channels are combined, processing pipeline 404 may also include Mel Spectrogram and Normalization operations 413 to convert the frequencies to the Mel spectrum and normalize.
[0077] Components 400 may also include a ML model 417. In the target sound removal mode of awareness, the ML model 417 may separate out sound content including ambient environmental noises from sound content including a non-user speaker, such as a noisy talker in the vicinity of the user. In some aspects, the components 400 may not receive any user cue input. In some examples, without user cue input the ML model 417 may have a fixed set of sources it can extract / separate (e.g. cancel voices and keep all other sound content hushed). In some aspects, the ambient noise may remain unprocessed, depending on user selection for sound content to adjust.
[0078] Post-processing pipeline 418 begins with conversion operations occurring at 419 which converts the mask from the Mel spectrum to the FFT binaural output. In some examples, the conversion may derive speech energy indicators based on speech energy to send to near term nontarget average for left and right channels at 421. Additionally, in some examples, this conversion may produce a common mask for the left and right channels to send to the mask application operation 425. In some embodiments, the near term non-target average for the left and right channels 421 may average the magnitude only for the left and right channels and send the output to the operations for removing near term non-speech average FFT for left and right channels 423. In some cases 423 may serve as a smoothing stage which may ensure the signal is not jumping too much between a user’s two ears. The mask application operation 425 may serve as a treatment toconvert to the spectral domain. The conversion operations 419 may also send the derived speech energy indicators based on speech energy to operations that utilize an energy-based threshold to determine a gain value based on the ML output 437.
[0079] Post-processing pipeline 418, the wearable device may include a mixing operation left 427 and mixing operation right 429, wherein the mixing occurs with smooth gain. In some embodiments, after the mask application operation 425, a user parameter may be taken in to determine the level of which a user wants to mix the audio. At Inversion Asymmetric Window (“InvAsym Window”) left 431 and Asymmetric Window (“InvAsym Window”) right 433, the signal may be subsequently inversely transformed to return to the time domain.
[0080] Post-processing pipeline 418, the wearable device may include a limiter 435, which may avoid clipping and signal anomaly. The output of post-processing pipeline 418 may include a two-channel output to maintain spatial awareness for the user.
[0081] FIG. 5 is a diagram illustrating example components 500 of a wearable device 110, according to aspects of the present disclosure. The components 500 provide more details of the components illustrated in FIG. 2. For example the components 500 may be part of the processor 221, audio hardware 223, feedback sensor 111 and / or feedforward sensor 112 of the wearable device 110. The processing pipeline illustrated using components 500 is one exemplary embodiment for implementing user selected, pre-trained Al aware modes.
[0082] According to certain aspects, the processing pipeline using components 500 are representative of the conversation boost mode of awareness. In some aspects, the conversation boost mode of awareness may be layered on to the target sound removal mode of awareness processing pipelines illustrated in FIG. 4. Specifically, a pre-processing pipeline 514 may be introduced. As such, in some cases, based on the content of the sounds and the selected modes of awareness, processing pipelines may be selectively used.
[0083] The outside microphones 503 receive audio input from the outside environment. As described above, audio input may comprise of self-voice, non-user speech, environmental noise, and / or any combination of user selected sounds.
[0084] Processing pipeline 504, including components 505, 507, 509, 511, and 513 may be similar to processing pipeline 404 including components 405, 407, 409, 411, and 413.
[0085] In addition to the pre-processing pipeline 504, the components 500 may also include an audio pre-processing pipeline 514. Relative to FIG. 4, this additional pre-processing pipeline 514 may be layered on to enable a conversation boost awareness mode in addition to the target sound removal awareness mode illustrated in FIG. 4. Specifically, the pre-processing pipeline 514 identifies a user’s voice through microphones. Pre-processing pipeline 514 includes a far-end signal 515 and a feedback microphone 517. The far-end signal 515 may be sent to the Asymwindow and FFT 519. The feedback microphone may be sent to the Asymwindow and FFT 521. In some aspects, the feedback microphone may enable control over self-voice pass-through. Subsequently, acoustic echo cancellation occurs at 523 and Mel Spectrogram and Normalization operations 525 to convert the frequencies to the Mel spectrum and normalize. The pre-processing pipeline 514 serves to enable conversation boost operations by determining the user’s voice, without requiring an enrollment phase. This allows for a user’s own voice to be suppressed while enhancing or amplifying another target voice.
[0086] Components 500 may also include a ML model 527. In the conversation boost awareness mode, the ML model 527 may separate out the own user’s voice, a target speaker’s voice, and ambient noise. In some aspects, other ambient noise may remain unprocessed by pipeline 528, depending on user selection for sound content to adjust.
[0087] Post-processing pipeline 528 may be triggered after the ML model 527 identifies and separates the various sounds based on content of the sound. Components 529, 531, 533, 535, 537, 539, 541, 543, 545, and 547 of the post-processing pipeline 528 perform similar operations as the components illustrated in post-processing pipeline 418 illustrated in FIG. 4.
[0088] FIG. 6 is a diagram illustrating example components 600 of a wearable device 110, according to aspects of the present disclosure. The components 600 provide more details of the components illustrated in FIG. 2. For example the components 600 may be part of the processor 221, audio hardware 223, feedback sensor 111 and / or feedforward sensor 112 of the wearable device 110. The processing pipeline illustrated using components 600 is one exemplary embodiment for implementing user selected, pre-trained Al aware modes.
[0089] According to certain aspects, the processing pipeline using components 600 are representative of a multi-task model, wherein a user has selected a target sound removal mode of awareness and removal of wind turbulence noise. Relative to the pipeline processing performed by components 400 illustrated in FIG. 4, the components 600 may additionally select content to remove such as wind turbulence. As with FIG. 5, based on the content of the sounds and the selected modes of awareness, certain processing pipelines (or components) may or may not be used.
[0090] The pre-processing pipeline 604, using inputs from 603 and components 605, 607, 609, 611, and 613 perform similar operations to the pre-processing pipelines 404 in FIG. 4 and 504 in FIG. 5.
[0091] The components 600 may also include a first ML model 615 to convert user selections into a cue which provides the cue into a second ML model 617. The second ML model 617 separates the audio sources based on content of the sounds. In certain aspects, the first ML model 615 may serve as a target source conditioner, wherein the first ML model 615 takes in user input instead of the same audio stream as the second ML model 617 source separator. In other words, in some examples, the first ML model 615 may behave as an interface between the user and the target source extraction system (e.g. the second ML model 617). In an example, the first ML model 615 may receive user input, such as a text or voice command. In some aspects, the user input may be a list of sources, wherein the first ML model 615 may map discrete sources into a cue vector. In other examples, the input may be voice commands, wherein the first ML model 615 may behave as an ASR system. In other examples, the input may be sampled fragments of unwanted sounds, wherein the first ML model 615 may convert the audio into a cue for the second ML model 617. In yet another example, the input may be text commands or a cue from an LLM-style virtual assistant, wherein the first ML model 615 may convert the audio into a cue for the second ML model 617. In an example, the first ML model 615 identifies any one or more of the own user’s speech, speech from the target speaker, and wind noise. In an example, the first ML model 615 output may be in an encoded condition. In certain aspects, the first ML model 615 may act to convert the user input into a cue for the second ML model 617. The second ML model 617 separates the identified sounds based on content. In some aspects, other ambient noise may remain unprocessed, depending on user selection for sound content to adjust. In an exemplary case, thefixed sources passed to the second ML model 617 may be more relaxed than in those described in Figs. 4-5. Furthermore, in aspects, the second ML model 617 may allow for separation of any set of sources requested by the user.
[0092] Post-processing state 618 including components 619, 621, 623, 625, 627, 629, 631, 633, 635, and 637 may perform similar operations as post-processing states 418 in FIG. 4 and 528 in FIG. 5.
[0093] In some cases, a feedback microphone may be used to control the self-voice pass- through; however, it is not strictly necessary.
[0094] It is noted that the processing related to ambient noise management as discussed in aspects of the present disclosure may be performed natively in the wearable device, by the computing device, or a combination thereof.Additional Considerations
[0095] It is noted that, descriptions of aspects of the present disclosure are presented above for purposes of illustration, but aspects of the present disclosure are not intended to be limited to any of the disclosed aspects. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described aspects.
[0096] In the preceding, reference is made to aspects presented in this disclosure. However, the scope of the present disclosure is not limited to specific described aspects. Aspects of the present disclosure can take the form of an entirely hardware aspect, an entirely software aspect (including firmware, resident software, micro-code, etc.) or an aspect combining software and hardware aspects that can all generally be referred to herein as a “component,” “circuit,” “module” or “system.” Furthermore, aspects of the present disclosure can take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied thereon.
[0097] Any combination of one or more computer readable medium(s) can be utilized. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, anelectronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a computer readable storage medium include: an electrical connection having one or more wires, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable readonly memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the current context, a computer readable storage medium can be any tangible medium that can contain, or store a program.
[0098] The flowchart and block diagrams in the Figures illustrate the architecture, functionality and operation of possible implementations of systems, methods and computer program products according to various aspects. In this regard, each block in the flowchart or block diagrams can represent a module, segment or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations the functions noted in the block can occur out of the order noted in the figures. For example, two blocks shown in succession can, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. Each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
Claims
WHAT IS CLAIMED IS:
1. A method for adjusting output of an audio output device, comprising: detecting, by one or more feedforward microphones, one or more spatialized sounds proximate a user; separating, using at least one pre-trained machine learned (ML) model, at least a portion of the spatialized sounds, wherein the separation is based on content of the sound; processing at least a portion of the spatialized sounds; and outputting, by the audio output device, an adaptive mix of the processed spatialized sounds.
2. The method of claim 1, further comprising: detecting, by one or more feedback microphones, self-voice from the user, wherein separating, by the at least one, pre-trained ML model comprises identifying and separating at least a portion of non-user speech from the one or more spatialized sounds.
3. The method of claim 2, wherein the at least one pre-trained ML model comprises a first pre-trained ML model that converts user selections into a cue, wherein the cue is passed to a second pre-trained ML model and the second pre-trained ML model identifies and separates a target sound selected by the user from the one or more spatialized sounds.
4. The method of claim 3, wherein the identifying, by the second pre-trained ML model for sound separation, occurs without engaging the user in an enrolment phase.
5. The method of claim 1, wherein the content comprises one or more of: wind noise, nonuser speech, and user-selected sounds.
6. The method of claim 1, further comprising: receiving user input regarding selection of one or more modes of awareness, wherein the processing and the outputting is further based on the selected one or more modes of awareness.
7. The method of claim 6, wherein the processing comprises adjusting mixing of any combination of self-voice, non-user speech, environmental noise, and streaming audio based on the selected more one or more modes of awareness.
8. The method of claim 6, wherein a first mode of the one or more modes uses a front facing beamformer targeting face to face conversation and a second mode of the one or more modes applies an omnidirectional beamformer.
9. The method of claim 1, wherein: the one or more spatialized sounds comprise non-user speech, and the processing comprises muting or decreasing a gain applied to the non-user speech.
10. The method of claim 9, wherein the processing further comprises adjusting ambient noise in accordance with a user-selection.
11. The method of claim 9, wherein the adaptive mix output by the audio output device includes unprocessed ambient noise from the spatialized sounds.
12. The method of claim 1, wherein: the one or more spatialized sounds comprise self-voice from the user, and the processing comprises ducking ambient sounds of the spatialized sounds.
13. The method of claim 12, wherein: the at least one pre-trained ML model identifies and separates non-user speech by a target speaker from the ambient sounds.
14. The method of claim 12, wherein the ducking gradually occurs when energy levels of the self-voice exceeds an energy threshold range for a first period of time.
15. The method of claim 14, wherein the processing further comprises gradually increasing a gain applied to the ambient sounds when the energy level of the self-voice drops below the threshold range for a second period of time.
16. A wearable audio output device comprising: at least one memory for storing a plurality of modes of awareness; one or more feedforward microphones to detect one or more spatialized sounds proximate a user; one or more pre-trained machine learned (ML) models for separating at least a portion of the spatialized sounds, wherein the separation is based on content of the sounds; one or more circuits for processing at least a portion of the spatialized sounds; and one or more speakers for outputting an adaptive mix of the processed spatialized sounds.
17. The wearable audio output device of claim 16, further comprising: one or more feedback microphones for detecting self-voice from the user, wherein the separating, by the one or more pre-trained ML models, comprises identifying and separating at least a portion of non-user speech from the one or more spatialized sounds.
18. The wearable audio output device of claim 17, the a first pre-trained ML model of the pre-trained ML models converts user selections into a cue, wherein the cue is passed to a second pre-trained ML model and the second pre-trained ML model of the ML models identifies and separates a target sound selected by the user from the one or more spatialized sounds.
19. The wearable audio output device of claim 18, wherein the identifying, by the second pre-trained ML model for sound separation, occurs without engaging the user in an enrolment phase.
20. The wearable audio output device of claim 16, wherein the processing and outputting is based on a received user input regarding selection of one or more modes of awareness.
Citation Information
Patent Citations
Context aware hearing optimization engine
US20180247646A1
Content-based audio stream separation
US20190206417A1
Fully customizable ear worn devices and associated development platform
US20230300532A1
Cited By
Automatic keyword pass-through system
US12621598B2