Known-speaker-based and conditional selective pass-through for wearable audio components with active noise control
A classification-based selective pass-through system using neural networks in ANC systems allows important ambient sounds to be detected and passed through, addressing the challenge of user interaction in conventional ANC systems, enhancing user engagement and comfort.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional active noise control (ANC) systems in wearable audio components struggle to distinguish and allow through ambient sounds that are important for user interaction, such as second-party-initiated conversations or critical alerts, requiring manual intervention by the user to disable ANC.
Implementing a classification-based selective pass-through system using artificial neural networks to detect and classify attention-seeking audio, allowing selective pass-through of important ambient sounds while maintaining ANC functionality.
Enables users to engage in conversations and hear critical alerts without removing the wearable audio components or disabling ANC, improving user comfort and awareness.
Smart Images

Figure US2024062385_09042026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 094021-1474769KNOWN-SPEAKER-BASED AND CONDITIONAL SELECTIVE PASSTHROUGH FOR WEARABLE AUDIO COMPONENTS WITH ACTIVE NOISE CONTROLCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of and priority to Indian Provisional Application No. 202441075217, filed on October 4, 2024, and titled ‘KNOWN-SPEAKER-BASED AND CONDITIONAL SELECTIVE PASS-THROUGH FOR WEARABLE AUDIO COMPONENTS WITH ACTIVE NOISE CONTROL,’' the content of which is herein incorporated by reference in its entirety for all purposes.BACKGROUND
[0002] Active noise control (ANC) is a common feature of headsets and earbuds. It operates by generating an anti-noise signal via a speaker that is approximately equal in magnitude, but opposite in phase to the ambient sound (e.g., ambient noise and other sounds in the vicinity). The ambient sound and anti-noise signal cancel each other acoustically, allowing the user to hear only a desired audio signal. Typically, signal processing in ANC includes two paths: an ambient sound signal from a reference microphone is taken as the input of a feed-forward ANC filter (FFANC); and an error microphone signal is taken as the input of a feedback ANC filter (FBANC).
[0003] When a user is listening to music or other desired audio through a wearable audio component (e.g., earbuds or on-ear headphones), ANC works to suppress any ambient sound. However, in some instances, ambient sound that is intended for the user can be very important to the user’s connectivity with others. For example, while a user is listening to music with ANC. it may be very difficult for an attention seeker to get the user’s attention, such as to alert the user of something important and / or to engage the user in a desired conversation. Conventionally, in such instances, the attention seeker gets the user’s attention by tapping the user on the shoulder, gesturing in front of the user’s face, or the like; and the user can then manually disable the ANC, pause the desired audio, and / or remove the wearable audio component.SUMMARY
[0004] Systems and methods are described herein for classification-based selective pass-through for wearable audio components with active noise control (ANC). Embodiments receive ambient noise with a first artificial neural network pre-trained to classify attention-seeking audio (ASA) components present in the ambient noise. A conditional signal is output based at least on the classification. For example, the first pre-trained network generates an embedded space representation (ESR) characterizing the ASA component, and the conditional signal is based onthe ESR. A second artificial neural network is pre-trained to generate a pass-through audio output based on the same ambient audio and the conditional signal, so that the pass-through audio output corresponds to the ASA component. The pass-through audio output is delayed past an ANC processing window so that essentially all the ambient noise except for the ASA component is cancelled by the ANC.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] A further understanding of the nature and advantages of various embodiments may be realized by reference to the following figures. In the appended figures, similar components or features may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.
[0006] FIG. 1 shows an audio management system for integration in a wearable audio component (WAC), according to embodiments described herein.
[0007] FIG. 2 shows a simplified diagram of a conventional active noise cancellation (ANC) environment for added context.
[0008] FIGS. 3 A and 3B show a wearable audio environment including a pair of WACs.
[0009] FIG. 4 shows a block diagram of an automated attention handling (AAH) system, according to embodiments described herein.
[0010] FIG. 5 shows a block diagram of another AAH system, focused on classification-based selective pass-through (CSPT) embodiments, according to embodiments described herein.
[0011] FIG. 6 shows a simplified block diagram of stages of an illustrative implementation of a name / conversation detector and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein.
[0012] FIGS. 7A and 7B show a first training stage and a second training stage, respectively, for training a machine learning model as an attention seeking audio detector and classifier (ASADC) model, according to various embodiments.
[0013] FIG. 8 shows an embodiment of a CSPT controller, according to embodiments described herein.
[0014] FIG. 9 shows a flow diagram of an illustrative method for AAH in a WAC, according to embodiments described herein.
[0015] FIG. 10 provides a schematic illustration of an illustrative computational system that can implement various system components and / or perform various steps of methods provided by various embodiments.
[0016] FIG. 11 shows a block diagram of an illustrative directional selective pass-through environment, according to embodiment described herein.
[0017] FIG. 12 shows an illustrative embodiment of a known speaker detector implemented by a machine learning model, according to embodiments described herein.
[0018] FIG. 13 shows a block diagram of an illustrative speaker embedded space representation (ESR) generation environment, according to embodiments described herein.
[0019] FIG. 14 shows a flow diagram of a method for automated attention handling (AAH) for a wearable audio component (WAC), including known speaker detection, according to embodiments described herein.
[0020] FIG. 15 shows a block diagram of an illustrative configurable AAH environment, according to some embodiments described herein.DETAILED DESCRIPTION
[0021] When a user is listening to music or other desired audio through a wearable audio component (e.g., earbuds or on-ear headphones), active noise control (ANC) works to suppress any ambient sound. However, in some instances, ambient sound that is intended for or otherwise important to the user can be very important to the user’s connectivity with the world. For example, although the user desires to suppress undesirable ambient sound, the user may still desire to be able on occasion to enter into desired conversations and / or to hear certain relevant or critical sounds. In general, three categories of "desired" audio are considered: first-party-initiated conversations; second-party-initiated conversations; and alert audio.
[0022] In first-party-initiated conversations, the user desires to start a conversation and may begin by trying to get someone’s attention. In such cases, some conventional ANC systems are adapted to detect that the user has begun speaking (e.g., by detecting the user’s speech via a beamforming microphone directed to the user’s mouth, accelerometer, or combination thereof), and the ANC system can turn off, switch to transparency mode, pause audio playback, etc. in response to detecting the user’s speaking. Because it tends to be relatively easy for the ANCsystem to distinguish the user’s own speech from ambient sound, such approaches tend to be effective for first-party-initiated conversations.
[0023] In second-party-initiated conversations, however, a second-party attention seeker is trying to get the user’s attention, and the attention seeker’s voice may be difficult to distinguish from other ambient sound. For example, while a user is listening to music with ANC, it may be very difficult for an attention seeker to get the user’s attention, such as to alert the user of something important and / or to engage the user in a desired conversation. Conventionally, in such instances, the attention seeker gets the user’s attention by tapping the user on the shoulder, gesturing in front of the user’s face, or the like. Before the user can participate in the conversation, the user conventionally must notice the interruption and then manually disable the ANC, pause the desired audio, remove the wearable audio component, etc.
[0024] In addition to conversations, users may desire to have effective ANC without the fear of missing important audible information. For example, a user may wish to get the benefits of good ANC (i.e., canceling of unwanted noise), while still being able to hear critical sounds, like their baby cry ing, a kitchen timer, oncoming traffic, or the like. From the user’s perspective, such audio is used to alert the user of something that the user wants to be alerted to, which maybe something critical or something non-critical. Such audio is referred to herein, accordingly, as "‘alert audio,” and more particularly as '‘critical alert audio” and “non-critical alert audio.”
[0025] Indeed, many users of wearable audio components enjoy the feeling of being in their “bubble” and the ability7to focus on their media that comes with effective ANC. However, as ANC continues to improve, the same users often feel increasingly unaware, not present, and fearful about missing out. Embodiments described herein seek to provide users with the ability to better stay aware and engage in desired conversations, while being able to continue wearing their wearable audio components and otherwise to take advantage of ANC. This can provide several benefits, including helping to improve user comfort and ear health.
[0026] Embodiments described herein are concerned with second-party-initiated conversations. As used herein, the term “user” refers to a wearer of a wearable audio component (i.e.. the first party). The term “attention seeker” is used herein generally to refer to any ambient party7trying to get the user’s attention while the user is wearing the wearable audio component (and presumably is listening to desired audio with ANC turned on). Typically, the attention seeker is a person.However, the attention seeker can also be a computational platform with a deterministic manner of seeking the user’s attention, such as a smart speaker programmed to call out the user’s name. Theterm “wearable audio component,'’ or “WAC” is used herein to generally refer to earbuds, on-ear headphones, over-ear headphones, or any type of wearable audio output device that includes ANC.
[0027] In general, at least four categories of audio are of interest herein. The term “desired audio,” or “media audio” is used herein to generally refer to any recorded or streaming audio signal that is being played to the user through the WAC, such as music, an audiobook, a podcast, a radio broadcast, a live event broadcast, an audio portion of audiovisual media (e.g., a video, movie, etc.) playing on a connected media device (e.g., a smartphone, laptop, etc.), etc. The term “ambient sound,” “ambient noise,” or “ambient audio” is used herein to generally refer to any audio in the vicinity of the WAC, other than the desired audio. It is generally the goal of the ANC system to suppress as much of the ambient sound as possible. The term “attention seeking audio,” or “AS audio,” is used herein to generally refer to any audio originating from an attention seeker and / or to alert audio detected while a user’s ANC system is active. AS audio will naturally be part of the ambient sound; it is audio that the ANC system would seek to cancel or reduce, if not for embodiments described herein. Finally, the term “playback audio” is used herein to generally refer to the audio signals played out to the user (e.g., reaching the user’s eardrum) and fed back into the ANC system. ANC systems typically use both the playback audio and the ambient audio to determine what to cancel.
[0028] FIG. 1 shows an audio management system 100 for integration in a wearable audio component (WAC), according to embodiments described herein. As illustrated, the audio management system 100 can include an active noise control (ANC) system 140, an attention handling system (AHS) 150. and an audio processing system 160. In general, the purpose of the WAC is to deliver desired audio 165 to a user’s ear or ears via one or more ear speakers, such as speaker 105. Embodiments of the audio processing system 160 are designed to process the desired audio 165 for output to the user. For example, the audio processing system 160 can include amplifiers, filters, and / or other audio components; and / or any other suitable components for receiving, processing, and / or outputting the desired audio 165.
[0029] Typically, while listening to the desired audio 165, the user is also in the presence of ambient audio 155. When in its ambient sound suppression mode, the ANC system 140 seeks to suppress as much of the ambient audio 155 as possible to enhance the user's experience of listening to the desired audio 165. As illustrated, the ANC system 140 includes a feed-forward ANC (FFANC) filter 120, a feedback ANC (FBANC) filter 125, a summer 130, and an ANC output control block 135. The ANC system 140 is also coupled with the speaker 105 and at least a reference microphone 110 and an error microphone 115. The reference microphone 110 issometimes referred to as the feed-forward microphone, and the error microphone 115 1s sometimes referred to as the feedback microphone. Reference to a "microphone" can include a single microphone or a group of microphones (e.g., a microphone array).
[0030] Embodiments of the speaker 105 generally convert an electrical audio signal into sound waves that are delivered to the ear of the wearer of the wearable audio component. Embodiments of the reference microphone 110 can be an omnidirectional microphone typically integrated with an outer casing of the wearable audio component. The reference microphone 110 generally captures at least the ambient audio 155 around the WAC, which is delivered as a reference audio signal (illustrated as x(n)) to the FFANC filter 120. Embodiments of the error microphone 115 are ty pically integrated with the inner casing of the w earable audio component to be positioned inside the ear canal or very close to it when the wearable audio component is being worn. The error microphone 115 captures the audio that reaches the eardrum, which includes the desired audio signal and any remaining ambient sound after suppression. The error microphone 115 outputs an error signal (illustrated as e(n)) to the FBANC filter 125.
[0031] The illustrated ANC system 140 includes a feed-forward noise control path and a feedback noise control path. The feed-forward noise control path includes the FFANC filter 120, which is a digital or analog filter designed to process the audio signal from the reference microphone 110. The FFANC filter 120 applies a specific frequency response to x(n) to adaptively cancel out noise. The specific frequency response is produced by continuously adjusting coefficients of the FFANC filter 120 to minimize the difference between the desired audio signal and the reference signal. The output of the FFANC filter 120 is illustrated as y (n).
[0032] The feedback noise control path includes the FBANC filter 125, which is a digital or analog filter designed to process the audio signal from the error microphone 115. The FBANC filter 125 applies a specific frequency response to e(n), and continuously adjusts coefficients of the FBANC filter 125 to minimize the difference between the desired audio signal and remaining ambient sound in the signal that reaches the eardrum. The output of the FFANC filter 120 is illustrated as y2(n)- In general, both the FFANC filter 120 and the FBANC filter 125 can adapt their respective filters (e.g., their coefficients) in real-time to a changing audio environment. For example, filter coefficients are iteratively adjusted using least mean squares (LMS), normalized LMS (NLMS). and / or other suitable adaptation algorithms.
[0033] Embodiments of the summer 130 combine the filtered output signals from the FFANC filter 120 and the FBANC filter 125. For example, the summer 130 calculates a sum of these signals. If tuned properly, the output of the summer 130 is an “anti -noise” signal that closelyrepresents the ambient sound at opposite polarity. Embodiments of the ANC output control block 135 control how and / or whether the anti-noise signal is output by ANC system 140. In some implementations, the ANC output control block 135 includes an amplifier to provide a controllable amount of gain (G) to the signal at the output of the summer 130, resulting in an output signal, y(n) = G(y (n~) + y2(n))- hr effect, the ANC gain block 135 adjusts the overall amplitude (i.e., corresponding to volume) of the combined filtered signal at the output of the summer 130. The output signal is sent to the speaker 105. In some implementations, as illustrated, the desired audio 165 can also be mixed in (e.g., by mixer 145) prior to sending the output to the speaker 105, such that what reaches the eardrum is almost entirely the desired audio signal with minimal ambient sound. Alternatively, the desired audio 165 is mixed into the output signal at the summer 130, such that the output of the ANC system 140 is an audio signal that is mostly the desired audio 165 with minimal residual ambient audio 155.
[0034] Embodiments of the ANC output control block 135 control the operating mode of the ANC system 140. For example, as described herein, the ANC system 140 can operate selectively in at least an active mode (i.e., an ambient sound suppression mode) or a conversation mode. Some implementations of the conversation mode correspond to an inactive mode (i.e., the ANC system 140 is turned off) or a transparency mode. Other implementations of the conversation mode are configured to pass through conversationally relevant audio from the ambient audio 155, while continuing to perform ANC functions to suppress other portions of the ambient audio 155. In some such implementations, a bandpass or notch filter is used to segregate out a range of frequencies typical for human speech and to treat the segregated audio as conversationally relevant audio. As one example, a filter can pass through portions of the ambient audio 155 only in the range of 75 to 300 Hertz and to suppress higher and lower frequency components of the ambient audio 155; thereby continuing to filter out white noise and other portions of ambient audio 155 that can interfere with a user’s abi 1 i ty to hear the passed-through conversationally relevant audio. Similarly, some implementations continue to pass through some desired audio 165 (e.g., at a reduced volume) while in conversation mode.
[0035] As described herein, embodiments of the AHS system 150 seek to detect and classify attention seeking (AS) audio 157 within the ambient audio 155 while the user is wearing the WAC and is listening to desired audio 165 with the ANC system 140 in the active mode. For example, the AHS system 150 listens for presence of alert audio and / or audio from a second-party attention seeker trying to get the attention of a user (second-party-initiated (SPI) conversation audio). Different embodiments can provide attention handling features based on detecting and responding to the detection of one or more types of such AS audio 157. As described herein, someembodiments of the AHS system 150 can additionally or alternatively provide attention handling features based on detecting and responding to the detection of first-person (FP) audio. For example, embodiments can detect that the user has begun talking and can provide features, accordingly.
[0036] Notably, some conventional ANC implementations can be said to selectively pass through certain audio. For example, some conventional ANC implementations only cancel out certain frequency ranges, thereby essentially passing through a select range or ranges of frequencies. In contrast to those conventional approaches, novel selective pass-through approaches described herein use trained machine learning models to detect and classify AS audio 157 from within ambient audio 155 and to selectively pass that audio through responsive to that classification. For the sake of clarify, such novel approaches are referred to herein as classification-based selective pass-through (CSPT).
[0037] FIG. 2 shows a simplified diagram of a conventional active noise cancellation (ANC) environment 200 for added context. During active use, it can be assumed that a user is listening to some desired audio 165 being played by some media player 230. Meanwhile, there is ambient audio 155 in the environment that is potentially frustrating the user’s ability to listen to the desired audio 165. Thus, the goal of the ANC 140 is to output a signal (playback signal 220) that seeks effectively to maximize the ratio of desired audio 165 to ambient audio 155 as experienced by the user. For example, the playback signal 220 ideally sounds to the user as if only the desired audio 165 is playing, even in the presence of ambient audio 155.
[0038] To that end, the ANC 140 seeks to generate an anti-noise signal 210 that cancels the ambient audio 155. As illustrated, the anti-noise signal 210 and the desired audio 165 are mixed (e.g., by mixer 145) to generate an output signal to deliver to the user via one or more speakers of the WAC. In fact, both that output signal and at least some of the ambient audio 155 are concurrently entering the user's ear canal. What is heard by the user, then, is a combination (e.g., a superposition) of those signals, such that the audio reaching the user’s eardrum (the "playback signal” 220) is effectively the desired audio 165 with whatever residual portion of the ambient audio 155 is not canceled by the anti-noise signal 210. For example, an ideal anti-noise signal 210 would perfectly cancel the ambient audio 155. such that the playback signal 220 is only the desired audio 165.
[0039] As illustrated, conventional ANCs 140 can receive the ambient audio 155 via a reference microphone 110 and can receive the playback signal 220 via an error microphone 115. When the ANC 140 receives the playback signal 220 via the error microphone 115, it correlates the signalagainst the ambient audio 155 received via the reference microphone 110 to determine how much residual ambient noise is still present in the playback signal 220 (i. e. , if the ANC 140 is perfectly canceling the ambient audio 155, the ANC will not find any residual in the playback signal 220). The ANC 140 then uses adaptive feedback control (e.g., one or more control loops) to adjust parameters of the ANC 140 in an attempt to iteratively improve the correlation (i.e., to further reduce the presence of ambient audio 155 in the playback signal 220).
[0040] Typically, the ANC 140 analyzes a time window (e.g., a sliding window) of the ambient audio 155 and playback signal 220 to compute the correlation. A typical conventional ANC 140 can use a sliding window of anywhere between fractions of a millisecond and tens of milliseconds. Using a shorter window tends to provide faster response to changes in the noise environment and can be more effective for high-frequency noise and transient sounds, but it tends to involve more processing power and more sophisticated algorithms. Using a longer window tends to be better for canceling steady, low-frequency noises (e g., engine hum, air conditioning, etc.) and is generally easier to implement with less computational overhead, but can tend to introduce more delay.Some modem ANCs 140 can use adaptive algorithms to adjust the window length dynamically.
[0041] As described herein, embodiments of the audio management system 100 are configured for integration in any suitable WAC. FIGS. 3A and 3B show a wearable audio environment 300 including a pair of WACs 310. Each WAC 310 is illustrated as an earbud. Alternatively, the pair of earbuds can be considered as a single WAC. In other embodiments, the WAC 310 can be implemented as over-ear headphones, or any other suitable wearable audio component that incorporates ANC. In the illustrated embodiments, each WAC (i.e., each earbud) has a respective instance of an audio management system 100, such as the audio management system 100 of FIG.1, and each instance of the audio management system 100 includes a respective instance of at least an ANC system 140 and an AHS system 150. Though not explicitly shown, each WAC 310 also has, integrated therein, an instance of the speaker 105, the reference microphone 110, the error microphone 115. one or more processors, and non-transitory processor-readable storage. Some implementations of the WAC 310 include additional components, such as instances of the audio processing system 160, one or more additional microphones (e.g., a beamforming microphone), one or more additional speakers, interface controls (e.g., one or more buttons), one or more power sources (e.g.. a rechargeable battery), one or more ports (e.g.. physical ports for charging and / or wired communication, logical ports for wireless charging and / or wireless communication), one or more antennas, etc.
[0042] In some embodiments, the one or more processors integrated in the WAC 310 implement components of the respective audio management system 100 instance. For example, a non- transitory processor-readable medium integrated therein has processor-executable instructions stored thereon, which, when executed, cause the set of processors to implement at least features of the respective ANC system 140 and / or AHS system 150 instances. As described herein, embodiments of the AHS system 150 include one or more types of artificial neural networks, corresponding trained network models, or the like. In some embodiment, such networks and / or models are implemented using specialized hardware, such as neuromorphic chips. In other embodiments, such networks and / or models are implemented by using processor-readable instructions to reconfigure general-purpose computing hardware (e.g., a central processing unit, CPU), specialized Al accelerators.
[0043] Turning specifically to FIG. 3A. a first type of wearable audio environment 300a is shown in which one or both WACs 310 is in communication with a cloud computing environment (“cloud”) 340. For example, the cloud 340 includes a server, or several distributed servers, accessible via the Internet. Though the WAC 310 is shown as directly in communication with the cloud 340, such a connection can be facilitated by any suitable intermediary devices, such as routers, hubs, etc. As described further herein, automated attention handling features described herein rely on generation of an inference model that includes several neural netw orks and / or models. In the illustrated embodiments of FIG. 3 A, the inference model is generated by the local computation environment of the WAC 310 and / or based on information ported to the WAC 310 from the cloud 340.
[0044] As described herein, some embodiments include features that involve user interaction. As one example, some embodiments involve name detection based on pre-enrollment by a user of invocation names. As another example, embodiments involve known speaker detection based on user-identified speakers. In those and / or other cases, users engage with physical and / or virtual interfaces, such as buttons, switches, light sensors, force sensors, proximity sensors, etc.: and / or with user interface applications, such as enrollment applications, configuration applications, graphical user interfaces, etc. In some embodiments, such as embodiments of FIG. 3 A, all the user interaction features are integrated with the WAC 310. For example, the WAC 310 includes physical interface elements, interface applications installed on its local computational environment, etc.
[0045] In some embodiments, such as embodiments of FIG. 3B, the wearable audio environment 300b further includes a user computational device 320 separate from the WAC 310. For example,the user computational device 320 can be a smartphone, laptop computer, tablet computer, smart watch, portable audio player, or any other suitable device that is separate from the WAC 310 and includes its own one or more processors and its own one or more non-transitory storage media for storing processor-readable instructions. The user computational device 320 can be in communication with each WAC 310 via any suitable wired and / or wireless communication link, such as via an audio cable (e.g., via a 3.5 -millimeter or 1 / 4-inch analog audio jack), a universal wired connection (e.g.. universal serial bus (USB)), a short-range universal wireless connection (e.g., Bluetooth, short-range radiofrequency, near field communication (NFC)), an optical connection (e.g., infrared), a proprietary' connector, a multi -pin connector, an intermediary' component or platform (e.g.. a docking starion or dongle), etc. In such embodiments, some or all of the user interaction features are integrated with the user computational device 320. For example, the user computational device 320 includes physical interface elements, interface applications installed on its local computational environment, etc. In some embodiments, both the WAC 310 and the user computational device 320 include user interaction features.
[0046] As described herein, embodiments use various machine learning models to support automated attention handling (AAH) features, illustrated generally as AAH model(s) 350. In some embodiments, the AAH model(s) 350 include one or more foundation models. In some embodiments, the AAH model(s) 350 include some or all of one or more inference models (e.g., generated based on the one or more foundation models). Each inference model can include some or all of a name embedding model, a deep image, a relation network, a false rejection network, etc. In some embodiments, the AAH model(s) 350 include one or more attention seeking audio detection and classification (ASADC) models. In some embodiments, the AAH model(s) 350 include a classification-based selective pass-through (CSPT) model. In some embodiments (e.g. as illustrated in FIG. 3A). some or all of the AAH model(s) 350 are stored in a cloud-based environment, so that they are accessible to the WAC 310 via the cloud 340. In some embodiments, some or all of the AAH model(s) 350 are stored in a local computational environment of the WAC 310. In some embodiments (e.g. as illustrated in FIG. 3B), some or all of the AAH model(s) 350 are stored in a cloud-based environment, so that they are accessible to the user computational device 320 via the cloud 340. In some embodiments, some or all of the AAH model(s) 350 are stored in a local computational environment of the user computational device 320. In some embodiments, the AAH model(s) 350 are distributed between two or more of the cloud-based environment, the user computational device 320, and the WAC 310.
[0047] FIG. 4 shows a block diagram of an automated attention handling (AAH) system 400, according to embodiments described herein. As illustrated, the AAH system 400 includes anattention seeking audio detection and classification (ASADC) processor 410, an AAH processor 420, and an audio output processor 440. For context, the AAH system 400 is illustrated in communication with a reference microphone 110, an error microphone 115, an output transducer (speaker) 105, an ANC system 140, and a media player 460.
[0048] Embodiments of the ASADC processor 410 receive an ambient audio 155 signal via the reference microphone 110. Various techniques can be used by the ASADC processor 410 to detect when the ambient audio 155 includes one or more types of AS audio (i.e., the AS audio 157 described above, but not explicitly shown in FIG. 4). In some cases, the ASADC processor 410 detects that the ambient audio 155 includes first-person-initiated (FPI) conversation audio. For example, the ASADC processor 410 detects that the user of the WAC has begun speaking. In some cases, the ASADC processor 410 detects that the ambient audio 155 includes second-person- initiated (SPI) conversation audio. For example, the ASADC processor 410 detects that a known attention seeker is speaking to the user, that an enrolled name of the user has been uttered, etc. In some cases, the ASADC processor 410 detects that the ambient audio 155 includes alert audio. For example, the ASADC processor 410 detects a particular type of sound (e g., traffic, a hom, an angry voice, etc.) is directed toward the user, detects any of a set of critical sounds, detects any of a set of enrolled sounds, etc.
[0049] As described herein, the ASADC processor 410 classifies any AS audio 157 that is detected. For example, the ASADC processor 410 classifies the AS audio 157 as an alert sound, as speech from a known speaker, as speech directed at the user from an unknown speaker, etc. The classification of the AS audio 157 by the ASADC processor 410 results in the ASADC processor 410 outputting an embedded space representation (ESR) signal 415 of the detected and classified AS audio 157. The ESR signal 415 can be a vector that represents the most salient features for classifying the one or more types of AS audio 157 detected within the ambient audio 155. If the ASADC processor 410 does not detect any AS audio 157 in the ambient audio 155, it will not generate any ESR signal 415 (or it will generate a signal that indicates the absence of any AS audio 157, such as a vector of all ‘0’s).
[0050] As illustrated, the AAH processor 420 can include one or more controllers for handling different types of AS audio 157. In some embodiments, the AAH processor 420 includes only a classification-based selective pass-through (CSPT) controller 430 configured to handle all types of AS audio 157. In some embodiments, the AAH processor 420 includes an CSPT controller 430 to handle second-person (e.g., known-speaker and / or unknown-speaker) types of AS audio 157, and an alert audio controller 435 to handle alert-related AS audio 157. Some embodiments do notconsider first-person (FP) audio (i.e., audio initiated by the user of the WAC) to be AS audio 157. Embodiments that do consider FP conversation audio to be AS audio 157 can include a separate FP audio controller 425 to handle that audio. In general, the signal path through the ASADC processor 410 and the CSPT controller 430 generates latency. The latency is considered to be acceptable for second-person audio and for alert audio, but not for the FP audio. As such, embodiments of the CSPT controller 430 are not configured to handle FP audio, and it is generally assumed herein that the AS audio 157 does not include FP audio.
[0051] Although three controllers are shown as part of the AAH processor 420, other implementations of the AAH processor 420 can include more or fewer controllers. Also, one or more of the controllers shown in FIG. 4 can be combined or segmented in any suitable manner. In one implementation, the CSPT controller 430 and the alert audio controller 435 are combined into a single controller. In one implementation, the CSPT controller 430 is segmented into more than one controller, each tailored to selectively pass through a particular subset of AS audio 157 classifications.
[0052] Embodiments of the AAH processor 420 can also include (or be in communication with) a configuration processor 428. As described herein, various features relating to the manner in which automated attention handling is implemented are user-configurable. As one example, some embodiments permit a user to determine whether, upon detecting AS audio 157, any currently playing media audio 465 should be paused, muted, lowered in volume, etc. As another example, embodiments permit a user to define different configurations for different locations, different priorities for different types of alerts, etc. The configuration processor 428 can be in communication with one or more user interfaces (I / F) 427 and can be configured to receive, update, maintain, and otherwise handle the configurations of the various AAH features.
[0053] As noted above, upon detecting AS audio 157 in the ambient audio 155, the ASADC processor 410 outputs an ESR signal 415 to the CSPT controller 430 in the AAH processor 420. As illustrated, the ambient audio 155 is also passed directly to the CSPT controller 430. As described herein, the ESR signal 415 is received by the CSPT controller 430 as a conditional signal for dynamically configuring an embedding (bottleneck) layer of the CSPT controller 430 model, thereby dynamically controlling selective pass-through of the detected AS audio 157.
[0054] As noted above, some embodiments of the AAH processor 420 have a separate alert audio controller 435. In such embodiments, the ambient audio 155 can be directly passed to the alert audio controller 435. In some such embodiments, all alert audio is processed by the alert audio controller 435. For example, the ASADC processor 410 is not configured to detect orclassify any alert audio, and the CSPT controller 430 is not configured to pass through alert audio. Instead, the alert audio controller 435 is configured to detect alert audio and to take one or more predetermined responsive actions accordingly. For example, such embodiments can use detected alert audio as a trigger to pause media audio 465, to disable the ANC system 140, to switch the ANC system 140 to a transparency mode, to generate corresponding alert audio, and / or perform other actions. In other such embodiments, certain types of alert audio is processed by the alert audio controller 435. and other types of alert audio is handled by the CSPT controller 430. In some implementations, different types of alert audio are categorized into different priority levels (e.g., by default, by user configuration, etc.), and different priority levels are handled by different controller. In some implementations, different types of alert audio are treated differently, ignored, treated by different controllers, etc. based on the present location of the user. In some implementations, different types of alert audio are treated differently, ignored, treated by different controllers, etc. based on other characteristic of the audio, such as its directionality7, its volume, etc.
[0055] As noted above, first-person audio is generally not considered AS audio 157 and is also not generally considered to be part of the ambient audio 155 (even though it will be picked up by the reference microphone 110). As such, FIG. 4 does not show the ambient audio 155 being passed to the reference microphone 110. Still, in some embodiments, the FP audio, or a processed version thereof, can be handled by the FP audio controller 425.
[0056] The AAH processor 420 outputs an AAH audio signal 424. In embodiments having a FP audio controller 425, in cases where the FP audio controller 425 is presently handling FP audio, the AAH audio signal 424 includes the FP audio, or processed FP audio, as output by the FP audio controller 425. In embodiments having an alert audio controller 435, in cases where the alert audio controller 435 is presently handling alert audio, the AAH audio signal 424 includes the alert audio, or processed (e.g., generated) alert audio, as output by the alert audio controller 435. In embodiments having an CSPT controller 430, in cases where the CSPT controller 430 is presently handling detected and classified AS audio 157, the AAH audio signal 424 includes the AS audio 157 as passed through (e.g., and processed) by the CSPT controller 430, indicated as a pass- through audio signal 432.
[0057] Embodiments of the audio output processor 440 can be implemented in or by the audio processing system 160 of FIG. 1. In general, the audio output processor 440 receives the AAH audio signal 424 (e.g., the pass-through audio signal 432), the output of the ANC 140 (i.e.. the anti-noise signal), and the output of the media player 460 (e.g., the media audio 465). These audio signals can be used to generate the playback audio to be provided via the WAC speakers 105 to theuser. As illustrated, the different audio signals can be received by a mixing controller 450 in the audio output processor 440. For example, the mixing controller 450 is a volumetric mixer that seeks to balance the audio signals to provide a comfortable listening experience for the user.
[0058] Some embodiments of the AAH processor 420 also output an AAH control signal 422 to control operation of the audio output processor 440. In some embodiments, the AAH control signal 422 is generated by each controller of the AAH processor 420 based on configured (e.g., by default, by a user, etc.) responses of the controller. As illustrated, the audio output processor 440 can include a gating controller 445, and the AAH control signal 422 can be used to control operation of the gating controller 445. The gating controller 445 operates to selectively pause operation of the ANC 140 and / or the media player 460 based on the AAH control signal 422.
[0059] Embodiments of the audio output processor 440 can operate differently based on the AAH audio signal 424 and / or the AAH control signal 422 using the mixing controller 450 and / or gating controller 445. In some implementations, the FP audio controller 425 and / or the CSPT controller 430, upon detecting the start of a conversation, directs the gating controller 445 to toggle the ANC 140 to a conversation mode and directs the gating controller 445 to pause the media player 460. The mixing controller 450 operates to volumetrically mix one or more sources of audio to generate the output to the WAC speaker or speakers. In some implementations, the alert audio controller 435, upon detecting alert audio, directs the mixing controller 450 to mix alert audio (and / or audio generated by the alert audio controller 435 in response to detecting the alert audio) with the media audio 465 being played by the media player 465, while also mixing in the anti-noise signal generated by the ANC 140. In some implementations, the FP audio controller 425 and / or the CSPT controller 430, upon detecting the start of a conversation, directs the mixing controller 450 to reduce the volume of the media player 460 output and to mix in the second party’s voice.
[0060] The audio output processor 440 (including the gating controller 445 and the mixing controller 450) can be configured to operate in any desired manner in response to detecting different types of AS audio 157. In some implementations, the audio output processor 440 has a static configuration. In other implementations, behavior of the audio output processor 440 (e.g., the gating controller 445 and / or the mixing controller 450) is user-configurable. In such implementations, I / F 427 can be used by the user to configure what actions should be taken in response to one or more types of AS audio 157. For example, the user can decide whether to pause the media player 460 or to reduce its volume, whether to disable the ANC system 140, etc.
[0061] The output from the audio output processor 440 is played through the WAC speaker(s) 105 to the user. A corresponding audio signal is received by the ANC 140 as the playback signal 475. Typically, the WAC speaker(s) 105 are located close to the user’s eardrum while the WAC is being worn, such as in the ear cavity. Thus, the playback signal 475 represents the audio being received by the user’s eardrum. As described above, the goal of the ANC 140 is to adaptively generate an anti-noise signal to mix into the playback signal 475, so that the playback signal 475 has maximum correlation with the media audio 465 (i. e. , desired audio 165) and minimum correlation with the ambient audio 155.
[0062] Conventionally, because the AS audio 157 is part of the ambient audio 155, the ANC 140 will actively tr\' to cancel the AS audio 157. One conventional way for the user to hear the AS audio 157 is to rely on the user to notice the presence of AS audio 157 in the environment and to either remove the WAC and / or to turn off the ANC 140. Such an approach can be inconvenient for the user and may not be particularly helpful when the user does not notice the presence of AS audio 157.
[0063] Another conventional way for the user to hear the AS audio 157 is for the ANC 140 to be configured not to cancel a certain range of frequencies. For example, some conventional "‘transparency” or “pass through” modes act in a similar manner to an advanced high-pass or notch filter, which actively cancels lower frequency noise (e.g., much of the energy of engine noise, wind noise, air conditioning noise, etc.), while allowing typical speech-related frequencies and other higher frequency sounds to pass through. Such approaches have several drawbacks. One drawback is that the noise-canceling of the ANC 140 with such approaches is limited to only a particular subset of audible frequencies to cancel, which appreciably reduces its effectiveness. Another drawback is that the partial processing and other features of such an approach tend to cause users to experience distortion or a hollow sound.
[0064] Selective pass-through embodiments described herein seek to keep the ANC 140 completely active (e.g., across at least the audible range of frequencies), while still selectively passing through AS audio 157 detected in the ambient audio 155. As illustrated, each of the FPI conversation controller 425, the SPI conversation controller 430, and the alert audio controller 435 outputs at least a processed AS audio signal 422. As described herein, the processed AS audio signal 422 corresponds to the detected AS audio 157 and is delayed by at least a predetermined ANC processing window duration and less than a maximum lip-sync delay. Because of the delay, the processed AS audio signal 422 will not correlate (e.g., will have close to zero correlation) withthe ambient audio 155 that is being canceled by the ANC 140, and it will effectively pass through to the user.
[0065] As noted above, ANCs 140 process a moving window of audio information to characterize the ambient audio 155 and perform noise canceling. The moving window has a particular duration (i.e., the ANC processing window duration). In some implementations, the ANC processing window duration is static, and the processed AS audio signal 422 is delayed by a predetermined amount of time that is greater than the ANC processing window duration. In other implementations, the ANC processing window duration is dynamically adjusted. For example, the ANC dynamically adjusts the ANC processing window duration, such as between fractions of a millisecond and up to 50 milliseconds. In some such implementations, the processed AS audio signal 422 is delayed by a predetermined amount of time that is greater than a predetermined maximum ANC processing window duration. In other such implementations, the ANC 140 informs the AAH processor 420 of the present ANC processing window duration, and the processed AS audio signal 422 is delayed by an amount of time that is greater than the present ANC processing window duration.
[0066] A lip-sync delay is the audio latency between a user seeing a source of sound (e.g., a speaker’s lips moving, etc.) and hearing the resulting sound. Relatively small lip-sync delays are generally imperceptible to humans. As such, a maximum lip-sync delay can be predetermined as within the generally imperceptible range. In one implementation, the maximum lip-sync delay is 100 milliseconds (i.e., humans are generally unable to perceive an audio latency of less than that amount). Based on the above, the processed AS audio signal 422 is delayed by enough time to be outside the ANC processing w indow , but by little enough time so that the delay is essentially imperceptible.
[0067] In some embodiments, the total processing time of the AS ADC processor 410 and the AAH processor 420 is greater than the ANC processing window duration and less than the maximum lip-sync delay. In such embodiments, there may be no need for inserting additional delay to ensure that the processed AS audio signal 422 is outside the ANC processing window. In other embodiments, the total processing time of the AS ADC processor 410 and the AAH processor 420 is less than the ANC processing window' duration, or the total processing time of the ASADC processor 410 and at least one of the controller paths of the AAH processor 420 is less than the ANC processing window duration. Such embodiments introduce additional delay to ensure that the processed AS audio signal 422 is outside the ANC processing window (but still less than the maximum lip-sync delay).
[0068] FIG. 5 shows a block diagram of another automated attention handling (AAH) system 500, focused on classification-based selective pass-through (CSPT) embodiments, according to embodiments described herein. The AAH system 500 can be an implementation of the AAH system 400. As in FIG. 4, the AAH system 400 includes an attention seeking audio detection and classification (AS ADC) processor 410 and an audio output processor 440. Rather than showing the AAH processor 420 as in FIG. 4, FIG. 5 shows only the CSPT controller 430. For context, the AAH system 500 is illustrated in communication with a reference microphone 110 (to receive ambient audio 155, which may include AS audio 157 as any given time), an output transducer (speaker) 105, and a media player 460 (to play media audio 465). Although FIG. 5 does not explicitly show the ANC system 140, error microphone 115, and certain other components, other embodiments, such as those illustrated in FIG. 4, illustrate how those components would be incorporated with embodiments of FIG. 5.
[0069] Embodiments of the ASADC processor 410 receive an ambient audio 155 signal via the reference microphone 110. The implementation illustrated in FIG. 5 includes a sound analytics block 510. The sound analytics block 510 includes detectors for different types of AS audio 157 within the ambient audio 155 signal. In some implementations, each of the detectors is implemented separately. In some implementations, some or all of the illustrated detectors can be combined. In some implementations, additional detectors can be included. In the illustrated implementation, the sound analytics block 510 includes a known speaker detector 512, a charged sound detector 514, a name / conversation detector 518, and an audio event detector 516.
[0070] Turning first to the name / conversation detector 518. embodiments generally detect when the ambient audio 155 includes a pre-enrolled invoked name, or other tag taking the place of a name. For example, the user enrolls a set of names that can be used to get the user’s attention (e.g., “John,” “Johnny,” “Mr. Smith.” etc.), and the name / conversation detector 518 uses one or more techniques to determine whether the ambient audio 155 includes one of those enrolled names. Various approaches have been previously described by inventors of the present disclosure for such name detection, for example, in International Patent Application No.PCT / US2024 / 014606, titled “AUTOMATED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS BASED ON LINGUISTIC NAME EMBEDDING,” filed on February 6, 2024; International Patent Application No. PCT / US2024 / 014788. titled “AUTOMATED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS BASED ON UNIVERSAL SOUND CONVERSION,” filed on February 7, 2024; Indian Patent Application No. IN202441019923, titled “SYSTEM AND METHOD FOR DETECTING A PREDEFINED SET OF NAMES FROM THE AUDIO SIGNAL USING AN ACOUSTIC SOUNDSEGMENTATION MODEL,” filed on March 18, 2024; and International Patent Application No. PCT / US2024 / 014820, titled “NAME-DETECTION BASED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS,” filed on February 7, 2024.
[0071] FIG. 6 shows a simplified block diagram of stages of an illustrative implementation of a name / conversation detector 600 and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein. The name / conversation detector 600 can be an implementation of the name / conversation detector 518 of FIG. 5. It is assumed that the name / conversation detector 600 is implemented in the context of a WAC 310.
[0072] Embodiments can include an enrollment stage 610, an identification stage 630, and a verification stage 640. In general, the enrollment stage 610 occurs outside of normal operation of the WAC 310, such as when a user first sets up and / or uses the WAC 310, when the user first registers the WAC 310, when the user first configures the WAC 310 for automated attention handling, etc. The identification stage 630 and the verification stage 640 are real-time blocks that occur during normal operation of the WAC 310 to facilitate real-time automated attention handling. Embodiments generally use the enrollment stage 610 to obtain any information needed to set up an inference model for storing at the WAC 310 to enable the WAC 310 to perform automated attention handling features during normal operation. As shown, the inference model includes at least an embedding model 615, a relation network 635, and a false rejection (FR) network 645.
[0073] In the enrollment stage 610, the embedding model 615 can be used to convert an enrollment audio stream 605 into a set of reference embeddings based on a foundation model 625. The foundation model 625 can be stored in and accessed via the cloud 340 (e.g., and / or any other suitable communication network). The reference embeddings can be stored as a deep image 620. The deep image 620 can also be considered as part of the inference model.
[0074] During normal operation, in the identification stage 630, a real-time (RT) audio stream (i.e., ambient audio 155) is received. The embedding model 615 (i.e., the same embedding model 615 generated in the enrollment stage 610) is used to generate a RT embedding from the received ambient audio 155. The relation network 635 can then compare the RT embedding with each of the reference embeddings in the deep image 620 to determine if there is a match. For example, the relation network 635 is configured to compute a similarity score for each comparison (e.g., corresponding to a mathematical correlation, or the like). The relation network 635 can determine if any similarity scores meet or exceed a predetermined threshold; if so, the reference embedding with the highest similarity score can be selected as a candidate matching embedding.
[0075] In the verification stage, the false rejection network 645 can then confirm the match by transforming the RT embedding and the candidate matching embedding into different mathematical spaces (e.g., domains) and determining whether the embeddings can be reliably discriminated from each other in any of the mathematical spaces. For example, the deep image 620 is configured to compute a discrimination score for each mathematical space. The deep image 620 can determine if any discrimination scores meet or exceed a predetermined threshold; if so, the match determined by the relation network 635 is determined to be a false match and is ignored. If the embeddings cannot be sufficiently discriminated in any mathematical spaces, the false rejection network 645 can output a signal that attention seeking audio has been detected.
[0076] As described herein, embodiments of the name / conversation detector 600 are configured to detect an invocation name as part of an attention handling system. The embedding model 615 is a name embedding model (e.g., a processor-executable name embedding model) that generates an output embedding from an audio sample. As used herein, the terms “audio sample” or an “audio signal” are used interchangeably in the context of an input to a component of the inference model; such an “audio sample” or an “audio signal” can be represented in any suitable manner, such as by any suitable number of digital samples. For example, reference to an input as an “audio sample” means an audio signal of a duration, or a sampled duration of an audio signal, at a sampling rate resulting in a large number of digital samples (e.g., one second of audio sampled at 16 kHz to yield 16,000 samples). As described above, the audio sample can be from an enrollment audio stream 605 in an enrollment stage 610, and the audio sample can be from the ambient audio 155 during normal operation. The name embedding model 615 is trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications. The deep image 620 (e.g., a processor-readable deep image) includes reference name embeddings generated by the name embedding model 615 based on a set of invocation names provided by a user during an enrollment procedure. The relation network 635 (e.g., a processor-executable relation network) is coupled with the deep image 620 and the name embedding model 615 to output a candidate name embedding responsive to determining that one of the reference name embeddings has a highest similarity with a RT embedding generated by the name embedding model 615. The RT embedding can be from the ambient audio 155 as received from a reference microphone associated with an ANC system of the WAC 310. The false rejection network 645 (e.g., a processorexecutable false rejection network) is coupled with the relation network 635 to output a name invoked signal 650 responsive to determining that the real-time embedding and the candidate name embedding cannot be reliably discriminated.
[0077] As described above, the interference model used for automated attention handling is based on a foundation model 350. Embodiments of the foundation model 350 are generated from a large speech-audio corpus of diversified words spoken by diversified speakers. “Diversified words” refers herein to the speech-audio corpus including a wide variety of at least phonemes and linguistic information. “Diversified speakers” refers to the speech-audio corpus representing a wide variety of at least accents and prosody. The speakers can be further diversified with respect to age, gender, geography, etc. For example, the speech-audio corpus includes tens of thousands of words (i.e., classifications) spoken multiple times (e.g., 10 - 15 times) by hundreds of speakers from around the w orld.
[0078] The term “suprasegmental” is used herein as an umbrella term to encompass properties of a speaker’s influence when speaking words, such as accent, prosody, intonation, rhythm, and other non-segmental aspects of speech. Such suprasegmental features can be contrasted with segmental features pertaining to individual speech sounds or segments, such as vowels and consonants, and can span multiple segments or an entire utterance. Examples of suprasegmental features of an utterance (e.g., a word, name, etc.) can include accent (including accent-influenced variations in pitch, loudness, and duration), prosody (including rhythm, intonation, and melody of speech), intonation (i.e., the rise and fall of pitch in speech), rhythm and / or rate (e.g., the temporal patterns of speech, such as duration and timing of sounds, syllables, and pauses), and stress (e.g., emphasis placed on a particular syllable). For example, a large speech-audio corpus of diversified words spoken by diversified speakers may include hundreds or thousands of samples of a particular word being spoken with wide suprasegmental variance over the samples.
[0079] Training of the foundation model 350 can begin w ith an auto-encoder architecture, which is a type of neural network architecture designed to leam compact representations of data, such as so-called “latent features.” In the context of embodiments described herein, the auto-encoder architecture is used to extract meaningful features from raw audio data to be used for automatic speech recognition (ASR). In general, the auto-encoder architecture includes three high-level stages: an encoder, a bottleneck layer, and a decoder. The encoder receives a high-dimensionality input (i.e., the raw7audio data) and includes a multi-layer netw ork to progressively reduce the dimensionality of the data using transformations. For example, the ASR information is represented as a sequence of audio frames, each including features that can be mathematically described as Mel-frequency cepstral coefficients (MFCCs), or in some other manner. Each layer of transformations (e.g., linear operations followed by non-linear operations, such as a rectified linear unit (ReLU) function) seeks to extract increasingly abstract and higher-level features from the input data.
[0080] The botleneck layer is so called because it ty pically includes significantly lower dimensionality than the layers of the encoder before it and or the decoder after it. This reduced dimensionality effectively forces the network to leam a compressed and informative representation of the input data. The decoder operates in essentially a reverse manner to that of the encoder. It takes the output of the botleneck layer as its input data (i.e., a lowest-dimensionality representation) and applies multiple layers of transformations to generate increasingly higherdimensional representations of the data. In effect, the decoder progressively increases the dimensionality with a goal of reconstructing the original input data from the highly compressed representation in the botleneck layer.
[0081] Although the decoder output ideally matches the encoder input (i.e., the raw audio data), there will practically be some difference, referred to as reconstruction loss. Training of the autoencoder seeks to optimize the auto-encoder network (e.g.. the bottleneck layer) to minimize reconstruction loss (e g., to minimize mean squared error loss, or the like). Effective training results in the botleneck layer producing a highly compact, but highly meaningful representation of the input data; effectively extracting the most salient features for ASR. The compact representation in the botleneck layer can be referred to as the ‘"latent space,” and the autoencoder’s trained (learned) knowledge of how to encode raw input data into the latent space is embedded in a set of weights. The set of weights can be represented as a feature vector, a set of feature vectors, or in any suitable manner. For example, one second of audio can be represented by 16,000 samples (i.e., at a 16 kHz sampling rate), and the 16,000 samples can effectively be represented by a set of 256 weights (e.g., or 128 weights, or another suitable number). The set of weights can represent the embedding from the auto-encoder, the filter bank energies, the MFCCs, etc.
[0082] In some embodiments, the name embedding model 615 is generated by applying transfer learning from the large speech-audio corpus of phonetically diversified words used in training the foundation model 350 (the source dataset) to a smaller corpus of real-world name data (the target dataset). For example, the foundation model 350 can use a large number (e.g., 11,000) of classifications to generate the set of weights, where each linguistically distinct word in the speechaudio corpus is classified into one of the classes. Transfer learning can then apply the trained foundation model 350 to generate the name embedding model 615 (deep feature generation model, or DFGNet) for a smaller number (e.g., 500 - 1000) of classifications, each associated with a linguistically distinct name. Implementations of the name embedding model 615 include only the encoder and botleneck layer as trained through the transfer learning.
[0083] In effect, the name embedding model 615 can be characterized as a tuned and reduced version of the foundation model 350 developed specifically for name detection, such as by using few-shot learning (FSL). Real audio samples used to train the name embedding model 615 can correspond to people's names from various regions and languages, and sample names can be chosen to cover most phonetic usage in each region. Implementations of the name embedding model 615 can be trained to recognize any suitable number of name classifications. For example, it may be impractical impossible to train the model for all possible names and their variants everywhere in the world, and a practical number of more common names (e.g., 1,000) can be chosen instead. In some implementations, different versions of the name embedding model can be generated 615 and / or trained differently for different user groupings (e.g., geographical regions, ethnicities, etc.) to capture the most popular names for the corresponding groupings. For example, grouping information can be entered by the user as part of enrollment, obtained for the user from account information, assumed for the user based on location or other demographic information, etc. Further, the name embedding model 615 can be designed with as much complex as needed to discriminate between the name classifications used for training with enough reliability to be suitable for name-detection-based attention handling. For example, a higher-dimensional model (e g., where the number of weighting vector dimensions, N, is larger) may be able to more reliably discriminate among a larger number of name classifications, but use of such a model in real time will involve more computational resources (e.g., which may correspond to more processing time, more battery usage, more heat generation, etc.). The name embedding model 615 can then generate a reference embedding for each invocation name.
[0084] Returning to FIG. 5, the name / conversation detector 518 is essentially one possible path by which AS audio 157 is detected as part of the ambient audio 155. In the name / conversation detector 518 context, it is the invocation of the pre-enrolled tag (e.g., name or other word or phrase, such as “excuse me,” “boss,” “mom,” etc.) that triggers the detection of relevant AS audio 157. In some embodiments, the sound analytics block 510 generally looks for any of three categories of AS audio 157: speech from a known speaker, speech from an unknow n speaker, and alert audio. The name / conversation detector 518 can be a manner of detecting speech from an unknown speaker. As used herein, an “unknown speaker” can be any speaker that is either unknown to the user or is not yet identified in the AAH system 500 (e.g., no representation has been previously stored for that speaker, as described below). Detection of the pre-enrolled tag acts as a w ay for the sound analytics block 510 to know that, although this is an unknown speaker, the speech is potentially relevant to the user and should be treated as AS audio 157.
[0085] Other detectors of the sound analytics block 510 include the known speaker detector 512, the charged sound detector 514, and the audio event detector 516. Embodiments train a machine learning model (e.g., an artificial neural network) to detect and classify ambient audio 155 (including training AS audio 157) into a number of classifications. The classification-based training can be used to formulate corresponding representation vectors that can help direct downstream behavior of the sound analytics block 510 and other portions of the ASADC processor 410 and CSPT controller 430.
[0086] FIGS. 7A and 7B show a first training stage and a second training stage, respectively, for training a machine learning model as an attention seeking audio detector and classifier (ASADC) model 700, according to various embodiments. Turning first to FIG. 7A, the ASADC model 700a includes an encoder 710 and a decoder 730 separated by a bottleneck layer (BNL) 720. Training data 705 is input to the encoder 710. and a set of attention seeking audio (ASA) classifications 735 is output by the decoder 730. As described with reference to the name / conversation detector 600 of FIG. 4, the encoder 710 is trained to generate successively compact representations of the training data 705 until the bottleneck layer 720 represents the most salient features for classifying the training data 705 into the desired set of ASA classifications 735.
[0087] Though not explicitly shown, the ASADC model 700a is trained based on iteratively inputting large amounts of diverse types of training data 705 into the encoder 710. In each iteration, it is known what expected output should be generated by the decoder 730 in response to receiving particular training data 705 at the encoder 710, so that the expected output is an N- dimensional vector to represent the set of N ASA classifications 735 (i.e., each of the N elements of the vector corresponds to a respective one of N ASA classifications 735). Once the ASADC model 700a is fully trained, inputting training data 705 that has an nth one of the set of ASA classifications 735 should result in the decoder 730 generating an output vector with energy7at the nth element. To that end, in each training iteration, the training data 705 includes one, several, or none of the types of AS audio 157 to be represented by the set of ASA classifications 735. The generated output vector is compared to a vector representing the expected output vector (i.e., what the ASADC model 700a should produce when fully trained), thereby generating an error. The error is back-propagated to adjust parameters of the encoder 710, bottleneck layer 720. and / or decoder 730. until the error falls below a predetermined training error threshold. In some embodiments, the encoder 710, bottleneck layer 720, and decoder 730 blocks are all trained together.
[0088] In some embodiments, the ASADC model 700a is trained to classify the training data 705 into one or more ASA classifications 735 that indicate when human speech is part of the training data 705. In such embodiments, the training data 705 includes a speech-audio corpus of large numbers of spoken word samples. In some such embodiments, the speech-audio corpus includes samples of words spoken by multiple different people with different accents, different prosody, and / or different suprasegmental features. In some such embodiments, the speech-audio corpus includes samples of words from different languages, from different regional dialects, etc. In some embodiments, the ASADC model 700a is trained so that all the different types of speech-audio samples (e.g., including all the different types of augmentations, variations, distortions, etc.) result in a same ASA classification 735 (i.e.. generally representing the presence of human speech).
[0089] In some other embodiments, the ASADC model 700a is trained so that different types of speech-audio samples result in different ASA classifications 735. For example, in some such embodiments, the speech-audio corpus includes samples of words spoken at different volumes, with different emotions, etc. In one such embodiment, at least one ASA classification 735 indicates that typical human speech is present in the training data 705, while at least another ASA classification 735 indicates that loud or angry speech is present in the training data 705. In general, embodiments of the ASADC model 700a are not trained to recognize linguistic information, acoustic segmentations, or the like. Rather, in relation to human speech, the ASADC model 700a is trained to recognize whether such human speech is present and / or whether it has some non-linguistic character (e g., particularly loud speech, angry -sounding or urgent-sounding speech, etc.) that indicates its salience as AS audio 157.
[0090] In some embodiments, the ASADC model 700a is trained to classify the training data 705 into one or more ASA classifications 735 that indicate when alert audio part of the training data 705. In such embodiments, the training data 705 includes a corpus of alert audio samples. For example, the corpus includes recorded samples of doorbells, alarms, appliance beeps, tea kettle whistles, dogs barking, babies crying, vehicle horns, etc. In some embodiments, the ASADC model 700a is trained so that any alert audio detected in the training data 705 results in a same ASA classification 735 (i.e., the ASA classification 735 indicates presence or absence of any- trained alert sound). In other embodiments, the ASADC model 700a is trained so that each of several categories of alert audio (e.g., baby-related alert sounds, pet-related alert sounds, alarm- related alert sounds, etc.) detected in the training data 705 results in a corresponding ASA classification 735 (i.e., each of the set of ASA classifications 735 indicates presence or absence of its corresponding category of alert audio). In such embodiments, one or more (or even all) of the categories can include a single respective sound.
[0091] In some embodiments, the ASADC model 700a is trained concurrently on multiple types of ASA classifications 735. In some such embodiments, the training data 705 is one or more large corpuses including large numbers and types of both speech and alert sound samples. In other such embodiments, the training data 705 includes combinations of different ASA classifications 735. For example, each combination can include multiple types of speech audio, multiple types of alert audio, or multiple types of speech and alert audio. In some such embodiments, the combinations are designed to represent likely combinations of AS audio 157 types and / or of ASA classifications 735. For example, one combination may include multiple types of sounds likely to be heard in a user’s home (e.g., combining a doorbell, a baby crying, and a dog barking), while another combination may include multiple types of sounds likely to be heard in an airport (e.g., loudspeaker announcement speech, airport transport sounds, and normal human speech), etc.
[0092] In any of the above embodiments, the training data 705 can be generated further by adding one or more noise profiles (e.g., Gaussian noise, typical ambient conversational noise, ty pical ambient street noise, wind noise, etc.) to the speech-audio corpus samples. In any of the above embodiments, the training data 705 can additionally or alternatively be generated by adding one or more audio transformations (e.g.. to change the speed, to change the pitch, to add a common type of distortion, etc.) to the speech-audio corpus samples.
[0093] When the first training phase is complete, the bottleneck layer 720 (e.g., or one block or layer of multiple blocks or layers of the bottleneck layer 720) can output a representation vector 725 that is a latent space, or embedded space representation of the features of the training data 705 signal that are most salient with respect to classification of AS audio 157 into desired ASA classifications 735. In some embodiments, only the first training phase is needed. Although not explicitly shown in FIG. 7A, the bottleneck layer 720 of the ASADC model 700a can output the representation vector 725.
[0094] Some embodiments include a second training phase. Turning to FIG. 7B, the ASADC model 700b takes the ASADC model 700a from FIG. 7A (i.e., the first training phase), and freezes the encoder 710 and bottleneck layer 720 blocks using the weights obtained in the first training phase. As illustrated, the decoder 730 of FIG. 7A is replaced with a new decoder 740 (e.g., and / or transfer learning is used from decoder 730 to decoder 740), which is again trained to generate a set of ASA classifications 745. The set of ASA classifications 745 can be the same or different from the set of ASA classifications 735 trained in FIG. 7A. As illustrated, the bottleneck layer 720 of the ASADC model 700b is configured to output the representation vector 725.
[0095] In some embodiments, the decoder 730 in the first training phase is trained and configured to classify each alert sound event and each type of speaker as a separate corresponding classification. The decoder 740 of the second training phase is then trained more categorically, such as to classify the input signal into a normal human voice classification, an angry or loud voice classification, and then separate classifications for different alert sound events. Using such a two- phase approach, the first training phase is used to carefully and completely train the encoder 710 and bottleneck layer 720 to be able to find the most salient features for recognizing all of the desired classes of speech, alert audio events, etc. By freezing those layers and blocks in the second training phase, the resulting ASADC model 700b is effectively trained to output a more generalized set of ASA classifications 745, while still preserving less generalized production of the representation vector 725. Thus, in some embodiments, the set of ASA classifications 745 can be used to categorically trigger certain automated attention handling behaviors (e.g., CSPT, gating, mixing, and / or other features) based on the categorical class of AS audio 157 that is detected; while the representation vector 725 can be used to capture a more nuanced embedded space representation of the detected and classified AS audio 157.
[0096] In some implementations, one of the ASA classifications 745 indicates generally that human speech has been detected. At the same time, the embedded space representation of that speech captured by the representation vector 725 can be used (as described below) to determine whether the speech is from a known speaker (e.g., matches a previously stored speaker representation vector). In some cases, if not, the representation vector 725 can be further used to generate a new speaker representation vector to assign to the unknown speaker, which can subsequently be stored in relation to a new known speaker.
[0097] In some implementations, the same ASA classification 745 indicating the general presence of human speech can be used to trigger other models. For example, rather than keeping the name / conversation detector 518 running at all times, embodiments can use the ASADC model 700b to detect and classify the presence of human speech in the ambient audio 155. Upon such detection, the name / conversation detector 518 can be triggered to process relevant portions of the ambient audio 155 to detect whether a pre-enrolled name (or word or phrase) has been invoked.
[0098] In some implementations, one or more ASA classifications 745 indicates a type of speech (e.g., angry, loud, emotionally charged, etc.) that suggests that the speech may be attention seeking and therefor considered as AS audio 157. In some embodiments, detection and classification of such charged speech (i.e., energy at the corresponding element of the vector representing the set of ASA classifications 745) can be used as a trigger to disable ANC, to mute or reduce the volume ofmedia playback, etc. In other embodiments, detection and classification of such charged speech can be used along with the representation vector 725 (e.g., by one or more automated attention handling features) so that even if the speaker of the charged speech is determined to be an unknown speaker, the speech should be selectively passed through by CSPT features, as described herein.
[0099] In some implementations, one or more of the ASA classifications 745 indicates a priority level or category of alert audio event, rather than a particular event. In some such implementations, there is a category of alert that all trigger generation of a predefined alert audio message, which is transmitted to the user (e.g., ‘'warning”). In such an embodiment, the ASA classification 745 itself can be used as a trigger for generating that message without any need for the ASADC model 700b to further specific the type of alert audio, or to use the representation vector 725. In some implementations, several of the ASA classifications 745 each indicates a particular corresponding alert audio event (e.g., one indicates a doorbell sound, one indicates a baby crying, one indicates, a fire alarm, etc.). In some such implementations, each type of alert audio event is used to trigger generation of a corresponding predefined alert audio message, which is transmitted to the user. For example, output by the ASADC model 700b of an ASA classification 745 indicating a doorbell sound is present in the ambient audio 155 triggers an audio message, such as a pre-recorded doorbell sound, or a simulated voice, saying “Someone is at your door.”
[0100] Returning again to FIG. 5, the ASADC model 700 (the ASADC model 700b of FIG. 7B, or in some cases, the ASADC model 700a of FIG. 7A) can be used as the basis for some or all of the detectors of the sound analytics block 510. In some embodiments, the ASA classifications 745 output by the ASADC model 700 can be used to indicate detection of human speech and / or charged (e.g., loud, angry, etc.) speech, which can be part of the known speaker detector 512, the part of name / conversation detector 518, and / or part of the charged sound detector 514; and the ASA classifications 745 output by the ASADC model 700 can be used to indicate detection of alert audio events, which can be part of the audio event detector 516.
[0101] As noted above, when human speech is present in the ambient audio 155 the bottleneck layer 720 of the trained ASADC model 700b generates a representation vector 725 that is an embedded (latent) space representation of the human speaker of the present speech. As part of configuring the WAC for AAH. and / or over time through use of AAH features, embedded space representations can be generated and stored for known speakers. Such representations can be referred to as known speaker representations and can be stored in a representation data store 525.The representation data store 525 can be implemented as or in any suitable type of non-transitory, computer-readable memory. The known speaker representations are generated using the same encoder 710 and bottleneck layer 720 of the ASADC model 700b, so that subsequent speech by those same speakers will tend to generate a representation vector 725 that matches the representations stored in the representation data store 525.
[0102] Returning again to FIG. 5, it can be seen that the sound analytics block 510 effectively analyzes the ambient sound 155 to detect whether any types of AS audio 157 is present, and. if so. the sound analytics block 510 classifies the type of AS audio 157 that is detected. For example, as described with reference to FIGS. 6 - 8, an ASADC model can be trained to output ASA classifications based on detecting and classifying presence, in the ambient sound 155, human speech, charged (e.g., angry, emotional, loud, etc.) speech, and / or alert audio.
[0103] Based on the classification, the ASADC processor 410 can take certain action. In some embodiments, there is a sequential or hierarchical set of actions. For example, in the illustrated implementation of FIG. 5, the sound analytics block 510 can initially use an ASADC model (not explicitly shown) to detect whether there is any human speech or any alert audio event present in the ambient sound 155. If human speech is detected, the known speaker detector 512 can determine whether the detected speech is from a speaker that has a corresponding stored embedded space representation (ESR) stored in the representation data store 525. If so, the classified speech is considered to be AS audio 157, and the corresponding ESR can be used as an active ESR 520. The active ESR 520 can be passed as an ESR signal 415 (as a conditional signal) to the CSPT controller 430, as described herein. In some embodiments, as described herein, even though the identified speaker already has an ESR stored in the representation data store 525, the representation vector 725 being produced by the ASADC model can be used to refine and / or update the stored ESR for that speaker.
[0104] If no known speaker is detected, embodiments can then use the name / conversation detector 518 to determine whether the classified speech (now considered as speech from an unknown speaker) includes an invocation of any pre-enrolled names or phrases. If so, the classified speech is still considered to be AS audio 157, and a speaker ESR engine 530 can be engaged. The speaker ESR engine 530 can use the representation vector 725 being generated by the ASADC model to produce a new ESR in association with the as-yet-unknown speaker. Once a stable new ESR has been generated, the new ESR becomes an active ESR 520. which can be used as an ESR signal 415 for the CSPT controller 430.
[0105] If no enrolled name or phrase is detected in the classified speech, embodiments can then use the charged speech detector 514 to determine whether the classified speech (speech from an unknown speaker and not explicitly invoking the user) has charged characteristics that indicate its relevance as AS audio 157. For example, the ASADC model can classify the ambient sound 155 as having angry' speech, particularly loud speech, particularly emotional speech, etc. In some implementations, rather than using the ASADC model to classify charged speech, the charged speech detector 514 includes different detection mechanisms. For example, one implementation of the charged speech detector 514 is an envelope detector that effectively detects when generally classified speech is above a predetermined threshold volume (i.e., the envelope detector detects a certain amplitude envelope), when generally classified speech includes particular predefined spectral characteristics (i.e., the envelope detector detects a certain spectral envelope), etc. If so, the classified speech is still considered to be AS audio 157, and the speaker ESR engine 530 can be engaged. The speaker ESR engine 530 can use the representation vector 725 being generated by the ASADC model to produce a new ESR in association with the as -yet- unknown speaker. Once a stable new ESR has been generated, the new ESR becomes an active ESR 520, which can be used as an ESR signal 415 for the CSPT controller 430.
[0106] Typically, classifying human speech as part of the ambient sound 155 indicates that a conversation will ensue, or at least that further human speech is likely to be received. As such, passing the present ESR 520 as a conditional signal to the CSPT controller 430 causes the speech from that speaker to be selectively passed through and not cancelled by the ANC. As illustrated, the human speech AS audio 157 passed through by the CSPT controller 430 (i.e., delayed beyond the ANC processing window) is mixed by the audio output processor 330 into the output audio heard by the user via the WAC audio transducer(s) 105 (one or more speakers).
[0107] For AS audio 157 relating to alert audio, some embodiments assume that alert audio is a discrete audio event, rather than a potentially ongoing engagement with an audio stream. As one example, detecting and characterizing a ringing doorbell can be captured as “doorbell ring.” regardless of the duration of the doorbell ring, the specific tones of the doorbell ring, etc. For such alert audio, there may be no significant benefit to passing through the actual audio of that alert, in comparison to generating a corresponding alert. For example, instead of hearing the actual doorbell sound, a user may obtain essentially the same result from hearing a generic recorded doorbell sound, or a recorded or synthesized voice saying “doorbell,” or “your doorbell is ringing,” or “someone is at your door.”
[0108] In some embodiments, as described above, each different type of alert audio is classified as a respective ASA classification by the ASADC model. In some such embodiments, all alert audio is treated as discrete events, and the detected and classified alert audio event triggers an alert generator 540 to generate a corresponding alert. In some implementations, the alert generator 540 includes a selectable set of stored alert outputs. For example, each stored alert output is representative stored audio, such as pre-recorded speech, pre-recorded sample audio (e.g., sample audio of a doorbell, an alarm, a dog bark, etc.). In some implementations, the alert generator 540 additionally or alternatively includes a selectable set of pre-configured options for synthesis into alert outputs. For example, one of the selectable alerts is stored as the text string “Oven timer is beeping’"; when selected, the text string is passed through a text-to-speech synthesizer, which generates a spoken version of the string.
[0109] Some embodiments are configured to handle alert audio that is not a discrete event, including AS audio 157 that is neither human speech nor a discrete alert audio event. Some or all AS audio 157 classified by the ASADC model as alert audio can also generate a representation vector 725 that is an ESR of the AS audio 157, regardless of whether the AS audio 157 is human speech or a discrete alert audio event. For example, the ASADC model classified a beeping alarm in ambient sound 155 as a corresponding ASA classification and generates a representation vector 725 that becomes a present ESR 520 for the beeping alarm, which is then used as an ESR signal 415 for the CSPT controller 430. According to such an approach, the alert audio is selectively- passed through in the same manner as speech-related AS audio 157. Some embodiments can be configured, using a similar approach, to handle any other desired types of AS audio 157. such as non-human (e g., robotic, or synthesized) speech.
[0110] Some embodiments can be configured to classify certain human speech audio tags (e.g., words or phrases) as alert audio events. In such embodiments, one or more models, such as those used for the name / conversation detector 518, are trained to detect and classify particular spoken audio tags as alert audio, rather than as conversation starters. As one example, an approaching cyclist says, “on your left.” In one implementation, the ASADC model classifies this as human speech from an unknow n speaker, and another trained classifier model classifies the speech as including a known spoken audio tag. In another implementation, the ASADC model classifies the detected audio as the spoken audio tag. Either way. the spoken audio tag is not passed through by the CSPT controller 430 as speech, but is rather passed to the alert generator 540, which generates a corresponding alert output (e.g., a recorded sound of a bicycle bell, a synthesized voice saying “watch out behind you,” etc.).
[0111] In some embodiments, alert audio includes different priority levels. For example, some of the ASA classifications correspond to higher-priority (e.g.. critical) alert audio, and other ASA classifications correspond to lower-priority (e.g., non-critical) alert audio. Some embodiments permit the user to define which types of alerts are assigned to which priority levels and / or how different priority levels are treated. For example, critical alert audio can include the sound of a fire alarm, oncoming traffic, a beeping oven timer, etc. In some such embodiments, different prioritylevels can be assigned to control the audio output processor in different ways. For example, a non- critical (or medium-critical) alert audio event may be mixed into the output audio without any ANC gating, but with a slight temporary- reduction in the volume of the media audio 465 playback; while a critical alert audio event may cause any media audio 465 playback to be paused, the ANC to be disabled, and only the alert audio output to be played through as the output audio.
[0112] In some embodiments, as illustrated in FIG. 5. only alert audio above a certain priority level is treated as AS audio 157, and any alert audio below that priority level is logged in an event log 535. Some such embodiments can monitor whether the same logged audio alert is detected repeatedly some number of times within some duration, in which case the alert audio can be subsequently treated as higher priority. Some such embodiments can alert the user (e.g.. by outputting a notification, by displaying on a graphical user interface at the user’s request, etc.) as to the logged events. Each event in the event log 535 can be logged in any suitable manner with any suitable related information. For example, each event can be logged with a corresponding timestamp, etc.
[0113] FIG. 8 shows an embodiment of a CSPT controller 800, according to embodiments described herein. The CSPT controller 800 is an implementation of the CSPT controller 430 of FIGS. 4 and 5. For added context, FIG. 8 shows the CSPT controller 800 in communication with the AS ADC processor 410. In particular, FIG. 8 illustrates the AS ADC processor 410 generating one or more present ESRs 520, which can be (or from which can be derived) an ESR signal 415. As described herein, detection and classification of AS audio 157 (e.g. or at least of certain speech- related ty pes of AS audio 157) by the ASADC processor 410 results in generation of the ESR signal 415, and the ESR signal 415 is a conditional signal used to control selective pass through by the CSPT controller 800.
[0114] The CSPT controller 800 is implemented as a generative machine learning (artificial neural) network. As illustrated, the generative network includes an encoder 810. a bottleneck layer 820, and a decoder 830. The encoder 810 processes the ambient sound 155, extracting and compressing the most salient features into a lower-dimensional representation within thebotleneck layer 820. As described in context of other models herein, the botleneck layer 820 serves as a compressed latent space that encodes the essence of the ambient sound 155 while discarding irrelevant components. The decoder 830 reconstructs an output from the latent space, expanding the compressed information back into a form that resembles the salient portion of the ambient sound 155 (i.e., any detected and classified AS audio 157).
[0115] Although not explicitly shown, the ambient sound 155 can first be processed through a short-time Fourier transform (STFT) to convert the audio signal into a time-frequency representation. The encoder 810 can include one or more conformer. Each conformer is a type of transformer architecture optimized for sequential data, combining multi-headed self-atention with convolutional layers to capture both global and local dependencies within the audio signal. The encoder 810 can be trained to efficiently capture nuanced temporal paterns and spectral features of the ambient sound 155, enhanced by local context and cross-atention mechanisms, to accurately represent the audio in the latent space. The bottleneck layer 820 then condenses this representation in a manner that preserves salient audio features while minimizing dimensionality. The decoder 830 can also include one or more conformers. The decoder conformers can use the latent space representation to synthesize an output audio signal (i.e.. the passed-through AS audio 157).
[0116] As illustrated, embodiments of the CSPT controller 800 include a skip connection 840 between the encoder 810 and the decoder 830. The skip connection 840 allows the network to pas information directly from one or more layers of the encoder 810 to one or more layers of the decoder 830, bypassing the bottleneck layer 820 (e.g.. and one or more internal layers of the encoder 810 and / or decoder 830). In effect, the audio information from the ambient sound 155 can both pass through the botleneck layer 820 and pass around the bypass layer 820. This can appreciably enhance the quality of the generated output, such as by preserving finer details that might be lost in additional network layers (e.g., through additional compression). For example, the botleneck layer 820 is trained to capture the most salient features for determining what portions of the ambient sound 155 to pass through based on the ESR signal 415, but in doing so. the botleneck layer 820 may also omit details needed to reconstruct those passed through portions of the ambient sound 155 with high fidelity and a natural sound.
[0117] As described herein, embodiments of the CSPT controller 800 use the ESR signal 415 as a conditional signal. After the ambient sound 155 undergoes STFT and is processed by the encoder 81 (e.g., by multiple conformer layers), the botleneck layer 820 generates an ESR (a latent space representation) of the ambient sound 155. The ESR signal 415 (representing the ESR of the classified AS audio 157) is configured to modulate the ESR of the ambient sound 155 in thebotleneck layer 820. When the decoder 830 reconstructs the audio to generate the pass-through audio signal 432. it does so based on the modulated ESR, which is based on the ESR signal 415. In this way, the generated pass-through audio signal 432 at the output of the CSPT controller 800 is a regeneration of only (or substantially only) the classified AS audio 157 portions of the ambient sound 155.
[0118] The above illustrates at least two features of CSPT embodiments described herein. One such feature is that the selective passing through of AS audio 157 is based on classifying the ambient sound 155 into one or more ASA classifications by a trained artificial neural classification model. Another such feature is that the ‘‘passed through” audio is a generated pass-through audio signal 432 that is produced by a trained artificial neural generation model by converting the ambient sound 155 into a latent space representation and then regenerating a particular AS audio 157 portion of the ambient sound 155 by classification-based modulation of the latent space representation.
[0119] FIG. 9 shows a flow diagram of an illustrative method 900 for automated atention handling (AAH) in a wearable audio component (WAC), according to embodiments described herein. Embodiments of the method 900 can be implemented using any of the systems described herein, and / or any other suitable system. Embodiments begin at stage 904 by receiving an ambient audio signal. For example, the ambient audio signal is received via a reference microphone of a WAC. As described herein, it is assumed that, at least some of the time, the ambient audio signal includes one or more atention seeking audio (ASA) component signals.
[0120] At stag 908, embodiments can classify the ASA component into one or more ASA classifications of multiple predetermined ASA classifications. The classifying is responsive to presence of the ASA component of the ambient audio signal (e.g., the classifying can effectively include detection of the ASA component). At stage 912, embodiments can output a conditional signal based on the ASA component and the one or more ASA classifications. In some embodiments, the classifying at stage 908 and the outputing at stage 912 use a pre-trained artificial neural classifier network, such as described herein. In some such embodiments, the artificial neural classifier network includes an encoder pre-trained to encode the ambient speech into an embedded space representation (ESR) in a botleneck layer, and a decoder pre-trained to decode the ESR into the plurality of ASA classifications. In such embodiments, the conditional signal output in stage 912 is based on the ESR for at least one of the plurality of ASA classifications. In some cases (e.g., for one or more of the ASA classifications), the conditionalsignal is the ESR. In other cases, the conditional signal is generated from, derived from, or otherwise based on the ESR.
[0121] At stage 916. embodiments can encode the ambient audio signal into an ESR (e.g., another ESR). In some embodiments, the encoding uses a pre-trained artificial neural generator network, such as described herein. For example, the ambient audio signal is encoded into a first ESR in stage 916 and is encoded into a second ESR in stage 908. At stage 920, embodiments can modulate the ESR into a modulated ESR based on the conditional signal. At stage 924, embodiments can decode the modulated ESR to generate a pass-through audio signal corresponding to the ASA component. At stage 928, embodiments can output the pass-through audio signal. For example, the outputting can be to an output audio transducer (e.g., a speaker) of the WAC. Some embodiments, at stage 932, further include generating an audio output signal by volumetrically mixing the pass-through audio signal with media audio being played by a media player. Such embodiments, at stage 936, can output the audio output signal to the audio output transducer (e.g., of the WAC).
[0122] In some embodiments, generating the audio output signal in stage 932 is further by volumetrically mixing the pass-through audio signal and the media audio with an anti-noise signal generated by an active noise control (ANC) system. As described herein, the anti-noise signal is to mitigate the ambient audio signal based on a moving ANC window having an ANC window duration. In such embodiments, the outputting the pass-through audio signal at stage 928 comprises controlling delay of the outputting to ensure that the pass-through audio signal is delayed relative to the ambient audio signal by at least the ANC window duration. In some such embodiments, the controlling the delay is to ensure that the pass-through audio signal is delayed relative to the ambient audio signal by more than the ANC window duration and less than a predetermined maximum lip-sync delay.
[0123] In some embodiments, the method 900 further includes, at stage 940, generating an AAH control signal based on the ASA classification. Such embodiments, at stage 944. can control operation of the ANC system and / or the media player based on the AAH control signal. In some implementations, the control signal includes a gating signal that selectively pauses media playback based on the ASA classification. In some embodiments, the control signal includes a volumetric control signal that adjusts the volumetric mixing of output signals based on the ASA classification. In some embodiments, the control signal includes a gating signal that disables the ANC system based on the ASA classification. In some embodiments, the control signal includes a mode controlsignal that changes the mode of the ANC system based on the ASA classification, for example, between an active canceling mode, a conversation mode, a transparency mode, etc.
[0124] FIG. 10 provides a schematic illustration of an illustrative computational system 1000 that can implement various system components and / or perform various steps of methods provided by various embodiments. Embodiments of the computational system 1000 can be integrated in a WAC, such as an earbud, headset, etc. Embodiments of the computational system 1000 can implement some or all of the audio management system 100 of FIG. 1, including embodiments of the AHS 150, the ANC 140, and / or the audio processing system (APS) 160 described herein. Additionally or alternatively, embodiments can implement any of the training environments described herein, and / or can execute any of the machine learning models described herein. FIG. 10 is meant only to provide a generalized illustration of various components, any or all of which may be utilized as appropriate. FIG. 10. therefore, broadly illustrates how individual system elements may be implemented in a relatively separated or relatively more integrated manner.
[0125] The computational system 1000 is shown including hardware elements that can be electrically coupled via a bus 1005 (or may otherwise be in communication, as appropriate). The hardware elements may include one or more processors 1010, including, without limitation, one or more general-purpose processors and / or one or more special-purpose processors (such as digital signal processing chips, graphics acceleration processors, video decoders, and / or the like); and one or more input / output (I / O) devices 1015. In the WAC context, the I / O devices 1015 can include wired and / or wireless ports, buttons, switches, microphones, touch interfaces, indicator lights, displays, speakers, and / or any other suitable input and / or output devices.
[0126] The computational system 1000 may further include (and / or be in communication with) one or more non-transitory storage devices 1025, which can comprise, without limitation, local and / or network accessible storage, and / or can include, without limitation, a disk drive, a drive array, an optical storage device, a solid-state storage device, such as a random access memory ("RAM"), and / or a read-only memory ("ROM"), which can be programmable, flash-updateable and / or the like. Such storage devices may be configured to implement any appropriate data stores, including, without limitation, various file systems, database structures, and / or the like. In some embodiments, the storage devices 1025 include the AAH model(s) 350, representation data store 525, event log 535, and / or any other suitable storage for storing any other suitable data to support embodiments herein. For example, embodiments can store an inference model, which can include one or more types of name embedding models, relation networks, false rejection networks, etc. for implementing name detection-based attention handling features.
[0127] The computational system 1000 can also include a communications subsystem 1030, which can include, without limitation, a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication device, etc.), and / or the like. As described herein, the communications subsystem 1030 supports multiple communication technologies. Further, as described herein, the communications subsystem 1030 can provide communications with one or more networks. For example, embodiments of the communications subsystem 1030 can communicate via the cloud 350. Though not explicitly shown, some embodiments interface via the communications subsystem 1030, and / or via input devices 1015 and output devices 1020, with one or more user computational devices 320.
[0128] In many embodiments, the computational system 1000 will further include a working memory 1035, which can include a RAM or ROM device, as described herein. The computational system 1000 also can include software elements, shown as currently being located within the working memory 1035, including an operating system 1040, device drivers, executable libraries, and / or other code, such as one or more application programs 1045, which may include computer programs provided by various embodiments, and / or may be designed to implement methods, and / or configure systems, provided by other embodiments, as described herein. Merely by way of example, one or more procedures described with respect to the method(s) discussed herein can be implemented as code and / or instructions executable by a computer (and / or a processor within a computer); in an aspect, then, such code and / or instructions can be used to configure and / or adapt a general-purpose computer (or other device) to perform one or more operations in accordance with the described methods. In some embodiments, the operating system 1040 and the working memory' 1035 are used in conjunction with the one or more processors 1010 to implement some or all of the audio management system 100 components, such as the ANC 140, AHS 150. and / or APS 160. In particular, as illustrated, the one or more processors 1010 can implement the AS ADC processor 410, the AAH processor 420, and / or the audio output processor; and the operating system 1040 and the working memory 1035 can be used in conjunction with the one or more processors 1010 to implement some or all of the sound analytics block 510 components (e.g., the know'll speaker detector 512. the charged sound detector 514. the audio event detector 516. and / or the name / conversation detector 518), the speaker ESR engine 530, the alert generator 540, etc.
[0129] A set of these instructions and / or codes can be stored on a non-transitory computer- readable storage medium, such as the non-transitory storage device(s) 1025 described above. In some cases, the storage medium can be incorporated within a computer system, such as computer system 1000. In other embodiments, the storage medium can be separate from a computer system(e.g., a removable medium, such as a compact disc), and / or provided in an installation package, such that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the instruct ons / code stored thereon. These instructions can take the form of executable code, which is executable by the computational system 1000 and / or can take the form of source and / or installable code, which, upon compilation and / or installation on the computational system 1000 (e.g., using any of a variety of generally available compilers, installation programs, compression / decompression utilities, etc.), then takes the form of executable code.
[0130] It will be apparent to those skilled in the art that substantial variations may be made in accordance with specific requirements. For example, customized hardware can also be used, and / or particular elements can be implemented in hardw are, softw are (including portable softw are, such as applets, etc.), or both. Further, connection to other computing devices, such as network input / output devices, may be employed.
[0131] As mentioned above, in one aspect, some embodiments may employ a computer system (such as the computer system 1000) to perform methods in accordance with various embodiments of the invention. According to a set of embodiments, some or all of the procedures of such methods are performed by the computational system 1000 in response to processor 1010 executing one or more sequences of one or more instructions (which can be incorporated into the operating system 1040 and / or other code, such as an application program 1045) contained in the working memory 1035. Such instructions may be read into the working memory 1035 from another computer-readable medium, such as one or more of the non-transitory storage device(s) 1025. Merely by way of example, execution of the sequences of instructions contained in the working memory 1035 can cause the processor(s) 1010 to perform one or more procedures of the methods described herein.
[0132] The terms "machine-readable medium,” “computer-readable storage medium” and “computer-readable medium,” as used herein, refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. These mediums may be non-transitory. In an embodiment implemented using the computer system 1000, various computer-readable media can be involved in providing instructions / code to processor(s) 1010 for execution and / or can be used to store and / or carry such instructions / code. In many implementations, a computer- readable medium is a physical and / or tangible storage medium. Such a medium may take the form of a non-volatile media or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as the non-transitory storage device(s) 1025. Volatile media include, without limitation, dynamic memory', such as the w orking memory' 1035. Common forms of physicaland / or tangible computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, any other physical medium with patterns of marks, a RAM, a PROM, EPROM, a FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.
[0133] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to the processor(s) 1010 for execution. Merely by way of example, the instructions may initially be carried on a magnetic disk and / or optical disc of a remote computer. A remote computer can load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received and / or executed by the computer system 1000. The communications subsystem 1030 (and / or components thereof) generally will receive signals, and the bus 1005 then can carry the signals (and / or the data, instructions, etc., carried by the signals) to the working memory 1035, from which the processor(s) 1010 retrieves and executes the instructions. The instructions received by the working memory 1035 may optionally be stored on a non-transitory storage device 1025 either before or after execution by the processor(s) 1010.
[0134] Directional Selective Pass Through
[0135] In some automated attention handling embodiments, the CSPT controller 430 is configured to selectively pass through components of the ambient sound 155 determined to be directed at the user. As one example, office conversations are generally muted to the user by the ANC system 140, but those sounds are passed through when a component of the sound is determined to be someone speaking directly at the user. As another example, road sounds are generally muted to the user by the ANC system 140, but a critical type of sound (a car horn) is passed through when sufficiently directed toward (e.g., approaching and sufficiently close as to be relevant to) the user.
[0136] FIG. 11 shows a block diagram of an illustrative directional selective pass through (SPT) environment 1100. according to embodiment described herein. The directional SPT environment 1 100 is intended to fit within the environment of FIG. 4 or 5, including at least the ASADC processor 410 and CSPT controller 430. The ASADC processor 410 includes a directionality detector 1110. In some embodiments, the directionality detector 1110 is implemented by the charged sound detector 514 of FIG. 5. such that the directionality detector 1110 can be part of the sound analytics block 510. In other embodiments, the directionality detector 1 110 is a separatedetector of the sound analytics block 510. In other embodiments, the directionality detector 1110 is separate from the sound analytics block 510.
[0137] As illustrated, the directionality detector 1110 includes a directional analyzer 1120 and an identification of crucial directional sounds (ICDS) block 1 130. Embodiments of the directional analyzer 1120 include a machine learning network (e.g., an artificial neural network) trained to ascertain the direction from which components of the ambient sound 155 originate, including dominant directionality with respect to the user. Embodiments assume that the ambient sound 155 is received by multiple microphones. Some such embodiments exploit the presence of multiple WACs, such as when the user is wearing earbuds, and each earbud has a physically separated reference microphone 110. Other such embodiments, additionally or alternatively, implement the reference microphone 110 as a microphone array (e g., multiple separate microphones, a transducer array, etc.). As such, substantially the same ambient sound 155 is received concurrently by multiple microphones (e g., with relative phase shift, difference in amplitude, etc.), and embodiments of the directional analyzer 1120 can determine directionality based on variances in arrival times, variances in intensity levels across, wavelet propagation characteristics, etc.
[0138] Embodiments of the ICDS block 1130 continuously monitor the directionality of ambient audio components to determine when there is a component exhibiting directionality relevant to the user. As illustrated, the ICDS block 1130 can communicate with the sound analytics block 510 to obtain components of the ambient sound 155, such as those assigned to one or more ASA classifications. In some embodiments, the ICDS block 1130 outputs a directionality trigger signal in response to detecting that one or more components of the ambient sound 155 is directed at the user. In some embodiments, and / or in certain instances (e.g., dependent on default or user- controlled settings), any components of the ambient sound 155 for which the ICDS block 1130 generates the directionality trigger signal are treated as AS audio 157 (regardless of ASA classification). In such cases, embodiments can automatically generate a corresponding active ESR 520 from which to output a corresponding ESR signal 415 for causing the CSPT controller 430 to pass through the associated directional component of the ambient sound 155. In some embodiments, the directionality trigger signal is used as an additional condition for controlling the CSPT controller 430 after a component of the ambient sound 155 is determined to be AS audio 157. For example, one or more components of the sound analytics block 510 determines that a component of the ambient sound 155 should be treated as candidate AS audio 157, but that AS audio 157 component is only passed through by the CSPT controller 430 when it is also determined to be directionally relevant to the user.
[0139] In some embodiments, all components of ambient sound 155 classified as speech from a known speaker are passed through by the CSPT controller 430; in other embodiments, such speech is only passed through if also determined by the directi onality detector 1110 to be directed at the user. In some embodiments, all components of ambient sound 155 classified as speech, regardless of whether it is from a know n or unknow n speaker, is passed through by the CSPT controller 430 if determined by the directionality detector 1110 to be directed at the user (i.e., any speech is treated as AS audio 157 if directed at the user). In some embodiments, all components of ambient sound 155 classified as alert audio are passed through by the CSPT controller 430; in other embodiments, such alert audio is only passed through if also determined by the directionality detector 1110 to be directed at the user.
[0140] In some embodiments, the directionality trigger signal can be combined with other conditions to determine when the CSPT controller 430 should pass through corresponding AS audio 157 components of the ambient sound 155, and / or otherwise how to behave in response to the presence of those components. For example, the directionality trigger signal can operate, so that AS audio 157 components of different ASA classifications are treated differently based on the directionality of the AS audio 157 component along with one or more of the user's present location, the volume or envelope of the AS audio 157 component, whether the AS audio 157 component is human speech, whether the AS audio 157 component is from a known speaker, whether the AS audio 157 component invokes the user’s name or other enrolled tag, whether the AS audio 157 component includes a particular priority level of alert audio, etc.
[0141] Additionally or alternatively, embodiments can use the directionality trigger signal to impact the AAH audio signal 424 and / or the AAH control signal 422. In some such embodiments, the directionality trigger signal can participate in control of the mixing controller 450 of the audio output processor 440, such as by increasing the relative volume of AS audio 157 components determined to be more directionally relevant. In other such embodiments, the directionality trigger signal can participate in control of the gating controller 445 of the audio output processor 440. such as by disabling ANC as soon as a directionally relevant AS audio 157 component is detected with a certain level of criticality, and / or muting playback of media audio 465 w hen a directionally relevant AS audio 157 component is detected.
[0142] Known Speaker Detection
[0143] FIG. 12 shows an illustrative embodiment of a known speaker detector 1200 implemented by a machine learning model (an artificial neural network), according to embodiments described herein. The known speaker detector 1200 is an implementation of theknown speaker detector 512 of FIG. 5. As illustrated, the known speaker detector 1200 includes the encoder 710 and bottleneck layer 720 trained in FIGS. 7A and 7B, such as the encoder 710 and bottleneck layer 720 of the AS ADC model 700b of FIG. 7B. As described above, the encoder 710 and bottleneck layer 720 are trained so that when human speech is present in the ambient sound 155 to produce a representation vector 725 that is characteristic of the speaker of the human speech.
[0144] As illustrated, the known speaker detector 1200 can also include a matching engine 1250. Embodiments of the matching engine 1250 include an identification network 1210 and a verification network 1220. Embodiments of the known speaker detector 1200 can operate in a similar manner to the name / conversation detector 600 described in FIG. 6. The encoder 710 and bottleneck layer 720 of the ASADC model 700b of FIG. 7B can be considered as an embedding model (similar to the embedding model 615 of FIG. 6) for generating the representation vector 725 from the ambient sound 155. The identification network 1220 can be implemented in a similar manner to the relation network 635 in the identification stage 630 of FIG. 6. The verification network 1220 can be implemented in a similar manner to the false rejection network 645 in the verification stage 640 of FIG. 6. The representation data store 525 can operate in a similar manner to the deep image 620 of FIG. 6. In some embodiments, where both the known speaker detector 1200 and the name / conversation detector 600 are included in the sound analytics block 510 of FIG. 5, the representation data store 525 and the deep image 620 of FIG. 6 can be stored in a same storage medium.
[0145] In runtime, the ambient sound 155 signal is received by the encoder 710 of the known speaker detector 1200, and the bottleneck layer 720 generates a representation vector 725. When human speech is present in the ambient sound 155, the representation vector 725 characterizes the speaker of the speech. The identification network 1210 compares the representation vector 725 to embedded space representations for known speakers stored in the representation data store 525 to select a most likely candidate speaker. The verification network 1220 verifies the selection (e.g., by transforming the vectors into several spaces) to confirm or reject the selection. The output from the verification network 1220 is a know n speaker signal 1265 that identifies the known speaker matching the detected speech.
[0146] In general, when a speech-related audio component is classified as part of the ambient sound 155. three general categories of possibility can be explored. The first category is where the speech is determined by the known speaker detector 1200 to be from a known speaker, such that the speech is considered to be AS audio 157. The second category is where the speech isdetermined (e.g., by the known speaker detector 1200) to be from an unknown speaker, and the speech is further determined (e.g., by the charged sound detector 514, the name / conversation detector 518, or otherwise) to be relevant to the user and thus to be AS audio 157. The third category is where the speech is determined (e.g., by the known speaker detector 1200) to be from an unknow n speaker, and the speech is not determined (e.g., by any of the charged sound detector 514, the name / conversation detector 518, or otherwise) to be relevant to the user, such that the speech is not considered to be AS audio 157. Embodiments herein generally ignore the third category.
[0147] As described in FIG. 5, with respect to the second category, embodiments seek to generate (e g., by the speaker ESR engine 530) an ESR to characterize the unknown speaker. The generated ESR can then be used as an active ESR 520, which can be passed as an ESR signal 415 (i.e., a conditional signal) to cause the CSPT controller 430 to effectively pass through (or. more accurately, to selectively regenerate) the portion of the ambient sound 155 corresponding to the AS audio 157 speech of the unknown speaker. In some embodiments, a similar technique can be applied to first category' speech from a known user. Even though the known speaker has a stored ESR in the representation data store 525, embodiments can use the new speech samples from the ambient sound 155 to update and / or refine the stored ESR for that speaker.
[0148] FIG. 13 shows a block diagram of an illustrative speaker embedded space representation (ESR) generation environment 1300, according to embodiments described herein. The speaker ESR generation environment 1300 includes a feedback loop from the CSPT controller 430 back to the ASADC processor 410 via an embodiment of the speaker ESR engine 530. Embodiments of FIG. 5 include this feedback loop, even though it is not explicitly shown in FIG. 5.
[0149] As described with reference to FIG. 5, ambient sound 155 is received by the ASADC processor 410. The sound analytics block 510 characterizes the ambient sound 155 to look for one or more ASA classifications. In the event that speech-related AS audio 157 is classified, the sound analytics block 510 can generate and output an ESR that seeks to characterize the speaker of the classified speech. This becomes one of one or more active ESR(s) 520, which can be passed as an ESR signal 415 (i.e., a conditional signal) to cause the CSPT controller 430 to effectively pass through (or, more accurately, to selectively regenerate) the portion of the ambient sound 155 corresponding to the AS audio 157 speech. As a result, the CSPT controller 430 outputs a pass- through audio signal 432 corresponding to the speech-related AS audio 157 classified from (i.e.. the AS audio component of) the ambient sound 155.
[0150] As illustrated, the pass-through audio signal 432 can also be fed back to the speaker ESR engine 530. In the speaker ESR engine 530, the pass-through audio signal 432 can be processed by a remove silence block 1310 to isolate only the portions of the pass-through audio signal 432 that represent the speaker’s speech. The signal can then be passed to an implementation of the AS ADC model 1320. In some embodiments, the AS ADC model 1320 is, or is implemented as an instance of, the ASADC model 700b of FIG. 7B. As described with reference to FIG. 7B, the AS ADC model 1320 is configured to output a representation vector 1325 (e.g.. representation vector 725 of FIG. 7B) that is an ESR that captures the most salient features of the input signal for characterizing the speaker.
[0151] In some embodiments, a stability analyzer block 1330 is used to determine whether the ASADC model 1320 appears to be generating the representation vector 1325 in a stable manner. For example, the stability analyzer block 1330 analyzes the output of the ASADC model 1320 over time (e g., using a moving window, or the like) to determine whether the output appears to be sufficiently consistent to be considered as a stable ESR. In general, particularly when starting with the ESR for an unknown speaker, there may be some period of time during which the ASADC model 1320 is adapting to leam the most salient features. During that period of time, the stability analyzer block 1330 can determine that the generated representation vector 1325 is not stable.
[0152] As illustrated, while the generated representation vector 1325 is not stable, the speaker ESR generation environment 1300 will remain in a holding pattern. In this state, the ASADC model 1320 will continue to receive the passed-through speech-related AS audio 157 component of the ambient sound 155 until it is able to generate a stable representation vector 1325. Once the stability analyzer block 1330 determines that the ASADC model 1320 is generating a stable representation vector 1325, the representation vector 1325 can be used to update the ESR (block 1335), and the updated ESR can be fed back to update the active ESR 520 and thereby to update the ESR signal 415 being used to control the CSPT controller 430.
[0153] Some embodiments provide the user with the opportunity to capture or confirm the identity of a classified speaker as a known speaker using a speaker enroll / confirm block 1345. In some such embodiments, subsequent to classification of speech as from a known user, the speaker enroll / confirm block 1345 prompts the user (e.g., via any suitable user interface on any suitable user device) to confirm that the known speaker was identified correctly (e.g., by the known speaker detector 1200). In some implementations, identification of the known speaker includes generating a confidence score that indicates the statistical likelihood that the identification is correct; and such confirmation of a know n speaker is only performed if the confidence score isbelow a predetermined threshold. In some embodiments, the ESR for the know n speaker is updated, refined, replaced, or otherwise modified based on the confirmation. For example, the feedback ESR updating process (e.g., by ESR block 1335) can determine how heavily to weight the present ESR for a known speaker (e.g., the stored ESR) relative to the new ESR generated by the ASADC model 1320 based on whether there is a positive or negative confirmation of the identified speaker from the user. Alternatively or additionally, each stored ESR can be associated with a confidence score that increases or decreases over time based on each instance of the user agreeing or disagreeing with a speaker identification based on that stored ESR.
[0154] In unknown speaker cases, the speaker enroll / confirm block 1345 can prompt the user to enroll the unknow n speaker as a new' know n speaker. Embodiments can include a user interface (e.g., any suitable user interface on any suitable user device) that allows users to manage their know'll speakers. In some cases, the indicates to the user that it seems the user just had a conversation with an unknown user and asks whether the user wants to enroll that user for future identification as a know n user. In some implementations, the interface also allows a user to prioritize known speakers, associate known speakers with particular locations, delete previously enrolled speakers, etc. If the user opts to enroll the speaker as a known speaker, the generated ESR can be stored in the representation data store 525.
[0155] Different implementations can handle new users in different ways, such as in the case of a user first w earing a WAC. In some implementations, the set of known speakers is empty; all attention-seeking speakers are treated as unknown users until they are detected and subsequently enrolled by the user. In other implementations, the set of known speakers can be populated by previously generated ESRs. As one example, a user replacing or upgrading their WAC can reload their previously enrolled known speakers, so that the WAC is initially populated with known speaker ESRs. As another example, enrolled speaker ESRs may be shared between users, such as among members of a family, roommates or officemates, etc.
[0156] In some embodiments, confirmation or enrollment of speakers by the speaker enroll / confirm block 1345 only occurs after a ’‘conversation” is detected. Such embodiments can include a detect conversation end block 1340 to determine at least when the classified speech from the unknow n user has ended. Some implementations of the conversation end block 1340 detect an end of the “conversation” by determining that there has not been additional speech from that speaker for some predetermined threshold amount of time (e.g.. two minutes). Some implementations of the conversation end block 1340 detect an end of the “conversation” by determining that there has not been additional speech from that speaker and from any otherspeaker (including the user) for some predetermined threshold amount of time (e g., to account for situations in which the conversation is continuing with turn taking). Some implementations of the conversation end block 1340 detect an end of the ‘'conversation” by further determining that there has been a conversation start prior to determining that there is a conversation end. For example, such implementations determine a conversation start after a predetermined threshold amount of speech has been received from the speaker (e.g., the speaker is only considered relevant enough to be a candidate known speaker if there has been a sufficiently long conversation with that speaker).
[0157] FIG. 14 shows a flow diagram of a method 1400 for automated attention handling (AAH) for a wearable audio component (WAC), including known speaker detection, according to embodiments described herein. Embodiments of the method 1400 begin at stage 1404 by receiving an ambient audio signal. At stage 1408. embodiments encode the ambient audio signal using an artificial neural classifier network into an embedded space representation (ESR) (referred to as the “second” ESR in this context). As described herein, the encoding is such that, when the ambient audio signal comprises a human speech component, the second ESR characterizes a speaker of the human speech component (whether that speaker is known or not).
[0158] At stage 1412, embodiments compare the second ESR with a set of first ESRs to determine whether the human speech component is being spoken by one of a set of known speakers. As described herein, the set of first ESRs is stored in a non-transient representation data store, and each of the set of first ESRs is previously enrolled by a user of the WAC as characterizing an associated one of the set of known speakers. In some embodiments, the comparing includes determining and outputting a candidate ESR as one of the set of first ESRs having above a predetermined threshold similarity with the second ESR. In some such embodiments, the comparing further includes determining whether to confirm the candidate ESR based on determining whether the second ESR can be discriminated from the candidate ESR in excess of a predetermined discrimination threshold in any of multiple mathematical spaces. Such embodiments can output the second ESR as the conditional signal is responsive to determining to confirm the candidate ESR.
[0159] At stage 1416 embodiments can output the second ESR as a conditional signal responsive to determining (i.e., in cases when it is determined) that the human speech component is being spoken by one of the set of known speakers. As described herein, some embodiments can still output the second ESR as a conditional signal when the speech component of the ambient sound signal is otherwise determined to be attention seeking audio.
[0160] At stage 1420, embodiments can encode the ambient audio signal with an artificial neural generator network into a third ESR. As described herein, at stage 1424, embodiments can modulate the third ESR in the artificial neural generator network into a fourth ESR based on the conditional signal. For example, the third ESR is the latent space representation of the ambient audio in a bottleneck layer of the artificial neural generator network. At stage 1428, embodiments can decode the fourth ESR by the artificial neural generator network to generate a pass-through audio signal, such that the pass-through audio signal corresponds to the human speech component of the ambient audio signal responsive to the human speech component being spoken by one of the set of know n speakers.
[0161] In some embodiments, the method 1400 further includes outputting the pass-through audio signal by the artificial neural generator network to an audio output processor in stage 1432.
[0162] As illustrated, some embodiments can further include executing a speaker confirmation / enrollment procedure 1414 based on the comparing at stage 1412. For example, the speaker confirmation / enrollment procedure 1414 can be an implementation of FIG. 13. In some embodiments, the speaker confirmation / enrollment procedure 1414 includes outputting a speaker identity message to the user via a user interface responsive to the comparing at stage 1412. The speaker identity message indicates a proposed identity of the one of the set of known speakers when the comparing determines that the human speech component is being spoken by one of the set of know n speakers, and the speaker identity message indicates an unknown identity' (e.g., that the identity is unknown) when the comparing determines that the human speech component is not being spoken by one of the set of known speakers. When the comparing determines that the human speech component is being spoken by one of the set of known speakers, the speaker identity message can include a prompt to the user to confirm or deny the proposed identity. In such cases, embodiments can receive a user message responsive to the user interacting with the prompt indicating confirmation or denial of the proposed identity by the user and can update at least the artificial neural classifier network and / or the representation data store (e.g., and the comparing in stage 1412) based on the confirmation or denial of the proposed identity by the user. When the comparing determines that the human speech component is not being spoken by one of the set of known speakers, the speaker identity message can include a prompt to the user whether to enroll the unknown identity. In such cases, embodiments can receive a user message responsive to the user interacting with the prompt indicating whether to enroll the unknown identity. If not, embodiments can ignore (i.e., not treat it as AS audio) or continue to treat the speech as AS audio from an unknown speaker (e.g., if determined to invoke an enrolled name or tag, if determined to be directed or angry, etc.). If the user indicates a desire to enroll the unknown identity, a ewspeaker enrollment can commence, including at least storing the second ESR in the representation data store to subsequently be one of the set of first ESRs.
[0163] Location-Based CSPT and Multi-Conditionality
[0164] Some embodiments of AAH systems described herein can account for additional present conditions to determine the manner in which to perform AAEI features, including CSPT. FIG. 15 shows a block diagram of an illustrative configurable AAH environment 1500, according to some embodiments described herein. The configurable AAH environment 1500 can be implemented by the configuration processor 428 (also shown in FIG. 4). The configuration processor 428 can be implemented by one of the one or more processors 1010 of the computational environment 1000 of FIG. 10.
[0165] The configuration processor 428 is generally configured to manage configuration settings in accordance with a stored AAH condition matrix 1510, determine a present attention handling condition based on one or more condition detectors 1520 and / or present outputs of the AS ADC processor 410, generate a present configuration 1530 based on the present attention handling condition, and control one or more attention handling features, accordingly. As illustrated, the AAH condition matrix 1510 can represent any suitable number and type of settable conditions 1504 for any of the ASA classifications 1502. As described herein, the ASADC processor 410 detects and classifies ASA components of ambient audio 155 into ASA classifications, such as speech from a known speaker, speech directed at the user, critical alert audio, non-critical alert audio, etc. For each of these ASA classifications 1502, the AAH condition matrix 1510 can define whether and how to implement AAH features based on values (i.e., settings) of the settable conditions 1504.
[0166] In some embodiments, the values of some or all settable conditions 1504 are set by default. For example, configuration processor 428 is factory set to accommodate one or more default configurations stored in the AAH condition matrix 1510. In some embodiments, as described with reference to FIG. 4, the configuration processor 428 can be in communication with a user interface 427. In such embodiments, the user interface 427 allows the user to configure or reconfigure some or all of the values of the settable conditions 1504 for some or all of the ASA classifications 1502, thereby modifying the AAH condition matrix 1510. Some embodiments can use machine learning to update the AAH condition matrix 1510 dynamically based on learning over time what changes to the AAH condition matrix 1510 would seem to be more desirable to the user. In some such embodiments, the machine learning automatically updates the AAH condition matrix 1510 based on learned user behaviors, learned condition tendencies, etc. In other suchembodiments, the machine learning automatically suggests updates to the AAH condition matrix 1510 by prompting the user accordingly (e.g., via the user interface 427). and the user can determine whether to accept or reject those suggested changes.
[0167] Embodiments include one or more condition detectors 1520. In the illustrated embodiments, the condition detectors 1520 includes one or more of a location detector 1522, a time detector 1524, and a mode detector 1526. The location detector 1522 determines a present location of the WAC(s)s as in use by the user. The location detector 1522 can use any suitable locating components integrated in the WAC and / or in a mobile device or other device of the user in local communication with the WAC. For example, the location detector 1522 can include or exploit a global positioning satellite (GPS) receiver, accelerometer, and / or other relevant component. Some implementations can determine the present location as geographic coordinates, or as an approximate street address. Some implementations can determine the present location categorically (e.g., outdoors, busy urban, indoor commercial, indoor residential, street, etc.). Some implementations can determine the present location based on being within some range of one of a set of user-configured locations (e.g., home, work, school, etc.).
[0168] The time detector 1524 determines a present time for the user. The time detector 1524 can use any suitable locating components integrated in the WAC and / or in a mobile device or other device of the user in local communication with the WAC, such as any suitable clock. Some implementations can determine the present time as a system time or clock time (e.g., 13:45). Some implementations can determine the present time categorically (e.g., morning, middle of the night, mid-afternoon, etc.). The mode detector 1526 determines a present configuration mode for the WAC being used by the user. The mode detector 1526 can determine the present mode as one of a set of default (preset) modes and / or as one of a set of user-defined modes. In some cases, the modes represent a defined combination and / or range of conditions as detected by the location detector 1522 and / or the time detector 1524. In other cases, the mode definitions facilitate other types of settings. For example, a user can define a “meditation mode’7that seeks to cancel all ambient sound 155, except for critical alert audio. In some embodiments, in some or all conditions, different ones of the condition detectors 1520 can operate hierarchically (e.g., in certain locations, the corresponding location-based configuration takes priority over any timebased configuration).
[0169] As described herein, the ASADC processor 410 generally receives the ambient sound 155 and outputs a conditional signal to control the CSPT controller 430 based on detecting and classifying one or more ASA components of the ambient sound 155. In some embodiments, asillustrated in FIG. 13, the manner in which the CSPT controller 430 is controlled is further based on the present configuration 1530. In some embodiments, additionally or alternatively, the manner in which the audio output controller 440 is controlled is further based on the present configuration 1530.
[0170] In one set of implementations, subsets of known speakers are associated with different conditions (e.g., with different locations, times of day, modes, etc.). For example, certain speakers are “known speakers" (e.g., speech from those speakers is identified as AS audio 157) only at the office and / or during the workday. In another set of implementations, different enrolled names are associated with different conditions. For example, “mom” or “Jane” may be considered as an enrolled name for name / conversation detection while at home, while “Mrs. Smith” or “boss” may be considered as an enrolled name for name / conversation detection while at work. In another set of implementations, charged speech is treated differently in different conditions. For example, approaching voices may be treated as AS audio 157 when the user is in a relatively unpopulous area (e.g., at work, at home, etc.), but not when the user is in a busy urban area. In another set of implementations, subsets of alert audio are associated with different conditions. For example, a horn honking is classified as AS audio 157 only when the user is in or near a street, but is not identified as such when the user is indoors or otherwise away from traffic. In another set of implementations, subsets of alert audio are assigned to different priorities in different conditions.
[0171] In some embodiments, the present configuration 1530 is used to impact operation of the audio output processor 440 (e.g., by affecting the AAH audio signal 424 and / or the AAH control signal 422). In some such embodiments, the present configuration 1530 can participate in control of the mixing controller 450 of the audio output processor 440, such as by determining an appropriate relative volume of different AS audio 157 components based on present location, time of day, etc. In other such embodiments, the present configuration 1530 can participate in control of the gating controller 445 of the audio output processor 440, such as by determining whether to disable ANC, to mute playback of media audio 465. etc. based on the present location, time of day, etc.
[0172] Having described several example configurations, various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the disclosure. For example, the above elements may be components of a larger system, wherein other rules may take precedence over or otherwise modify the application of the invention. Also, a number of steps may be undertaken before, during, or after the above elements are considered.
Claims
WHAT IS CLAIMED IS:
1. An automated attention handling (AAH) system for integration with a wearable audio component (WAC), the AAH system comprising: a non-transient representation data store having a set of first embedded space representations (ESRs) stored thereon, each of the set of first ESRs previously enrolled as characterizing an associated one of a set of known speakers; a known speaker detector comprising: a first artificial neural network trained to encode a received ambient audio signal into a second ESR, such that, when the ambient audio signal comprises a human speech component, the second ESR characterizes a speaker of the human speech component; and a matching engine configured to compare the second ESR with the set of first ESRs to determine whether the human speech component is being spoken by one of the set of known speakers, and to output the second ESR as a conditional signal responsive to determining that the human speech component is being spoken by one of the set of known speakers; and a classification-based selective pass-through (CSPT) controller comprising a second artificial neural network trained to: encode the ambient audio signal into a third ESR and modulate the third ESR into a fourth ESR based on the conditional signal; and decode the fourth ESR to generate a pass-through audio signal, such that the pass-through audio signal corresponds to the human speech component of the ambient audio signal responsive to the human speech component being spoken by one of the set of known speakers.
2. The AAH system of claim 1, wherein the CSPT controller is further configured to output the pass-through audio signal to an audio output processor.
3. The AAH system of claim 1, wherein the matching engine comprises: an identification netw ork coupled with the first artificial neural netw ork and the representation data store and configured to output a candidate ESR as one of the set of ESRs having above a predetermined threshold similarity with the second ESR.
4. The AAH system of claim 3, wherein:the matching engine further comprises a verification network coupled with the identification network to determine whether to confirm the candidate ESR based on determining whether the second ESR can be discriminated from the candidate ESR in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces; and the matching engine is configured to output the second ESR as the conditional signal responsive to determining to confirm the candidate ESR.
5. The AAH system of claim 1, wherein the matching engine is further configured to output a speaker identification to a user interface responsive to determining that the human speech component is being spoken by one of the set of known speakers, the speaker identification indicating an identity of the one of the set of known speakers.
6. The AAH system of claim 5, wherein: the output of the speaker identification to the user interface comprises a prompt to confirm or deny whether the speaker identification correctly identifies the one of the set of known speakers that is presently speaking the human speech component; the known speaker detector is further configured to receive a message responsive to the user interacting with the prompt and to update the artificial neural network, the matching engine, and / or the representation data store based on whether the user indicated to confirm or deny.
7. A method for automated attention handling (AAH) for a wearable audio component (WAC), the method comprising: receiving an ambient audio signal; encoding, with an artificial neural classifier network , the ambient audio signal into a second embedded space representation (ESR), such that, when the ambient audio signal comprises a human speech component, the second ESR characterizes a speaker of the human speech component; comparing the second ESR with a set of first ESRs to determine whether the human speech component is being spoken by one of a set of known speakers, the set of first ESRs stored in a non-transient representation data store, each of the set of first ESRs previously enrolled as characterizing an associated one of the set of known speakers; outputting the second ESR as a conditional signal responsive to determining that the human speech component is being spoken by one of the set of known speakers; encoding the ambient audio signal with an artificial neural generator network into a third ESR;modulating the third ESR in the artificial neural generator network into a fourth ESR based on the conditional signal; and decoding the fourth ESR by the artificial neural generator network to generate a pass-through audio signal, such that the pass-through audio signal corresponds to the human speech component of the ambient audio signal responsive to the human speech component being spoken by one of the set of known speakers.
8. The method of claim 7, further comprising: outputting the pass-through audio signal by the artificial neural generator network to an audio output processor.
9. The method of claim 7, wherein the comparing comprises determining and outputting a candidate ESR as one of the set of ESRs having above a predetermined threshold similarity with the second ESR.
10. The method of claim 9, wherein: the comparing further comprises determining whether to confirm the candidate ESR based on determining whether the second ESR can be discriminated from the candidate ESR in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces; and the outputting the second ESR as the conditional signal is responsive to determining to confirm the candidate ESR.
11. The method of claim 7, further comprising: outputting a speaker identity message via a user interface responsive to the comparing, such that the speaker identity message indicates a proposed identity of the one of the set of known speakers when the comparing determines that the human speech component is being spoken by one of the set of known speakers, and the speaker identity message indicates an unknown identity when the comparing determines that the human speech component is not being spoken by one of the set of know n speakers.
12. The method of claim 11, wherein the speaker identity message includes a prompt to confirm or deny the proposed identity when the comparing determines that the human speech component is being spoken by one of the set of known speakers, and further comprising: receiving a response message responsive to the prompt indicating confirmation or denial of the proposed identity; andupdating at least the artificial neural classifier netw ork and / or the representation data store based on the confirmation or denial of the proposed identity.
13. The method of claim 11, wherein the speaker identity message includes a prompt to enroll the unknown identity when the comparing determines that the human speech component is not being spoken by one of the set of known speakers, and further comprising: receiving a response message responsive to the prompt indicating whether to enroll the unknown identity; and executing, responsive to the response message indicating to enroll the unknown identity, a new speaker enrollment comprising storing the second ESR in the representation data store to subsequently be one of the set of first ESRs.
14. A computer program comprising instructions for implementing the method of claim 7.
15. A system for automated attention handling (AAH) for a wearable audio component (WAC), the system comprising: one or more processors to execute at least an artificial neural classifier network and an artificial neural generator network; a non-transitory computer-readable storage medium having instructions stored thereon which, when executed, cause the one or more processors to perform steps comprising: receiving an ambient audio signal; encoding, with the artificial neural classifier network, the ambient audio signal into a second embedded space representation (ESR), such that, when the ambient audio signal comprises a human speech component, the second ESR characterizes a speaker of the human speech component; comparing the second ESR with a set of first ESRs to determine whether the human speech component is being spoken by one of a set of known speakers, the set of first ESRs stored in a non-transient representation data store, each of the set of first ESRs previously enrolled as characterizing an associated one of the set of known speakers; outputting the second ESR as a conditional signal responsive to determining that the human speech component is being spoken by one of the set of known speakers; encoding the ambient audio signal with the artificial neural generator network into a third ESR; modulating the third ESR in the artificial neural generator netw ork into a fourth ESR based on the conditional signal; anddecoding the fourth ESR by the artificial neural generator network to generate a pass-through audio signal, such that the pass-through audio signal corresponds to the human speech component of the ambient audio signal responsive to the human speech component being spoken by one of the set of known speakers.
16. The system of claim 15, further comprising: responsive to determining that the human speech component is not being spoken by one of the set of known speakers.
17. The system of claim 15, further comprising: outputting the pass-through audio signal by the artificial neural generator network to an audio output processor.
18. The system of claim 15, wherein the comparing comprises determining and outputting a candidate ESR as one of the set of ESRs having above a predetermined threshold similarity with the second ESR.
19. The system of claim 18, wherein: the comparing further comprises determining whether to confirm the candidate ESR based on determining whether the second ESR can be discriminated from the candidate ESR in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces; and the outputting the second ESR as the conditional signal is responsive to determining to confirm the candidate ESR.
20. The system of claim 15, further comprising: outputting a speaker identity message via a user interface responsive to the comparing, such that the speaker identity message indicates a proposed identity of the one of the set of known speakers when the comparing determines that the human speech component is being spoken by one of the set of known speakers, and the speaker identity7message indicates an unknown identity when the comparing determines that the human speech component is not being spoken by one of the set of known speakers.
Citation Information
Patent Citations
Automated attention handling in active noise control systems based on linguistic name embedding
WO2025128138A1
Automated attention handling in active noise control systems based on universal sound conversion
WO2025128139A1
Name-detection based attention handling in active noise control systems
WO2025128140A1
Interrupt for noise-cancelling audio devices
US20220020387A1
Electronic device including speaker and microphone and method for operating the same
US20220261218A1