Name-detection based attention handling in active noise control systems

Automated attention handling systems in wearable audio components, using linguistic name embedding or universal sound conversion, address the challenge of detecting attention-seeking sounds, allowing users to easily engage in conversations without manually disabling noise cancellation.

WO2025128140A1PCT designated stage expired Publication Date: 2025-06-19GOOGLE LLC

Patent Information

Application Number
PCT/US2024/014820
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-02-07
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing active noise control systems in wearable audio components struggle to automatically detect attention-seeking sounds, making it difficult for users to engage in conversations without manually disabling the noise cancellation feature.

Method used

The implementation of automated attention handling systems that utilize linguistic name embedding, universal sound conversion, or a hybrid approach to detect attention-seeking sounds in real-time, automatically switching the ANC system from ambient sound suppression mode to conversation mode.

Benefits of technology

Enables users to seamlessly transition into conversations by automatically switching the ANC system, improving user experience and comfort while maintaining effective noise cancellation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024014820_19062025_PF_FP_ABST
    Figure US2024014820_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Automated attention handling techniques are described herein for use with wearable audio components with active noise control (ANC) to suppress ambient sound. A name embedding model is trained automatically to convert name audio samples into acoustic segments based on a knowledge distillation model. The name embedding model is used to generate reference embeddings for each of a user-enrolled set of names, and a relation network and a false rejection network are also trained. In real-time operation, the name embedding model converts real-time audio samples to real-time embeddings, the relation network compared the real-time embeddings to the reference embeddings to look for candidate matches, and the false rejection network validates the candidate matches to detect when one of the user-enrolled names has been invoked. Detecting such an invocation automatically triggers the ANC to switch to a conversation mode.
Need to check novelty before this filing date? Find Prior Art

Description

NAME-DETECTION BASED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMSCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of and priority to Indian Provisional Application No. 202341085855, filed on December 15, 2023, and titled “AUTOMATED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS BASED ON LINGUISTIC NAME EMBEDDING,” the content of which is herein incorporated by reference in its entirety for all purposes.BACKGROUND

[0002] Active noise control (ANC) is a common feature of headsets and earbuds. It operates by generating an anti-noise signal via a speaker that is approximately equal in magnitude, but opposite in phase to the ambient sound (e.g., ambient noise and other sounds in the vicinity). The ambient sound and anti-noise signal cancel each other acoustically, allowing the user to hear only a desired audio signal. Typically, signal processing in ANC includes two paths: an ambient sound signal from a reference microphone is taken as the input of a feed-forward ANC filter (FFANC); and an error microphone signal is taken as the input of a feedback ANC filter (FBANC).

[0003] When a user is listening to music or other desired audio through a wearable audio component (e.g., earbuds or on-ear headphones), ANC works to suppress any ambient sound. However, in some instances, ambient sound that is intended for the user can be very important to the user’s connectivity with others. For example, while a user is listening to music with ANC, it may be very difficult for an attention seeker to get the user’s attention, such as to alert the user of something important and / or to engage the user in a desired conversation. Conventionally, in such instances, the attention seeker gets the user’s attention by tapping the user on the shoulder, gesturing in front of the user’s face, or the like; and the user can then manually disable the ANC, pause the desired audio, and / or remove the wearable audio component.SUMMARY

[0004] Systems and methods are described herein for automated attention handling.Embodiments operate in the context of a user wearing a wearable audio component (e.g., in-ear headphones, on-ear headphones, etc.) and having active noise control (ANC) turned on to suppress ambient sound. Various novel techniques are described for detecting that an attention seeker is trying audibly to get the attention of the user and for automatically switching the ANC into a conversation mode in response to such detection.

[0005] In one category of embodiments, the automated attention handling is based on linguistic name embedding (LNE). A name embedding model is trained automatically to classify name audio samples into linguistically distinct name classifications. The name embedding model is used to generate reference embeddings for each of a user-enrolled set of names, and a relation network and a false rejection network are also trained. In real-time operation, the name embedding model converts real-time audio samples to real-time embeddings, the relation network compared the realtime embeddings to the reference embeddings to look for candidate matches, and the false rejection network validates the candidate matches to detect when one of the user-enrolled names has been invoked. Detecting such an invocation automatically triggers the ANC to switch to the conversation mode.

[0006] In another category of embodiments, the automated attention handling is based on universal sound conversion (USC). In such embodiments, spoken audio of a class (i.e., a word) is converted into a unified sound that mimics what would be generated by a speech synthesizer for that class (i.e., the unified sound is stripped of the speaker’s influence on suprasegmental features of the audio). An embedding model essentially converts enrolled invocation names into an enrolled set of unified sound names and stores them as reference embeddings. The same embedding model can then convert real-time audio into a unified sound to produce a real-time embedding. Other models are then used to determine whether the real-time embedding matches any of the reference embeddings, which would indicate detection of one of the enrolled invocation names in the real-time audio and can be used automatically to switch ANC into the conversation mode.

[0007] In another category of embodiments, the automated attention handling is based on a hybrid between LNE and USC, referred to as unified LNE, or ULNE. In such embodiments, USC techniques are used at the front-end to normalize suprasegmental audio features (i.e., to remove speaker influence from the received audio) prior to performing LNE, so that the LNE is performed on a unified version of the name stripped of any particular suprasegmental influence by the attention seeker. This can enable LNE-based classification to generate well-defined and linguistically discriminative name embedding with sustained and effective performance across a wide variety of suprasegmental audio features. As such, performance of ULNE-based embodiments tend to be better than either of the LNE-only or USC-only embodiments.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] A further understanding of the nature and advantages of various embodiments may be realized by reference to the following figures. In the appended figures, similar components orfeatures may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.

[0009] FIG. 1 shows an audio management system for integration in a wearable audio component (WAC), according to embodiments described herein.

[0010] FIG. 2 shows a conceptual circuit block diagram of a partial audio management system environment with an automated attention handling system (AHS).

[0011] FIGS. 3A and 3B show a wearable audio environment including a pair of WACs.

[0012] FIG. 4 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein.

[0013] FIGS. 5A and 5B show a block diagrams of illustrative uses of the name embedding model to generate the deep image.

[0014] FIG. 6 shows several example screenshots from an example enrollment application running on a user device.

[0015] FIG. 7 shows a flow diagram of an illustrative method for audio management that includes automated attention handling in a WAC, according to embodiments described herein.

[0016] FIG. 8 shows a flow diagram of an illustrative method for an enrollment phase.

[0017] FIG. 9 shows a flow diagram of an illustrative method for conversation end detection.

[0018] FIG. 10 shows a flow diagram of an illustrative method for audio management that includes automated attention handling in a WAC using linguistic name embedding (LNE) techniques, according to embodiments described herein.

[0019] FIGS. 11A and 11B show block diagrams of a training environment for training a foundation model to support unified sound conversion (USC) for automated name-detection-based attention handling, according to some embodiments described herein.

[0020] FIG. 12 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein.

[0021] FIG. 13 shows a flow diagram of an illustrative method for automated audio management based on name detection, according to USC-based embodiments described herein.

[0022] FIG. 14 shows a flow diagram of an illustrative method for training a USC inference model.

[0023] FIG. 15 shows a flow diagram of an illustrative method for using the trained USC inference model during an operation time.

[0024] FIG. 16 shows a simplified block diagram of an embedding environment incorporating an illustrative hybrid name embedding model for use in a unified linguistic name embedding attention handling system (ULNE-AHS).

[0025] FIG. 17 provides a schematic illustration of an illustrative computational system that can implement various system components and / or perform various steps of methods provided by various embodiments.DETAILED DESCRIPTION

[0026] When a user is listening to music or other desired audio through a wearable audio component (e.g., earbuds or on-ear headphones), active noise control (ANC) works to suppress any ambient sound. However, in some instances, ambient sound that is intended for the user can be very important to the user’s connectivity with others. For example, although the user desired to suppress undesirable ambient sound, the user may still desire to be able on occasion to enter into desired conversations. In general, two types of desired conversation can be considered: first-party- initiated; and second-party-initiated.

[0027] In first-party -initiated conversations, the user desires to start a conversation and may begin by trying to get someone’s attention. In such cases, some conventional ANC systems are adapted to detect that the user has begun speaking (e.g., by detecting the user’s speech via a beamforming microphone directed to the user’s mouth, accelerometer, or combination thereof), and the ANC system can turn off, switch to transparency mode, pause audio playback, etc. in response to detecting the user’s speaking. Because it tends to be relatively easy for the ANC system to distinguish the user’s own speech from ambient sound, such approaches tend to be effective for first-party-initiated conversations.

[0028] In second-party-initiated conversations, however, a second-party attention seeker is trying to get the user’s attention, and the attention seeker’s voice may be difficult to distinguish from other ambient sound. For example, while a user is listening to music with ANC, it may be very difficult for an attention seeker to get the user’s attention, such as to alert the user ofsomething important and / or to engage the user in a desired conversation. Conventionally, in such instances, the attention seeker gets the user’s attention by tapping the user on the shoulder, gesturing in front of the user’s face, or the like. Before the user can participate in the conversation, the user conventionally must notice the interruption and then manually disable the ANC, pause the desired audio, remove the wearable audio component, etc.

[0029] Indeed, many users of wearable audio components enjoy the feeling of being in their “bubble” and the ability to focus on their media that comes with effective ANC. However, as ANC continues to improve, the same users often feel increasingly unaware, not present, and fearful about missing out. Embodiments described herein seek to provide users with the ability to better stay aware and engage in desired conversations, while being able to continue wearing their wearable audio components and otherwise to take advantage of ANC. This can provide several benefits, including helping to improve user comfort and ear health.

[0030] Embodiments described herein are concerned with second-party-initiated conversations.As used herein, the term “user” refers to a wearer of a wearable audio component (i.e., the first party). The term “attention seeker” is used herein generally to refer to any ambient party trying to get the user’s attention while the user is wearing the wearable audio component (and presumably is listening to desired audio with ANC turned on). Typically, the attention seeker is a person.However, the attention seeker can also be a computational platform with a deterministic manner of seeking the user’s attention, such as a smart speaker programmed to call out the user’s name. The term “wearable audio component,” or “WAC” is used herein to generally refer to earbuds, on-ear headphones, over-ear headphones, or any type of wearable audio output device that includes ANC. The term “desired audio” is used herein to generally refer to any recorded or streaming audio signal that is being played to the user through the WAC, such as music, an audiobook, a podcast, a radio broadcast, a live event broadcast, etc. The term “ambient sound,” or “ambient audio” is used herein to generally refer to any audio in the vicinity of the WAC, other than the desired audio. It is generally the goal of the ANC system to suppress as much of the ambient sound as possible.Audio originating from an attention seeker while a user’s ANC system is active is part of the ambient sound.

[0031] FIG. 1 shows an audio management system 100 for integration in a wearable audio component (WAC), according to embodiments described herein. As illustrated, the audio management system 100 can include an active noise control (ANC) system 140, an attention handling system (AHS) 150, and an audio processing system 160. In general, the purpose of the WAC is to deliver desired audio 165 to a user’s ear or ears via one or more ear speakers, such asspeaker 105. Embodiments of the audio processing system 160 are designed to process the desired audio 165 for output to the user. For example, the audio processing system 160 can include amplifiers, filters, and / or other audio components; and / or any other suitable components for receiving, processing, and / or outputting the desired audio 165.

[0032] Typically, while listening to the desired audio 165, the user is also in presence of ambient audio 155. When in its ambient sound suppression mode, the ANC system 140 seeks to suppress as much of the ambient audio 155 as possible to enhance the user’s experience of listening to the desired audio 165. As illustrated, the ANC system 140 includes a feed-forward ANC (FFANC) filter 120, a feedback ANC (FBANC) filter 125, a summer 130, and an ANC output control block 135. The ANC system 140 is also coupled with the speaker 105 and at least a reference microphone 110 and an error microphone 115. Embodiments of the speaker 105 generally convert an electrical audio signal into sound waves that are delivered to the ear of the wearer of the wearable audio component. Embodiments of the reference microphone 110 can be an omnidirectional microphone typically integrated with an outer casing of the wearable audio component. The reference microphone 110 generally captures at least the ambient audio 155 around the WAC, which is delivered as a reference audio signal (illustrated as x(n)) to the FFANC filter 120. Embodiments of the error microphone 115 are typically integrated with the inner casing of the wearable audio component to be positioned inside the ear canal or very close to it when the wearable audio component is being worn. The error microphone 115 captures the audio that reaches the eardrum, which includes the desired audio signal and any remaining ambient sound after suppression. The error microphone 115 outputs an error signal (illustrated as e(ri)) to the FBANC filter 125.

[0033] The illustrated ANC system 100 includes a feed-forward noise control path and a feedback noise control path. The feed-forward noise control path includes the FFANC filter 120, which is a digital or analog filter designed to process the audio signal from the reference microphone 110. The FFANC filter 120 applies a specific frequency response to x(n) to adaptively cancel out noise. The specific frequency response is produced by continuously adjusting coefficients of the FFANC filter 120 to minimize the difference between the desired audio signal and the reference signal. The output of the FFANC filter 120 is illustrated as(n). The feedback noise control path includes the FBANC filter 125, which is a digital or analog filter designed to process the audio signal from the error microphone 115. The FBANC filter 125 applies a specific frequency response to e( ), and continuously adjusts coefficients of the FBANC filter 125 to minimize the difference between the desired audio signal and remaining ambient sound in the signal that reaches the eardrum. The output of the FFANC filter 120 is illustrated asy2(n). In general, both the FFANC filter 120 and the FBANC filter 125 can adapt their respective filters (e.g., their coefficients) in real-time to a changing audio environment. For example, filter coefficients are iteratively adjusted using least mean squares (LMS), normalized LMS (NLMS), and / or other suitable adaptation algorithms.

[0034] Embodiments of the summer 130 combine the filtered output signals from the FFANC filter 120 and the FBANC filter 125. For example, the summer 130 calculates a sum of these signals. If tuned properly, the output of the summer 130 is an “anti-noise” signal that closely represents the ambient sound at opposite polarity. Embodiments of the ANC output control block 135 control how and / or whether the anti-noise signal is output by ANC system 140. In some implementations, the ANC output control block 135 includes an amplifier to provide a controllable amount of gain (G) to the signal at the output of the summer 130, resulting in an output signal, y(n) = G(y (n) + y2(n)). Ineffect, the ANC gain block 135 adjusts the overall amplitude (i.e., corresponding to volume) of the combined filtered signal at the output of the summer 130. The output signal is sent to the speaker 105. In some implementations, as illustrated, the desired audio 165 can also be mixed in (e.g., by mixer 145) prior to sending the output to the speaker 105, such that what reaches the eardrum is almost entirely the desired audio signal with minimal ambient sound. Alternatively, the desired audio 165 is mixed into the output signal at the summer 130, such that the output of the ANC system 140 is an audio signal that is mostly the desired audio 165 with minimal residual ambient audio 155.

[0035] Embodiments of the ANC output control block 135 control the operating mode of the ANC system 140. For example, as described herein, the ANC system 140 can operate selectively in at least an active mode (i.e., an ambient sound suppression mode) or a conversation mode. Some implementations of the conversation mode correspond to an inactive mode (i.e., the ANC system 140 is turned off) or a transparency mode. Other implementations of the conversation mode are configured to pass through conversationally relevant audio from the ambient audio 155, while continuing to perform ANC functions to suppress other portions of the ambient audio 155. In some such implementations, a bandpass or notch filter is used to segregate out a range of frequencies typical for human speech and to treat the segregated audio as conversationally relevant audio. As one example, a filter can pass through portions of the ambient audio 155 only in the range of 75 to 300 Hertz and to suppress higher and lower frequency components of the ambient audio 155; thereby continuing to filter out white noise and other portions of ambient audio 155 that can interfere with a user’s ability to hear the passed-through conversationally relevant audio. Similarly, some implementations continue to pass through some desired audio 165 (e.g., at a reduced volume) while in conversation mode.

[0036] As described herein, embodiments of the AHS system 150 seek to detect when a second- party attention seeker is trying to get the attention of a user while the user is wearing the WAC and is listening to desired audio 165 with the ANC system 140 in the active mode. In particular, the AHS system 150 listens for presence of attention seeking (AS) audio 157 within the ambient audio 155. Some embodiments of the AHS system 150 described herein are implemented as a linguistic name-embedded attention handling system (LNE-AHS). Other embodiments of the AHS system 150 described herein are implemented as a universal sound conversion attention handling system (USC-AHS). Other embodiments of the AHS system 150 described herein are implemented as a hybrid universal LNE attention handling system (ULNE-AHS). Any of the types of AHS system 150 described herein can be configured specifically to listen for AS audio 157 corresponding to a previously enrolled invocation name (e.g., a name of the user). When the AS audio 157 is detected by the AHS system 150, the AHS system 150 automatically directs the ANC system 140 to switch from the active mode to the conversation mode. In some embodiments, when the AS audio 157 is detected, the AHS system 150 also directs the audio processing system 160 to enter a conversation enhancement mode. As described above with reference to the ANC system 140, the conversation enhancement mode as implemented by the audio processing system 160 can include segregating conversationally relevant audio from the ambient audio 155, adapting equalization of passed through audio to enhance speech, muting or reducing the volume of playback of the desired audio 165, pausing playback of the desired audio 165, etc. Typically, in response to the AS audio 157, the user will begin to engage in a conversation with the attention seeker. Such a conversation can involve the user speaking, and embodiments of the conversation mode of the ANC system 140 and / or the conversation enhancement mode of the audio processing system 160 can include using techniques to help ensure that the user’s own speech is not fed back in a manner that results in an apparent echo, feedback noise, or the like. For example, the user’s own speech may be captured by a separate beamforming microphone as a user speech audio stream, while ambient audio 155 is being received by the reference microphone 110. The user speech audio stream can be subtracted from the ambient audio 155 prior to passing the signal through other blocks of the system, so that the fed-back audio stream includes only ambient audio other than the user’s own speech.

[0037] Some embodiments of the AHS system 150, after having detected AS audio 157 and directing the ANC system 140 into conversation mode, can further detect when the conversation ends. Such embodiments of the AHS system 150 can automatically direct the ANC system 140 to return to the active mode, accordingly. As part of returning to the active mode, some such embodiments also return settings (e.g., in the ANC system 140 and / or the audio processing system160) to those appropriate for listening to the desired audio 165 and suppressing all of the ambient audio 155 (e.g., all frequencies of the ambient audio 155).

[0038] FIG. 2 shows a conceptual circuit block diagram of a partial audio management system environment 200 with an automated attention handling system (AHS) 150. The illustrated environment 200 can be an illustrative portion of the audio management system 100 of FIG. 1, and the AHS 150 can be an illustrative implementation of the AHS 150 of FIG. 1. As illustrated, the AHS 150 includes an attention seeking (AS) trigger detection block 210 and a conversation end detection block 220. Some implementations further include a conversation enhancement block 230. The AHS 150 is illustrated in context of a desired audio 165 stream, a reference microphone 110 that receives ambient audio 155 and outputs an ambient audio stream, and a speaker 105. For the sake of simplicity, the AHS 150 is illustrated without other components of the audio management system 100 of FIG. 1, such as without the ANC system 140 and the audio processing system 160.

[0039] The role of the AHS 150 can be generally described as to toggle the audio environment between an active mode and a conversation mode based on whether a desired conversation is detected, as represented by a switch network 215. In the active mode, the user is listening to the desired audio 165 via the speaker 105, and the ANC system 140 (not shown) is suppressing as much of the ambient audio 155 as possible. This is conceptually represented by the switches of the switch network 215 being in the solid-line position, whereby the desired audio 165 passes through to the speaker 105 and the ambient audio 155 does not. When attention seeking audio (i.e., audio associated with getting the user’s attention) is detected by the AS trigger detection block 210, the AHS system 150 switches the switch network 215 to the dashed-line position, whereby the ambient audio 155 passes through to the speaker 105 and the desired audio 165 does not. In some embodiments, while in the conversation mode, the passed-through ambient audio 155 (e.g., either all of the ambient audio 155, or a conversationally relevant portion of the ambient audio 155) is passed through the conversation enhancement block 230 in line with the speaker 105. As described above, the conversation enhancement block 230 can use various techniques to enhance conversationally relevant portions of the ambient audio 155. Further, as described above, the conversation enhancement block 230 can be implemented in the AHS system 150, in the ANC system 140, in the audio processing system 160, and / or in any suitable location.

[0040] When the end of the conversation is detected by the conversation end detection block 220, the AHS system 150 switches the switch network 215 back to the solid-line position, whereby the desired audio 165 again passes through to the speaker 105 and the ambient audio 155 againdoes not. In some embodiments, as described above, the end of the conversation is detected based on detecting the user’s own speech, such as detecting that the user is no longer speaking for some time, or that the user has issued an audio cue (e.g., “resume ANC”). This user speech can be detected via the reference microphone 110 as part of the ambient audio 155, or detected through a separate microphone 240, such as a beamforming microphone with its beam directed toward the user’s mouth. Additionally, or alternatively, some embodiments of the conversation end detection block 220 detect the end of a conversation based on detecting user interfacing with an interface element, such as detecting that the user pressed a play / pause button 245 on the WAC 210, or the like.

[0041] Although the AHS 150 is illustrated as directly coupled with the microphones, the desired audio 165 is directly coupled with the speaker 105 in active mode, the ambient audio 155 is directly coupled with the speaker 105 in conversation mode, etc., some or all of such connections can be through other components that are not shown in FIG. 2. For example, embodiments herein assume that the ambient audio 155 is passing through the ANC system 140, that the desired audio 165 is mixed with an anti -noise signal in active mode, etc.

[0042] As described herein, embodiments of the audio management system 100 are configured for integration in any suitable WAC. FIGS. 3A and 3B show a wearable audio environment 300 including a pair of WACs 310. Each WAC 310 is illustrated as an earbud. Alternatively, the pair of earbuds can be considered as a single WAC. In other embodiments, the WAC 310 can be implemented as over-ear headphones, or any other suitable wearable audio component that incorporates ANC. In the illustrated embodiments, each WAC (i.e., each earbud) has a respective instance of an audio management system 100, such as the audio management system 100 of FIG.1, and each instance of the audio management system 100 includes a respective instance of at least an ANC system 140 and an AHS system 150. Though not explicitly shown, each WAC 310 also has, integrated therein, an instance of the speaker 105, the reference microphone 110, the error microphone 115, one or more processors, and non-transitory processor-readable storage. Some implementations of the WAC 310 include additional components, such as instances of the audio processing system 160, one or more additional microphones (e.g., a beamforming microphone), one or more additional speakers, interface controls (e.g., one or more buttons), one or more power sources (e.g., a rechargeable battery), one or more ports (e.g., physical ports for charging and / or wired communication, logical ports for wireless charging and / or wireless communication), one or more antennas, etc.

[0043] In some embodiments, the one or more processors integrated in the WAC 310 implement components of the respective audio management system 100 instance. For example, a non- transitory processor-readable medium integrated therein has processor-executable instructions stored thereon, which, when executed, cause the set of processors to implement at least features of the respective ANC system 140 and / or AHS system 150 instances. As described herein, embodiments of the AHS system 150 include one or more types of artificial neural networks, corresponding trained network models, or the like. In some embodiment, such networks and / or models are implemented using specialized hardware, such as neuromorphic chips. In other embodiments, such networks and / or models are implemented by using processor-readable instructions to reconfigure general-purpose computing hardware (e.g., a central processing unit, CPU), specialized Al accelerators.

[0044] Turning specifically to FIG. 3 A, a first type of wearable audio environment 300a is shown in which one or both WACs 310 is in communication with a cloud computing environment (“cloud”) 340. For example, the cloud 340 includes a server, or several distributed servers, accessible via the Internet. Though the WAC 310 is shown as directly in communication with the cloud 340, such a connection can be facilitated by any suitable intermediary devices, such as routers, hubs, etc. As described further herein, automated attention handling features described herein rely on generation of an inference model that includes several neural networks and / or models. In the illustrated embodiments of FIG. 3 A, the inference model is generated by the local computation environment of the WAC 310 and / or based on information ported to the WAC 310 from the cloud 340.

[0045] As described herein, some embodiments involve enrollment of invocation names for use in name detection. In the embodiments of FIG. 3 A, an enrollment application 330 is downloaded to the local computational environment of the WAC 310, and the enrollment application 330 is used for such enrollment. The enrollment application 330 can facilitate generation of the inference model, based on a foundation model 350. As illustrated, the foundation model 350 can be stored in the cloud 340. In some embodiments, the foundation model 350 is accessible to the WAC 310 (e.g., to the enrollment application 330) via the cloud 340. The inference model can include some or all of a name embedding model, a deep image, a relation network, and a false rejection network. Each is described more fully below.

[0046] Turning to FIG. 3B, some embodiments of the wearable audio environment 300b further include a user computational device 320 separate from the WAC 310. For example, the user computational device 320 can be a smartphone, laptop computer, tablet computer, smart watch,portable audio player, or any other suitable device that is separate from the WAC 310 and includes its own one or more processors and its own one or more non-transitory storage media for storing processor-readable instructions. The user computational device 320 can be in communication with each WAC 310 via any suitable wired and / or wireless communication link, such as via an audio cable (e.g., via a 3.5 -millimeter or 1 / 4-inch analog audio jack), a universal wired connection (e.g., universal serial bus (USB)), a short-range universal wireless connection (e.g., Bluetooth, short- range radiofrequency, near field communication (NFC)), an optical connection (e.g., infrared), a proprietary connector, a multi-pin connector, an intermediary component or platform (e.g., a docking station or dongle), etc.

[0047] Similar to the embodiments of FIG. 3 A, the embodiments of FIG. 3B can include an enrollment application 330, communications with the cloud 340, use of a foundation model 350, etc. As illustrated in FIG. 3B, these features can be facilitated via the user computational device 320 (rather than directly by the WAC 310). For example (as described more fully below), when a user first registers the WAC 310 (e.g., first attempts to pair the earbuds with the user computational device 320), the user computational device 320 automatically accesses and / or downloads the enrollment application 330, or prompts the user to access and / or download the enrollment application 330. The enrollment application 330 can be downloaded from the cloud 340, or from any other suitable environment. The enrollment application 330 can then access the foundation model 350 via the cloud 340 and can use the foundation model 350 to generate an inference model for automated attention handling. For example, to facilitate LNE-AHS, the inference model includes some or all of a name embedding model, a deep image, a relation network, and a false rejection network. The generated inference model can then be ported to the WAC 310. For example, the inference model can be ported from the user computational device 320 to both earbuds; ported from the user computational device 320 to a master earbud, and from the master earbud to a slave earbud; etc.

[0048] Embodiments generally build an inference model for storing at the WAC 310 to enable the WAC 310 to subsequently use the inference model to perform automated attention handling, as described herein. FIG. 4 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block 400 and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein. The AS trigger detection block 400 can be an implementation of the AS trigger detection block 210 of FIG. 2. It is assumed that the AS trigger detection block 400 is implemented in the context of a WAC 310.

[0049] Embodiments can include an enrollment stage 410, an identification stage 430, and a verification stage 440. In general, the enrollment stage 410 occurs outside of normal operation of the WAC 310, such as when a user first sets up and / or uses the WAC 310, when the user first registers the WAC 310, when the user first configures the WAC 310 for automated attention handling, etc. The identification stage 430 and the verification stage 440 are real-time blocks that occur during normal operation of the WAC 310 to facilitate real-time automated attention handling. Embodiments generally use the enrollment stage 410 to obtain any information needed to set up an inference model for storing at the WAC 310 to enable the WAC 310 to perform automated attention handling features during normal operation. As shown, the inference model includes at least an embedding model 415, a relation network 435, and a false rejection network 445.

[0050] In the enrollment stage 410, the embedding model 415 can be used to convert an enrollment audio stream 405 into a set of reference embeddings based on a foundation model 350. The foundation model 350 can be stored in and accessed via the cloud 340 (e.g., and / or any other suitable communication network). The reference embeddings can be stored as a deep image 420. The deep image 420 can also be considered as part of the inference model.

[0051] During normal operation, in the identification stage 430, a real-time (RT) audio stream 407 is received. The embedding model 415 (i.e., the same embedding model 415 generated in the enrollment stage 410) is used to generate a RT embedding from the received RT audio stream 407. The relation network 435 can then compare the RT embedding with each of the reference embeddings in the deep image 420 to determine if there is a match. For example, the relation network 435 is configured to compute a similarity score for each comparison (e.g., corresponding to a mathematical correlation, or the like). The relation network 435 can determine if any similarity scores meet or exceed a predetermined threshold; if so, the reference embedding with the highest similarity score can be selected as a candidate matching embedding.

[0052] In the verification stage, the false rejection network 445 can then confirm the match by transforming the RT embedding and the candidate matching embedding into different mathematical spaces (e.g., domains) and determining whether the embeddings can be reliably discriminated from each other in any of the mathematical spaces. For example, the deep image 420 is configured to compute a discrimination score for each mathematical space. The deep image 420 can determine if any discrimination scores meet or exceed a predetermined threshold; if so, the match determined by the relation network 435 is determined to be a false match and is ignored. Ifthe embeddings cannot be sufficiently discriminated in any mathematical spaces, the false rejection network 445 can output a signal that attention seeking audio has been detected.

[0053] As described herein, embodiments of the attention seeker (AS) trigger detection block 400 are configured to detect an invocation name as part of an attention handling system. The embedding model 415 is a name embedding model (e.g., a processor-executable name embedding model) that generates an output embedding from an audio sample. As used herein, the terms “audio sample” or an “audio signal” are used interchangeably in the context of an input to a component of the inference model; such an “audio sample” or an “audio signal” can be represented in any suitable manner, such as by any suitable number of digital samples. For example, reference to an input as an “audio sample” means an audio signal of a duration, or a sampled duration of an audio signal, at a sampling rate resulting in a large number of digital samples (e.g., one second of audio sampled at 16 kHz to yield 16,000 samples). As described above, the audio sample can be from an enrollment audio stream 405 in an enrollment stage 410, and the audio sample can be from a RT audio stream 407 during normal operation. The name embedding model 415 is trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications. The deep image 420 (e.g., a processor-readable deep image) includes reference name embeddings generated by the name embedding model 415 based on a set of invocation names provided by a user during an enrollment procedure. The relation network 435 (e.g., a processor-executable relation network) is coupled with the deep image 420 and the name embedding model 415 to output a candidate name embedding responsive to determining that one of the reference name embeddings has a highest similarity with a RT embedding generated by the name embedding model 415. The RT embedding can be from a RT audio stream 407 received from a reference microphone associated with an ANC system of the WAC 310. The false rejection network 445 (e.g., a processor-executable false rejection network) is coupled with the relation network 435 to output a name invoked signal 450 responsive to determining that the real-time embedding and the candidate name embedding cannot be reliably discriminated. As described herein, the name invoked signal 450 can direct an ANC system of the WAC 310 automatically to enter a conversation mode.

[0054] As described above, the interference model used for automated attention handling is based on a foundation model 350. Embodiments of the foundation model 350 are generated from a large speech-audio corpus of diversified words spoken by diversified speakers. “Diversified words” refers herein to the speech-audio corpus including a wide variety of at least phonemes and linguistic information. “Diversified speakers” refers to the speech-audio corpus representing a wide variety of at least accents and prosody. The speakers can be further diversified with respectto age, gender, geography, etc. For example, the speech-audio corpus includes tens of thousands of words (i.e., classifications) spoken multiple times (e.g., 10 - 15 times) by hundreds of speakers from around the world.

[0055] The term “suprasegmental” is used herein as an umbrella term to encompass properties of a speaker’s influence when speaking words, such as accent, prosody, intonation, rhythm, and other non-segmental aspects of speech. Such suprasegmental features can be contrasted with segmental features pertaining to individual speech sounds or segments, such as vowels and consonants, and can span multiple segments or an entire utterance. Examples of suprasegmental features of an utterance (e.g., a word, name, etc.) can include accent (including accent-influenced variations in pitch, loudness, and duration), prosody (including rhythm, intonation, and melody of speech), intonation (i.e., the rise and fall of pitch in speech), rhythm and / or rate (e.g., the temporal patterns of speech, such as duration and timing of sounds, syllables, and pauses), and stress (e.g., emphasis placed on a particular syllable). For example, a large speech-audio corpus of diversified words spoken by diversified speakers may include hundreds or thousands of samples of a particular word being spoken with wide suprasegmental variance over the samples.

[0056] Training of the foundation model 350 can begin with an auto-encoder architecture, which is a type of neural network architecture designed to learn compact representations of data, such as so-called “latent features.” In the context of embodiments described herein, the auto-encoder architecture is used to extract meaningful features from raw audio data to be used for automatic speech recognition (ASR). In general, the auto-encoder architecture includes three high-level stages: an encoder, a bottleneck layer, and a decoder. The encoder receives a high-dimensionality input (i.e., the raw audio data) and includes a multi-layer network to progressively reduce the dimensionality of the data using transformations. For example, the ASR information is represented as a sequence of audio frames, each including features that can be mathematically described as Mel-frequency cepstral coefficients (MFCCs), or in some other manner. Each layer of transformations (e.g., linear operations followed by non-linear operations, such as a rectified linear unit (ReLU) function) seeks to extract increasingly abstract and higher-level features from the input data.

[0057] The bottleneck layer is so called because it typically includes significantly lower dimensionality than the layers of the encoder before it and or the decoder after it. This reduced dimensionality effectively forces the network to learn a compressed and informative representation of the input data. The decoder operates in essentially a reverse manner to that of the encoder. It takes the output of the bottleneck layer as its input data (i.e., a lowest-dimensionalityrepresentation) and applies multiple layers of transformations to generate increasingly higherdimensional representations of the data. In effect, the decoder progressively increases the dimensionality with a goal of reconstructing the original input data from the highly compressed representation in the bottleneck layer.

[0058] Although the decoder output ideally matches the encoder input (i.e., the raw audio data), there will practically be some difference, referred to as reconstruction loss. Training of the autoencoder seeks to optimize the auto-encoder network (e.g., the bottleneck layer) to minimize reconstruction loss (e.g., to minimize mean squared error loss, or the like). Effective training results in the bottleneck layer producing a highly compact, but highly meaningful representation of the input data; effectively extracting the most salient features for ASR. The compact representation in the bottleneck layer can be referred to as the ’’latent space,” and the autoencoder’s trained (learned) knowledge of how to encode raw input data into the latent space is embedded in a set of weights. The set of weights can be represented as a feature vector, a set of feature vectors, or in any suitable manner. For example, one second of audio can be represented by 16,000 samples (i.e., at a 16 kHz sampling rate), and the 16,000 samples can effectively be represented by a set of 256 weights (e.g., or 128 weights, or another suitable number). The set of weights can represent the embedding from the auto-encoder, the filter bank energies, the MFCCs, etc.

[0059] In the context of linguistic name embedding embodiments described herein (i.e., LNE- AHS), the foundation model 350 is a classification model that takes features of an audio signal as its input and classifies the audio signal in accordance with the salient linguistic content of the audio signal. The foundation model 350 seeks to operate so that any audio signal representing a same linguistic word (e.g., whether from the same or different users, whether having the same suprasegmental features, etc.) will be grouped into the same classification, such that each classification represents a linguistically unique word. As used in this context, a “word” can include any type of word, name, or utterance that could reasonably be used to get someone’s attention, such as “John,” “mister,” “hey,” “excuse me,” etc. Embodiments of the foundation model 350 can be trained (e.g., using transfer learning) based on the trained bottleneck layer (e.g., and encoder) from the auto-encoder.

[0060] For example, in the case of linguistic name embedding embodiments described herein, the auto-encoder is trained to figure out how to encode (compress) an input audio signal into a highly compact representation of the input data that represents the most linguistically salient features of the input audio signal. The linguistically salient features are stripped of speakerinfluence, so that any particular word will always be encoded to the same set of weights, regardless of the speaker’s suprasegmental influence (e.g., influence on accent, prosody, speaking speed, pitch, etc.). Thus, at the bottleneck layer, an input audio signal is represented by a set of weights (e.g., auto-encoder embeddings, filter bank energies, MFCCs, etc.) that capture the salient linguistic features of the word represented by the audio signal. The foundation model 350 is based on an encoder-decoder architecture. The dimensions of the input and output of the foundation model 350 may not be the same, and they may not be symmetric. The input labels of the foundation model 350 are the set of weights trained to represent the most salient linguistic features of an input audio signal provided to the model, and the output of the foundation model 350 is a classification. In some embodiments, the output of the model may be large number of layers, each corresponding to a respective one of the classifications; and there may be substantial energy only at whichever of the layers corresponds to the classification appropriate for the classifying the input audio signal. Thus, the foundation model 350 can generate an output label that represents a linguistic classification identifier, such as using one-hot encoding, a numeric representation, or any other suitable representation of the classification identifier.

[0061] Thus, the foundation model 350 is trained to remove speaker influence from words in order to generate classifications (the terms classifications and classes are used interchangeably herein), so that all audio samples representing the same linguistic information (i.e., the same word, name, etc.) are in a same class. For example, hundreds of speakers can provide audio samples of them saying the word “EXAMPLE” with different accents, prosody, etc., and the foundation model 350 is trained to pull out all the salient features so that all those audio samples will be classified into a single classification associated with the word “EXAMPLE.” The classification can, in some implementations, be further associated with corresponding text, such as text of the word “EXAMPLE.”

[0062] The name embedding model 415 is generated by applying transfer learning from the large speech-audio corpus of phonetically diversified words used in training the foundation model 350 (the source dataset) to a smaller corpus of real-world name data (the target dataset). For example, the foundation model 350 can use a large number (e.g., 11,000) of classifications to generate the set of weights, where each linguistically distinct word in the speech-audio corpus is classified into one of the classes. Transfer learning can then apply the trained foundation model 350 to generate the name embedding model 415 (deep feature generation model, or DFGNet) for a smaller number (e.g., 500 - 1000) of classifications, each associated with a linguistically distinct name. Implementations of the name embedding model 415 include only the encoder and bottleneck layer as trained through the transfer learning.

[0063] In effect, the name embedding model 415 can be characterized as a tuned and reduced version of the foundation model 350 developed specifically for name detection, such as by using few-shot learning (FSL). Real audio samples used to train the name embedding model 415 can correspond to people's names from various regions and languages, and sample names can be chosen to cover most phonetic usage in each region. Implementations of the name embedding model 415 can be trained to recognize any suitable number of name classifications. For example, it may be impractical impossible to train the model for all possible names and their variants everywhere in the world, and a practical number of more common names (e.g., 1,000) can be chosen instead. In some implementations, different versions of the name embedding model can be generated 415 and / or trained differently for different user groupings (e.g., geographical regions, ethnicities, etc.) to capture the most popular names for the corresponding groupings. For example, grouping information can be entered by the user as part of enrollment, obtained for the user from account information, assumed for the user based on location or other demographic information, etc. Further, the name embedding model 415 can be designed with as much complex as needed to discriminate between the name classifications used for training with enough reliability to be suitable for name-detection-based attention handling. For example, a higher-dimensional model (e.g., where the number of weighting vector dimensions, N, is larger) may be able to more reliably discriminate among a larger number of name classifications, but use of such a model in real time will involve more computational resources (e.g., which may correspond to more processing time, more battery usage, more heat generation, etc.).

[0064] The name embedding model 415 can then generate a reference embedding for each invocation name. For example, FIGS. 5A and 5B show a block diagrams 500 of illustrative uses of the name embedding model 415 to generate the deep image 420. Turning first to FIG. 5 A, a user 505 provides a set of J invocation names 510 via an enrollment application 330 (J is a positive integer). For example, the user 505 speaks each name one or more times, types each name using its proper spelling, types each name phonetically, etc. The J invocation names 510 are passed to the name embedding model 415, which generates J corresponding reference embeddings.

[0065] In some embodiments, each reference embedding is an N-dimensional vector corresponding to a set of N weights in the name embedding model 415 that represents the invocation name that yielded that reference embedding. N can be any suitable integer number of weights to provide sufficiently reliable classification. In one implementation, N is 128. In another implementation, N is 256. For example, the name embedding model 415 generates each reference embedding as the N-dimensional vector, and stores the vectors in the deep image 420. Forexample, the deep image 420 stores J N-dimensional vectors, a single J-by-N-dimensional matrix, or the like.

[0066] Turning first to FIG. 5B, a user 505 again provides a set of J invocation names 510 via an enrollment application 330 (J is a positive integer). Unlike in FIG. 5 A, the J invocation names 510 are passed to a name augmenter 520, which augments the user-provided set of invocation names 510 to generate an augmented set of invocation names 510’. The name augmenter 520 can include, or be in communication with, an augmentation model 515. Embodiments of the augmentation model 515 include mathematical transformations to apply to each of some or all of the invocation names 510. The name augmenter 520 can generate K augmentations (K is a positive integer) for each of the J invocation names 510, so that the augmented set of invocation names 510’ includes J * K names. For example, a user 505 enrolls four invocation names, nine augmentations are applied to each invocation name to generate ten total names for each invocation name, or forty total entries in the augmented set of invocation names 510’.

[0067] In some implementations, the name augmenter 520 adds time-based augmentations to each of some or all of the invocation names 510, such as by time-stretching and / or timecompressing a user-provided audio sample of the invocation name. In some implementations, the name augmenter 520 adds accent-based augmentations to each of some or all of the invocation names 510, such as by mathematically applying different vowel changes, regional variations, pronunciations, etc. to the invocation name. In some implementations, the name augmenter 520 adds suprasegmental augmentations to each of some or all of the invocation names 510, such as by mathematically applying different syllable accenting, intonation, volume, pitch, etc. Other augmentations can account for differences across genders, ages, etc. Other augmentations can account for noise models, such as models of ambient background noise, television or music noise, traffic noise, road noise, engine noise, air conditioning noise, running water noise, etc. The J * K invocation names 510’ are passed to the name embedding model 415, which generates J * K corresponding reference embeddings. For example, the name embedding model 415 generates each reference embedding as an N-dimensional vector for storage in the deep image 420 (e.g., as J * K N-dimensional vectors, as a (J * K)-by-N-dimensional matrix, or the like). In some implementations, the name augmenter 520 applies different augmentations to different invocation names, and / or different numbers of augmentations to different invocation names. As one example, different augmentations can be applied based on whether the invocation name is characterized more by its vowel content, or more by its consonant content. As another example, a more common term enrolled as an invocation name (e.g., “boss,” “mom”), or a shorter name enrolled asan invocation name (e.g., “Max,” “Tim”) may be augmented differently than less common terms, longer names, etc.

[0068] Returning to FIG. 4, the relation network 435 is trained with the linear and non-linear features that characterize the name embedding model 415. Embodiments of the relation network 435 are trained to output a similarity score (e.g., a probability of a match) responsive to two inputs: one of the reference embeddings from the deep image 420, and a real-time embedding generated from a real-time audio sample received via the reference microphone. For example, the one of the reference embeddings from the deep image 420 is a N-dimensional vector previously generated by the name embedding model 415 during enrollment, and the real-time embedding is an N- dimensional vector generated by the name embedding model 415 in real-time. As described above, the reference embedding vector essentially represents salient linear and non-linear features of the associated classification, and the real-time embedding vector essentially represents the same salient linear and non-linear features of the real-time audio sample. The relation network can map those same linear and non-linear features between real-time embeddings and reference embeddings to find candidate matches. For example, in a scenario where 40 classifications are generated (i.e., the deep image 420 is a 40-by-N matrix), the relation network 435 can compute a correspondence between the real-time embedding and each of the 40 reference embeddings. This can be performed as 40 serial computations (e.g., iterative), 40 parallel computations, or in any suitable manner. Some embodiments of the relation network 435 are implemented as a two-dimensional convolutional neural network (CNN). Some other embodiments of the relation network 435 are implemented as a one-dimensional CNN, a time-delay neural network (TDNN), or another suitable neural network. Some other embodiments of the relation network 435 are implemented using simple cosine similarity or equilidian distance estimation. For example, thresholding is performed based on the measured metric, and either a high value of the cosine similarity score represents a high relationship (for simple cosine similarity), or a minimum score represents a high relationship (for equlidian distance).

[0069] Embodiments compute a similarity score (e.g., a mathematical correlation) between a present real-time embedding (RTE) and each of the reference embeddings and determine whether the similarity score exceeds a predetermined matching threshold (e.g., 0.3) for any one or more of the reference embeddings. If none of the reference embeddings yields a similarity score exceeding the predetermined matching threshold, embodiments determine that there is no name match and ignore the analyzed portion of the real-time audio signal (i.e., discards the RTE). If one of the reference embeddings yields a similarity score exceeding the predetermined matching threshold, the class associated with that reference embodiment is selected as a candidate matching name (i.e.,that reference embedding is selected as the candidate matching reference embedding, or CMRE). If multiple reference embeddings yield similarity scores exceeding the predetermined matching threshold, the reference embedding associated with the highest similarity score is selected as the CMRE.

[0070] Embodiments of the relation network 435 are trained to output a similarity score (e.g., a probability of a match) responsive to two inputs: one of the reference embeddings from the deep image, and a real-time embedding generated from a real-time audio sample received via the reference microphone. During training of the relation network 435, a training audio sample can be used as the real-time audio sample, which is fed into the name embedding model 415 to generate a training embedding (corresponding to the real-time embedding generated during normal operation). For example, the reference embedding for a particular invoked name classification is input to the relation network 435, and a training embedding is generated by the name embedding model 415 for a training audio sample: if the training audio sample is known to correspond to a particular invocation name, the relation network 435 is trained to output ‘1’, ‘ 100 percent’, etc. when fed the corresponding reference and training embeddings; if the training audio sample is known not to correspond to a particular invocation name, the relation network 435 is trained to output ‘O’, ‘0 percent’, etc. when fed the corresponding reference and training embeddings. In some embodiments, the relation network 435 is trained based on the same corpus of audio samples (or a portion thereof) used to train the foundation model 350. In other embodiments, the relation network 435 is trained based on the same corpus of audio samples (or a portion thereof) used to train the name embedding model 415.

[0071] The false rejection network (FRNet) 445 seeks to determine whether the CMRE and the RTE can be discriminated. In effect, the relation network 435 seeks to find a candidate match, and the false rejection network 445 seeks to determine whether the candidate match is a false match. Embodiments of the false rejection network 445 apply multiple mathematical transformations (e.g., rotations), each transformation designed to transform both the CMRE and the RTE into a corresponding domain and / or space to see whether the two datasets continue to match. For example, suppose a user has enrolled the invocation name, “Jonathan,” and the real-time audio signal includes the phrase “on a thin.” In such a scenario, the relation network 435 may find a candidate match (i.e., a similarity score exceeding the threshold), but the false rejection network 445 may determine that the candidate match is likely not a match and can be rejected. Some embodiments of the false rejection network 445 are implemented as a progressive layered extraction (PLE) neural network. Some other embodiments of the false rejection network 445 are implemented as a probabilistic linear discriminant analysis (PLDA) network.

[0072] Embodiments of the false rejection network 445 are trained to output a discrimination score (e.g., a likelihood ratio representing probability of a false match) responsive to two inputs: one of the reference embeddings from the deep image 420, and a real-time embedding generated from a real-time audio sample received via the reference microphone. The training of the false rejection network 445 can be similar to the training of the relation network 435. For example, during training of the false rejection network 445, a training audio sample can be used as the realtime audio sample, which is fed into the name embedding model 415 to generate a training embedding (corresponding to the real-time embedding generated during normal operation).Unlike the relation network 435, the false rejection network 445 is trained to apply transformations to the two inputs to look for a particular domain or space in which the two can be discriminated. For example, the training can use some training audio samples that are similar to a particular invoked name classification and other audio samples that are completely different (e.g., effectively linguistically orthogonal) to the invoked name classification. The false rejection network 445 is trained to find transformations that reliably discriminate involved name classifications from audio samples that sound like those invoked names but actually carry a different linguistic meaning. In some embodiments, the false rejection network 445 is trained based on the same corpus of audio samples (or a portion thereof) used to train the foundation model 350. In other embodiments, the false rejection network 445 is trained based on the same corpus of audio samples (or a portion thereof) used to train the name embedding model 415.

[0073] The name embedding model 415, the relation network 435, and the false rejection network 445 can all be trained together (e.g., in parallel, or serially). As noted above, the name embedding model 415 is trained by transfer learning from the foundation model 350 using a corpus of real-world name data. The input is an audio sample, and the output is a classification (e.g., M-dimensional vector, where M is the number of classifications, such as J or J * K). The specific invocation names (e.g., including augmentations) are used to generate reference embeddings for each of a set of invocation name classifications, which are stored as the deep image 420. Embodiments of the relation network 435 are trained to output a respective similarity score between a real-time embedding generated from a real-time audio sample received via the reference microphone and each of the reference embeddings from the deep image 420. Embodiments of the false rejection network 445 are trained to output a respective discrimination score between a real-time embedding generated from a real-time audio sample received via the reference microphone and each of the reference embeddings from the deep image 420. Because the relation network 435 only computes similarity scores and matches them to a threshold, the relation network 435 can be very lightweight (e.g., resource-efficient). For example, even in thecontext of a small processor and a small battery, such as in an earbud), the relation network 435 can run continuously without using excessively processor computation cycles, without draining excessive power, without generating excessive heat, etc. Embodiments of the false rejection network 445, which may use appreciably more resources to perform transformations, etc., only run when a candidate match has been identified. Alternative embodiments can combine the functionality of the relation network 435 and the false rejection network 445, such as in contexts where resources are not as limited (e.g., implemented in over-ear headphones that include wired power).

[0074] As described above, some or all of the name embedding model 415, the relation network 435, the false rejection network 445, and the deep image 420 can be treated as a single inference model (or a “name detection model”). For example, the foundation model 350 is a common model that is computed in and / or stored in the cloud 340. When invocation names are first enrolled, an enrollment application 330 is downloaded to a user device. For example, the application is downloaded to the user’s laptop computer, tablet computer, smartphone, smart watch, portable audio player, headset, etc. In some implementations, the WAC 310 is associated with a case, such as for storage and / or charging; and the enrollment application 330 can be downloaded to a computational environment stored in the case.

[0075] For the sake of illustration, FIG. 6 shows several example screenshots from an example enrollment application 330 running on a user device. At a first screen 610, the user begins a name enrollment process. By clicking “NEXT” using a user interface of the user device (e.g., a touchscreen), the user can proceed to a second screen 620. At the second screen 620, the user is prompted to enroll an invocation name. For example, the second screen 620 includes a button to activate a microphone of the user device by which to receive an audio sample from the user representing the invocation name being enrolled. Additionally or alternatively, the second screen 620 (or another screen) can include interface elements for receiving text, etc. Proceeding to a third screen 630 (e.g., by clicking “NEXT”), the user is presented with several options, such as an option to re-record the enrollment name, to enroll another name, or to end the enrollment process. Some implementations can present additional options, such as permitting the user to select any previously enrolled name to re-record, to delete, etc. In some cases, opting to re-record or to enroll another name can bring the user back to the second screen 620, or another similar screen. Opting to end the enrollment can bring the user to a fourth screen 640, which indicates to the user that the enrollment is complete.

[0076] In some implementations, conclusion of the user enrollment of invocation names automatically triggers the enrollment application 330 to compute (generate) some or all of the name detection model. In other implementations, subsequent to the user enrollment of invocation names, the user is prompted to continue with generation of some or all of the name detection model. In some implementations, some or all of the name detection model is generated separately from the enrollment application 330. After the name detection model is generated, the name detection model can be ported to the WAC 310 for local execution. Some embodiments of the enrollment application 330 permit the user, at any suitable time, to enroll additional invocation names, delete enrolled invocation names, etc.

[0077] Some embodiments described herein assume joint participation of a cloud-based computational platform, a local computational platform separate from the WAC 310 (e.g., a smartphone), and the computational platform integrated in the WAC 310. Different arrangements of features, components, etc. can be implemented depending on the computing, power, storage, and / or other resources of these computational platforms. In one implementation, the application is downloaded directly to the WAC 310 (or is previously loaded to the WAC 310), and the name detection model is computed directly by the WAC 310 (i.e., there is no need for a separate computational platform. In another implementation, enrollment information is exchanged with cloud-based processing resources to generate some or all of the name detection model. For example, audio samples corresponding to the invocation names (e.g., including augmentations thereof) are sent to the cloud, cloud-based resources are used to compute the name detection model, and the name detection model is ported (e.g., directly from the cloud, or via one or more intermediary devices) to the WAC 310. In other implementations, the application is directly ported to the WAC 310, and it is then downloaded to, or installed on, the local computational platform separate from the WAC 310 (e.g., the smartphone, etc.), if the local computational platform does not already have it while pairing.

[0078] FIG. 7 shows a flow diagram of an illustrative method 700 for audio management that includes automated attention handling in a wearable audio component (WAC), according to embodiments described herein. Embodiments of the method 700 can be implemented using any of the system implementations described herein, or any other suitable variation thereof. Some embodiments begin at stage 704 by receiving a real-time audio signal (e.g., from a reference microphone associated with an active noise control (ANC) system). The real-time audio signal is ambient audio received via the WAC while a user is listening to desired audio with ANC in an active (ambient sound suppression) mode.

[0079] At stage 708, embodiments can detect whether the real-time audio signal includes attention seeking (AS) audio. For example, as described above with reference to FIG. 1, an AHS system 150 can be used to detect when a second-party attention seeker is trying to get the attention of a user while the user is wearing the WAC and is listening to desired audio 165 with the ANC system 140 in the active mode. In particular, the AHS system 150 listens for presence of attention seeking (AS) audio 157 within the ambient audio 155. For example, embodiments of the AHS system 150 are configured to listen for AS audio 157 corresponding to a previously enrolled invocation name (e.g., a name of the user). As described herein, the detection in stage 708 can be performed using automated attention handling based on linguistic name embedding, based on universal sound conversion, or based on a hybrid thereof. As illustrated, the detection at stage 708 can rely on prior training of a name detection (i.e., inference) model at stage 720, generation and storage of reference embeddings based on a set of enrolled invocation names using the name detection model at stage 722, generation of real-time embeddings from the real-time audio signal using the name detection model at stage 724, and comparison of the real-time embeddings with the reference embeddings to determine whether the AS audio is present at stage 726.

[0080] A determination block at stage 712 represents the result of the determination at stage 708. If no AS audio is detected, embodiments of the method 700 return to stage 704. For example, embodiments continue to listen to the real-time audio signal, and the ANC system remains in active mode. If AS audio is detected, embodiments proceed to stage 716 by triggering the ANC system automatically to switch to a conversation mode. For example, referring back to FIG. 1, when the AS audio 157 is detected by the AHS system 150, the AHS system 150 automatically directs the ANC system 140 to switch from the active mode to the conversation mode. As described herein, the conversation mode can include disabling ANC, lowering the volume of desired audio, pausing the desired audio, enhancing conversationally relevant audio, suppressing feedback of the user’s own speech, etc.

[0081] As illustrated by off-page reference “A”, some embodiments of the method 700 include an enrollment phase prior to stage 704. FIG. 8 shows a flow diagram of an illustrative method 800 for such an enrollment phase. In some embodiments, the method 800 begins at stage 804 when a new WAC is detected. Such a detection can occur when a WAC is first paired with a user device, first paired with a network, first set to perform linguistic name-embedded attention handling, etc. For example, the term “new” in this context can simply indicate that the WAC is new with respect to features of automated attention handling described herein. In response to the detection at stage 804, embodiments can obtain an enrollment application (e.g., from the cloud) at stage 808.

[0082] At stage 812, embodiments can receiving a set of invocation names (e.g., see stage 712 of FIG. 7) from the user. At stage 816, embodiments can generate a reference name embedding for each of the invocation names by the processor-executable name embedding model. Stage 816 can correspond to stage 722 of FIG. 7. In some embodiments, generating the reference name embedding at stage 816 includes applying a plurality of augmentation transformations to each of the set of invocation names to generate an augmented set of invocation names and generating a reference name embedding for each of the augmented set of invocation names by the processorexecutable name embedding model. At stage 820, embodiments can storing the reference name embeddings in a non-transitory deep image.

[0083] Returning to FIG. 7, embodiments of the method 800 can also include conversation end detection subsequent to stage 716, as indicated by off-page reference “B.” FIG. 9 shows a flow diagram of an illustrative method 900 for such conversation end detection. For example, it is assumed that the output of the name invoked signal at stage 716 of FIG. 7 indicates the beginning of a conversation involving the user and a second party. At stage 904, embodiments detect a conversation end trigger subsequent to stage 716 (i.e., after the name invoked signal directed the ANC system automatically to enter the conversation mode). At stage 908, in response to the detection at stage 904, embodiments can output a conversation end signal responsive to detecting the conversation end trigger. The name invoked signal directs the ANC system automatically to switch from an ambient sound suppression mode to a conversation mode, and the conversation end signal directs the ANC system automatically to switch from the conversation mode to the ambient sound suppression mode.

[0084] FIG. 10 shows a flow diagram of an illustrative method 1000 for audio management that includes automated attention handling in a wearable audio component (WAC) using linguistic name embedding techniques, according to embodiments described herein. Embodiments of the method 1000 can be implemented using any of the system implementations described herein, or any other suitable variation thereof. Embodiments of the method 1000 can be a linguistic-name- embedding-specific implementation of the method 700 of FIG. 7. Similar to stage 704 of FIG. 7, some embodiments of the method 1000 begin at stage 1004 by receiving a real-time audio signal (e.g., from a reference microphone associated with an active noise control (ANC) system). The real-time audio signal is ambient audio received via the WAC while a user is listening to desired audio with ANC in an active (ambient sound suppression) mode.

[0085] At stage 1008, embodiments can generate a real-time embedding from the real-time audio signal by a processor-executable name embedding model trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications. In some embodiments, prior to receiving the real-time audio signal at stage 1008, the name embedding model is trained at stage 1006. For example, the name embedding model is trained by transfer learning from a foundation model that is an artificial neural network trained to linguistically classify a speech-audio corpus of phonetically diversified words spoken by a plurality of accent- diversified speakers.

[0086] At stage 1012, embodiments can select a candidate name embedding based on determining, by a processor-executable relation network, that one or more of a stored plurality of reference name embeddings has at least a threshold similarity with the real-time embedding, the stored plurality of reference name embeddings previously generated by the name embedding model based on a set of invocation names provided by a user during an enrollment procedure. In some embodiments, the selecting at stage 1012 includes computing a similarity score between each of the reference name embeddings and the real-time embedding, the similarity score indicating a mathematical correlation. Such embodiments can determine whether the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold and can output the candidate name embedding as the one of the reference name embeddings yielding a highest similarity score when it is determined that the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold.

[0087] At stage 1016, embodiments can output a name invoked signal to the ANC system based on determining that the real-time embedding and the candidate name embedding cannot be discriminated by a processor-executable false rejection network in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces. The name invoked signal can direct the ANC system automatically to enter a conversation mode. In some embodiments, the outputting at stage 1016 includes transforming the real-time embedding and the candidate name embedding into a plurality of mathematical spaces. For each mathematical space, a discrimination score can be computed indicating a likelihood of discriminating between the real-time embedding and the candidate name embedding in the mathematical space. The name invoked signal can be output responsive to the discrimination score computed for at least one of the mathematical spaces falling below the predetermined discrimination threshold. Such a determination can also be performed by computing a likelihood ratio, such that the name invoked signal can be output responsive to determining that the likelihood ratio exceeds the discrimination threshold.

[0088] Embodiments of the method 1000 can include preceding and / or subsequent stages, such an indicated by off-page references. As in the method 700 of FIG. 7, off-page reference “A” optionally references a preceding enrollment phase. An example of such an enrollment phase is described with reference to FIG. 8. Also as in the method 700 of FIG. 7, off-page reference “B” optionally references a subsequent conversation end detection phase. An example of such a conversation end detection phase is described with reference to FIG. 9.

[0089] Unified Sound Conversion

[0090] Some of the embodiments described above automatically classify spoken audio into one of a number of linguistically distinct classifications (i.e., each classification corresponds to a unique word, regardless of suprasegmental variations). An embedding model essentially converts enrolled invocation names into an enrolled set of classifications and stores associated reference embeddings. The same embedding model can then try to classify real-time audio to produce a real-time embedding. Other models are then used to determine whether the real-time embedding matches any of the reference embeddings, which would indicate detection of one of the enrolled invocation names in the real-time audio and can be used automatically to switch ANC into a conversation mode.

[0091] Other embodiments convert spoken audio of a class (i.e., a word) into a unified sound that mimics what would be generated by a speech synthesizer for that class (i.e., the unified sound is stripped of the speaker’s suprasegmental influence on the audio). An embedding model essentially converts enrolled invocation names into an enrolled set of unified sound names and stores them as reference embeddings. The same embedding model can then convert real-time audio into a unified sound to produce a real-time embedding. Other models are then used to determine whether the real-time embedding matches any of the reference embeddings, which would indicate detection of one of the enrolled invocation names in the real-time audio and can be used automatically to switch ANC into a conversation mode.

[0092] FIGS. 11A and 11B show block diagrams of a training environment 1100 for training a foundation model 350 to support unified sound conversion (USC) for automated name-detectionbased attention handling, according to some embodiments described herein. As illustrated, the training environment 1100 can begin with a spoken word audio repository 1110. The repository 1110 can include one or more large corpuses of spoken audio data. Preferably, the corpuses include many words (classes), and many diversified samples in each class, so that each class includes many versions of the same word spoken with wide suprasegmental variance. For example, a given word may be spoken 10,000 times by different speakers from around the world.Thus, for each class, the repository 1110 can output a large number of diversified audio samples for the class, which can be referred to as the “repository spoken audio samples” for the class.

[0093] The repository spoken audio samples for each class form the entire set of spoken audio samples for that class, referred to as the “class spoken audio samples.” For example, if the repository 1110 includes S samples for a particular class (S is a positive integer), there are S class spoken audio samples. In other embodiments, additional variation in the audio samples for each class is created by passing the class spoken audio samples through an augmenter 1120. The augmenter 1120 uses one or more augmentation models 1115 to generate an augmented set of class spoken audio samples with variations in features, such as speech speed (e.g., lengthening or shortening of the audio sample, lengthening or shortening of some or all vowel sounds, etc.), modeled suprasegmental variations, models of noise profiles and / or ambient noise features (e.g., traffic sounds, background conversation sounds, etc.), etc. Some embodiments of the augmenter 1120 are implemented in the same manner as the name augmenter 520 of FIG. 5B (e.g., and the augmentation models 1115 can be the same as, or different from the augmentation models 515 of FIG. 5B). Other embodiments of the augmenter 1120 introduce more and / or different types of variation using more and / or other augmentation models 1115. In embodiments that include the augmenter 1120, the augmented set of audio samples for the class is used as the class spoken audio samples. For example, if the augmenter 1120 produces A augmentations for each of the S repository spoken audio samples (A is a positive integer), there are A * S class spoken audio samples. In some cases, the repository 1110 may include different numbers of samples for different classes, and / or the augmenter 1120 may apply different types of augmentations for different classes, such that the values of A and / or S may be class dependent.

[0094] Embodiments also generate a unified audio sample for each class, referred to as a “class unified audio sample.” As illustrated in FIG. 11 A, in some embodiments, the repository 1110 includes a lexical entry for each of some or all of the classes, which can be used directly as “class text.” For example, the term “INDEPENDENCE” can have hundreds of diversified spoken audio samples for the word, all stored in association with a lexical entry (i.e., the text) for the word. In other embodiments, the repository 1110 may not include lexical entries for classes, or may not include a lexical entry for one or more classes. In such embodiments, for any class that does not have an associated lexical entry, one or more of the repository spoken audio samples is fed to a speech -to-text (STT) engine 1130, which generates the class text from the repository spoken audio sample(s). The class text, whether derived from a lexical entry in the repository 1110 or generated by the STT engine 1130, is passed to a text-to-speech (TTS) synthesizer 1135 to produce the class unified audio sample. In some implementations, the class unified audio sample is stripped of allaccent, prosody, etc. For example, the class unified audio sample is a purely phonetic representation of the class in a standardized set of phonemes. In other implementations, the TTS synthesizer 1135 produces the class unified audio sample to have synthesized suprasegmental features (i.e., the class unified audio sample is stripped of the speaker’s suprasegmental influence, but it may still have synthesized suprasegmental features).

[0095] FIG. 1 IB shows alternative embodiments for generating the class unified audio sample. Rather than generating class unified audio sample from TTS synthesis, the class unified audio sample is generated using a selected audio sample for the class. As illustrated, embodiments can include a selector 1170. The repository 1110 includes spoken audio samples from many different speakers for each class (i.e., the class spoken audio samples), and the class unified audio sample for each class is generated by the selector 1170 by selecting one of the class spoken audio samples, and / or selecting one of the speakers of the class spoken audio samples. In one implementation, the selected spoken audio sample and / or speaker is the same for all classes. For example, the selector 1170 uses a particular identifier (e.g., index, etc.) to select a same speaker’s audio contribution as the unified sound for all classes. In other implementations, the selector 1170 randomly (or otherwise) selects a selected speaker from among the available speakers and / or samples for each class. For example, a particular speaker may not have provided audio contributions for all classes. In other implementations, the selection is based on quality features. For example, if the repository 1110 aggregates data from multiple corpuses, some corpuses may provide better quality audio samples than others; or within a particular one or more corpuses, some audio samples may be of better quality than others. In such cases, the selector 1170 is configured to select one of the class spoken audio samples as the class unified audio sample for each class based on the quality of the sample. In some cases, a particular corpus or portion of a corpus may provide better data for a particular geographic region in which the user resides. In such cases, the selector 1170 is configured to select one of the class spoken audio samples as the class unified audio sample for each class based on regional factors, such as suprasegmental features more common to a particular geographic region. For example, some implementations may ignore certain very strong regional accents for the purposes of generating the class unified audio sample, except in cases where the user resides in geographic proximity to speakers with that regional accent.

[0096] As illustrated in both FIGS. 11 A and 1 IB, embodiments include a feature extractor 1140, which takes an audio waveform as its input and outputs a computed a set of features, referred to as a feature vector. In some implementations, the computed feature vectors are sets of filter bank energies, which are a set of values representing the respective energy content at each of a set of different frequency bands in the input audio signal. For example, the input audio signal is passedthrough a set of bandpass filters (e.g., each designed to respond to a specific range of frequencies that mimics the frequency response of the human auditory system). The filtered signals can then be rectified (e.g., to preserve magnitude information and remove phase information) and smoothed (e.g., averaged), resulting in the so-called filter bank energies. As illustrated, passing the class spoken audio samples through the feature extractor 1140 produces spoken feature vectors 1142, and passing the class unified audio sample through the feature extractor 1140 produces a unified feature vector 1144. In some implementations, the feature vectors can be generated with a representation other than filter bank energies. In some such implementations, a complex fast Fourier transform (FFT) is used to generate the feature vectors as an array of complex numbers. Each complex number in the array represents a frequency component in the audio signal, where the magnitude of the complex number represents an amplitude of a corresponding frequency component, and the phase of the complex number represents a phase shift of the corresponding component relative to a reference. In other such embodiments, MFCCs can be used (e.g., or the filter bank energies can be converted to MFCCs) for a more compact representation, but such representations may not yield sufficient accuracy for all applications (e.g., the filter resolution may effectively be too low for successful training of the foundation model 1150). In other such embodiments, sample-to-sample modeling can be used, but such implementations may be too bulky to be practical in all applications.

[0097] As illustrated, during training, a foundation model 1150 takes the spoken feature vectors 1142 as its input labels and the unified feature vector 1144 as its output label. The foundation model 1150 is iteratively trained based on the input and output labels, so that inputting any of the spoken feature vectors 1142 for a particular class into the foundation model 1150 will cause the foundation model 350 to output the same unified feature vector 1144 for that class. In some embodiments, the feature extractor 1140 is integrated with the foundation model 1150 so that the class spoken audio samples can be input to the foundation model 1150, and the foundation model 1150 will try to generate an output unified audio sample that mimics the class synthesized audio sample (or to output a feature vector that mimics the unified feature vector 1144).

[0098] As one example, each audio sample at the input to the feature extractor 1140 is one second of audio sampled at 16 kHz (i.e., 16,000 samples). The sample is split into 100 frames of ten milliseconds each. The feature extractor 1140 extracts 30 Mel frequency bins (e.g., using bandpass filters tuned to the Mel scale) from each frame, thereby compressing each audio input of 16,000 samples into 30 * 100, or 3,000, filter bank energies. In this example, the feature extractor 1140 has an input dimension of 16,000 and an output dimension of 3,000. Suppose the repository 1110 has 10,000 repository spoken audio samples for a particular word, and the augm enter 1120generates ten augmentations for each repository spoken audio sample, resulting in 100,000 class spoken audio samples. The feature extractor 1140 converts each of the 100,000 class spoken audio samples into a respective 3,000 filter bank energies (i.e., the spoken feature vectors 1142), and converts the single class unified audio sample into another respective 3,000 filter bank energies (i.e., the unified feature vector 1144); and the foundation model 1150 uses all those filter bank energies to figure out how to convert any of the spoken feature vectors 1142 into the same unified feature vector 1144.

[0099] Ultimately, the foundation model 1150 is trained so that inputting any input audio sample 1160 will cause the foundation model 1150 to generate an output unified feature vector 1165 that is stripped of any speaker influence on suprasegmental features. For example, the generated output unified feature vector 1165 mimics the unified feature vector 1144 that would be generated if the input audio sample were to represent a class from the repository 1110, class text corresponding to the class were passed through the TTS synthesizer 1135 to generate a class synthesized audio sample, and the class synthesized audio sample were passed through the feature extractor 1140 to generate the unified feature vector 1144. In other words, the foundation model 1150 is trained to convert any input audio sample 1160 into a unified sound representation.

[0100] Having trained the foundation model 1150 to generate a unified sound version of an input audio sample, use of the foundation model 1150 in automated attention handling can be similar to what is described with reference to LNE-AHS above. FIG. 12 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block 1200 and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein. The AS trigger detection block 1200 can be implemented in a similar manner to the AS trigger detection block 400 of FIG. 4, and consistent reference designators are used to indicate such similarities. For example, both include an enrollment stage 410 that serves the same overall purpose with respect to the system and has is shown with the same reference designator, accordingly; but each enrollment stage 410 yields a different type of embedding model, and the embedding models are shown with different reference designators, accordingly. Like the AS trigger detection block 400 of FIG. 4, the AS trigger detection block 1200 of FIG. 12 can be an implementation of the AS trigger detection block 210 of FIG. 2. It is assumed that the AS trigger detection block 400 is implemented in the context of a WAC 310.

[0101] Embodiments can include an enrollment stage 410, an identification stage 430, and a verification stage 440. In general, the enrollment stage 410 occurs outside of normal operation ofthe WAC 310, such as when a user first sets up and / or uses the WAC 310, when the user first registers the WAC 310, when the user first configures the WAC 310 for automated attention handling, etc. The identification stage 430 and the verification stage 440 are real-time blocks that occur during normal operation of the WAC 310 to facilitate real-time automated attention handling. Embodiments generally use the enrollment stage 410 to obtain any information needed to set up an inference model for storing at the WAC 310 to enable the WAC 310 to perform automated attention handling features during normal operation. As shown, the inference model includes at least an embedding model 1215, a relation network 1235, and a false rejection network 1245.

[0102] In the enrollment stage 410, the embedding model 1215 can be used to convert an enrollment audio stream 405 into a set of reference embeddings based on a foundation model 1150. The foundation model 1150 can be stored in and accessed via the cloud 340 (e.g., and / or any other suitable communication network). As described with reference to FIG. 11, in the context of USC-based embodiments, the foundation model 1150 can be trained to convert an input audio stream into either a unified output audio sample (i.e., corresponding to the input audio stream stripped of any speaker influence on accent, prosody, etc.), or an output unified feature vector 1165 (e.g., a set of filter bank energies that describe the unified output audio sample). In this context, the embedding model 1215 is trained by transfer learning from the foundation model 1150 on a smaller corpus of name audio samples. For example, the encoder portion of the foundation model 1150 is taken by the embedding model 1215, and a new decoder is trained on the smaller name audio corpus.

[0103] As described above (e.g., with reference to FIGS. 4, 5A, 5B, 6, and 8), during the enrollment stage 410, the enrollment audio stream 405 can include several user-enrolled audio samples of invocation names and / or a set of augmentations to those invocation names. In this context, the “reference embeddings” refer to either the set of unified output audio samples generated from the set of invocation names (e.g., and their augmentations), or the set of output unified feature vectors 1165 generated from the set of invocation names (e.g., and their augmentations). The reference embeddings can be stored as a deep image 1220. The deep image 1220 can also be considered as part of the inference model.

[0104] During normal operation, in the identification stage 430, a real-time (RT) audio stream 407 is received. The embedding model 1215 (i.e., the same embedding model 1215 generated in the enrollment stage 410) is used to generate a RT embedding from the received RT audio stream 407. In this context, the “RT embedding” refers to either a unified output audio sample generatedfrom the RT audio stream 407, or an output unified feature vector 1165 generated from the RT audio stream 407. The relation network 1235 can then compare the RT embedding with each of the reference embeddings in the deep image 1220 to determine if there is a match. For example, the relation network 1235 is configured to compute a similarity score for each comparison (e.g., corresponding to a mathematical correlation, or the like). The relation network 1235 can determine if any similarity scores meet or exceed a predetermined threshold; if so, the reference embedding with the highest similarity score can be selected as a candidate matching embedding. Further options and / or features of the relation network 1235 can be understood with reference to descriptions of the relation network 435 above.

[0105] In the verification stage, the false rejection network 1245 can then confirm the match by transforming the RT embedding and the candidate matching embedding into different mathematical spaces (e.g., domains) and determining whether the embeddings can be reliably discriminated from each other in any of the mathematical spaces. For example, the deep image 1220 is configured to compute a discrimination score for each mathematical space. The deep image 1220 can determine if any discrimination scores meet or exceed a predetermined threshold; if so, the match determined by the relation network 1235 is determined to be a false match and is ignored. If the embeddings cannot be sufficiently discriminated in any mathematical spaces, the false rejection network 1245 can output a signal that attention seeking audio has been detected (i.e., a name invoked signal 450). As described herein, the name invoked signal 450 can direct an ANC system of the WAC 310 automatically to enter a conversation mode. As noted with reference to FIG. 4 above, embodiments of the embedding model 1215 are a processor-executable name embedding model, embodiments of the deep image 1220 are a processor-readable deep image, embodiments of the relation network 1235 are a processor-executable relation network, and embodiments of the false rejection network 1245 are a processor-executable false rejection network. Further options and / or features of the false rejection network 1245 can be understood with reference to descriptions of the false rejection network 445 above.

[0106] It can be seen that, although the stages used in unified sound conversion embodiments are similar to those used in LNE-AHS embodiments, the underlying modeling approach used for name detection is different. For example, in both categories of embodiments, input audio can be encoded using types of language modeling, linguistic embedding, etc. However, the LNE-AHS embodiments use a classifier architecture that essentially converts the encoded linguistically salient features into a classification, while the unified sound conversion embodiments use an encoder-decoder architecture that essentially converts the encoded linguistically salient features into a unified audio output. For example, as described above, the LNE-AHS embodiments canrely on architectures, such as two-dimensional convolution with long short-term memory (LSTM) networks, etc.; while unified sound conversion embodiments can rely on architectures, such as two-dimensional convolution with recurrent neural networks (e.g., a gated recurrent unit (GRU) network).

[0107] As described above, some USC-based embodiments implement the embedding model 1215 to include both an encoder and a decoder, so that the output of the embedding model 1215 is the output unified feature vector 1165 (e.g., synthesized filter bank energies). As illustrated in FIGS. 11 A and 1 IB, in other embodiments, the embedding model 1215 is implemented only to include the encoder portion of the model (i.e., all layers of the model after the bottleneck layer are removed after training). As described above, encoder-decoder architectures (e.g., auto-encoder architectures) include a neural network architecture designed to learn compact representations of data, referred to as “latent features.” The encoder portion receives a high-dimensionality input (i.e., the raw audio data) and includes a multi-layer network to progressively reduce the dimensionality of the data using transformations, until a lowest dimensionality (i.e., most compressed) representation is reached at the bottleneck layer. The vector space that is represented at the bottleneck layer can be referred to as a latent space embedding of the audio sample. In some embodiments, the latent space embedding generated by the encoder (e.g., at the bottleneck layer) is used as a unified sound code word (USCW) 1167.

[0108] In such embodiments, the embedding model 1215 is trained to generate respective USCWs 1167 as the reference embeddings stored in the deep image 1220 and to generate a respective USCW 1167 as the RT embedding, the relation network 1235 is trained to determine candidate matches based on comparing the USCWs 1167, and the false rejection network 1245 is trained to reject false matches by discriminating between the USCWs 1167. With proper training and implementation, the performance of the AS trigger detection block 1200 can be almost the same for embodiments of the inference model that include a decoder and are based on output unified feature vectors 1165 as for embodiments of the inference model that do not include a decoder and are based on UCSWs 1167.

[0109] FIG. 13 shows a flow diagram of an illustrative method 1300 for automated audio management based on name detection, according to USC-based embodiments described herein. Embodiments of the method 1300 begin at stage 1304 by obtaining (e.g., by one or more processors) spoken audio samples for each word of a large number of phonetically diversified words from a spoken word audio repository. As described herein, the spoken word audio repository includes one or more speech-audio corpuses of suprasegmentally diversified speech-audio samples of the phonetically diversified words. As such, the spoken audio samples for each word of the phonetically diversified words includes those of the suprasegmentally diversified speech-audio samples representing respective instances of the word.

[0110] At stage 1308, embodiments can obtain (e.g., by the one or more processors) a respective text representation of each word of the phonetically diversified words. For some or all the phonetically diversified words, some embodiments can convert one or more of the class spoken audio samples associated with the word into the respective text representation of the word. For example, a class spoken audio sample is speech-to-text converted in stage 1306 to generate the textual representation. In some implementations, for each of at least some of the phonetically diversified words, the spoken word audio repository includes a lexical entry for the word (e.g., stored in association with at least one of the class spoken audio samples for the word). In such implementations, embodiments of the method 1300 can obtain the respective text representation of each of some or all the phonetically diversified words in stage 1308 by obtaining the lexical entry for the word from the spoken word audio repository.[OHl] At stage 1312, embodiments can synthesize (e.g., by the one or more processors) a respective class unified audio sample for each word from a respective text representation of the word. At stage 1316, embodiments can convert, for each word (e.g., by the one or more processors), the respective class unified audio sample for the word into respective output labels representing a unified feature vector of the respective class unified audio sample. At stage 1320, embodiments can convert, for each word (e.g., by the one or more processors), the class spoken audio samples associated with the word into respective input labels representing spoken feature vectors of the class spoken audio samples. In some embodiments, the spoken feature vectors represent filter bank energies of the plurality of class spoken audio samples, and the unified feature vector represents filter bank energies of the respective class unified audio sample.

[0112] Some embodiments, at stage 1318, prior to the converting in stage 1320, can add class augmented audio samples to the class spoken audio samples associated with the word by applying one or more augmenter models to one or more of the class spoken audio samples from the spoken word audio repository (i.e., the class spoken audio samples for the word include both the class spoken audio samples from the spoken word audio repository and the class augmented audio samples). For example, each of the augmenter models mathematically represents a respective one of several noise models and / or suprasegmental feature models (e.g., models of different accents, prosody, etc.).

[0113] At stage 1324, embodiments can train a foundation model. As described herein, the foundation model is an artificial neural network trained, based on the input labels and the output labels for each word, to convert any particular one of the class spoken audio samples for the word into the respective class unified audio sample for the word by automatically removing suprasegmental differences between the particular class spoken audio sample and the class unified audio sample. For example, once properly trained, the foundation model is able to remove a speaker’s suprasegmental influence on spoken audio received from the speaker as the input stream, thereby generating a unified version of the audio that mimics what would be generated by text-to- speech synthesis.

[0114] FIG. 14 shows a flow diagram of an illustrative method 1400 for training a unified sound conversion (USC) inference model. As represented by off-page reference “C,” the method 1400 can proceed based on and subsequent to the training of the foundation model in stage 1324 of FIG. 13. At stage 1404, embodiments train the USC inference model (e.g., a processor-executable USC inference model) based on transfer learning from the foundation model. The training is such that the USC inference model can automatically: receive a real-time audio stream; generate a real-time embedding representing the real-time audio stream stripped of suprasegmental features; and output a name invoked signal based on matching, with at least a predetermined threshold confidence level, the real-time embedding to one of a stored plurality of reference name embeddings generated to represent unified audio representations of each of a set of invocation names stripped of the suprasegmental features, the name invoked signal to direct an ANC system automatically to enter a conversation mode.

[0115] In some embodiments, the training in stage 1404 can proceed according to stages 1408 - 1416. At stage 1408, embodiments can train a name embedding model, by transfer learning from the foundation model based on a corpus of real-world name audio samples, automatically to remove the suprasegmental features from an input audio sample to generate a corresponding unified audio output. In such embodiments, the real-time embedding is generated by the name embedding model from the real-time audio signal received during an operational time, and the reference name embeddings are generated by the name embedding model based on enrollment audio samples received during an enrollment time.

[0116] At stage 1412, embodiments can train a relation network to output a candidate name embedding responsive to determining that one of the reference name embeddings has a highest similarity with the real-time embedding and that the highest similarity exceeds a predetermined similarity threshold. For example, the relation network is trained to output the candidate name by:computing a similarity score between each of the reference name embeddings and the real-time embedding, the similarity score indicating a mathematical correlation; determining whether the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold; and outputting the candidate name embedding as the one of the reference name embeddings yielding a highest similarity score when it is determined that the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold.

[0117] At stage 1416, embodiments can train a false rejection network to output the name invoked signal responsive to determining that the real-time embedding and the candidate name embedding cannot be discriminated in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces. For example, the false rejection network is trained to output the name invoked signal by: transforming the real-time embedding and the candidate name embedding into a plurality of mathematical spaces; computing, for each mathematical space, a discrimination score indicating a likelihood of discriminating between the real-time embedding and the candidate name embedding in the mathematical space; and outputting the name invoked signal responsive to the discrimination score computed for at least one of the mathematical spaces falling below the predetermined discrimination threshold. In embodiments that perform stages 1412 and 1416, the predetermined confidence level in stage 1404 can be based on a combination of the predetermined similarity threshold and the predetermined discrimination threshold.

[0118] FIG. 15 shows a flow diagram of an illustrative method 1500 for using the trained unified sound conversion (USC) inference model during an operation time. As represented by off- page reference “D,” the method 1500 can proceed based on and subsequent to the training of the USC inference model in stage 1404 of FIG. 14. Embodiments of the method 1500 begin at stage 1504 by receiving a real-time audio signal by the USC inference model (e.g., from a reference microphone associated with the ANC system). At stage 1508, embodiments can generate a realtime embedding from the real-time audio signal by the USC inference model. At stage 1512, embodiments can output the name invoked signal in response to (e.g., only after determining) determining (e.g., only after determining), by the USC inference model, that the real-time embedding matches any of the stored plurality of reference name embeddings with at least the predetermined threshold confidence level.

[0119] Similar to the method 1000 of FIG. 10, embodiments of the method 1500 can include additional preceding and / or subsequent stages, such an indicated by off-page references “A” and “B”. As described above, off-page reference “A” optionally references a preceding enrollmentphase. An example of such an enrollment phase is described with reference to FIG. 8. Also as described above, off-page reference “B” optionally references a subsequent conversation end detection phase. An example of such a conversation end detection phase is described with reference to FIG. 9.

[0120] Hybrid Unified Linguistic Name Embedding (ULNE)

[0121] Some embodiments described above perform automated name detection-based attention handing using linguistic name embedding (LNE). LNE-based embodiments rely on a classification model to convert names into linguistically distinct classifications based on linguistically salient features. Such a classification model tends to have a performance tradeoff between a true positive rate and a false acceptance rate for names and sounds that are linguistically similar. Additionally, to ensure good performance of such LNE-based approaches when there is discrepancy between the prosody or accent of enrolled names and those of microphone-captured names uttered by attention seekers, such LNE-based approaches tend to involve augmentation of audio samples with many types of time scaling, pitch shifting, etc., which tends to increase the computation load of the identification stage.

[0122] Other embodiments described above perform automated name detection-based attention handing using unified sound conversion (USC). USC-based embodiments rely on an encoderdecoder model to convert names into unified sound representations that stripped of any speaker influence on suprasegmental features of the audio. This suprasegmental unification (e.g., normalization) can help USC-based embodiments avoid some of the limitations of LNE-based embodiments. However, some USC-based embodiments can have their own limitations, such as challenges in classifying similar sounding names from the resulting types of embedding or normalized features.

[0123] Other embodiments perform automated name detection-based attention handing using a hybrid unified linguistic name embedding (ULNE) approach, which combines features of USC- based embodiments with features of LNE-based embodiments. FIG. 16 shows a simplified block diagram of an embedding environment 1600 incorporating an illustrative hybrid name embedding model 1610 for use in a unified linguistic name embedding attention handling system (ULNE- AHS). As illustrated, the hybrid name embedding model 1610 includes two encoding stages. In a first stage, a USC-trained embedding model includes a USC encoder 1620 and a USC decoder 1625, which can be components of embedding model 1215, as described with reference to FIGS. 11 A and 1 IB. The USC-trained embedding model is trained to convert an input audio stream 1605 into a unified sound representation as a set of unified features. For example, the unifiedfeatures are an output unified feature vector 1165. In a second stage of the hybrid name embedding model 1610, an LNE-trained embedding model includes an LNE encoder 1630 and an LNE classifier 1635, which can be components of embedding model 415, as described with reference to FIG. 4. The LNE-trained embedding model is trained to convert the unified features at the output of the USC-trained embedding model into an output classification 1640.

[0124] Ultimately, the output classification 1640 at the output of the hybrid name embedding model 1610 can be the same classification output as in the LNE-based embodiments described above (e.g., at the output of name embedding model 415). As such, ULNE -based embodiments can use the same relation network 435 and the same false rejection network 445 used in the LNE- based embodiments. For example, referring to FIG. 4, the embedding model 415 can be replaced with the hybrid name embedding model 1610 from FIG. 16. In that context, during the enrollment stage 410, the hybrid name embedding model 1610 converts the enrollment audio stream 405 into corresponding classifications for storage in the deep image 420 by converting each name audio sample into unified features, and converting the unified features into an appropriate classification. As such, during the enrollment stage 410, the output classifications 1640 generated by the hybrid name embedding model 1610 are used as the reference embeddings. During real-time operation, the hybrid name embedding model 1610 converts the RT audio stream 407 into a corresponding classification by converting the received audio into unified features, and converting the unified features into an appropriate classification. As such, during the enrollment stage 410, the output classification 1640 generated by the hybrid name embedding model 1610 is used as the RT embedding. The classifications represented by the reference and RT embeddings are then used by the relation network 435 and the false rejection network 445 to determine whether to output the name invoked signal 450.

[0125] In effect, unified sound conversion is used at the front-end to normalize prosody and accent prior to performing linguistic name embedding, so that the linguistic name embedding is performed on a unified version of the name stripped of any particular accent and / or prosody influence by the attention seeker. This can enable LNE-based classification to generate well- defined and linguistically discriminative name embedding with sustained and effective performance across a wide variety of accents and prosody. As such, performance of ULNE-based embodiments can be better than embodiments based on either LNE alone or USC alone.

[0126] Referring back to the method 700 of FIG. 7, any of the LNE-based, USC-based, or ULNE-based approaches described herein can be used for automated attention handling. For example, in stage 704, a real-time audio signal is received by a reference microphone associatedwith an active noise control (ANC) system while the ANC system is in an ambient sound suppression mode. It can be assumed that the ANC system is integrated into a wearable audio component being worn by a first party (e.g., the user). At stage 708, a pre-trained inference model is used to detect whether the real-time audio signal includes attention-seeking audio spoken by a second party (i.e., the attention seeker, who is someone other than the user). The attention-seeking audio is predetermined to indicate that the second party is seeking attention of the first party. At stages 712 and 716, embodiments can output a name invoked signal to the ANC system automatically in response to determining that the real-time audio signal includes the attentionseeking audio, the name invoked signal directing the ANC system automatically to switch from the ambient sound suppression mode to a conversation mode. As described there, some embodiments can further include enrollment (e.g., according to the method 800 of FIG. 8) and / or detection of a conversation end trigger (e.g., according to the method 900 of FIG. 9).

[0127] Generally, the attention-seeking audio can be predetermined to indicate that the second party is seeking attention of the first party based on a set of invocation names previously enrolled by the first party that is previously converted by the pre-trained inference model into a set of reference embeddings. As such, the detecting can generally involve: converting the real-time audio signal to a real-time embedding by the pre-trained inference model; and determining whether the real-time embedding matches any of the reference embeddings with at least a threshold confidence level. The converting of the real-time audio signal to the real-time embedding can be performed by a processor-executable name embedding model.

[0128] For example, in LNE-based embodiments, the name embedding model is an LNE-based name embedding model trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications by transfer learning from a foundation model, and the foundation model is an artificial neural network trained to linguistically classify a speech-audio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers. In USC-based embodiments, the name embedding model is a USC-based name embedding model trained, by transfer learning from a foundation model based on a corpus of real- world name audio samples, automatically to remove suprasegmental features from an input audio sample to generate a corresponding unified audio output; and the foundation model is an artificial neural network trained, based on a speech-audio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers, to generate class unified audio samples from class spoken audio samples by removing the suprasegmental features.

[0129] In ULNE-based embodiments, the name embedding model is a ULNE-based name embedding model (i.e., hybrid name embedding model 1610) that can include: a first embedding stage trained automatically to remove suprasegmental features from a corpus of real-world name audio samples to generate corresponding unified audio outputs; and a second embedding stage coupled with the first embedding stage and trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications. In such embodiments, the converting of the real-time audio signal to the real-time embedding by the name embedding model can involve converting the real-time audio signal, by the first embedding stage, into a unified realtime audio signal stripped of suprasegmental influence by the second party; and converting the unified real-time audio signal, by the second embedding stage, into the real-time embedding, the real-time embedding representing a name classification.

[0130] FIG. 17 provides a schematic illustration of an illustrative computational system 1700 that can implement various system components and / or perform various steps of methods provided by various embodiments. Embodiments of the computational system 1700 can be integrated in a WAC, such as an earbud, headset, etc. Embodiments of the computational system 1700 can implement some or all of the audio management system 100 of FIG. 1, including embodiments of the AHS 150, the ANC 140, and / or the audio processing system (APS) 160 described herein. FIG. 17 is meant only to provide a generalized illustration of various components, any or all of which may be utilized as appropriate. FIG. 17, therefore, broadly illustrates how individual system elements may be implemented in a relatively separated or relatively more integrated manner.

[0131] The computational system 1700 is shown including hardware elements that can be electrically coupled via a bus 1705 (or may otherwise be in communication, as appropriate). The hardware elements may include one or more processors 1710, including, without limitation, one or more general-purpose processors and / or one or more special-purpose processors (such as digital signal processing chips, graphics acceleration processors, video decoders, and / or the like); one or more input devices 1715; and one or more output devices 1720. In the WAC context, the input devices 1715 can include wired and / or wireless ports, buttons, switches, microphones, touch interfaces, and / or any other suitable input device 1715; and the output devices 1720 can include indicator lights, displays, speakers, and / or any other suitable output devices 1720.

[0132] The computational system 1700 may further include (and / or be in communication with) one or more non-transitory storage devices 1725, which can comprise, without limitation, local and / or network accessible storage, and / or can include, without limitation, a disk drive, a drive array, an optical storage device, a solid-state storage device, such as a random access memory(“RAM”), and / or a read-only memory (“ROM”), which can be programmable, flash-updateable and / or the like. Such storage devices may be configured to implement any appropriate data stores, including, without limitation, various file systems, database structures, and / or the like. In some embodiments, the storage devices 1725 include the deep image 420 and / or an inference model 1727. As described herein, the inference model can include one or more types of name embedding models, relation networks, false rejection networks, etc. for implementing name detection-based attention handling.

[0133] The computational system 1700 can also include a communications subsystem 1730, which can include, without limitation, a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication device, etc.), and / or the like. As described herein, the communications subsystem 1730 supports multiple communication technologies. Further, as described herein, the communications subsystem 1730 can provide communications with one or more networks 140, and / or other networks. For example, embodiments of the communications subsystem 1730 can communicate with a foundation model 350 via the cloud 350. Though not explicitly shown, some embodiments interface via the communications subsystem 1730, and / or via input devices 1715 and output devices 1720, with one or more user computational devices 320.

[0134] In many embodiments, the computational system 1700 will further include a working memory 1735, which can include a RAM or ROM device, as described herein. The computational system 1700 also can include software elements, shown as currently being located within the working memory 1735, including an operating system 1740, device drivers, executable libraries, and / or other code, such as one or more application programs 1745, which may include computer programs provided by various embodiments, and / or may be designed to implement methods, and / or configure systems, provided by other embodiments, as described herein. Merely by way of example, one or more procedures described with respect to the method(s) discussed herein can be implemented as code and / or instructions executable by a computer (and / or a processor within a computer); in an aspect, then, such code and / or instructions can be used to configure and / or adapt a general -purpose computer (or other device) to perform one or more operations in accordance with the described methods. In some embodiments, the operating system 1740 and the working memory 1735 are used in conjunction with the one or more processors 1710 to implement some or all of the audio management system 100 components, such as the ANC 140, AHS 150, and / or APS 160.

[0135] A set of these instructions and / or codes can be stored on a non-transitory computer- readable storage medium, such as the non-transitory storage device(s) 1725 described above. In some cases, the storage medium can be incorporated within a computer system, such as computer system 1700. In other embodiments, the storage medium can be separate from a computer system (e.g., a removable medium, such as a compact disc), and / or provided in an installation package, such that the storage medium can be used to program, configure, and / or adapt a general -purpose computer with the instructions / code stored thereon. These instructions can take the form of executable code, which is executable by the computational system 1700 and / or can take the form of source and / or installable code, which, upon compilation and / or installation on the computational system 1700 (e.g., using any of a variety of generally available compilers, installation programs, compression / decompression utilities, etc.), then takes the form of executable code.

[0136] It will be apparent to those skilled in the art that substantial variations may be made in accordance with specific requirements. For example, customized hardware can also be used, and / or particular elements can be implemented in hardware, software (including portable software, such as applets, etc.), or both. Further, connection to other computing devices, such as network input / output devices, may be employed.

[0137] As mentioned above, in one aspect, some embodiments may employ a computer system (such as the computer system 1700) to perform methods in accordance with various embodiments of the invention. According to a set of embodiments, some or all of the procedures of such methods are performed by the computational system 1700 in response to processor 1710 executing one or more sequences of one or more instructions (which can be incorporated into the operating system 1740 and / or other code, such as an application program 1745) contained in the working memory 1735. Such instructions may be read into the working memory 1735 from another computer-readable medium, such as one or more of the non-transitory storage device(s) 1725. Merely by way of example, execution of the sequences of instructions contained in the working memory 1735 can cause the processor(s) 1710 to perform one or more procedures of the methods described herein.

[0138] The terms “machine-readable medium,” “computer-readable storage medium” and “computer-readable medium,” as used herein, refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. These mediums may be non-transitory. In an embodiment implemented using the computer system 1700, various computer-readable media can be involved in providing instructions / code to processor(s) 1710 for execution and / or can be used to store and / or carry such instructions / code. In many implementations, a computer-readable medium is a physical and / or tangible storage medium. Such a medium may take the form of a non-volatile media or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as the non-transitory storage device(s) 1725. Volatile media include, without limitation, dynamic memory, such as the working memory 1735. Common forms of physical and / or tangible computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, any other physical medium with patterns of marks, a RAM, a PROM, EPROM, a FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.

[0139] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to the processor(s) 1710 for execution. Merely by way of example, the instructions may initially be carried on a magnetic disk and / or optical disc of a remote computer. A remote computer can load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received and / or executed by the computer system 1700. The communications subsystem 1730 (and / or components thereof) generally will receive signals, and the bus 1705 then can carry the signals (and / or the data, instructions, etc., carried by the signals) to the working memory 1735, from which the processor(s) 1710 retrieves and executes the instructions. The instructions received by the working memory 1735 may optionally be stored on a non-transitory storage device 1725 either before or after execution by the processor(s) 1710.

[0140] Having described several example configurations, various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the disclosure. For example, the above elements may be components of a larger system, wherein other rules may take precedence over or otherwise modify the application of the invention. Also, a number of steps may be undertaken before, during, or after the above elements are considered.

Claims

WHAT IS CLAIMED IS:

1. A method for audio management in a wearable audio component, the method comprising: receiving a real-time audio signal by a reference microphone associated with an active noise control (ANC) system while the ANC system is in an ambient sound suppression mode, the ANC system integrated into a wearable audio component being worn by a first party; detecting, using a pre-trained inference model, whether the real-time audio signal includes attention-seeking audio spoken by a second party, the attention-seeking audio predetermined to indicate that the second party is seeking attention of the first party; and outputting a name invoked signal to the ANC system automatically in response to determining that the real-time audio signal includes the attention-seeking audio, the name invoked signal directing the ANC system automatically to switch from the ambient sound suppression mode to a conversation mode.

2. The method of claim 1, further comprising: detecting a conversation end trigger subsequent to the name invoked signal directing the ANC system automatically to switch from the ambient sound suppression mode to a conversation mode; outputting a conversation end signal responsive to detecting the conversation end trigger, wherein the conversation end signal directs the ANC system automatically to switch from the conversation mode to the ambient sound suppression mode.

3. The method of claim 1, wherein: the attention-seeking audio is predetermined to indicate that the second party is seeking attention of the first party based on a set of invocation names previously enrolled by the first party and previously converted by the pre-trained inference model into a set of reference embeddings; and the detecting comprises: converting the real-time audio signal to a real-time embedding by the pretrained inference model; and determining whether the real-time embedding matches any of the reference embeddings with at least a threshold confidence level.

4. The method of claim 3, wherein: the pre-trained inference model comprises a processor-executable name embedding model trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications by transfer learning from a foundation model, the foundation model being an artificial neural network trained to linguistically classify a speechaudio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers; and the converting the real-time audio signal to the real-time embedding is by the processor-executable name embedding model.

5. The method of claim 3, wherein: the pre-trained inference model comprises a processor-executable name embedding model trained, by transfer learning from a foundation model based on a corpus of real-world name audio samples, automatically to remove suprasegmental features from an input audio sample to generate a corresponding unified audio output, wherein the foundation model is an artificial neural network trained, based on a speech-audio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers, to generate class unified audio samples from class spoken audio samples by removing the suprasegmental features; and the converting the real-time audio signal to the real-time embedding is by the processor-executable name embedding model.

6. The method of claim 3, wherein: the pre-trained inference model comprises a processor-executable name embedding model comprising: a first embedding stage trained automatically to remove suprasegmental features from a corpus of real-world name audio samples to generate corresponding unified audio outputs; and a second embedding stage coupled with the first embedding stage and trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications; and the converting the real-time audio signal to the real-time embedding is by the processor-executable name embedding model comprising:converting the real-time audio signal, by the first embedding stage, into a unified real-time audio signal stripped of suprasegmental influence by the second party; and converting the unified real-time audio signal, by the second embedding stage, into the real-time embedding, the real-time embedding representing a name classification.

7. The method of claim 1, wherein the detecting comprises: converting the real-time audio signal to a real-time embedding by the pre-trained inference model; selecting a candidate name embedding based on determining, by the pre-trained inference model, that one or more of a stored plurality of reference name embeddings meets at least a predetermined similarity threshold with the real-time embedding, the stored plurality of reference name embeddings previously generated by the name embedding model based on a set of invocation names provided by a user during an enrollment procedure; and determining whether the real-time embedding and the candidate name embedding cannot be discriminated by the pre-trained inference model in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces, wherein the outputting the name invoked signal is only responsive to the determining that the real-time embedding and the candidate name embedding cannot be discriminated in excess of the predetermined discrimination threshold.

8. The method of claim 7, wherein the selecting the candidate name comprises: computing a similarity score between each of the reference name embeddings and the real-time embedding, the similarity score indicating a mathematical correlation; determining whether the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold; and outputting the candidate name embedding as the one of the reference name embeddings yielding a highest similarity score when it is determined that the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold.

9. The method of claim 7, wherein the determining whether the real-time embedding and the candidate name embedding cannot be discriminated comprises: transforming the real-time embedding and the candidate name embedding into a plurality of mathematical spaces;computing, for each mathematical space, a discrimination score indicating a likelihood of discriminating between the real-time embedding and the candidate name embedding in the mathematical space; and outputting the name invoked signal responsive to the discrimination score computed for at least one of the mathematical spaces falling below the predetermined discrimination threshold.

10. The method of claim 1, further comprising: receiving a set of invocation names from a user during an enrollment procedure; and generating a reference name embedding for each of the invocation names by the pre-trained inference model, wherein the attention-seeking audio is predetermined to indicate that the second party is seeking attention of the first party based on the reference embeddings.

11. The method of claim 10, wherein the generating the reference name embedding for each of the invocation names comprises: applying a plurality of augmentation transformations to each of the set of invocation names to generate an augmented set of invocation names; and generating a reference name embedding for each of the augmented set of invocation names by the pre-trained inference model.

12. An automatic attention handling system for integration into a wearable audio component, the automatic attention handling system comprising: an audio input to receive a real-time audio signal from a reference microphone associated with an active noise control (ANC) system of the wearable audio component; one or more processors coupled with the audio input; a non-transitory processor-readable storage medium having, stored thereon, a pretrained inference model and instructions, which, when executed, cause the one or more processors to perform steps comprising: receiving the real-time audio signal while the ANC system is in an ambient sound suppression mode and the wearable audio component is being worn by a first party; detecting, using the pre-trained inference model, whether the real-time audio signal includes attention-seeking audio spoken by a second party, the attention-seeking audio predetermined to indicate that the second party is seeking attention of the first party; andoutputting a name invoked signal to the ANC system automatically in response to determining that the real-time audio signal includes the attention-seeking audio, the name invoked signal directing the ANC system automatically to switch from the ambient sound suppression mode to a conversation mode.

13. The automatic attention handling system of claim 12, wherein the instructions further comprise: detecting a conversation end trigger subsequent to the name invoked signal directing the ANC system automatically to switch from the ambient sound suppression mode to a conversation mode; and outputting a conversation end signal responsive to detecting the conversation end trigger, wherein the conversation end signal directs the ANC system automatically to switch from the conversation mode to the ambient sound suppression mode.

14. The automatic attention handling system of claim 12, wherein: the attention-seeking audio is predetermined to indicate that the second party is seeking attention of the first party based on a set of invocation names previously enrolled by the first party and previously converted by the pre-trained inference model into a set of reference embeddings; and the detecting comprises: converting the real-time audio signal to a real-time embedding by the pretrained inference model; and determining whether the real-time embedding matches any of the reference embeddings with at least a threshold confidence level.

15. The automatic attention handling system of claim 14, wherein: the pre-trained inference model comprises a processor-executable linguistic name embedding (LNE) based name embedding model trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications by transfer learning from a foundation model, the foundation model being an artificial neural network trained to linguistically classify a speech-audio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers; and the converting the real-time audio signal to the real-time embedding is by the processor-executable LNE-based name embedding model.

16. The automatic attention handling system of claim 14, wherein: the pre-trained inference model comprises a processor-executable unified sound conversion (USC) based name embedding model trained, by transfer learning from a foundation model based on a corpus of real-world name audio samples, automatically to remove suprasegmental features from an input audio sample to generate a corresponding unified audio output; the foundation model is an artificial neural network trained, based on a speechaudio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers, to generate class unified audio samples from class spoken audio samples by removing the suprasegmental features; and the converting the real-time audio signal to the real-time embedding is by the processor-executable USC-based name embedding model.

17. The automatic attention handling system of claim 14, wherein:The pre-trained inference model comprises a processor-executable hybrid name embedding model comprising: a first embedding stage trained automatically to remove suprasegmental features from a corpus of real-world name audio samples to generate corresponding unified audio outputs; and a second embedding stage coupled with the first embedding stage and trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications; and the converting the real-time audio signal to the real-time embedding is by the processor-executable hybrid name embedding model comprising: converting the real-time audio signal, by the first embedding stage, into a unified real-time audio signal stripped of suprasegmental influence by the second party; and converting the unified real-time audio signal, by the second embedding stage, into the real-time embedding, the real-time embedding representing a name classification.

18. The automatic attention handling system of claim 12, wherein the detecting comprises: converting the real-time audio signal to a real-time embedding by the pre-trained inference model;selecting a candidate name embedding based on determining, by the pre-trained inference model, that one or more of a stored plurality of reference name embeddings meets at least a predetermined similarity threshold with the real-time embedding, the stored plurality of reference name embeddings previously generated by the name embedding model based on a set of invocation names provided by a user during an enrollment procedure; and determining whether the real-time embedding and the candidate name embedding cannot be discriminated by the pre-trained inference model in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces, wherein the outputting the name invoked signal is only responsive to the determining that the real-time embedding and the candidate name embedding cannot be discriminated in excess of the predetermined discrimination threshold.

19. The automatic attention handling system of claim 18, wherein the selecting the candidate name comprises: computing a similarity score between each of the reference name embeddings and the real-time embedding, the similarity score indicating a mathematical correlation; determining whether the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold; and outputting the candidate name embedding as the one of the reference name embeddings yielding a highest similarity score when it is determined that the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold.

20. The automatic attention handling system of claim 18, wherein the determining whether the real-time embedding and the candidate name embedding cannot be discriminated comprises: transforming the real-time embedding and the candidate name embedding into a plurality of mathematical spaces; computing, for each mathematical space, a discrimination score indicating a likelihood of discriminating between the real-time embedding and the candidate name embedding in the mathematical space; and outputting the name invoked signal responsive to the discrimination score computed for at least one of the mathematical spaces falling below the predetermined discrimination threshold.

Citation Information

Patent Citations

  • Automated control of noise reduction or noise masking

    US20200329297A1

  • Context aware hearing optimization engine

    US20200380979A1

  • Electronic device including speaker and microphone and method for operating the same

    US20220261218A1

Cited By

  • Name-detection based attention handling in active noise control systems based on automated acoustic segmentation

    WO2025198624A1

  • Speech signal repair and enhancement using an integrated network based on progressive learning

    WO2025250160A1

  • Architecture and network topology for acoustic segmentation of speech

    WO2026010638A1

  • Classification-based selective pass-through architecture for wearable audio components with active noise control

    WO2026075674A1

  • Known-speaker-based and conditional selective pass-through for wearable audio components with active noise control

    WO2026075675A1