Architecture and network topology for universal sound conversion

Machine learning networks with convolutional encoders and axial self-attention blocks enhance active noise control systems by automatically switching to conversation mode upon detecting attention-seeking audio, addressing the challenge of engaging users in conversations while maintaining effective noise cancellation.

WO2026039056A1PCT designated stage Publication Date: 2026-02-19GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/051040
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2024-10-11
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Conventional active noise control systems in wearable audio components struggle to effectively distinguish ambient sound intended for the user, making it difficult for attention seekers to engage the user in conversations, requiring manual intervention to disable ANC and remove the device.

Method used

Implementing machine learning network architectures with convolutional encoders, axial self-attention blocks, and deconvolutional decoders to convert spoken audio into unified sounds, automatically switching ANC to conversation mode upon detecting attention-seeking audio.

Benefits of technology

Enables users to engage in conversations without removing their wearable audio components, improving user comfort and awareness while maintaining effective noise cancellation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024051040_19022026_PF_FP_ABST
    Figure US2024051040_19022026_PF_FP_ABST
Patent Text Reader

Abstract

Machine learning network topologies, and training systems and methods therefor, are described for implementing unified sound conversion (USC), such as for automated name detection. Such automated name detection can support automated attention handling in wearable audio components with active noise control (ANC) to suppress ambient sound. Embodiments of USC network topologies include a feature generator comprising a convolutional encoder, an axial self-attention (ASA) block, and a deconvolutional decoder. A MEL converter generates input filter bank energies (FBEs) from an audio sample of a spoken word. The feature generator is trained to estimate output FBEs from the input FBEs, such that the output FBEs represent the linguistic information of the spoken word absent any speaker-specific suprasegmental features.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.094021-1470322 ARCHITECTURE AND NETWORK TOPOLOGY FOR UNIVERSAL SOUND CONVERSION CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of and priority to Indian Provisional Application No. 202441062106, filed on August 16, 2024, and titled “ARCHITECTURE AND NETWORK TOPOLOGY FOR UNIVERSAL SOUND CONVERSION,” the content of which is herein incorporated by reference in its entirety for all purposes. BACKGROUND

[0002] Active noise control (ANC) is a common feature of headsets and earbuds. It operates by generating an anti-noise signal via a speaker that is approximately equal in magnitude, but opposite in phase to the ambient sound (e.g., ambient noise and other sounds in the vicinity). The ambient sound and anti-noise signal cancel each other acoustically, allowing the user to hear only a desired audio signal. Typically, signal processing in ANC includes two paths: an ambient sound signal from a reference microphone is taken as the input of a feed-forward ANC filter (FFANC); and an error microphone signal is taken as the input of a feedback ANC filter (FBANC).

[0003] When a user is listening to music or other desired audio through a wearable audio component (e.g., earbuds or on-ear headphones), ANC works to suppress any ambient sound. However, in some instances, ambient sound that is intended for the user can be very important to the user’s connectivity with others. For example, while a user is listening to music with ANC, it may be very difficult for an attention seeker to get the user’s attention, such as to alert the user of something important and / or to engage the user in a desired conversation. Conventionally, in such instances, the attention seeker gets the user’s attention by tapping the user on the shoulder, gesturing in front of the user’s face, or the like; and the user can then manually disable the ANC, pause the desired audio, and / or remove the wearable audio component. SUMMARY

[0004] Architectures and network topologies are described for machine learning networks universal sound conversion (USC) machine learning networks. For example, in the context of a user wearing a wearable audio component (e.g., in-ear headphones, on-ear headphones, etc.) and having active noise control (ANC) turned on to suppress ambient sound, automated attention handling techniques can be used to detecting when an attention seeker is trying audibly to get the attention of the user. The ANC can be automatically switched into a conversation mode in response to such detection. One technique for automated attention handling is based on universal sound conversion (USC), by which spoken audio of a class (i.e., a word) is converted into a unifiedsound that mimics what would be generated by a speech synthesizer for that class and / or by a normalized speaker (i.e., the unified sound is stripped of the speaker’s influence on suprasegmental features of the audio).

[0005] Embodiments are described herein for machine learning network architectures and topologies, training systems and methods therefor, to implement such universal sound conversion. Embodiments of such USC network topologies include a feature generator comprising a convolutional encoder, an axial self-attention (ASA) block, and a deconvolutional decoder. A MEL converter generates input filter bank energies (FBEs) from an audio sample of a spoken word. The feature generator is trained to estimate output FBEs from the input FBEs, such that the output FBEs represent the linguistic information of the spoken word absent any speaker-specific suprasegmental features. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] A further understanding of the nature and advantages of various embodiments may be realized by reference to the following figures. In the appended figures, similar components or features may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.

[0007] FIG.1 shows an audio management system for integration in a wearable audio component (WAC), according to embodiments described herein.

[0008] FIG.2 shows a conceptual circuit block diagram of a partial audio management system environment with an automated attention handling system (AHS).

[0009] FIGS.3A and 3B show a wearable audio environment including a pair of WACs.

[0010] FIG.4 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein.

[0011] FIG.5 shows an embodiment of a universal sound conversion (USC) machine learning network, according to embodiments described herein.

[0012] FIG.6 shows an illustrative USC machine learning network that can be a more detailed implementation of the USC machine learning network of FIG.5.

[0013] FIGS.7A and 7B show architectures for illustrative implementations of an encoder block and a decoder block, respectively.

[0014] FIG.8 shows a block diagram of an illustrative architecture for an axial self-attention (ASA) block, according to some embodiments described herein.

[0015] FIG.9 shows an architecture for an illustrative implementation of a convolution block of the ASA block.

[0016] FIG.10 shows an illustrative training environment for training a USC machine learning model, such as the one described in FIGS.5 – 9.

[0017] FIG.11 shows a flow diagram of an illustrative method for training an artificial neural network (ANN) for universal sound conversion (USC), according to embodiments described herein.

[0018] FIGS.12A and 12B show a block diagrams of illustrative uses of the name embedding model to generate the deep image.

[0019] FIG.13 shows several example screenshots from an example enrollment application running on a user device.

[0020] FIG.14 shows a flow diagram of an illustrative method for audio management that includes automated attention handling in a WAC, according to embodiments described herein.

[0021] FIG.15 shows a flow diagram of an illustrative method for an enrollment phase.

[0022] FIG.16 shows a flow diagram of an illustrative method for conversation end detection.

[0023] FIGS.17A and 17B show block diagrams of a training environment for training a foundation model to support unified sound conversion (USC) for automated name-detection-based attention handling, according to some embodiments described herein.

[0024] FIG.18 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein.

[0025] FIG.19 shows a flow diagram of an illustrative method for automated audio management based on name detection, according to USC-based embodiments described herein.

[0026] FIG.20 shows a flow diagram of an illustrative method for training a USC inference model.

[0027] FIG.21 shows a flow diagram of an illustrative method for using the trained USC inference model during an operation time.

[0028] FIG.22 provides a schematic illustration of an illustrative computational system that can implement various system components and / or perform various steps of methods provided by various embodiments. DETAILED DESCRIPTION

[0029] When a user is listening to music or other desired audio through a wearable audio component (e.g., earbuds or on-ear headphones), active noise control (ANC) works to suppress any ambient sound. However, in some instances, ambient sound that is intended for the user can be very important to the user’s connectivity with others. For example, although the user desired to suppress undesirable ambient sound, the user may still desire to be able on occasion to enter into desired conversations. In general, two types of desired conversation can be considered: first-party- initiated; and second-party-initiated.

[0030] In first-party-initiated conversations, the user desires to start a conversation and may begin by trying to get someone’s attention. In such cases, some conventional ANC systems are adapted to detect that the user has begun speaking (e.g., by detecting the user’s speech via a beamforming microphone directed to the user’s mouth, accelerometer, or combination thereof), and the ANC system can turn off, switch to transparency mode, pause audio playback, etc. in response to detecting the user’s speaking. Because it tends to be relatively easy for the ANC system to distinguish the user’s own speech from ambient sound, such approaches tend to be effective for first-party-initiated conversations.

[0031] In second-party-initiated conversations, however, a second-party attention seeker is trying to get the user’s attention, and the attention seeker’s voice may be difficult to distinguish from other ambient sound. For example, while a user is listening to music with ANC, it may be very difficult for an attention seeker to get the user’s attention, such as to alert the user of something important and / or to engage the user in a desired conversation. Conventionally, in such instances, the attention seeker gets the user’s attention by tapping the user on the shoulder, gesturing in front of the user’s face, or the like. Before the user can participate in the conversation, the user conventionally must notice the interruption and then manually disable the ANC, pause the desired audio, remove the wearable audio component, etc.

[0032] Indeed, many users of wearable audio components enjoy the feeling of being in their “bubble” and the ability to focus on their media that comes with effective ANC. However, as ANC continues to improve, the same users often feel increasingly unaware, not present, andfearful about missing out. Embodiments described herein seek to provide users with the ability to better stay aware and engage in desired conversations, while being able to continue wearing their wearable audio components and otherwise to take advantage of ANC. This can provide several benefits, including helping to improve user comfort and ear health.

[0033] Embodiments described herein are concerned with second-party-initiated conversations. As used herein, the term “user” refers to a wearer of a wearable audio component (i.e., the first party). The term “attention seeker” is used herein generally to refer to any ambient party trying to get the user’s attention while the user is wearing the wearable audio component (and presumably is listening to desired audio with ANC turned on). Typically, the attention seeker is a person. However, the attention seeker can also be a computational platform with a deterministic manner of seeking the user’s attention, such as a smart speaker programmed to call out the user’s name. The term “wearable audio component,” or “WAC” is used herein to generally refer to earbuds, on-ear headphones, over-ear headphones, or any type of wearable audio output device that includes ANC. The term “desired audio” is used herein to generally refer to any recorded or streaming audio signal that is being played to the user through the WAC, such as music, an audiobook, a podcast, a radio broadcast, a live event broadcast, etc. The term “ambient sound,” or “ambient audio” is used herein to generally refer to any audio in the vicinity of the WAC, other than the desired audio. It is generally the goal of the ANC system to suppress as much of the ambient sound as possible. Audio originating from an attention seeker while a user’s ANC system is active is part of the ambient sound.

[0034] FIG.1 shows an audio management system 100 for integration in a wearable audio component (WAC), according to embodiments described herein. As illustrated, the audio management system 100 can include an active noise control (ANC) system 140, an attention handling system (AHS) 150, and an audio processing system 160. In general, the purpose of the WAC is to deliver desired audio 165 to a user’s ear or ears via one or more ear speakers, such as speaker 105. Embodiments of the audio processing system 160 are designed to process the desired audio 165 for output to the user. For example, the audio processing system 160 can include amplifiers, filters, and / or other audio components; and / or any other suitable components for receiving, processing, and / or outputting the desired audio 165.

[0035] Typically, while listening to the desired audio 165, the user is also in presence of ambient audio 155. When in its ambient sound suppression mode, the ANC system 140 seeks to suppress as much of the ambient audio 155 as possible to enhance the user’s experience of listening to the desired audio 165. As illustrated, the ANC system 140 includes a feed-forward ANC (FFANC)filter 120, a feedback ANC (FBANC) filter 125, a summer 130, and an ANC output control block 135. The ANC system 140 is also coupled with the speaker 105 and at least a reference microphone 110 and an error microphone 115. Embodiments of the speaker 105 generally convert an electrical audio signal into sound waves that are delivered to the ear of the wearer of the wearable audio component. Embodiments of the reference microphone 110 can be an omnidirectional microphone typically integrated with an outer casing of the wearable audio component. The reference microphone 110 generally captures at least the ambient audio 155 around the WAC, which is delivered as a reference audio signal (illustrated as x(n)) to the FFANC filter 120. Embodiments of the error microphone 115 are typically integrated with the inner casing of the wearable audio component to be positioned inside the ear canal or very close to it when the wearable audio component is being worn. The error microphone 115 captures the audio that reaches the eardrum, which includes the desired audio signal and any remaining ambient sound after suppression. The error microphone 115 outputs an error signal (illustrated as e(n)) to the FBANC filter 125.

[0036] The illustrated ANC system 100 includes a feed-forward noise control path and a feedback noise control path. The feed-forward noise control path includes the FFANC filter 120, which is a digital or analog filter designed to process the audio signal from the reference microphone 110. The FFANC filter 120 applies a specific frequency response to x(n) to adaptively cancel out noise. The specific frequency response is produced by continuously adjusting coefficients of the FFANC filter 120 to minimize the difference between the desired audio signal and the reference signal. The output of the FFANC filter 120 is illustrated as(^^^^). The feedback noise control path includes the FBANC filter 125, which is a digital or analog filter designed to process the audio signal from the error microphone 115. The FBANC filter 125 applies a specific frequency response to e(n), and continuously adjusts coefficients of the FBANC filter 125 to minimize the difference between the desired audio signal and remaining ambient sound in the signal that reaches the eardrum. The output of the FFANC filter 120 is illustrated as ^^^^2(^^^^). In general, both the FFANC filter 120 and the FBANC filter 125 can adapt their respective filters (e.g., their coefficients) in real-time to a changing audio environment. For example, filter coefficients are iteratively adjusted using least mean squares (LMS), normalized LMS (NLMS), and / or other suitable adaptation algorithms.

[0037] Embodiments of the summer 130 combine the filtered output signals from the FFANC filter 120 and the FBANC filter 125. For example, the summer 130 calculates a sum of these signals. If tuned properly, the output of the summer 130 is an “anti-noise” signal that closely represents the ambient sound at opposite polarity. Embodiments of the ANC output control block135 control how and / or whether the anti-noise signal is output by ANC system 140. In some implementations, the ANC output control block 135 includes an amplifier to provide a controllable amount of gain (G) to the signal at the output of the summer 130, resulting in an output signal,^^^^(^^^^) = ^^^^(^^^^1(^^^^) + ^^^^2(^^^^)). In effect, the ANC gain block 135 adjusts the overall amplitude (i.e.,corresponding to volume) of the combined filtered signal at the output of the summer 130. The output signal is sent to the speaker 105. In some implementations, as illustrated, the desired audio 165 can also be mixed in (e.g., by mixer 145) prior to sending the output to the speaker 105, such that what reaches the eardrum is almost entirely the desired audio signal with minimal ambient sound. Alternatively, the desired audio 165 is mixed into the output signal at the summer 130, such that the output of the ANC system 140 is an audio signal that is mostly the desired audio 165 with minimal residual ambient audio 155.

[0038] Embodiments of the ANC output control block 135 control the operating mode of the ANC system 140. For example, as described herein, the ANC system 140 can operate selectively in at least an active mode (i.e., an ambient sound suppression mode) or a conversation mode. Some implementations of the conversation mode correspond to an inactive mode (i.e., the ANC system 140 is turned off) or a transparency mode. Other implementations of the conversation mode are configured to pass through conversationally relevant audio from the ambient audio 155, while continuing to perform ANC functions to suppress other portions of the ambient audio 155. In some such implementations, a bandpass or notch filter is used to segregate out a range of frequencies typical for human speech and to treat the segregated audio as conversationally relevant audio. As one example, a filter can pass through portions of the ambient audio 155 only in the range of 75 to 300 Hertz and to suppress higher and lower frequency components of the ambient audio 155; thereby continuing to filter out white noise and other portions of ambient audio 155 that can interfere with a user’s ability to hear the passed-through conversationally relevant audio. Similarly, some implementations continue to pass through some desired audio 165 (e.g., at a reduced volume) while in conversation mode.

[0039] As described herein, embodiments of the AHS system 150 seek to detect when a second- party attention seeker is trying to get the attention of a user while the user is wearing the WAC and is listening to desired audio 165 with the ANC system 140 in the active mode. In particular, the AHS system 150 listens for presence of attention seeking (AS) audio 157 within the ambient audio 155. Some embodiments of the AHS system 150 described herein are implemented as a linguistic name-embedded attention handling system (LNE-AHS). Other embodiments of the AHS system 150 described herein are implemented as a universal sound conversion attention handling system (USC-AHS). Other embodiments of the AHS system 150 described herein are implemented as ahybrid universal LNE attention handling system (ULNE-AHS). Any of the types of AHS system 150 described herein can be configured specifically to listen for AS audio 157 corresponding to a previously enrolled invocation name (e.g., a name of the user). When the AS audio 157 is detected by the AHS system 150, the AHS system 150 automatically directs the ANC system 140 to switch from the active mode to the conversation mode. In some embodiments, when the AS audio 157 is detected, the AHS system 150 also directs the audio processing system 160 to enter a conversation enhancement mode. As described above with reference to the ANC system 140, the conversation enhancement mode as implemented by the audio processing system 160 can include segregating conversationally relevant audio from the ambient audio 155, adapting equalization of passed through audio to enhance speech, muting or reducing the volume of playback of the desired audio 165, pausing playback of the desired audio 165, etc. Typically, in response to the AS audio 157, the user will begin to engage in a conversation with the attention seeker. Such a conversation can involve the user speaking, and embodiments of the conversation mode of the ANC system 140 and / or the conversation enhancement mode of the audio processing system 160 can include using techniques to help ensure that the user’s own speech is not fed back in a manner that results in an apparent echo, feedback noise, or the like. For example, the user’s own speech may be captured by a separate beamforming microphone as a user speech audio stream, while ambient audio 155 is being received by the reference microphone 110. The user speech audio stream can be subtracted from the ambient audio 155 prior to passing the signal through other blocks of the system, so that the fed-back audio stream includes only ambient audio other than the user’s own speech.

[0040] Some embodiments of the AHS system 150, after having detected AS audio 157 and directing the ANC system 140 into conversation mode, can further detect when the conversation ends. Such embodiments of the AHS system 150 can automatically direct the ANC system 140 to return to the active mode, accordingly. As part of returning to the active mode, some such embodiments also return settings (e.g., in the ANC system 140 and / or the audio processing system 160) to those appropriate for listening to the desired audio 165 and suppressing all of the ambient audio 155 (e.g., all frequencies of the ambient audio 155).

[0041] FIG.2 shows a conceptual circuit block diagram of a partial audio management system environment 200 with an automated attention handling system (AHS) 150. The illustrated environment 200 can be an illustrative portion of the audio management system 100 of FIG.1, and the AHS 150 can be an illustrative implementation of the AHS 150 of FIG.1. As illustrated, the AHS 150 includes an attention seeking (AS) trigger detection block 210 and a conversation end detection block 220. Some implementations further include a conversation enhancement block 230. The AHS 150 is illustrated in context of a desired audio 165 stream, a reference microphone110 that receives ambient audio 155 and outputs an ambient audio stream, and a speaker 105. For the sake of simplicity, the AHS 150 is illustrated without other components of the audio management system 100 of FIG.1, such as without the ANC system 140 and the audio processing system 160.

[0042] The role of the AHS 150 can be generally described as to toggle the audio environment between an active mode and a conversation mode based on whether a desired conversation is detected, as represented by a switch network 215. In the active mode, the user is listening to the desired audio 165 via the speaker 105, and the ANC system 140 (not shown) is suppressing as much of the ambient audio 155 as possible. This is conceptually represented by the switches of the switch network 215 being in the solid-line position, whereby the desired audio 165 passes through to the speaker 105 and the ambient audio 155 does not. When attention seeking audio (i.e., audio associated with getting the user’s attention) is detected by the AS trigger detection block 210, the AHS system 150 switches the switch network 215 to the dashed-line position, whereby the ambient audio 155 passes through to the speaker 105 and the desired audio 165 does not. In some embodiments, while in the conversation mode, the passed-through ambient audio 155 (e.g., either all of the ambient audio 155, or a conversationally relevant portion of the ambient audio 155) is passed through the conversation enhancement block 230 in line with the speaker 105. As described above, the conversation enhancement block 230 can use various techniques to enhance conversationally relevant portions of the ambient audio 155. Further, as described above, the conversation enhancement block 230 can be implemented in the AHS system 150, in the ANC system 140, in the audio processing system 160, and / or in any suitable location.

[0043] When the end of the conversation is detected by the conversation end detection block 220, the AHS system 150 switches the switch network 215 back to the solid-line position, whereby the desired audio 165 again passes through to the speaker 105 and the ambient audio 155 again does not. In some embodiments, as described above, the end of the conversation is detected based on detecting the user’s own speech, such as detecting that the user is no longer speaking for some time, or that the user has issued an audio cue (e.g., “resume ANC”). This user speech can be detected via the reference microphone 110 as part of the ambient audio 155, or detected through a separate microphone 240, such as a beamforming microphone with its beam directed toward the user’s mouth. Additionally, or alternatively, some embodiments of the conversation end detection block 220 detect the end of a conversation based on detecting user interfacing with an interface element, such as detecting that the user pressed a play / pause button 245 on the WAC 210, or the like.

[0044] Although the AHS 150 is illustrated as directly coupled with the microphones, the desired audio 165 is directly coupled with the speaker 105 in active mode, the ambient audio 155 is directly coupled with the speaker 105 in conversation mode, etc., some or all of such connections can be through other components that are not shown in FIG.2. For example, embodiments herein assume that the ambient audio 155 is passing through the ANC system 140, that the desired audio 165 is mixed with an anti-noise signal in active mode, etc.

[0045] As described herein, embodiments of the audio management system 100 are configured for integration in any suitable WAC. FIGS.3A and 3B show a wearable audio environment 300 including a pair of WACs 310. Each WAC 310 is illustrated as an earbud. Alternatively, the pair of earbuds can be considered as a single WAC. In other embodiments, the WAC 310 can be implemented as over-ear headphones, or any other suitable wearable audio component that incorporates ANC. In the illustrated embodiments, each WAC (i.e., each earbud) has a respective instance of an audio management system 100, such as the audio management system 100 of FIG. 1, and each instance of the audio management system 100 includes a respective instance of at least an ANC system 140 and an AHS system 150. Though not explicitly shown, each WAC 310 also has, integrated therein, an instance of the speaker 105, the reference microphone 110, the error microphone 115, one or more processors, and non-transitory processor-readable storage. Some implementations of the WAC 310 include additional components, such as instances of the audio processing system 160, one or more additional microphones (e.g., a beamforming microphone), one or more additional speakers, interface controls (e.g., one or more buttons), one or more power sources (e.g., a rechargeable battery), one or more ports (e.g., physical ports for charging and / or wired communication, logical ports for wireless charging and / or wireless communication), one or more antennas, etc.

[0046] In some embodiments, the one or more processors integrated in the WAC 310 implement components of the respective audio management system 100 instance. For example, a non- transitory processor-readable medium integrated therein has processor-executable instructions stored thereon, which, when executed, cause the set of processors to implement at least features of the respective ANC system 140 and / or AHS system 150 instances. As described herein, embodiments of the AHS system 150 include one or more types of artificial neural networks, corresponding trained network models, or the like. In some embodiment, such networks and / or models are implemented using specialized hardware, such as neuromorphic chips. In other embodiments, such networks and / or models are implemented by using processor-readable instructions to reconfigure general-purpose computing hardware (e.g., a central processing unit, CPU), specialized AI accelerators.

[0047] Turning specifically to FIG.3A, a first type of wearable audio environment 300a is shown in which one or both WACs 310 is in communication with a cloud computing environment (“cloud”) 340. For example, the cloud 340 includes a server, or several distributed servers, accessible via the Internet. Though the WAC 310 is shown as directly in communication with the cloud 340, such a connection can be facilitated by any suitable intermediary devices, such as routers, hubs, etc. As described further herein, automated attention handling features described herein rely on generation of an inference model that includes several neural networks and / or models. In the illustrated embodiments of FIG.3A, the inference model is generated by the local computation environment of the WAC 310 and / or based on information ported to the WAC 310 from the cloud 340.

[0048] As described herein, some embodiments involve enrollment of invocation names for use in name detection. In the embodiments of FIG.3A, an enrollment application 330 is downloaded to the local computational environment of the WAC 310, and the enrollment application 330 is used for such enrollment. The enrollment application 330 can facilitate generation of the inference model, based on a foundation model 350. As illustrated, the foundation model 350 can be stored in the cloud 340. In some embodiments, the foundation model 350 is accessible to the WAC 310 (e.g., to the enrollment application 330) via the cloud 340. The inference model can include some or all of a name embedding model, a deep image, a relation network, and a false rejection network. Each is described more fully below.

[0049] Turning to FIG.3B, some embodiments of the wearable audio environment 300b further include a user computational device 320 separate from the WAC 310. For example, the user computational device 320 can be a smartphone, laptop computer, tablet computer, smart watch, portable audio player, or any other suitable device that is separate from the WAC 310 and includes its own one or more processors and its own one or more non-transitory storage media for storing processor-readable instructions. The user computational device 320 can be in communication with each WAC 310 via any suitable wired and / or wireless communication link, such as via an audio cable (e.g., via a 3.5-millimeter or 1 / 4-inch analog audio jack), a universal wired connection (e.g., universal serial bus (USB)), a short-range universal wireless connection (e.g., Bluetooth, short- range radiofrequency, near field communication (NFC)), an optical connection (e.g., infrared), a proprietary connector, a multi-pin connector, an intermediary component or platform (e.g., a docking station or dongle), etc.

[0050] Similar to the embodiments of FIG.3A, the embodiments of FIG.3B can include an enrollment application 330, communications with the cloud 340, use of a foundation model 350,etc. As illustrated in FIG.3B, these features can be facilitated via the user computational device 320 (rather than directly by the WAC 310). For example (as described more fully below), when a user first registers the WAC 310 (e.g., first attempts to pair the earbuds with the user computational device 320), the user computational device 320 automatically accesses and / or downloads the enrollment application 330, or prompts the user to access and / or download the enrollment application 330. The enrollment application 330 can be downloaded from the cloud 340, or from any other suitable environment. The enrollment application 330 can then access the foundation model 350 via the cloud 340 and can use the foundation model 350 to generate an inference model for automated attention handling. For example, to facilitate automated attention handling, the inference model includes some or all of a name embedding model, a deep image, a relation network, and a false rejection network. The generated inference model can then be ported to the WAC 310. For example, the inference model can be ported from the user computational device 320 to both earbuds; ported from the user computational device 320 to a master earbud, and from the master earbud to a slave earbud; etc.

[0051] Embodiments generally build an inference model for storing at the WAC 310 to enable the WAC 310 to subsequently use the inference model to perform automated attention handling, as described herein. FIG.4 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block 400 and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein. The AS trigger detection block 400 can be an implementation of the AS trigger detection block 210 of FIG.2. It is assumed that the AS trigger detection block 400 is implemented in the context of a WAC 310.

[0052] Embodiments can include an enrollment stage 410, an identification stage 430, and a verification stage 440. In general, the enrollment stage 410 occurs outside of normal operation of the WAC 310, such as when a user first sets up and / or uses the WAC 310, when the user first registers the WAC 310, when the user first configures the WAC 310 for automated attention handling, etc. The identification stage 430 and the verification stage 440 are real-time blocks that occur during normal operation of the WAC 310 to facilitate real-time automated attention handling. Embodiments generally use the enrollment stage 410 to obtain any information needed to set up an inference model for storing at the WAC 310 to enable the WAC 310 to perform automated attention handling features during normal operation. As shown, the inference model includes at least an embedding model 415, a relation network 435, and a false rejection network 445.

[0053] In the enrollment stage 410, the embedding model 415 can be used to convert an enrollment audio stream 405 into a set of reference embeddings based on a foundation model 350. The foundation model 350 can be stored in and accessed via the cloud 340 (e.g., and / or any other suitable communication network). The reference embeddings can be stored as a deep image 420. The deep image 420 can also be considered as part of the inference model.

[0054] During normal operation, in the identification stage 430, a real-time (RT) audio stream 407 is received. The embedding model 415 (i.e., the same embedding model 415 generated in the enrollment stage 410) is used to generate a RT embedding from the received RT audio stream 407. The relation network 435 can then compare the RT embedding with each of the reference embeddings in the deep image 420 to determine if there is a match. For example, the relation network 435 is configured to compute a similarity score for each comparison (e.g., corresponding to a mathematical correlation, or the like). The relation network 435 can determine if any similarity scores meet or exceed a predetermined threshold; if so, the reference embedding with the highest similarity score can be selected as a candidate matching embedding.

[0055] In the verification stage, the false rejection network 445 can then confirm the match by transforming the RT embedding and the candidate matching embedding into different mathematical spaces (e.g., domains) and determining whether the embeddings can be reliably discriminated from each other in any of the mathematical spaces. For example, the deep image 420 is configured to compute a discrimination score for each mathematical space. The deep image 420 can determine if any discrimination scores meet or exceed a predetermined threshold; if so, the match determined by the relation network 435 is determined to be a false match and is ignored. If the embeddings cannot be sufficiently discriminated in any mathematical spaces, the false rejection network 445 can output a signal that attention seeking audio has been detected.

[0056] As described herein, embodiments of the attention seeker (AS) trigger detection block 400 are configured to detect an invocation name as part of an attention handling system. The embedding model 415 is a name embedding model (e.g., a processor-executable name embedding model) that generates an output embedding from an audio sample. As used herein, the terms “audio sample” or an “audio signal” are used interchangeably in the context of an input to a component of the inference model; such an “audio sample” or an “audio signal” can be represented in any suitable manner, such as by any suitable number of digital samples. For example, reference to an input as an “audio sample” means an audio signal of a duration, or a sampled duration of an audio signal, at a sampling rate resulting in a large number of digital samples (e.g., one second of audio sampled at 16 kHz to yield 16,000 samples). As described above, the audio sample can befrom an enrollment audio stream 405 in an enrollment stage 410, and the audio sample can be from a RT audio stream 407 during normal operation. The name embedding model 415 is trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications. The deep image 420 (e.g., a processor-readable deep image) includes reference name embeddings generated by the name embedding model 415 based on a set of invocation names provided by a user during an enrollment procedure. The relation network 435 (e.g., a processor-executable relation network) is coupled with the deep image 420 and the name embedding model 415 to output a candidate name embedding responsive to determining that one of the reference name embeddings has a highest similarity with a RT embedding generated by the name embedding model 415. The RT embedding can be from a RT audio stream 407 received from a reference microphone associated with an ANC system of the WAC 310. The false rejection network 445 (e.g., a processor-executable false rejection network) is coupled with the relation network 435 to output a name invoked signal 450 responsive to determining that the real-time embedding and the candidate name embedding cannot be reliably discriminated. As described herein, the name invoked signal 450 can direct an ANC system of the WAC 310 automatically to enter a conversation mode.

[0057] As described above, the interference model used for automated attention handling is based on a foundation model 350. Embodiments of the foundation model 350 are generated from a large speech-audio corpus of diversified words spoken by diversified speakers. “Diversified words” refers herein to the speech-audio corpus including a wide variety of at least phonemes and linguistic information. “Diversified speakers” refers to the speech-audio corpus representing a wide variety of at least accents and prosody. The speakers can be further diversified with respect to age, gender, geography, etc. For example, the speech-audio corpus includes tens of thousands of words (i.e., classifications) spoken multiple times (e.g., 10 – 15 times) by hundreds of speakers from around the world.

[0058] The term “suprasegmental” is used herein as an umbrella term to encompass properties of a speaker’s influence when speaking words, such as accent, prosody, intonation, rhythm, and other non-segmental aspects of speech. Such suprasegmental features can be contrasted with segmental features pertaining to individual speech sounds or segments, such as vowels and consonants, and can span multiple segments or an entire utterance. Examples of suprasegmental features of an utterance (e.g., a word, name, etc.) can include accent (including accent-influenced variations in pitch, loudness, and duration), prosody (including rhythm, intonation, and melody of speech), intonation (i.e., the rise and fall of pitch in speech), rhythm and / or rate (e.g., the temporal patterns of speech, such as duration and timing of sounds, syllables, and pauses), and stress (e.g., emphasisplaced on a particular syllable). For example, a large speech-audio corpus of diversified words spoken by diversified speakers may include hundreds or thousands of samples of a particular word being spoken with wide suprasegmental variance over the samples.

[0059] Training of the foundation model 350 can begin with an auto-encoder architecture, which is a type of neural network architecture designed to learn compact representations of data, such as so-called “latent features.” In the context of embodiments described herein, the auto-encoder architecture is used to extract meaningful features from raw audio data to be used for automatic speech recognition (ASR). In general, the auto-encoder architecture includes three high-level stages: an encoder, a bottleneck layer, and a decoder.

[0060] Embodiments described herein provide novel network topologies and architectures for implementing a machine learning network for use with unified sound conversion (USC) based approaches to automated attention handling. FIG.5 shows an embodiment of a universal sound conversion (USC) machine learning network 500, according to embodiments described herein. The network 500 includes a Mel converter 510, an encoder 520, an axial self-attention (ASA) block 530, and a decoder 540.

[0061] The Mel converter 510 converts input audio to a set of Mel filter bank energies 515. In one implementation, the set of Mel filter bank energies 515 is forty (40) Mel filter bank energies. Mel filter bank energies 515 are derived by converting an input audio signal into the frequency domain (e.g., using a Fourier Transform, typically a Fast Fourier Transform (FFT)), thereby converting the time-domain signal into a spectrum of frequencies. The frequency spectrum is segmented into bins corresponding to the frequencies captured by each FFT point. Typically, such frequency bins are not uniformly distributed in terms of human auditory perception. To mimic the non-linear human ear perception of sound, the frequencies can be converted into the Mel scale, which is a perceptual scale that relates perceived frequency (or pitch) of a pure tone to its actual measured frequency. Frequencies are converted to the Mel scale using a formula that approximates the human ear's response to different frequencies, typically using a logarithmic transformation that emphasizes the relevance of lower frequencies more than higher ones. Once frequencies are mapped onto the Mel scale, a set of triangular filters (known as a Mel filter bank) is applied to the spectrum. Each Mel filter in the bank covers a specific range of frequencies and is designed to pass a certain band of frequencies while attenuating others based on the Mel scale. The Mel filters can be overlapping and can cover the entire frequency spectrum obtained from the FFT. The energy in each Mel filter can be calculated by taking the sum of the squared values of the FFT coefficients that fall within the filter's range, which corresponds to the power spectraldensity in that frequency band. The resulting values are referred to as the Mel filter bank energies 515. Thus, the Mel filter bank energies provide a compact representation of the original sound spectrum in a manner that also aligns with human auditory perception.

[0062] The Mel filter bank energies 515 are input to the encoder 520. The encoder 520 converts the relatively high-dimensionality input, via a multi-layer network, to progressively reduce the dimensionality of the data using transformations. Each layer of transformations seeks to extract increasingly abstract and higher-level features from the input data. The output of the encoder can be referred to generally as an encoded representation 525 of the input data (i.e., of the Mel filter bank energies 515).

[0063] The ASA block 530 implements the bottleneck layer. In general, the bottleneck layer is so called because it typically includes significantly lower dimensionality than the layers of the encoder 520 before it and or the decoder 540 after it. This reduced dimensionality effectively forces the network to learn a compressed and informative representation of the input data, referred to herein as a bottleneck embedding 535. As illustrated, the ASA block 530 uses axial self- attention to perform the bottleneck layer features. Self-attention is a mechanism in machine learning that allows models to weigh the importance of different components of the input data in different ways. In general, self-attention allows inputs to interact with each other (“attend” to each other) and to determine how much focus should be placed on other parts of the input for generating a specific output. For a sequence of input values, the self-attention mechanism calculates a weight matrix for which each element determines how much each part of the input should contribute to every other part.

[0064] As described more fully below, self-attention computes three vectors: a query vector, a key vector, and a value vector. These vectors are derived by multiplying the input embeddings by the corresponding trainable weights. Attention scores can be computed by taking the dot product of the query vector of each input with the key vector of all other inputs, resulting in a score that represents how much each element of the sequence should attend to every other element. The scores can then be scaled down (e.g., by the square root of the dimension of the key vectors) to stabilize the gradients during training. The scores can then be normalized (e.g., by applying a softmax function to convert them into a probability distribution) and can be multiplied by the value vectors, which then sums to the output of the self-attention layer for each input. Axial self- attention is a variant of the self-attention mechanism designed to efficiently handle multi- dimensional data. For example, a non-axial approach can treat input data as a sequence, which can become computationally expensive as the sequence length increases. Axial self-attentiondecomposes the self-attention operation across different dimensions of the data. For example, attention can be computed across each dimension separately. This can significantly reduce computational complexity.

[0065] The axial self-attention described herein includes both inter-axial and intra-axial self- attention. Intra-axial self-attention refers to performing self-attention operations within a same axis of the data, so that the self-attention mechanism computes relationships and interactions among elements that lie on the same row or column (or however the axes are defined). Inter-axial self-attention refers to performing self-attention operations across different axes of the data. For example, after applying intra-axial self-attention separately along X and Y axes, inter-axial self- attention can integrate these results to allow for interactions between the features learned in each separate axis. This staging reduces computational while still allowing the model to consider the complete context of the data.

[0066] As illustrated, the decoder 540 receives both the bottleneck embedding 535 from the ASA block 530 and the encoded representation(s) 525 from the encoder 520 via a skip channel 550. The decoder 540 operates in essentially a reverse manner to that of the encoder. It applies multiple layers of transformations to generate increasingly higher-dimensional representations of the data, thereby progressively increasing the dimensionality with a goal of reconstructing the original input data from the highly compressed representation in the bottleneck layer. The output of the decoder 540 is a set of estimated Mel filter bank energies 545. The goal of training the network 500 can essentially be to get the estimated Mel filter bank energies 545 as close as practical to the Mel filter bank energies 515.

[0067] Although the decoder 540 output ideally matches the encoder 520 input, there will practically be some difference, referred to as reconstruction loss. Training of the auto-encoder seeks to optimize the network 500 to minimize reconstruction loss. Effective training results in the bottleneck layer producing a highly compact, but highly meaningful representation of the input data; effectively extracting the most salient features for ASR. The compact representation (i.e., the bottleneck embedding 535) in the bottleneck layer can be referred to as the ”latent space,” and the auto-encoder’s trained (learned) knowledge of how to encode raw input data into the latent space is embedded in a set of weights. The set of weights can be represented as a feature vector, a set of feature vectors, or in any suitable manner. For example, one second of audio can be represented by 16,000 samples (i.e., at a 16 kHz sampling rate), and the 16,000 samples can effectively be represented by a set of 256 weights (e.g., or 128 weights, or another suitable number) in the bottleneck embedding 535.

[0068] FIG.6 shows an illustrative USC machine learning network 600 that can be a more detailed implementation of the USC machine learning network 500 of FIG.5. As in FIG.5, the network 600 includes a Mel converter 510, an encoder 520, an ASA block 530, and a decoder 540. The encoder 520 can include ^^^^ encoder blocks 620 (^^^^ is a positive integer) and X decoder blocks 640. In the illustrated implementation, ^^^^ is 4. Each encoder block 620 can be implemented with generic encoder features to perform a stepwise reduction in dimensionality between its input and its output. As such, each encoder block 620 can be nominally identical (i.e., designed to have identical functionality, except for differences in input and output labeling, and other instance- specific identifiers). Similarly, each decoder block 640 can be implemented with generic decoder features to perform a stepwise increase in dimensionality between its input and its output. As such, each decoder block 640 can be nominally identical (i.e., designed to have identical functionality, except for differences in input and output labeling, and other instance-specific identifiers).

[0069] For example, FIGS.7A and 7B show architectures for illustrative implementations of an encoder block 620 and a decoder block 640, respectively. In FIG.7A, each encoder block 620 includes a two-dimensional convolution (Conv2D) block 710-E, a batch normalization (BN) block 720-E, and a leaky rectified linear unit (ReLU) block 730-E. The Conv2D block 710-E applies a set of learnable filters to the input frame received by the encoder. The input frame is represented in two-dimensional data, such as in rows and columns of a two-dimensional matrix. Each filter is responsible for learning relational (e.g., spatial) hierarchies between the elements of the input frame. The output of the Conv2D block 710-E is a feature map that represents the input frame in a more compact and abstract way. This is generated by using a convolution operation to effectively slide each filter across the two dimensions of the input frame and compute the dot product between the weights of the filter and the input elements at every spatial location.

[0070] The output of the Conv2D block 710-E is passed to the BN block 720-E, which is configured to increase the stability of the network by normalizing the output of the previous layer. For example, the BN block 720-E subtracts a batch mean and divides it by a batch standard deviation. This can stabilize the learning process and appreciably reduce the number of training epochs required to train the network.

[0071] The output of the BN block 720-E is passed through the leaky ReLU block 730-E. LeakyReLU is an activation function defined as: ^^^^(^^^^) = ^^^^^^^^^^^^(0.01^^^^, ^^^^), where ^^^^ is the input. Thefunction introduces a small slope to keep updates alive when the input is negative, thereby addressing a so-called “dying ReLU” problem present in traditional ReLU activation functions. Assuch, for every negative input, the function has a small negative slope rather than outputting zero. Essentially, the leaky ReLU function applies a non-linear transformation to the data, allowing the model to learn and represent more complex patterns. For positive inputs, the function returns the input itself; for negative inputs, it returns a small fraction of the input (thus, it is “leaky”).

[0072] Turning to FIG.7B, each decoder block 640 includes a transposed Conv2D block 710-D, a BN block 720-D, and a leaky ReLU block 730-D. The transposed Conv2D block 710-D performs deconvolution, which allows the spatial dimensions of the feature maps to be increased (i.e., in contrast to the Conv2D block 710-E, which reduce the spatial dimensions of the input frame). The transposed convolution operation effectively slides a learnable filter over the input frame and spreads the value of each input element across the filter, thereby resulting in a larger output. The manner in which the spreading occurs depends on the stride and padding used in the operation. The BN block 720-D and the leaky ReLU block 730-D can act in substantially the same manner as the corresponding blocks of the encoder block 620. The output of the transposed Conv2D block 710-D is normalized by the BN block 720-D, and the output of the BN block 720-D is activated by the leaky ReLU block 730-D.

[0073] Returning to FIG.6, the Mel converter 510 outputs Mel filter bank energies 515 based on an input audio signal. As illustrated, the dimensions of the Mel filter bank energies 515 are represented herein as [^^^^ x ^^^^ x ^^^^], where ^^^^ represents a frame, ^^^^ represents a feature, and ^^^^ represents a channel. The multi-dimensional set of [^^^^^^^^ x ^^^^^^^^ x ^^^^^^^^] is referred to herein as an ^^^^th data image. The Mel filter bank energies 515 are passed to the first encoder block 620 of the encoder 520.

[0074] The encoder 520 and decoder 540 are configured as a convolutional neural network. For example, as described with reference to FIGS.6A and 6B, each encoder block 620 includes a Conv2D block 710-E, and each decoder block 640 includes a transposed Conv2D block 710-D. The functionality of each of these convolution blocks is controlled by a set of parameters,including filters, kernel, and stride, represented in FIG. 5 as [^^^^^^^^,^^^^^^^^, ^^^^^^^^]. The filters parameterrefers to the number of unique learnable convolution kernels in a particular layer. Each filter is responsible for learning a different feature from the input data, and the number of filters determines the depth of the output feature map. Increasing the number of filters allows the model to learn more complex representations but increases computational complexity. The kernel parameter refers to the size of the window or receptive field that is convolved with the input data. It is typically a square matrix of weights that is slid across the input data to perform the convolution operation. The size of the kernel affects the level of detail the model can learn fromthe input data. For example, smaller kernels can capture finer-grained details, while larger kernels tend to capture more global features. The stride feature refers to the number of elements by which the kernel is shifted when sliding across the input data during the convolution operation (e.g., a stride of 1 indicates that the kernel slides by one element at a time). A larger stride yields a smaller output dimension (i.e., more down-sampling of the input data).

[0075] The input to each ^^^^th encoder block 620-I is the (^^^^ − 1)th data image, and the outputfrom each ^^^^th encoder block 620-I is the ^^^^th data image. The output from the Mel converter 510 (i.e., the Mel filter bank energies 515) can be represented as [^^^^0,^^^^0,^^^^0]. As illustrated, encoderblock 620-1 has an input of [^^^^0,^^^^0,^^^^0] and uses parameters [^^^^1^^^^ ,^^^^1^^^^ , ^^^^1^^^^] to generate an outputof [^^^1^ ,^^^^1,^^^^1]; encoder block 620-2 has an input of [^^^^1,^^^^1,^^^^1] and uses parameters[^^^^2^^^^ ,^^^^2^^^^ , ^^^^2^^^^] to generate an output of [^^^^2,^^^^2,^^^^2]; encoder block 620-3 has an input of[^^^^2,^^^^2,^^^^2] and uses parameters [^^^^3^^^^ ,^^^^3^^^^ , ^^^^3^^^^] to generate an output of [^^^^3,^^^^3,^^^^3]; and encoderblock 620-4 has an input of [^^^^3,^^^^3,^^^^3] and uses parameters [^^^^4^^^^ ,^^^^4^^^^ , ^^^^4^^^^] to generate an outputof [^^^4^ ,^^^4^ ,^^^^4]. The respective output of each encoder block 620 has lower dimensionality than its respective input.

[0076] The output of encoder block 620-4 (referred to in FIG.5 as the encoded representation 525) is passed to the ASA block 530, which is described in more detail below. The output of the ASA block 530 is referred to in FIG.5 as the bottleneck embedding 535. As illustrated, the input to each decoder block 640 is generated by a respective merging layer 645. Each merging layer 645 merges a preceding decoder output (or the bottleneck embedding 535 in the case of merging layer 645-4) with the output of a corresponding one of the encoder blocks 620 via a respective skip connection 650 of the skip channel 550. For example, each merging layer 645 can perform element-wise addition. In effect, the skip connections 650 and merging layers 645 form a nested filter architecture. For example, encoder block 620-4, the ASA block 530, and decoder block 640- 4 form an innermost filter via skip connection 650-4 and merging layer 645-4; that innermost filter is nested within another filter formed by encoder block 620-3, the ASA block 530, and decoder block 640-3 via skip connection 650-3 and merging layer 645-3; and so on.

[0077] The nested filters effectively form a hierarchical arrangement of interactions between the encoder 520, decoder 540, and ASA block 530. For example, each encoder block 620 processes input data and passes its output to both the next encoder block 620 and its corresponding decoder block 640 through a corresponding skip connection 650. This skip connection 650 ensures that each decoder block 640 receives both high-level encoded information from the ASA block 530 and lower-level features directly from the encoder block 620, thereby creating a composite featuremap that integrates multiple levels of abstraction. The ASA block 530 further refines the encoded representations before passing them to the decoder 540. In such a structure, each level of the encoder-decoder pair (e.g., encoder block 620-4, the ASA block 530, and decoder block 640-4) functions as a filter that processes the data at a particular level of abstraction. By nesting the filters in the illustrated manner, each higher-level filter (e.g., encoder block 620-3, the ASA block 530, and decoder block 640-3) encompassing the lower-level ones (e.g., encoder block 620-4, the ASA block 530, and decoder block 640-4). This facilitates capturing and reconstructing complex data representations in a multi-scale manner, allowing for more accurate and detailed outputs.

[0078] The output of each merging layer 645 (i.e., the input to each decoder block 640) is a corresponding estimated data image. As illustrated, merging layer 645-4 merges the output of encoder block 620-4 with the bottleneck embedding 535 from the ASA block 530 to generate an estimated frame image represented by [^^^^′4,^^^^′4,^^^^′4]. Once the network is trained, [^^^^′4,^^^^′4,^^^^′4] is ideally equal to (or as close as practical to) [^^^4^ ,^^^4^ ,^^^^4]. Indeed, the nested filter architecture seeks to reconstruct the input data as accurately as possible at each level of abstraction. For example, referring to the output of encoder block 4620-4 as “data image 4,” merging layer 4645-4 combines “data image 4” with the bottleneck embedding 535 from the ASA block 530 to generate an “estimated image 4,” which becomes the input to decoder block 4640-4. During training, using back-propagation and other optimization techniques, the network iteratively adjusting the weights of the encoder 520, decoder 540, and ASA block 530 until it minimizes the difference between “estimated image 4” and “data image 4.” In the illustrated architecture, the training process can ensure that the information captured and processed at each level of the network is preserved and accurately reconstructed, leading to more effective and precise overall encoding and decoding.

[0079] Decoder block 640 uses parameters [^^^^4^^^^ ,^^^^4^^^^ , ^^^^4^^^^] to generate an output of. Although FIG.6 represents the parameters for each ^^^^th encoder block 620 and each^^^^th decoder block 640 simply as [^^^^^^^^,^^^^^^^^ , ^^^^^^^^], this is not intended to require that the parameters foreach ^^^^th encoder block 620-I are the same as those for the corresponding ^^^^th decoder block 640-I. Indeed, that may be the case in some instances and / or implementations. For example,[^^^^4^^^^ ,^^^^4^^^^ ,^^^^4^^^^] can be the same as [^^^^4^^^^ ,^^^^4^^^^ , ^^^^4^^^^]. However, in other instances and / orimplementations, the values of [^^^^^^^^^^^^ ,^^^^^^^^^^^^ ,^^^^^^^^^^^^] can be different from those of [^^^^^^^^^^^^ ,^^^^^^^^^^^^ , ^^^^^^^^^^^^]. Forexample, encoder and decoder parameters can be independently tuned, independently frozen, forced to correspond, allowed to diverge, etc.

[0080] As illustrated, the remainder of the decoder 540 can proceed similarly. Merging layer 645-3 merges the output of encoder block 620-3 (via skip connection 650-3) with the output ofdecoder block 640-4 to generate an estimated frame image represented by [^^^^′3,^ decoder block 640-3 takes the output of merging layer 645-3 as its input and uses parameters [^^^^3^^^^,^^^^3^^^^,^^^^3^^^^] to generate an expanded image frame. Merging layer 645-2 merges the output of encoder block 620-2 (via skip connection 650-2) with the output of decoder block 640-3 to generate an estimated frame image represented by [^^^^′2,^^^^′2,^^^^′2]; decoder block 640-2 takes the output of merging layer 645-2 as its input and uses parameters [^^^^2^^^^,^^^^2^^^^,^^^^2^^^^] to generate a further expanded image frame [^^^^′1,^^^^′1,^^^^′1]. Finally, merging layer 645-1 merges the output of encoder block 620-1 (via skip connection 650-1) with the output of decoder block 640-2 to generate an estimated frame image represented by [^^^^′1,^^^^′1,^^^^′1]; decoder block 640-1 takes the output ofmerging layer 645-1 as its input and uses parameters [^^^^1^^^^ ,^^^^1^^^^ , ^^^^1^^^^] to generate a furtherexpanded image frame. The expanded image framecorresponds to the estimated Mel filter bank energies 545 of FIG.5.

[0081] The following table demonstrates example parameters, input data, and output data for blocks of the model. For the sake of simplicity, the example assumes that the model is perfectly trained.

[0082] FIG.8 shows a block diagram of an illustrative architecture for an axial self-attention (ASA) block 800, according to some embodiments described herein. The ASA block 800 can be an implementation of the ASA block 530 of FIGS.5 and 6. The input and output to the ASAblock 800 are both shown generally as having the dimensions [^^^^ x ^^^^ x ^^^^] (i.e., not as any particular [^^^^^^^^,^^^^^^^^,^^^^^^^^]), as such an ASA block 800 architecture can be used after any suitable number of encoder blocks 620 and before any suitable number of decoder blocks 640. As illustrated, the ASA block 800 includes three convolution stages, each with a respective convolution (Conv) block 810.

[0083] Each convolution block 810 of the ASA block 800 can be implemented using essentially the same architecture as the encoder blocks 620 of FIG.6A. For example, FIG.9 shows an architecture for an illustrative implementation of a convolution block 810 of the ASA block 800. The convolution block 810 includes a two-dimensional convolution (Conv2D) block 710-C, a batch normalization (BN) block 720-C, and a leaky rectified linear unit (ReLU) block 730-C. As described with reference to FIGS.7A and 7B, the Conv2D block 710-C applies a set of learnable filters to the input frame received at its input and outputs a feature map that represents the input frame in a more compact and abstract way. This is passed to the BN block 720-C, which normalizes the feature map. The normalized feature map is then passed to the leaky ReLU block 730-E for activation.

[0084] Returning to FIG.8, each convolution block 810 can have a dimensionality representedby [^^^^,^^^^, ^^^^, 1]. In the illustrated implementation, the parameters are the same for all theconvolution blocks 810: [^^^^ = 1,^^^^ = 1x1, ^^^^ = 1, 1]. In the first convolution stage, a firstconvolution block 810-1 uses an input data image (i.e., the encoded representation 525) to generate and output a first set of feature maps to three dense blocks 820. For example, the feature maps are the result of applying various filters across the input data, and each filter detects specific features at different spatial hierarchies according to the layer and depth of the convolution block 810. The dense blocks 820 refer to fully connected layers within the neural network. Each “neuron” in a dense block 820 receives input from all neurons of a previous layer (which is fundamentally different from the way the convolutional layers operate, where each output is only connected to a local region of the input through the convolution operation). The dense blocks 820 can perform high-level reasoning from the features extracted by the convolutional layers of the convolution block 810. Typically, outputting the feature maps to the dense blocks 820 can involve flattening the feature maps into single vectors.

[0085] As illustrated, dense block 820-11 generates one or more “query” output vectors, dense block 820-12 generates one or more “value” output vectors, and dense block 820-13 generates one or more “key” output vectors. In the ASA context, the dense blocks 820 are configured to generate these vector outputs to support a subsequent attention mechanism implemented by anattention block 830 (attention block 830-1 in this stage). The key vector(s) help the attention mechanism to determine how much attention to pay to different parts of the input data. The query output vector(s) are used by the attention mechanism to score each key. The score can be generated based on one or more outputs from one or more previous layers to retrieve information from the keys that are most relevant for answering the query. For example, a scoring function (e.g., a dot product) can be applied between the query and each key. Once the keys are scored by the query, the value vector(s) are used to construct the output of the attention block 830-1. Each value is associated with a key, and the attention mechanism essentially creates a weighted sum of these values based on the scores computed between the keys and the query. The weighted sum can be output by the attention layer as a representation of the most relevant information, as dictated by the query. As illustrated, the attention block 830-1 in the first convolution stage can output attention scores (e.g., the results of scoring each key by the query vector(s)) and an attention output (e.g., the weighted sum(s) generated by the value vector(s)).

[0086] In the second convolution stage, the input to the second convolution block 810-2 is generated by a fuser 835, which combines the attention scores from the first convolution stage with the initial input data. For example, the fuser 835 can used a cross-product to combine the data. Prior to passing the output of the fuser 835 to the second convolution block 810-2, the data is reshaped by pivoting the dimensions from [^^^^ x ^^^^ x ^^^^] to [^^^^ x ^^^^ x ^^^^]. The output of the second convolution block 810-2 is passed to two dense blocks 820. One of the dense blocks 820-21 generates one or more second query vectors as a first corresponding input to a second attention block 830-2, and the other of the dense blocks 820-22 generates one or more second key vectors as a second corresponding input to the second attention block 830-2. The value input to the second attention block 830-2 is taken from the attention output of the first attention block 830-1. The second attention block 830-2 outputs a second attention output.

[0087] In the third convolution stage, the input to the third convolution block 810-3 is generated from the second attention output. Prior to passing the second attention output to the third convolution block 810-3, the data is again reshaped by pivoting the dimensions back from [^^^^ x ^^^^ x ^^^^] to [^^^^ x ^^^^ x ^^^^]. The output of the third convolution block 810-3 is now back to the dimensions of the initial input data but is in a more compact form generated by inter-axial and intra-axial attention mechanisms. The output of the third convolution block 810-3 corresponds to the bottleneck embedding 535 of FIGS.5 and 6.

[0088] As described above, the topologies and architectures of the USC machine learning models described above are trainable with a goal of converting an input audio sample to an outputhas only speaker-specific influences removed. FIG.10 shows an illustrative training environment 1000 for training a USC machine learning model, such as the one described in FIGS.5 – 9. The USC machine learning model can be considered as the Mel converter 510 plus a feature generator 1010 consisting of the encoder 520, ASA block 530, and decoder 540. Alternatively, the USC machine learning model can be considered as only including the feature generator 1010 (i.e., the model begins with the receipt of Mel filter bank energies 515).

[0089] As illustrated, the environment includes a spoken word audio (SWA) repository 1005 and a USC repository 1015. In some implementations, the SWA repository 1005 is a large speech- audio corpus of phonetically diversified classes (e.g., 11,000), where each class (or classification) corresponds to a linguistically distinct word in the speech-audio corpus. The USC repository 1015 can include a single unified representation of each of the same classes. For example, as described below, the unified representations can be generated by converting the speech to text and back to synthesized speech. Alternatively, the unified representations can be generated by selecting a same speaker from each classification to use as a “normalized” or “generic” speaker.

[0090] In each of a large number of training epochs, an SWA sample for a particular class is selected (e.g., at random) from the SWA repository 1005, and the corresponding USC sample for the same class is selected from the USC repository 1015. The selected SWA sample is passed through a first instance of the Mel converter 510-1 to generate a first set of Mel filter bank energies 515-1, and the selected USC sample is passed through a second instance of the Mel converter 510- 2 to generate a second set of Mel filter bank energies 515-2. The first Mel filter bank energies 515-1 is passed through the feature generator 1010 to generate a corresponding set of estimated Mel filter bank energies 545.

[0091] Particularly during earlier training epochs, the feature generator 1010 may have no sense of how to remove speaker-specific influences from the first Mel filter bank energies 515-1 to make them look like the second Mel filter bank energies 515-2. As such, the estimated MEL filter bank energies 545 generated in those epochs are not expected to correlate well with the second Mel filter bank energies 515-2. A loss function 1020 (e.g., a mean square error (MSE) loss function) is used to compute an error between the estimated Mel filter bank energies 545 for the epoch and the second (i.e., USC) Mel filter bank energies 515-2. This error is back-propagated to the feature generator 1010 to tune parameters of the model. Thus, in each epoch, it is expected that the error will decrease as the estimated Mel filter bank energies 545 approach the Mel filter bank energies 515-2 generated from the USC samples. The training can be considered complete when the error falls below a predetermined threshold.

[0092] Once trained, it can be assumed that any SWA sample of any class with any suprasegmental features (e.g., any prosody, etc.) can be converted by the USC machine learning network into a set of estimated Mel filter bank energies 545 corresponding to a linguistic representation of the same class with predefined suprasegmental features (e.g., prosody corresponding to a synthesized speaker, a normalized speaker, etc.). Thus, the USC machine learning networks described in FIGS.5 – 10 can be used to implement USC approaches to automated attention handling. In some embodiments, additional layers and / or components can be added to the model to generate a different output representation. For example, the model can be expanded to output an audio signal corresponding to the estimated USC representation of the class. In other embodiments, the model can be expanded to classify the input audio, based on the estimated unified representation, into one of a predetermined set of output classifications. As used in automated attention handling contexts, a “word” can include any type of word, name, or utterance that could reasonably be used to get someone’s attention, such as “John,” “mister,” “hey,” “excuse me,” etc. Notably, the term “attention” used for automated attention relates to getting a human user’s attention (e.g., to engage them in a conversation). This should not be confused with use of the term “attention” in the context of self-attention mechanisms of the machine learning network itself, such as implemented by the ASA block 530.

[0093] FIG.11 shows a flow diagram of an illustrative method 1100 for training an artificial neural network (ANN) (i.e., a machine learning network) for universal sound conversion (USC), according to embodiments described herein. Embodiments of the method 1100 can iterate for a large number of training epochs, refining the ANN until an error is computed to be below a predefined threshold. Embodiments begin at stage 1104 by generating a first set of Mel filter bank energies from a spoken word audio (SWA) sample selected for the training epoch from a SWA repository. The SWA sample is a spoken representation of a word corresponding to a class, as described herein. At stage 1108, embodiments can generate a second set of Mel filter bank energies from a USC sample selected for the training epoch from a USC repository. The USC sample selected as a unified audio representation of the class absent speaker-specific suprasegmental features.

[0094] At stage 1112, embodiments can generate an estimated set of Mel filter bank energies for the training epoch from the first set of Mel filter bank energies using an ANN topology. As described herein, the ANN topology includes a convolutional encoder, an axial self-attention (ASA) block, and a deconvolutional decoder. In some embodiments, the generating in stage 1112 involves: generating one or more encoded representations by compressing the first set of Mel filter bank energies using the convolutional encoder according to a set of encoder parameters;generating a bottleneck embedding by compressing the one or more encoded representations using the ASA block according to a set of ASA parameters; and generating the estimated set of Mel filter bank energies by using the deconvolutional decoder to decompress a combination of the bottleneck embedding and the one or more encoded representations according to a set of decoder parameters. In some such embodiments, generating the bottleneck embedding involves: computing a first query vector, a first key vector, and a first value vector based on the encoded representation; computing attention scores and a first attention output vector based on the first query vector, the first key vector, and the first value vector; computing a second query vector and a second key vector based on fusing the encoded representation with the attention scores; computing a second attention output vector based on the second query vector, the second key vector, and the first attention output vector; and computing the bottleneck embedding based on the second attention output vector. In some implementations, the encoded representation has dimensions [T x F x C], wherein T represents a frame dimension, F represents a feature dimension, and C represents a channel dimension. Such implementations can further involve reshaping the dimensions to [F x T x C] prior to computing the second query vector and a second key vector and reshaping the dimensions back to [T x F x C] prior to computing the bottleneck embedding.

[0095] At stage 1116, embodiments can compute the error between the estimated set of Mel filter bank energies and the second set of Mel filter bank energies. For example, the error is computed as a loss function, such as a mean square error. At stage 1120, embodiments can determine whether the error computed in stage 1116 is below the predetermined threshold. If so, the ANN can be considered as trained, and embodiments can end. If not, at stage 1120, embodiments can back-propagate the error to the ANN topology to reduce the error in further iterations, and the method 1100 can return to stage 1104 for the next training epoch. In some embodiments, back-propagating the error involves tuning at least one of the encoder parameters, the ASA parameters, or the decoder parameters to reduce the error in subsequent epochs.

[0096] Referring back to FIG.4, the machine learning networks described in FIGS.5 – 11 can be used to implement a foundation model 350 and / or a name embedding model 415. For example, a foundation model 350 can be trained as per FIG.10. The name embedding model 415 can then be generated by applying transfer learning from the large speech-audio corpus of phonetically diversified words used in training the foundation model 350 (the source dataset) to a smaller corpus of real-world name data (the target dataset). For example, the foundation model 350 can use a large number (e.g., 11,000) of classifications to generate the set of weights, where each linguistically distinct word in the speech-audio corpus is classified into one of the classes. Transfer learning can then apply the trained foundation model 350 to generate the nameembedding model 415 (deep feature generation model, or DFGNet) for a smaller number (e.g., 500 – 1000) of classifications, each associated with a linguistically distinct name. Implementations of the name embedding model 415 include only the encoder and bottleneck layer as trained through the transfer learning.

[0097] In effect, the name embedding model 415 can be characterized as a tuned and reduced version of the foundation model 350 developed specifically for name detection, such as by using few-shot learning (FSL). Real audio samples used to train the name embedding model 415 can correspond to people's names from various regions and languages, and sample names can be chosen to cover most phonetic usage in each region. Implementations of the name embedding model 415 can be trained to recognize any suitable number of name classifications. For example, it may be impractical impossible to train the model for all possible names and their variants everywhere in the world, and a practical number of more common names (e.g., 1,000) can be chosen instead. In some implementations, different versions of the name embedding model can be generated 415 and / or trained differently for different user groupings (e.g., geographical regions, ethnicities, etc.) to capture the most popular names for the corresponding groupings. For example, grouping information can be entered by the user as part of enrollment, obtained for the user from account information, assumed for the user based on location or other demographic information, etc. Further, the name embedding model 415 can be designed with as much complex as needed to discriminate between the name classifications used for training with enough reliability to be suitable for name-detection-based attention handling. For example, a higher-dimensional model (e.g., where the number of weighting vector dimensions, N, is larger) may be able to more reliably discriminate among a larger number of name classifications, but use of such a model in real time will involve more computational resources (e.g., which may correspond to more processing time, more battery usage, more heat generation, etc.).

[0098] The name embedding model 415 can then generate a reference embedding for each invocation name. For example, FIGS.12A and 12B show block diagrams 1200 of illustrative uses of the name embedding model 415 to generate the deep image 420. Turning first to FIG.12A, a user 1205 provides a set of J invocation names 1210 via an enrollment application 330 (J is a positive integer). For example, the user 1205 speaks each name one or more times, types each name using its proper spelling, types each name phonetically, etc. The J invocation names 1210 are passed to the name embedding model 415, which generates J corresponding reference embeddings.

[0099] In some embodiments, each reference embedding is an N-dimensional vector corresponding to a set of N weights in the name embedding model 415 that represents the invocation name that yielded that reference embedding. N can be any suitable integer number of weights to provide sufficiently reliable classification. In one implementation, N is 128. In another implementation, N is 256. For example, the name embedding model 415 generates each reference embedding as the N-dimensional vector, and stores the vectors in the deep image 420. For example, the deep image 420 stores J N-dimensional vectors, a single J-by-N-dimensional matrix, or the like.

[0100] Turning first to FIG.12B, a user 1205 again provides a set of J invocation names 1210 via an enrollment application 330 (J is a positive integer). Unlike in FIG.12A, the J invocation names 1210 are passed to a name augmenter 1220, which augments the user-provided set of invocation names 1210 to generate an augmented set of invocation names 1210’. The name augmenter 1220 can include, or be in communication with, an augmentation model 1215. Embodiments of the augmentation model 1215 include mathematical transformations to apply to each of some or all of the invocation names 1210. The name augmenter 1220 can generate K augmentations (K is a positive integer) for each of the J invocation names 1210, so that the augmented set of invocation names 1210’ includes J * K names. For example, a user 1205 enrolls four invocation names, nine augmentations are applied to each invocation name to generate ten total names for each invocation name, or forty total entries in the augmented set of invocation names 1210’.

[0101] In some implementations, the name augmenter 1220 adds time-based augmentations to each of some or all of the invocation names 1210, such as by time-stretching and / or time- compressing a user-provided audio sample of the invocation name. In some implementations, the name augmenter 1220 adds accent-based augmentations to each of some or all of the invocation names 1210, such as by mathematically applying different vowel changes, regional variations, pronunciations, etc. to the invocation name. In some implementations, the name augmenter 1220 adds suprasegmental augmentations to each of some or all of the invocation names 1210, such as by mathematically applying different syllable accenting, intonation, volume, pitch, etc. Other augmentations can account for differences across genders, ages, etc. Other augmentations can account for noise models, such as models of ambient background noise, television or music noise, traffic noise, road noise, engine noise, air conditioning noise, running water noise, etc. The J * K invocation names 1210’ are passed to the name embedding model 415, which generates J * K corresponding reference embeddings. For example, the name embedding model 415 generates each reference embedding as an N-dimensional vector for storage in the deep image 420 (e.g., as J* K N-dimensional vectors, as a (J * K)-by-N-dimensional matrix, or the like). In some implementations, the name augmenter 1220 applies different augmentations to different invocation names, and / or different numbers of augmentations to different invocation names. As one example, different augmentations can be applied based on whether the invocation name is characterized more by its vowel content, or more by its consonant content. As another example, a more common term enrolled as an invocation name (e.g., “boss,” “mom”), or a shorter name enrolled as an invocation name (e.g., “Max,” “Tim”) may be augmented differently than less common terms, longer names, etc.

[0102] Returning to FIG.4, the relation network 435 is trained with the linear and non-linear features that characterize the name embedding model 415. Embodiments of the relation network 435 are trained to output a similarity score (e.g., a probability of a match) responsive to two inputs: one of the reference embeddings from the deep image 420, and a real-time embedding generated from a real-time audio sample received via the reference microphone. For example, the one of the reference embeddings from the deep image 420 is a N-dimensional vector previously generated by the name embedding model 415 during enrollment, and the real-time embedding is an N- dimensional vector generated by the name embedding model 415 in real-time. As described above, the reference embedding vector essentially represents salient linear and non-linear features of the associated classification, and the real-time embedding vector essentially represents the same salient linear and non-linear features of the real-time audio sample. The relation network can map those same linear and non-linear features between real-time embeddings and reference embeddings to find candidate matches. For example, in a scenario where 40 classifications are generated (i.e., the deep image 420 is a 40-by-N matrix), the relation network 435 can compute a correspondence between the real-time embedding and each of the 40 reference embeddings. This can be performed as 40 serial computations (e.g., iterative), 40 parallel computations, or in any suitable manner. Some embodiments of the relation network 435 are implemented as a two-dimensional convolutional neural network (CNN). Some other embodiments of the relation network 435 are implemented as a one-dimensional CNN, a time-delay neural network (TDNN), or another suitable neural network. Some other embodiments of the relation network 435 are implemented using simple cosine similarity or equilidian distance estimation. For example, thresholding is performed based on the measured metric, and either a high value of the cosine similarity score represents a high relationship (for simple cosine similarity), or a minimum score represents a high relationship (for equlidian distance).

[0103] Embodiments compute a similarity score (e.g., a mathematical correlation) between a present real-time embedding (RTE) and each of the reference embeddings and determine whetherthe similarity score exceeds a predetermined matching threshold (e.g., 0.3) for any one or more of the reference embeddings. If none of the reference embeddings yields a similarity score exceeding the predetermined matching threshold, embodiments determine that there is no name match and ignore the analyzed portion of the real-time audio signal (i.e., discards the RTE). If one of the reference embeddings yields a similarity score exceeding the predetermined matching threshold, the class associated with that reference embodiment is selected as a candidate matching name (i.e., that reference embedding is selected as the candidate matching reference embedding, or CMRE). If multiple reference embeddings yield similarity scores exceeding the predetermined matching threshold, the reference embedding associated with the highest similarity score is selected as the CMRE.

[0104] Embodiments of the relation network 435 are trained to output a similarity score (e.g., a probability of a match) responsive to two inputs: one of the reference embeddings from the deep image, and a real-time embedding generated from a real-time audio sample received via the reference microphone. During training of the relation network 435, a training audio sample can be used as the real-time audio sample, which is fed into the name embedding model 415 to generate a training embedding (corresponding to the real-time embedding generated during normal operation). For example, the reference embedding for a particular invoked name classification is input to the relation network 435, and a training embedding is generated by the name embedding model 415 for a training audio sample: if the training audio sample is known to correspond to a particular invocation name, the relation network 435 is trained to output ‘1’, ‘100 percent’, etc. when fed the corresponding reference and training embeddings; if the training audio sample is known not to correspond to a particular invocation name, the relation network 435 is trained to output ‘0’, ‘0 percent’, etc. when fed the corresponding reference and training embeddings. In some embodiments, the relation network 435 is trained based on the same corpus of audio samples (or a portion thereof) used to train the foundation model 350. In other embodiments, the relation network 435 is trained based on the same corpus of audio samples (or a portion thereof) used to train the name embedding model 415.

[0105] The false rejection network (FRNet) 445 seeks to determine whether the CMRE and the RTE can be discriminated. In effect, the relation network 435 seeks to find a candidate match, and the false rejection network 445 seeks to determine whether the candidate match is a false match. Embodiments of the false rejection network 445 apply multiple mathematical transformations (e.g., rotations), each transformation designed to transform both the CMRE and the RTE into a corresponding domain and / or space to see whether the two datasets continue to match. For example, suppose a user has enrolled the invocation name, “Jonathan,” and the real-time audiosignal includes the phrase “on a thin.” In such a scenario, the relation network 435 may find a candidate match (i.e., a similarity score exceeding the threshold), but the false rejection network 445 may determine that the candidate match is likely not a match and can be rejected. Some embodiments of the false rejection network 445 are implemented as a progressive layered extraction (PLE) neural network. Some other embodiments of the false rejection network 445 are implemented as a probabilistic linear discriminant analysis (PLDA) network.

[0106] Embodiments of the false rejection network 445 are trained to output a discrimination score (e.g., a likelihood ratio representing probability of a false match) responsive to two inputs: one of the reference embeddings from the deep image 420, and a real-time embedding generated from a real-time audio sample received via the reference microphone. The training of the false rejection network 445 can be similar to the training of the relation network 435. For example, during training of the false rejection network 445, a training audio sample can be used as the real- time audio sample, which is fed into the name embedding model 415 to generate a training embedding (corresponding to the real-time embedding generated during normal operation). Unlike the relation network 435, the false rejection network 445 is trained to apply transformations to the two inputs to look for a particular domain or space in which the two can be discriminated. For example, the training can use some training audio samples that are similar to a particular invoked name classification and other audio samples that are completely different (e.g., effectively linguistically orthogonal) to the invoked name classification. The false rejection network 445 is trained to find transformations that reliably discriminate involved name classifications from audio samples that sound like those invoked names but actually carry a different linguistic meaning. In some embodiments, the false rejection network 445 is trained based on the same corpus of audio samples (or a portion thereof) used to train the foundation model 350. In other embodiments, the false rejection network 445 is trained based on the same corpus of audio samples (or a portion thereof) used to train the name embedding model 415.

[0107] The name embedding model 415, the relation network 435, and the false rejection network 445 can all be trained together (e.g., in parallel, or serially). As noted above, the name embedding model 415 is trained by transfer learning from the foundation model 350 using a corpus of real-world name data. The input is an audio sample, and the output is a classification (e.g., M-dimensional vector, where M is the number of classifications, such as J or J * K). The specific invocation names (e.g., including augmentations) are used to generate reference embeddings for each of a set of invocation name classifications, which are stored as the deep image 420. Embodiments of the relation network 435 are trained to output a respective similarity score between a real-time embedding generated from a real-time audio sample received via thereference microphone and each of the reference embeddings from the deep image 420. Embodiments of the false rejection network 445 are trained to output a respective discrimination score between a real-time embedding generated from a real-time audio sample received via the reference microphone and each of the reference embeddings from the deep image 420. Because the relation network 435 only computes similarity scores and matches them to a threshold, the relation network 435 can be very lightweight (e.g., resource-efficient). For example, even in the context of a small processor and a small battery, such as in an earbud), the relation network 435 can run continuously without using excessively processor computation cycles, without draining excessive power, without generating excessive heat, etc. Embodiments of the false rejection network 445, which may use appreciably more resources to perform transformations, etc., only run when a candidate match has been identified. Alternative embodiments can combine the functionality of the relation network 435 and the false rejection network 445, such as in contexts where resources are not as limited (e.g., implemented in over-ear headphones that include wired power).

[0108] As described above, some or all of the name embedding model 415, the relation network 435, the false rejection network 445, and the deep image 420 can be treated as a single inference model (or a “name detection model”). For example, the foundation model 350 is a common model that is computed in and / or stored in the cloud 340. When invocation names are first enrolled, an enrollment application 330 is downloaded to a user device. For example, the application is downloaded to the user’s laptop computer, tablet computer, smartphone, smart watch, portable audio player, headset, etc. In some implementations, the WAC 310 is associated with a case, such as for storage and / or charging; and the enrollment application 330 can be downloaded to a computational environment stored in the case.

[0109] For the sake of illustration, FIG.13 shows several example screenshots from an example enrollment application 330 running on a user device. At a first screen 1310, the user begins a name enrollment process. By clicking “NEXT” using a user interface of the user device (e.g., a touchscreen), the user can proceed to a second screen 1320. At the second screen 1320, the user is prompted to enroll an invocation name. For example, the second screen 1320 includes a button to activate a microphone of the user device by which to receive an audio sample from the user representing the invocation name being enrolled. Additionally or alternatively, the second screen 1320 (or another screen) can include interface elements for receiving text, etc. Proceeding to a third screen 1330 (e.g., by clicking “NEXT”), the user is presented with several options, such as an option to re-record the enrollment name, to enroll another name, or to end the enrollment process. Some implementations can present additional options, such as permitting the user to select anypreviously enrolled name to re-record, to delete, etc. In some cases, opting to re-record or to enroll another name can bring the user back to the second screen 1320, or another similar screen. Opting to end the enrollment can bring the user to a fourth screen 1340, which indicates to the user that the enrollment is complete.

[0110] In some implementations, conclusion of the user enrollment of invocation names automatically triggers the enrollment application 330 to compute (generate) some or all of the name detection model. In other implementations, subsequent to the user enrollment of invocation names, the user is prompted to continue with generation of some or all of the name detection model. In some implementations, some or all of the name detection model is generated separately from the enrollment application 330. After the name detection model is generated, the name detection model can be ported to the WAC 310 for local execution. Some embodiments of the enrollment application 330 permit the user, at any suitable time, to enroll additional invocation names, delete enrolled invocation names, etc.

[0111] Some embodiments described herein assume joint participation of a cloud-based computational platform, a local computational platform separate from the WAC 310 (e.g., a smartphone), and the computational platform integrated in the WAC 310. Different arrangements of features, components, etc. can be implemented depending on the computing, power, storage, and / or other resources of these computational platforms. In one implementation, the application is downloaded directly to the WAC 310 (or is previously loaded to the WAC 310), and the name detection model is computed directly by the WAC 310 (i.e., there is no need for a separate computational platform. In another implementation, enrollment information is exchanged with cloud-based processing resources to generate some or all of the name detection model. For example, audio samples corresponding to the invocation names (e.g., including augmentations thereof) are sent to the cloud, cloud-based resources are used to compute the name detection model, and the name detection model is ported (e.g., directly from the cloud, or via one or more intermediary devices) to the WAC 310. In other implementations, the application is directly ported to the WAC 310, and it is then downloaded to, or installed on, the local computational platform separate from the WAC 310 (e.g., the smartphone, etc.), if the local computational platform does not already have it while pairing.

[0112] FIG.14 shows a flow diagram of an illustrative method 1400 for audio management that includes automated attention handling in a wearable audio component (WAC), according to embodiments described herein. Embodiments of the method 1400 can be implemented using any of the system implementations described herein, or any other suitable variation thereof. Someembodiments begin at stage 1404 by receiving a real-time audio signal (e.g., from a reference microphone associated with an active noise control (ANC) system). The real-time audio signal is ambient audio received via the WAC while a user is listening to desired audio with ANC in an active (ambient sound suppression) mode.

[0113] At stage 1408, embodiments can detect whether the real-time audio signal includes attention seeking (AS) audio. For example, as described above with reference to FIG.1, an AHS system 150 can be used to detect when a second-party attention seeker is trying to get the attention of a user while the user is wearing the WAC and is listening to desired audio 165 with the ANC system 140 in the active mode. In particular, the AHS system 150 listens for presence of attention seeking (AS) audio 157 within the ambient audio 155. For example, embodiments of the AHS system 150 are configured to listen for AS audio 157 corresponding to a previously enrolled invocation name (e.g., a name of the user). As described herein, the detection in stage 1408 can be performed using automated attention handling based on universal sound conversion. As illustrated, the detection at stage 1408 can rely on prior training of a name detection (i.e., inference) model at stage 1420, generation and storage of reference embeddings based on a set of enrolled invocation names using the name detection model at stage 1422, generation of real-time embeddings from the real-time audio signal using the name detection model at stage 1424, and comparison of the real-time embeddings with the reference embeddings to determine whether the AS audio is present at stage 1426.

[0114] A determination block at stage 1412 represents the result of the determination at stage 1408. If no AS audio is detected, embodiments of the method 1400 return to stage 1404. For example, embodiments continue to listen to the real-time audio signal, and the ANC system remains in active mode. If AS audio is detected, embodiments proceed to stage 1416 by triggering the ANC system automatically to switch to a conversation mode. For example, referring back to FIG.1, when the AS audio 157 is detected by the AHS system 150, the AHS system 150 automatically directs the ANC system 140 to switch from the active mode to the conversation mode. As described herein, the conversation mode can include disabling ANC, lowering the volume of desired audio, pausing the desired audio, enhancing conversationally relevant audio, suppressing feedback of the user’s own speech, etc.

[0115] As illustrated by off-page reference “A”, some embodiments of the method 1400 include an enrollment phase prior to stage 1404. FIG.15 shows a flow diagram of an illustrative method 1500 for such an enrollment phase. In some embodiments, the method 1500 begins at stage 1504 when a new WAC is detected. Such a detection can occur when a WAC is first paired with a userdevice, first paired with a network, first set to perform attention handling, etc. For example, the term “new” in this context can simply indicate that the WAC is new with respect to features of automated attention handling described herein. In response to the detection at stage 1504, embodiments can obtain an enrollment application (e.g., from the cloud) at stage 1508.

[0116] At stage 1512, embodiments can receiving a set of invocation names (e.g., see stage 1412 of FIG.14) from the user. At stage 1516, embodiments can generate a reference name embedding for each of the invocation names by the processor-executable name embedding model. Stage 1516 can correspond to stage 1422 of FIG.14. In some embodiments, generating the reference name embedding at stage 1516 includes applying a plurality of augmentation transformations to each of the set of invocation names to generate an augmented set of invocation names and generating a reference name embedding for each of the augmented set of invocation names by the processor- executable name embedding model. At stage 1520, embodiments can store the reference name embeddings in a non-transitory deep image.

[0117] Returning to FIG.14, embodiments of the method 1500 can also include conversation end detection subsequent to stage 1416, as indicated by off-page reference “B.” FIG.16 shows a flow diagram of an illustrative method 1600 for such conversation end detection. For example, it is assumed that the output of the name invoked signal at stage 1416 of FIG.14 indicates the beginning of a conversation involving the user and a second party. At stage 1604, embodiments detect a conversation end trigger subsequent to stage 1416 (i.e., after the name invoked signal directed the ANC system automatically to enter the conversation mode). At stage 1608, in response to the detection at stage 1604, embodiments can output a conversation end signal responsive to detecting the conversation end trigger. The name invoked signal directs the ANC system automatically to switch from an ambient sound suppression mode to a conversation mode, and the conversation end signal directs the ANC system automatically to switch from the conversation mode to the ambient sound suppression mode.

[0118] FIGS.17A and 17B show block diagrams of a training environment 1700 for training a foundation model 350 to support unified sound conversion (USC) for automated name-detection- based attention handling, according to some embodiments described herein. As illustrated, the training environment 1700 can begin with a spoken word audio repository 1710. The repository 1710 can include one or more large corpuses of spoken audio data. Preferably, the corpuses include many words (classes), and many diversified samples in each class, so that each class includes many versions of the same word spoken with wide suprasegmental variance. For example, a given word may be spoken 10,000 times by different speakers from around the world.Thus, for each class, the repository 1710 can output a large number of diversified audio samples for the class, which can be referred to as the “repository spoken audio samples” for the class.

[0119] The repository spoken audio samples for each class form the entire set of spoken audio samples for that class, referred to as the “class spoken audio samples.” For example, if the repository 1710 includes S samples for a particular class (S is a positive integer), there are S class spoken audio samples. In other embodiments, additional variation in the audio samples for each class is created by passing the class spoken audio samples through an augmenter 1720. The augmenter 1720 uses one or more augmentation models 1715 to generate an augmented set of class spoken audio samples with variations in features, such as speech speed (e.g., lengthening or shortening of the audio sample, lengthening or shortening of some or all vowel sounds, etc.), modeled suprasegmental variations, models of noise profiles and / or ambient noise features (e.g., traffic sounds, background conversation sounds, etc.), etc. Some embodiments of the augmenter 1720 are implemented in the same manner as the name augmenter 520 of FIG.5B (e.g., and the augmentation models 1715 can be the same as, or different from the augmentation models 515 of FIG.5B). Other embodiments of the augmenter 1720 introduce more and / or different types of variation using more and / or other augmentation models 1715. In embodiments that include the augmenter 1720, the augmented set of audio samples for the class is used as the class spoken audio samples. For example, if the augmenter 1720 produces A augmentations for each of the S repository spoken audio samples (A is a positive integer), there are A * S class spoken audio samples. In some cases, the repository 1710 may include different numbers of samples for different classes, and / or the augmenter 1720 may apply different types of augmentations for different classes, such that the values of A and / or S may be class dependent.

[0120] Embodiments also generate a unified audio sample for each class, referred to as a “class unified audio sample.” As illustrated in FIG.17A, in some embodiments, the repository 1710 includes a lexical entry for each of some or all of the classes, which can be used directly as “class text.” For example, the term “INDEPENDENCE” can have hundreds of diversified spoken audio samples for the word, all stored in association with a lexical entry (i.e., the text) for the word. In other embodiments, the repository 1710 may not include lexical entries for classes, or may not include a lexical entry for one or more classes. In such embodiments, for any class that does not have an associated lexical entry, one or more of the repository spoken audio samples is fed to a speech-to-text (STT) engine 1730, which generates the class text from the repository spoken audio sample(s). The class text, whether derived from a lexical entry in the repository 1710 or generated by the STT engine 1730, is passed to a text-to-speech (TTS) synthesizer 1735 to produce the class unified audio sample. In some implementations, the class unified audio sample is stripped of allaccent, prosody, etc. For example, the class unified audio sample is a purely phonetic representation of the class in a standardized set of phonemes. In other implementations, the TTS synthesizer 1735 produces the class unified audio sample to have synthesized suprasegmental features (i.e., the class unified audio sample is stripped of the speaker’s suprasegmental influence, but it may still have synthesized suprasegmental features).

[0121] FIG.17B shows alternative embodiments for generating the class unified audio sample. Rather than generating class unified audio sample from TTS synthesis, the class unified audio sample is generated using a selected audio sample for the class. As illustrated, embodiments can include a selector 1770. The repository 1710 includes spoken audio samples from many different speakers for each class (i.e., the class spoken audio samples), and the class unified audio sample for each class is generated by the selector 1770 by selecting one of the class spoken audio samples, and / or selecting one of the speakers of the class spoken audio samples. In one implementation, the selected spoken audio sample and / or speaker is the same for all classes. For example, the selector 1770 uses a particular identifier (e.g., index, etc.) to select a same speaker’s audio contribution as the unified sound for all classes. In other implementations, the selector 1770 randomly (or otherwise) selects a selected speaker from among the available speakers and / or samples for each class. For example, a particular speaker may not have provided audio contributions for all classes. In other implementations, the selection is based on quality features. For example, if the repository 1710 aggregates data from multiple corpuses, some corpuses may provide better quality audio samples than others; or within a particular one or more corpuses, some audio samples may be of better quality than others. In such cases, the selector 1770 is configured to select one of the class spoken audio samples as the class unified audio sample for each class based on the quality of the sample. In some cases, a particular corpus or portion of a corpus may provide better data for a particular geographic region in which the user resides. In such cases, the selector 1770 is configured to select one of the class spoken audio samples as the class unified audio sample for each class based on regional factors, such as suprasegmental features more common to a particular geographic region. For example, some implementations may ignore certain very strong regional accents for the purposes of generating the class unified audio sample, except in cases where the user resides in geographic proximity to speakers with that regional accent.

[0122] As illustrated in both FIGS.17A and 17B, embodiments include a feature extractor 1740, which takes an audio waveform as its input and outputs a computed a set of features, referred to as a feature vector. In some implementations, the computed feature vectors are sets of filter bank energies, which are a set of values representing the respective energy content at each of a set of different frequency bands in the input audio signal. For example, the input audio signal is passedthrough a set of bandpass filters (e.g., each designed to respond to a specific range of frequencies that mimics the frequency response of the human auditory system). The filtered signals can then be rectified (e.g., to preserve magnitude information and remove phase information) and smoothed (e.g., averaged), resulting in the so-called filter bank energies. As illustrated, passing the class spoken audio samples through the feature extractor 1740 produces spoken feature vectors 1742, and passing the class unified audio sample through the feature extractor 1740 produces a unified feature vector 1744. In some implementations, the feature vectors can be generated with a representation other than filter bank energies. In some such implementations, a complex fast Fourier transform (FFT) is used to generate the feature vectors as an array of complex numbers. Each complex number in the array represents a frequency component in the audio signal, where the magnitude of the complex number represents an amplitude of a corresponding frequency component, and the phase of the complex number represents a phase shift of the corresponding component relative to a reference. In other such embodiments, MFCCs can be used (e.g., or the filter bank energies can be converted to MFCCs) for a more compact representation, but such representations may not yield sufficient accuracy for all applications (e.g., the filter resolution may effectively be too low for successful training of the foundation model 1750). In other such embodiments, sample-to-sample modeling can be used, but such implementations may be too bulky to be practical in all applications.

[0123] As illustrated, during training, a foundation model 1750 takes the spoken feature vectors 1742 as its input labels and the unified feature vector 1744 as its output label. The foundation model 1750 is iteratively trained based on the input and output labels, so that inputting any of the spoken feature vectors 1742 for a particular class into the foundation model 1750 will cause the foundation model 350 to output the same unified feature vector 1744 for that class. In some embodiments, the feature extractor 1740 is integrated with the foundation model 1750 so that the class spoken audio samples can be input to the foundation model 1750, and the foundation model 1750 will try to generate an output unified audio sample that mimics the class synthesized audio sample (or to output a feature vector that mimics the unified feature vector 1744).

[0124] As one example, each audio sample at the input to the feature extractor 1740 is one second of audio sampled at 16 kHz (i.e., 16,000 samples). The sample is split into 100 frames of ten milliseconds each. The feature extractor 1740 extracts 40 Mel frequency bins (e.g., using bandpass filters tuned to the Mel scale) from each frame, thereby compressing each audio input of 16,000 samples into 40 * 100, or 4,000, filter bank energies. In this example, the feature extractor 1740 has an input dimension of 16,000 and an output dimension of 4,000. Suppose the repository 1710 has 10,000 repository spoken audio samples for a particular word, and the augmenter 1720generates ten augmentations for each repository spoken audio sample, resulting in 100,000 class spoken audio samples. The feature extractor 1740 converts each of the 100,000 class spoken audio samples into a respective 4,000 filter bank energies (i.e., the spoken feature vectors 1742), and converts the single class unified audio sample into another respective 3,000 filter bank energies (i.e., the unified feature vector 1744); and the foundation model 1750 uses all those filter bank energies to figure out how to convert any of the spoken feature vectors 1742 into the same unified feature vector 1744.

[0125] Ultimately, the foundation model 1750 is trained so that inputting any input audio sample 1760 will cause the foundation model 1750 to generate an output unified feature vector 1765 that is stripped of any speaker influence on suprasegmental features. For example, the generated output unified feature vector 1765 mimics the unified feature vector 1744 that would be generated if the input audio sample were to represent a class from the repository 1710, class text corresponding to the class were passed through the TTS synthesizer 1735 to generate a class synthesized audio sample, and the class synthesized audio sample were passed through the feature extractor 1740 to generate the unified feature vector 1744. In other words, the foundation model 1750 is trained to convert any input audio sample 1760 into a unified sound representation.

[0126] FIG.18 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block 1800 and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein. The AS trigger detection block 1800 can be implemented in a similar manner to the AS trigger detection block 400 of FIG.4, and consistent reference designators are used to indicate such similarities. For example, both include an enrollment stage 410 that serves the same overall purpose with respect to the system and has is shown with the same reference designator, accordingly; but each enrollment stage 410 yields a different type of embedding model, and the embedding models are shown with different reference designators, accordingly. Like the AS trigger detection block 400 of FIG.4, the AS trigger detection block 1800 of FIG.18 can be an implementation of the AS trigger detection block 210 of FIG.2. It is assumed that the AS trigger detection block 400 is implemented in the context of a WAC 310.

[0127] Embodiments can include an enrollment stage 410, an identification stage 430, and a verification stage 440. In general, the enrollment stage 410 occurs outside of normal operation of the WAC 310, such as when a user first sets up and / or uses the WAC 310, when the user first registers the WAC 310, when the user first configures the WAC 310 for automated attention handling, etc. The identification stage 430 and the verification stage 440 are real-time blocks thatoccur during normal operation of the WAC 310 to facilitate real-time automated attention handling. Embodiments generally use the enrollment stage 410 to obtain any information needed to set up an inference model for storing at the WAC 310 to enable the WAC 310 to perform automated attention handling features during normal operation. As shown, the inference model includes at least an embedding model 1815, a relation network 1835, and a false rejection network 1845.

[0128] In the enrollment stage 410, the embedding model 1815 can be used to convert an enrollment audio stream 405 into a set of reference embeddings based on a foundation model 1750. The foundation model 1750 can be stored in and accessed via the cloud 340 (e.g., and / or any other suitable communication network). As described with reference to FIG.17, in the context of USC-based embodiments, the foundation model 1750 can be trained to convert an input audio stream into either a unified output audio sample (i.e., corresponding to the input audio stream stripped of any speaker influence on accent, prosody, etc.), or an output unified feature vector 1765 (e.g., a set of filter bank energies that describe the unified output audio sample). In this context, the embedding model 1815 is trained by transfer learning from the foundation model 1750 on a smaller corpus of name audio samples. For example, the encoder portion of the foundation model 1750 is taken by the embedding model 1815, and a new decoder is trained on the smaller name audio corpus.

[0129] As described above (e.g., with reference to FIGS.4, 5A, 5B, 6, and 8), during the enrollment stage 410, the enrollment audio stream 405 can include several user-enrolled audio samples of invocation names and / or a set of augmentations to those invocation names. In this context, the “reference embeddings” refer to either the set of unified output audio samples generated from the set of invocation names (e.g., and their augmentations), or the set of output unified feature vectors 1765 generated from the set of invocation names (e.g., and their augmentations). The reference embeddings can be stored as a deep image 1820. The deep image 1820 can also be considered as part of the inference model.

[0130] During normal operation, in the identification stage 430, a real-time (RT) audio stream 407 is received. The embedding model 1815 (i.e., the same embedding model 1815 generated in the enrollment stage 410) is used to generate a RT embedding from the received RT audio stream 407. In this context, the “RT embedding” refers to either a unified output audio sample generated from the RT audio stream 407, or an output unified feature vector 1765 generated from the RT audio stream 407. The relation network 1835 can then compare the RT embedding with each of the reference embeddings in the deep image 1820 to determine if there is a match. For example,the relation network 1835 is configured to compute a similarity score for each comparison (e.g., corresponding to a mathematical correlation, or the like). The relation network 1835 can determine if any similarity scores meet or exceed a predetermined threshold; if so, the reference embedding with the highest similarity score can be selected as a candidate matching embedding. Further options and / or features of the relation network 1835 can be understood with reference to descriptions of the relation network 435 above.

[0131] In the verification stage, the false rejection network 1845 can then confirm the match by transforming the RT embedding and the candidate matching embedding into different mathematical spaces (e.g., domains) and determining whether the embeddings can be reliably discriminated from each other in any of the mathematical spaces. For example, the deep image 1820 is configured to compute a discrimination score for each mathematical space. The deep image 1820 can determine if any discrimination scores meet or exceed a predetermined threshold; if so, the match determined by the relation network 1835 is determined to be a false match and is ignored. If the embeddings cannot be sufficiently discriminated in any mathematical spaces, the false rejection network 1845 can output a signal that attention seeking audio has been detected (i.e., a name invoked signal 450). As described herein, the name invoked signal 450 can direct an ANC system of the WAC 310 automatically to enter a conversation mode. As noted with reference to FIG.4 above, embodiments of the embedding model 1815 are a processor-executable name embedding model, embodiments of the deep image 1820 are a processor-readable deep image, embodiments of the relation network 1835 are a processor-executable relation network, and embodiments of the false rejection network 1845 are a processor-executable false rejection network. Further options and / or features of the false rejection network 1845 can be understood with reference to descriptions of the false rejection network 445 above.

[0132] As described above, some USC-based embodiments implement the embedding model 1815 to include both an encoder and a decoder, so that the output of the embedding model 1815 is the output unified feature vector 1765 (e.g., synthesized filter bank energies). As illustrated in FIGS.17A and 17B, in other embodiments, the embedding model 1815 is implemented only to include the encoder portion of the model (i.e., all layers of the model after the bottleneck layer are removed after training). As described above, encoder-decoder architectures (e.g., auto-encoder architectures) include a neural network architecture designed to learn compact representations of data, referred to as “latent features.” The encoder portion receives a high-dimensionality input (i.e., the raw audio data) and includes a multi-layer network to progressively reduce the dimensionality of the data using transformations, until a lowest dimensionality (i.e., most compressed) representation is reached at the bottleneck layer. The vector space that is representedat the bottleneck layer can be referred to as a latent space embedding of the audio sample. In some embodiments, the latent space embedding generated by the encoder (e.g., at the bottleneck layer) is used as a unified sound code word (USCW) 1767.

[0133] In such embodiments, the embedding model 1815 is trained to generate respective USCWs 1767 as the reference embeddings stored in the deep image 1820 and to generate a respective USCW 1767 as the RT embedding, the relation network 1835 is trained to determine candidate matches based on comparing the USCWs 1767, and the false rejection network 1845 is trained to reject false matches by discriminating between the USCWs 1767. With proper training and implementation, the performance of the AS trigger detection block 1800 can be almost the same for embodiments of the inference model that include a decoder and are based on output unified feature vectors 1765 as for embodiments of the inference model that do not include a decoder and are based on UCSWs 1767.

[0134] FIG.19 shows a flow diagram of an illustrative method 1900 for automated audio management based on name detection, according to USC-based embodiments described herein. Embodiments of the method 1900 begin at stage 1904 by obtaining (e.g., by one or more processors) spoken audio samples for each word of a large number of phonetically diversified words from a spoken word audio repository. As described herein, the spoken word audio repository includes one or more speech-audio corpuses of suprasegmentally diversified speech- audio samples of the phonetically diversified words. As such, the spoken audio samples for each word of the phonetically diversified words includes those of the suprasegmentally diversified speech-audio samples representing respective instances of the word.

[0135] At stage 1908, embodiments can obtain (e.g., by the one or more processors) a respective text representation of each word of the phonetically diversified words. For some or all the phonetically diversified words, some embodiments can convert one or more of the class spoken audio samples associated with the word into the respective text representation of the word. For example, a class spoken audio sample is speech-to-text converted in stage 1906 to generate the textual representation. In some implementations, for each of at least some of the phonetically diversified words, the spoken word audio repository includes a lexical entry for the word (e.g., stored in association with at least one of the class spoken audio samples for the word). In such implementations, embodiments of the method 1900 can obtain the respective text representation of each of some or all the phonetically diversified words in stage 1908 by obtaining the lexical entry for the word from the spoken word audio repository.

[0136] At stage 1912, embodiments can synthesize (e.g., by the one or more processors) a respective class unified audio sample for each word from a respective text representation of the word. At stage 1916, embodiments can convert, for each word (e.g., by the one or more processors), the respective class unified audio sample for the word into respective output labels representing a unified feature vector of the respective class unified audio sample. At stage 1920, embodiments can convert, for each word (e.g., by the one or more processors), the class spoken audio samples associated with the word into respective input labels representing spoken feature vectors of the class spoken audio samples. In some embodiments, the spoken feature vectors represent filter bank energies of the plurality of class spoken audio samples, and the unified feature vector represents filter bank energies of the respective class unified audio sample.

[0137] Some embodiments, at stage 1918, prior to the converting in stage 1920, can add class augmented audio samples to the class spoken audio samples associated with the word by applying one or more augmenter models to one or more of the class spoken audio samples from the spoken word audio repository (i.e., the class spoken audio samples for the word include both the class spoken audio samples from the spoken word audio repository and the class augmented audio samples). For example, each of the augmenter models mathematically represents a respective one of several noise models and / or suprasegmental feature models (e.g., models of different accents, prosody, etc.).

[0138] At stage 1924, embodiments can train a foundation model. As described herein, the foundation model is an artificial neural network trained, based on the input labels and the output labels for each word, to convert any particular one of the class spoken audio samples for the word into the respective class unified audio sample for the word by automatically removing suprasegmental differences between the particular class spoken audio sample and the class unified audio sample. For example, once properly trained, the foundation model is able to remove a speaker’s suprasegmental influence on spoken audio received from the speaker as the input stream, thereby generating a unified version of the audio that mimics what would be generated by text-to- speech synthesis.

[0139] FIG.20 shows a flow diagram of an illustrative method 2000 for training a unified sound conversion (USC) inference model. As represented by off-page reference “C,” the method 2000 can proceed based on and subsequent to the training of the foundation model in stage 1924 of FIG. 19. At stage 2004, embodiments train the USC inference model (e.g., a processor-executable USC inference model) based on transfer learning from the foundation model. The training is such that the USC inference model can automatically: receive a real-time audio stream; generate a real-timeembedding representing the real-time audio stream stripped of suprasegmental features; and output a name invoked signal based on matching, with at least a predetermined threshold confidence level, the real-time embedding to one of a stored plurality of reference name embeddings generated to represent unified audio representations of each of a set of invocation names stripped of the suprasegmental features, the name invoked signal to direct an ANC system automatically to enter a conversation mode.

[0140] In some embodiments, the training in stage 2004 can proceed according to stages 2008 – 2016. At stage 2008, embodiments can train a name embedding model, by transfer learning from the foundation model based on a corpus of real-world name audio samples, automatically to remove the suprasegmental features from an input audio sample to generate a corresponding unified audio output. In such embodiments, the real-time embedding is generated by the name embedding model from the real-time audio signal received during an operational time, and the reference name embeddings are generated by the name embedding model based on enrollment audio samples received during an enrollment time.

[0141] At stage 2012, embodiments can train a relation network to output a candidate name embedding responsive to determining that one of the reference name embeddings has a highest similarity with the real-time embedding and that the highest similarity exceeds a predetermined similarity threshold. For example, the relation network is trained to output the candidate name by: computing a similarity score between each of the reference name embeddings and the real-time embedding, the similarity score indicating a mathematical correlation; determining whether the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold; and outputting the candidate name embedding as the one of the reference name embeddings yielding a highest similarity score when it is determined that the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold.

[0142] At stage 2016, embodiments can train a false rejection network to output the name invoked signal responsive to determining that the real-time embedding and the candidate name embedding cannot be discriminated in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces. For example, the false rejection network is trained to output the name invoked signal by: transforming the real-time embedding and the candidate name embedding into a plurality of mathematical spaces; computing, for each mathematical space, a discrimination score indicating a likelihood of discriminating between the real-time embedding and the candidate name embedding in the mathematical space; and outputting the name invokedsignal responsive to the discrimination score computed for at least one of the mathematical spaces falling below the predetermined discrimination threshold. In embodiments that perform stages 2012 and 2016, the predetermined confidence level in stage 2004 can be based on a combination of the predetermined similarity threshold and the predetermined discrimination threshold.

[0143] FIG.21 shows a flow diagram of an illustrative method 2100 for using the trained unified sound conversion (USC) inference model during an operation time. As represented by off- page reference “D,” the method 2100 can proceed based on and subsequent to the training of the USC inference model in stage 2004 of FIG.20. Embodiments of the method 2100 begin at stage 2104 by receiving a real-time audio signal by the USC inference model (e.g., from a reference microphone associated with the ANC system). At stage 2108, embodiments can generate a real- time embedding from the real-time audio signal by the USC inference model. At stage 2112, embodiments can output the name invoked signal in response to (e.g., only after determining) determining (e.g., only after determining), by the USC inference model, that the real-time embedding matches any of the stored plurality of reference name embeddings with at least the predetermined threshold confidence level.

[0144] Similar to the method 1400 of FIG.14, embodiments of the method 2100 can include additional preceding and / or subsequent stages, such an indicated by off-page references “A” and “B”. As described above, off-page reference “A” optionally references a preceding enrollment phase. An example of such an enrollment phase is described with reference to FIG.15. Also as described above, off-page reference “B” optionally references a subsequent conversation end detection phase. An example of such a conversation end detection phase is described with reference to FIG.16.

[0145] FIG.22 provides a schematic illustration of an illustrative computational system 2200 that can implement various system components and / or perform various steps of methods provided by various embodiments. Embodiments of the computational system 2200 can be integrated in a WAC, such as an earbud, headset, etc. Embodiments of the computational system 2200 can implement some or all of the audio management system 100 of FIG.1, including embodiments of the AHS 150, the ANC 140, and / or the audio processing system (APS) 160 described herein. FIG. 22 is meant only to provide a generalized illustration of various components, any or all of which may be utilized as appropriate. FIG.22, therefore, broadly illustrates how individual system elements may be implemented in a relatively separated or relatively more integrated manner.

[0146] The computational system 2200 is shown including hardware elements that can be electrically coupled via a bus 2205 (or may otherwise be in communication, as appropriate). Thehardware elements may include one or more processors 2210, including, without limitation, one or more general-purpose processors and / or one or more special-purpose processors (such as digital signal processing chips, graphics acceleration processors, video decoders, and / or the like); one or more input devices 2215; and one or more output devices 2220. In the WAC context, the input devices 2215 can include wired and / or wireless ports, buttons, switches, microphones, touch interfaces, and / or any other suitable input device 2215; and the output devices 2220 can include indicator lights, displays, speakers, and / or any other suitable output devices 2220.

[0147] The computational system 2200 may further include (and / or be in communication with) one or more non-transitory storage devices 2225, which can comprise, without limitation, local and / or network accessible storage, and / or can include, without limitation, a disk drive, a drive array, an optical storage device, a solid-state storage device, such as a random access memory (“RAM”), and / or a read-only memory (“ROM”), which can be programmable, flash-updateable and / or the like. Such storage devices may be configured to implement any appropriate data stores, including, without limitation, various file systems, database structures, and / or the like. In some embodiments, the storage devices 2225 include the deep image 420 and / or an inference model 2227. As described herein, the inference model can include one or more types of name embedding models, relation networks, false rejection networks, etc. for implementing name detection-based attention handling.

[0148] The computational system 2200 can also include a communications subsystem 2230, which can include, without limitation, a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth^ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication device, etc.), and / or the like. As described herein, the communications subsystem 2230 supports multiple communication technologies. Further, as described herein, the communications subsystem 2230 can provide communications with one or more networks 140, and / or other networks. For example, embodiments of the communications subsystem 2230 can communicate with a foundation model 350 via the cloud 350. Though not explicitly shown, some embodiments interface via the communications subsystem 2230, and / or via input devices 2215 and output devices 2220, with one or more user computational devices 320.

[0149] In many embodiments, the computational system 2200 will further include a working memory 2235, which can include a RAM or ROM device, as described herein. The computational system 2200 also can include software elements, shown as currently being located within the working memory 2235, including an operating system 2240, device drivers, executable libraries,and / or other code, such as one or more application programs 2245, which may include computer programs provided by various embodiments, and / or may be designed to implement methods, and / or configure systems, provided by other embodiments, as described herein. Merely by way of example, one or more procedures described with respect to the method(s) discussed herein can be implemented as code and / or instructions executable by a computer (and / or a processor within a computer); in an aspect, then, such code and / or instructions can be used to configure and / or adapt a general-purpose computer (or other device) to perform one or more operations in accordance with the described methods. In some embodiments, the operating system 2240 and the working memory 2235 are used in conjunction with the one or more processors 2210 to implement some or all of the audio management system 100 components, such as the ANC 140, AHS 150, and / or APS 160.

[0150] A set of these instructions and / or codes can be stored on a non-transitory computer- readable storage medium, such as the non-transitory storage device(s) 2225 described above. In some cases, the storage medium can be incorporated within a computer system, such as computer system 2200. In other embodiments, the storage medium can be separate from a computer system (e.g., a removable medium, such as a compact disc), and / or provided in an installation package, such that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the instructions / code stored thereon. These instructions can take the form of executable code, which is executable by the computational system 2200 and / or can take the form of source and / or installable code, which, upon compilation and / or installation on the computational system 2200 (e.g., using any of a variety of generally available compilers, installation programs, compression / decompression utilities, etc.), then takes the form of executable code.

[0151] It will be apparent to those skilled in the art that substantial variations may be made in accordance with specific requirements. For example, customized hardware can also be used, and / or particular elements can be implemented in hardware, software (including portable software, such as applets, etc.), or both. Further, connection to other computing devices, such as network input / output devices, may be employed.

[0152] As mentioned above, in one aspect, some embodiments may employ a computer system (such as the computer system 2200) to perform methods in accordance with various embodiments of the invention. According to a set of embodiments, some or all of the procedures of such methods are performed by the computational system 2200 in response to processor 2210 executing one or more sequences of one or more instructions (which can be incorporated into the operating system 2240 and / or other code, such as an application program 2245) contained in the workingmemory 2235. Such instructions may be read into the working memory 2235 from another computer-readable medium, such as one or more of the non-transitory storage device(s) 2225. Merely by way of example, execution of the sequences of instructions contained in the working memory 2235 can cause the processor(s) 2210 to perform one or more procedures of the methods described herein.

[0153] The terms “machine-readable medium,” “computer-readable storage medium” and “computer-readable medium,” as used herein, refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. These mediums may be non-transitory. In an embodiment implemented using the computer system 2200, various computer-readable media can be involved in providing instructions / code to processor(s) 2210 for execution and / or can be used to store and / or carry such instructions / code. In many implementations, a computer- readable medium is a physical and / or tangible storage medium. Such a medium may take the form of a non-volatile media or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as the non-transitory storage device(s) 2225. Volatile media include, without limitation, dynamic memory, such as the working memory 2235. Common forms of physical and / or tangible computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, any other physical medium with patterns of marks, a RAM, a PROM, EPROM, a FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.

[0154] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to the processor(s) 2210 for execution. Merely by way of example, the instructions may initially be carried on a magnetic disk and / or optical disc of a remote computer. A remote computer can load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received and / or executed by the computer system 2200. The communications subsystem 2230 (and / or components thereof) generally will receive signals, and the bus 2205 then can carry the signals (and / or the data, instructions, etc., carried by the signals) to the working memory 2235, from which the processor(s) 2210 retrieves and executes the instructions. The instructions received by the working memory 2235 may optionally be stored on a non-transitory storage device 2225 either before or after execution by the processor(s) 2210.

[0155] Having described several example configurations, various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the disclosure. Forexample, the above elements may be components of a larger system, wherein other rules may take precedence over or otherwise modify the application of the invention. Also, a number of steps may be undertaken before, during, or after the above elements are considered.

Claims

WHAT IS CLAIMED IS:

1. An artificial neural network (ANN) topology for universal sound conversion (USC) comprising: a Mel converter to convert an input spoken word audio (SWA) sample into a plurality of input Mel filter bank energies; a convolutional encoder coupled with the Mel converter to compress the Mel filter bank energies into an encoded representation according to a set of encoder parameters; an axial self-attention (ASA) block coupled with the convolutional encoder to further compress the encoded representation into a bottleneck embedding according to a set of ASA parameters; and a deconvolutional decoder coupled with the ASA block and coupled with the convolutional encoder via a skip channel, the deconvolutional decoder to decompress the bottleneck embedding into an estimated set of output Mel filter bank energies according to a set of decoder parameters, wherein the encoder parameters, the ASA parameters, and the decoder parameters are trained so that the output Mel filter bank energies represent linguistic features of the SWA sample absent speaker-specific suprasegmental features.

2. The ANN topology of claim 1, wherein the convolutional encoder comprises X encoder blocks, and the deconvolutional decoder comprises X decoder blocks, wherein X is a positive integer.

3. The ANN topology of claim 2, wherein X is 4.

4. The ANN topology of claim 2, wherein: each encoder block comprises a two-dimensional convolution (Conv2D) block, a batch normalization (BN) block, and a leaky rectified linear unit (ReLU) block; and each decoder block comprises a transposed Conv2D, a BN block, and a leaky ReLU block.

5. The ANN topology of claim 2, wherein: the deconvolutional decoder further comprises X merging layers; and each merging layer generates an input for a corresponding one of the X decoder blocks based on merging an output from a directly preceding block with an output from a corresponding one of the X encoder blocks received via a corresponding skip connection.

6. The ANN topology of claim 1, wherein the ASA block comprises: a first convolution stage configured to compute a first query vector, a first key vector, and a first value vector based on the encoded representation, and to compute attention scores and a first attention output vector based on the first query vector, the first key vector, and the first value vector; a second convolution stage configured to compute a second query vector and a second key vector based on fusing the encoded representation with the attention scores, and to compute a second attention output vector based on the second query vector, the second key vector, and the first attention output vector; and a third convolution stage to compute the bottleneck embedding based on the second attention output vector.

7. The ANN topology of claim 6, wherein: the first convolution stage comprises: a first convolution block to generate a first set of feature maps from the encoded representation; a first set of dense blocks, each to generate a respective one of the first query vector, the first value vector, and the first key vector based on the first set of feature maps; and a first attention block to generate the attention scores and the first attention output vector based on the first query vector, the first key vector, and the first value vector; the second convolution stage comprises: a fuser to fuse the encoded representation with the attention scores; a second convolution block to generate a second set of feature maps from the output of the fuser; a second set of dense blocks, each to generate a respective one of the second query vector and the second key vector based on the second set of feature maps; and a second attention block to generate the second attention output vector based on the second query vector, the second key vector, and the first attention output vector; and the third convolution stage comprises a third convolution block to compute the bottleneck embedding based on the second attention output vector.

8. The ANN topology of claim 7, wherein: each convolution block comprises a two-dimensional convolution (Conv2D) block, a batch normalization (BN) block, and a leaky rectified linear unit (ReLU) block.

9. The ANN topology of claim 6, wherein: the encoded representation has dimensions [T x F x C], wherein T represents a frame dimension, F represents a feature dimension, and C represents a channel dimension; the second convolution stage is further configured to reshape the dimensions to [F x T x C] prior to computing the second query vector and a second key vector; and the third convolution stage is further configured to reshape the dimensions back to [T x F x C] prior to computing the bottleneck embedding.

10. A method for training an artificial neural network (ANN) for universal sound conversion (USC), the method comprising: for each of a plurality of training epochs, until an error is below a predetermined threshold: generating a first set of Mel filter bank energies from a spoken word audio (SWA) sample selected for the training epoch from a SWA repository, the SWA sample being a spoken representation of a word corresponding to a class of a plurality of classes; generating a second set of Mel filter bank energies from a USC sample selected for the training epoch from a USC repository, the USC sample selected as a unified audio representation of the class absent speaker-specific suprasegmental features; generating an estimated set of Mel filter bank energies for the training epoch from the first set of Mel filter bank energies using an ANN topology comprising a convolutional encoder, an axial self-attention (ASA) block, and a deconvolutional decoder; computing the error between the estimated set of Mel filter bank energies and the second set of Mel filter bank energies; and back-propagating the error to the ANN topology to reduce the error.

11. The method of claim 10, wherein the convolutional encoder comprises X encoder blocks, and the deconvolutional decoder comprises X decoder blocks, wherein X is a positive integer.

12. The method of claim 11, wherein: each encoder block comprises a two-dimensional convolution (Conv2D) block, a batch normalization (BN) block, and a leaky rectified linear unit (ReLU) block; and each decoder block comprises a transposed Conv2D, a BN block, and a leaky ReLU block.

13. The method of claim 10, wherein the generating the estimated set of Mel filter bank energies comprises: generating one or more encoded representations by compressing the first set of Mel filter bank energies using the convolutional encoder according to a set of encoder parameters; generating a bottleneck embedding by compressing the one or more encoded representations using the ASA block according to a set of ASA parameters; and generating the estimated set of Mel filter bank energies by using the deconvolutional decoder to decompress a combination of the bottleneck embedding and the one or more encoded representations according to a set of decoder parameters, wherein back-propagating the error comprises tuning at least one of the encoder parameters, the ASA parameters, or the decoder parameters to reduce the error.

14. The method claim 13, wherein the generating the bottleneck embedding comprises: computing a first query vector, a first key vector, and a first value vector based on the encoded representation; computing attention scores and a first attention output vector based on the first query vector, the first key vector, and the first value vector; computing a second query vector and a second key vector based on fusing the encoded representation with the attention scores; computing a second attention output vector based on the second query vector, the second key vector, and the first attention output vector; and computing the bottleneck embedding based on the second attention output vector.

15. The method of claim 14, wherein the encoded representation has dimensions [T x F x C], wherein T represents a frame dimension, F represents a feature dimension, and C represents a channel dimension, and further comprising: reshaping the dimensions to [F x T x C] prior to computing the second query vector and a second key vector; and reshaping the dimensions back to [T x F x C] prior to computing the bottleneck embedding.

16. A computer program comprising instructions for implementing a method of claim 10.

17. A system comprising: one or more processors; a non-transitory memory having instructions stored thereon which, when executed, cause the one or more processors to train an artificial neural network (ANN) for universal sound conversion (USC) by performing steps comprising, for each of a plurality of training epochs, until an error is below a predetermined threshold: generating a first set of Mel filter bank energies from a spoken word audio (SWA) sample selected for the training epoch from a SWA repository, the SWA sample being a spoken representation of a word corresponding to a class of a plurality of classes; generating a second set of Mel filter bank energies from a USC sample selected for the training epoch from a USC repository, the USC sample selected as a unified audio representation of the class absent speaker-specific suprasegmental features; generating an estimated set of Mel filter bank energies for the training epoch from the first set of Mel filter bank energies using an ANN topology comprising a convolutional encoder, an axial self-attention (ASA) block, and a deconvolutional decoder; computing the error between the estimated set of Mel filter bank energies and the second set of Mel filter bank energies; and back-propagating the error to the ANN topology to reduce the error.

18. The system of claim 17, wherein the generating the estimated set of Mel filter bank energies comprises: generating one or more encoded representations by compressing the first set of Mel filter bank energies using the convolutional encoder according to a set of encoder parameters; generating a bottleneck embedding by compressing the one or more encoded representations using the ASA block according to a set of ASA parameters; and generating the estimated set of Mel filter bank energies by using the deconvolutional decoder to decompress a combination of the bottleneck embedding and the one or more encoded representations according to a set of decoder parameters, wherein back-propagating the error comprises tuning at least one of the encoder parameters, the ASA parameters, or the decoder parameters to reduce the error.

19. The system of claim 18, wherein the generating the bottleneck embedding comprises: computing a first query vector, a first key vector, and a first value vector based on the encoded representation;computing attention scores and a first attention output vector based on the first query vector, the first key vector, and the first value vector; computing a second query vector and a second key vector based on fusing the encoded representation with the attention scores; computing a second attention output vector based on the second query vector, the second key vector, and the first attention output vector; and computing the bottleneck embedding based on the second attention output vector.