Name detection against environmental interferences using a progressive learning
A progressive learning-based machine learning network in ANC systems detects user-specific invocation names, automatically switching to conversation mode, addressing the challenge of ambient sound distinction and improving user engagement and comfort.
Patent Information
- Application Number
- PCT/US2024/045118
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-10
- Filing Date
- 2024-09-04
- Publication Date
- 2026-01-15
AI Technical Summary
Conventional active noise control systems in wearable audio components struggle to distinguish ambient sound intended for the user from other noise, making it difficult for attention seekers to engage users in conversations, requiring manual intervention by the user to disable ANC and remove the device.
Implementing a progressive learning-based machine learning network for name detection and automated attention handling, which trains a robust name embedding model to identify and respond to user-specific invocation names amidst noise and competing speech, automatically switching the ANC system to conversation mode.
Enables users to remain aware of ambient conversations while wearing ANC-equipped devices, improving user comfort and engagement without manual intervention, enhancing the user experience and ear health.
Smart Images

Figure US2024045118_15012026_PF_FP_ABST
Abstract
Description
NAME DETECTION AGAINST ENVIRONMENTAL INTERFERENCES USING A PROGRESSIVE LEARNINGCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of and priority to Indian Provisional Application No. 202441052763, filed on July 10, 2024, and titled “NAME DETECTION AGAINST ENVIRONMENTAL INTERFERENCES USING A PROGRESSIVE LEARNING,” the content of which is herein incorporated by reference in its entirety for all purposes.BACKGROUND
[0002] Active noise control (ANC) is a common feature of headsets and earbuds. It operates by generating an anti-noise signal via a speaker that is approximately equal in magnitude, but opposite in phase to the ambient sound (e.g., ambient noise and other sounds in the vicinity). The ambient sound and anti-noise signal cancel each other acoustically, allowing the user to hear only a desired audio signal. Typically, signal processing in ANC includes two paths: an ambient sound signal from a reference microphone is taken as the input of a feed-forward ANC filter (FFANC); and an error microphone signal is taken as the input of a feedback ANC filter (FBANC).
[0003] When a user is listening to music or other desired audio through a wearable audio component (e.g., earbuds or on-ear headphones), ANC works to suppress any ambient sound. However, in some instances, ambient sound that is intended for the user can be very important to the user’s connectivity with others. For example, while a user is listening to music with ANC, it may be very difficult for an attention seeker to get the user’s attention, such as to alert the user of something important and / or to engage the user in a desired conversation. Conventionally, in such instances, the attention seeker gets the user’s attention by tapping the user on the shoulder, gesturing in front of the user’s face, or the like; and the user can then manually disable the ANC, pause the desired audio, and / or remove the wearable audio component.SUMMARY
[0004] Systems and methods are described herein for progressive training of a machine learning network for name embedding. Embodiments seek to refine embedding models so that automated name detection, automated attention handling, and other similar features can be applied to active noise control systems in a manner that is robust to noise and competing speech. Embodiments begin by training a foundation model based on a name detection paradigm. Progressive training is used, based initially on the foundation model, to teach progressive machine learning networks to generate unified embeddings for each of multiple linguistic classes robustly in the presence of noise and / or competing speech. Those networks are ultimately used to train a robust nameembedding (RNE) model to produce target outputs (e.g., classifications, acoustic segments, etc.) according to the name detection paradigm.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] A further understanding of the nature and advantages of various embodiments may be realized by reference to the following figures. In the appended figures, similar components or features may have the same reference label. Further, various components of the same type may be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.
[0006] FIG. 1 shows an audio management system for integration in a wearable audio component (WAC), according to embodiments described herein.
[0007] FIG. 2 shows a conceptual circuit block diagram of a partial audio management system environment with an automated attention handling system (AHS).
[0008] FIGS. 3A and 3B show a wearable audio environment including a pair of WACs.
[0009] FIG. 4 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein.
[0010] FIG. 5 shows a training environment for implementing a first phase of training a knowledge distillation model (KDM).
[0011] FIG. 6 shows an example of a posterior probability matrix (PPM).
[0012] FIG. 7 shows a training environment for implementing a second phase of training the KDM.
[0013] FIG. 8 shows an illustrative candidate segmentation and an illustrative corresponding PPM and ordered acoustic segmentation vector (OASV).
[0014] FIGS. 9A and 9B show example OASVs resulting from an illustrative automated orthosegmentation and an illustrative re-segmentation, respectively.
[0015] FIGS. 10A and 10B show block diagrams of illustrative uses of the name embedding model to generate the deep image.
[0016] FIG. 11 shows several example screenshots from an example enrollment application running on a user device.
[0017] FIG. 12 shows a flow diagram of an illustrative method for audio management that includes automated attention handling in a WAC, according to embodiments described herein.
[0018] FIG. 13 shows a flow diagram of an illustrative method for an enrollment phase.
[0019] FIG. 14 shows a flow diagram of an illustrative method for conversation end detection.
[0020] FIG. 15 shows a flow diagram of an illustrative method for training an automated acoustic segmentation (AAS) system for use with embodiments described herein.
[0021] FIG. 16 shows a flow diagram of an illustrative method for automated acoustic segmentation-based attention handling in a wearable audio component (WAC), according to embodiments described herein.
[0022] FIG. 17 shows a flow diagram of an illustrative method for audio management that includes automated attention handling in a WAC using linguistic name embedding (LNE) techniques, according to embodiments described herein.
[0023] FIGS. 18A and 18B show block diagrams of a training environment for training a foundation model to support unified sound conversion (USC) for automated name-detection-based attention handling, according to some embodiments described herein.
[0024] FIG. 19 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein.
[0025] FIG. 20 shows a flow diagram of an illustrative method for automated audio management based on name detection, according to USC-based embodiments described herein.
[0026] FIG. 21 shows a flow diagram of an illustrative method for training a USC inference model.
[0027] FIG. 22 shows a flow diagram of an illustrative method for using the trained USC inference model during an operation time.
[0028] FIG. 23 shows a simplified block diagram of an embedding environment incorporating an illustrative hybrid name embedding model for use in a unified linguistic name embedding attention handling system (ULNE-AHS).
[0029] FIG. 24 shows a schematic illustration of an illustrative computational system that can implement various system components and / or perform various steps of methods provided by various embodiments.
[0030] FIG. 25 shows a block diagram of a first refinement stage training environment for an embedding model, according to embodiments described herein.
[0031] FIG. 26 shows a block diagram of a second refinement stage training environment for an embedding model, according to embodiments described herein.
[0032] FIG. 27 shows a block diagram of a third refinement stage training environment for an embedding model, according to embodiments described herein.
[0033] FIG. 28 shows a block diagram of a final training environment for an embedding model, according to embodiments described herein.
[0034] FIG. 29 shows a flow diagram of an illustrative method for refinement of a name embedding model, according to embodiments described herein.DETAILED DESCRIPTION
[0035] When a user is listening to music or other desired audio through a wearable audio component (e.g., earbuds or on-ear headphones), active noise control (ANC) works to suppress any ambient sound. However, in some instances, ambient sound that is intended for the user can be very important to the user’s connectivity with others. For example, although the user desired to suppress undesirable ambient sound, the user may still desire to be able on occasion to enter into desired conversations. In general, two types of desired conversation can be considered: first-party- initiated; and second-party-initiated.
[0036] In first-party-initiated conversations, the user desires to start a conversation and may begin by trying to get someone’s attention. In such cases, some conventional ANC systems are adapted to detect that the user has begun speaking (e.g., by detecting the user’s speech via a beamforming microphone directed to the user’s mouth, accelerometer, or combination thereof), and the ANC system can turn off, switch to transparency mode, pause audio playback, etc. in response to detecting the user’s speaking. Because it tends to be relatively easy for the ANC system to distinguish the user’s own speech from ambient sound, such approaches tend to be effective for first-party-initiated conversations.
[0037] In second-party-initiated conversations, however, a second-party attention seeker is trying to get the user’s attention, and the attention seeker’s voice may be difficult to distinguish from other ambient sound. For example, while a user is listening to music with ANC, it may bevery difficult for an attention seeker to get the user’s attention, such as to alert the user of something important and / or to engage the user in a desired conversation. Conventionally, in such instances, the attention seeker gets the user’s attention by tapping the user on the shoulder, gesturing in front of the user’s face, or the like. Before the user can participate in the conversation, the user conventionally must notice the interruption and then manually disable the ANC, pause the desired audio, remove the wearable audio component, etc.
[0038] Indeed, many users of wearable audio components enjoy the feeling of being in their “bubble” and the ability to focus on their media that comes with effective ANC. However, as ANC continues to improve, the same users often feel increasingly unaware, not present, and fearful about missing out. Embodiments described herein seek to provide users with the ability to better stay aware and engage in desired conversations, while being able to continue wearing their wearable audio components and otherwise to take advantage of ANC. This can provide several benefits, including helping to improve user comfort and ear health.
[0039] Embodiments described herein are concerned with second-party-initiated conversations.As used herein, the term “user” refers to a wearer of a wearable audio component (i.e., the first party). The term “attention seeker” is used herein generally to refer to any ambient party trying to get the user’s attention while the user is wearing the wearable audio component (and presumably is listening to desired audio with ANC turned on). Typically, the attention seeker is a person.However, the attention seeker can also be a computational platform with a deterministic manner of seeking the user’s attention, such as a smart speaker programmed to call out the user’s name. The term “wearable audio component,” or “WAC” is used herein to generally refer to earbuds, on-ear headphones, over-ear headphones, or any type of wearable audio output device that includes ANC. The term “desired audio” is used herein to generally refer to any recorded or streaming audio signal that is being played to the user through the WAC, such as music, an audiobook, a podcast, a radio broadcast, a live event broadcast, etc. The term “ambient sound,” or “ambient audio” is used herein to generally refer to any audio in the vicinity of the WAC, other than the desired audio. It is generally the goal of the ANC system to suppress as much of the ambient sound as possible.Audio originating from an attention seeker while a user’s ANC system is active is part of the ambient sound.
[0040] FIG. 1 shows an audio management system 100 for integration in a wearable audio component (WAC), according to embodiments described herein. As illustrated, the audio management system 100 can include an active noise control (ANC) system 140, an attention handling system (AHS) 150, and an audio processing system 160. In general, the purpose of theWAC is to deliver desired audio 165 to a user’s ear or ears via one or more ear speakers, such as speaker 105. Embodiments of the audio processing system 160 are designed to process the desired audio 165 for output to the user. For example, the audio processing system 160 can include amplifiers, filters, and / or other audio components; and / or any other suitable components for receiving, processing, and / or outputting the desired audio 165.
[0041] Typically, while listening to the desired audio 165, the user is also in presence of ambient audio 155. When in its ambient sound suppression mode, the ANC system 140 seeks to suppress as much of the ambient audio 155 as possible to enhance the user’s experience of listening to the desired audio 165. As illustrated, the ANC system 140 includes a feed-forward ANC (FFANC) filter 120, a feedback ANC (FBANC) filter 125, a summer 130, and an ANC output control block 135. The ANC system 140 is also coupled with the speaker 105 and at least a reference microphone 110 and an error microphone 115. Embodiments of the speaker 105 generally convert an electrical audio signal into sound waves that are delivered to the ear of the wearer of the wearable audio component. Embodiments of the reference microphone 110 can be an omnidirectional microphone typically integrated with an outer casing of the wearable audio component. The reference microphone 110 generally captures at least the ambient audio 155 around the WAC, which is delivered as a reference audio signal (illustrated as x(n)) to the FFANC filter 120. Embodiments of the error microphone 115 are typically integrated with the inner casing of the wearable audio component to be positioned inside the ear canal or very close to it when the wearable audio component is being worn. The error microphone 115 captures the audio that reaches the eardrum, which includes the desired audio signal and any remaining ambient sound after suppression. The error microphone 115 outputs an error signal (illustrated as e(ri)) to the FBANC filter 125.
[0042] The illustrated ANC system 100 includes a feed-forward noise control path and a feedback noise control path. The feed-forward noise control path includes the FFANC filter 120, which is a digital or analog filter designed to process the audio signal from the reference microphone 110. The FFANC filter 120 applies a specific frequency response to x(n) to adaptively cancel out noise. The specific frequency response is produced by continuously adjusting coefficients of the FFANC filter 120 to minimize the difference between the desired audio signal and the reference signal. The output of the FFANC filter 120 is illustrated as(n). The feedback noise control path includes the FBANC filter 125, which is a digital or analog filter designed to process the audio signal from the error microphone 115. The FBANC filter 125 applies a specific frequency response to e(n), and continuously adjusts coefficients of the FBANC filter 125 to minimize the difference between the desired audio signal and remaining ambientsound in the signal that reaches the eardrum. The output of the FFANC filter 120 is illustrated as y2(n). In general, both the FFANC filter 120 and the FBANC filter 125 can adapt their respective filters (e.g., their coefficients) in real-time to a changing audio environment. For example, filter coefficients are iteratively adjusted using least mean squares (LMS), normalized LMS (NLMS), and / or other suitable adaptation algorithms.
[0043] Embodiments of the summer 130 combine the filtered output signals from the FFANC filter 120 and the FBANC filter 125. For example, the summer 130 calculates a sum of these signals. If tuned properly, the output of the summer 130 is an “anti-noise” signal that closely represents the ambient sound at opposite polarity. Embodiments of the ANC output control block 135 control how and / or whether the anti-noise signal is output by ANC system 140. In some implementations, the ANC output control block 135 includes an amplifier to provide a controllable amount of gain (G) to the signal at the output of the summer 130, resulting in an output signal, y(n) = G y- n) + y2(n))- Ineffect, the ANC gain block 135 adjusts the overall amplitude (i.e., corresponding to volume) of the combined filtered signal at the output of the summer 130. The output signal is sent to the speaker 105. In some implementations, as illustrated, the desired audio 165 can also be mixed in (e.g., by mixer 145) prior to sending the output to the speaker 105, such that what reaches the eardrum is almost entirely the desired audio signal with minimal ambient sound. Alternatively, the desired audio 165 is mixed into the output signal at the summer 130, such that the output of the ANC system 140 is an audio signal that is mostly the desired audio 165 with minimal residual ambient audio 155.
[0044] Embodiments of the ANC output control block 135 control the operating mode of the ANC system 140. For example, as described herein, the ANC system 140 can operate selectively in at least an active mode (i.e., an ambient sound suppression mode) or a conversation mode. Some implementations of the conversation mode correspond to an inactive mode (i.e., the ANC system 140 is turned off) or a transparency mode. Other implementations of the conversation mode are configured to pass through conversationally relevant audio from the ambient audio 155, while continuing to perform ANC functions to suppress other portions of the ambient audio 155. In some such implementations, a bandpass or notch filter is used to segregate out a range of frequencies typical for human speech and to treat the segregated audio as conversationally relevant audio. As one example, a filter can pass through portions of the ambient audio 155 only in the range of 75 to 300 Hertz and to suppress higher and lower frequency components of the ambient audio 155; thereby continuing to filter out white noise and other portions of ambient audio 155 that can interfere with a user’s ability to hear the passed-through conversationally relevant audio.Similarly, some implementations continue to pass through some desired audio 165 (e.g., at a reduced volume) while in conversation mode.
[0045] As described herein, embodiments of the AHS system 150 seek to detect when a second- party attention seeker is trying to get the attention of a user while the user is wearing the WAC and is listening to desired audio 165 with the ANC system 140 in the active mode. In particular, the AHS system 150 listens for presence of attention seeking (AS) audio 157 within the ambient audio 155. Various techniques can be used by AHS systems 150 to detect presence of such AS audio 157 within the ambient audio 155 and to perform automated attention handling, accordingly. Some such approaches have been previously described by inventors of the present disclosure, for example, in International Patent Application No. PCT / US2024 / 014606, titled “AUTOMATED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS BASED ON LINGUISTIC NAME EMBEDDING,” filed on February 6, 2024; International Patent Application No. PCT / US2024 / 014788, titled “AUTOMATED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS BASED ON UNIVERSAL SOUND CONVERSION”, filed on February 7, 2024; International Patent Application No. PCT / US2024 / 014820, titled “NAMEDETECTION BASED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS”, filed on February 7, 2024; and Indian Provisional Patent Application No. 202441019923, titled “NAME-DETECTION BASED ATTENTION HANDLING IN ACTIVE NOISE CONTROL SYSTEMS BASED ON AUTOMATED ACOUSTIC SEGMENTATION”, filed on March 18, 2024.
[0046] Embodiments of the AHS system 150 described herein can be configured specifically to listen for AS audio 157 corresponding to a previously enrolled invocation name (e.g., a name of the user). When the AS audio 157 is detected by the AHS system 150, the AHS system 150 automatically directs the ANC system 140 to switch from the active mode to the conversation mode. In some embodiments, when the AS audio 157 is detected, the AHS system 150 also directs the audio processing system 160 to enter a conversation enhancement mode. As described above with reference to the ANC system 140, the conversation enhancement mode as implemented by the audio processing system 160 can include segregating conversationally relevant audio from the ambient audio 155, adapting equalization of passed through audio to enhance speech, muting or reducing the volume of playback of the desired audio 165, pausing playback of the desired audio 165, etc. Typically, in response to the AS audio 157, the user will begin to engage in a conversation with the attention seeker. Such a conversation can involve the user speaking, and embodiments of the conversation mode of the ANC system 140 and / or the conversation enhancement mode of the audio processing system 160 can include using techniques to helpensure that the user’s own speech is not fed back in a manner that results in an apparent echo, feedback noise, or the like. For example, the user’s own speech may be captured by a separate beamforming microphone as a user speech audio stream, while ambient audio 155 is being received by the reference microphone 110. The user speech audio stream can be subtracted from the ambient audio 155 prior to passing the signal through other blocks of the system, so that the fed-back audio stream includes only ambient audio other than the user’s own speech.
[0047] Some embodiments of the AHS system 150, after having detected AS audio 157 and directing the ANC system 140 into conversation mode, can further detect when the conversation ends. Such embodiments of the AHS system 150 can automatically direct the ANC system 140 to return to the active mode, accordingly. As part of returning to the active mode, some such embodiments also return settings (e.g., in the ANC system 140 and / or the audio processing system 160) to those appropriate for listening to the desired audio 165 and suppressing all of the ambient audio 155 (e.g., all frequencies of the ambient audio 155).
[0048] FIG. 2 shows a conceptual circuit block diagram of a partial audio management system environment 200 with an automated attention handling system (AHS) 150. The illustrated environment 200 can be an illustrative portion of the audio management system 100 of FIG. 1, and the AHS 150 can be an illustrative implementation of the AHS 150 of FIG. 1. As illustrated, the AHS 150 includes an attention seeking (AS) trigger detection block 210 and a conversation end detection block 220. Some implementations further include a conversation enhancement block 230. The AHS 150 is illustrated in context of a desired audio 165 stream, a reference microphone 110 that receives ambient audio 155 and outputs an ambient audio stream, and a speaker 105. For the sake of simplicity, the AHS 150 is illustrated without other components of the audio management system 100 of FIG. 1, such as without the ANC system 140 and the audio processing system 160.
[0049] The role of the AHS 150 can be generally described as to toggle the audio environment between an active mode and a conversation mode based on whether a desired conversation is detected, as represented by a switch network 215. In the active mode, the user is listening to the desired audio 165 via the speaker 105, and the ANC system 140 (not shown) is suppressing as much of the ambient audio 155 as possible. This is conceptually represented by the switches of the switch network 215 being in the solid-line position, whereby the desired audio 165 passes through to the speaker 105 and the ambient audio 155 does not. When attention seeking audio (i.e., audio associated with getting the user’s attention) is detected by the AS trigger detection block 210, the AHS system 150 switches the switch network 215 to the dashed-line position,whereby the ambient audio 155 passes through to the speaker 105 and the desired audio 165 does not. In some embodiments, while in the conversation mode, the passed-through ambient audio 155 (e.g., either all of the ambient audio 155, or a conversationally relevant portion of the ambient audio 155) is passed through the conversation enhancement block 230 in line with the speaker 105. As described above, the conversation enhancement block 230 can use various techniques to enhance conversationally relevant portions of the ambient audio 155. Further, as described above, the conversation enhancement block 230 can be implemented in the AHS system 150, in the ANC system 140, in the audio processing system 160, and / or in any suitable location.
[0050] When the end of the conversation is detected by the conversation end detection block 220, the AHS system 150 switches the switch network 215 back to the solid-line position, whereby the desired audio 165 again passes through to the speaker 105 and the ambient audio 155 again does not. In some embodiments, as described above, the end of the conversation is detected based on detecting the user’s own speech, such as detecting that the user is no longer speaking for some time, or that the user has issued an audio cue (e.g., “resume ANC”). This user speech can be detected via the reference microphone 110 as part of the ambient audio 155, or detected through a separate microphone 240, such as a beamforming microphone with its beam directed toward the user’s mouth. Additionally, or alternatively, some embodiments of the conversation end detection block 220 detect the end of a conversation based on detecting user interfacing with an interface element, such as detecting that the user pressed a play / pause button 245 on the WAC 210, or the like.
[0051] Although the AHS 150 is illustrated as directly coupled with the microphones, the desired audio 165 is directly coupled with the speaker 105 in active mode, the ambient audio 155 is directly coupled with the speaker 105 in conversation mode, etc., some or all of such connections can be through other components that are not shown in FIG. 2. For example, embodiments herein assume that the ambient audio 155 is passing through the ANC system 140, that the desired audio 165 is mixed with an anti -noise signal in active mode, etc.
[0052] As described herein, embodiments of the audio management system 100 are configured for integration in any suitable WAC. FIGS. 3A and 3B show a wearable audio environment 300 including a pair of WACs 310. Each WAC 310 is illustrated as an earbud. Alternatively, the pair of earbuds can be considered as a single WAC. In other embodiments, the WAC 310 can be implemented as over-ear headphones, or any other suitable wearable audio component that incorporates ANC. In the illustrated embodiments, each WAC (i.e., each earbud) has a respective instance of an audio management system 100, such as the audio management system 100 of FIG.1, and each instance of the audio management system 100 includes a respective instance of at least an ANC system 140 and an AHS system 150. Though not explicitly shown, each WAC 310 also has, integrated therein, an instance of the speaker 105, the reference microphone 110, the error microphone 115, one or more processors, and non-transitory processor-readable storage. Some implementations of the WAC 310 include additional components, such as instances of the audio processing system 160, one or more additional microphones (e.g., a beamforming microphone), one or more additional speakers, interface controls (e.g., one or more buttons), one or more power sources (e.g., a rechargeable battery), one or more ports (e.g., physical ports for charging and / or wired communication, logical ports for wireless charging and / or wireless communication), one or more antennas, etc.
[0053] In some embodiments, the one or more processors integrated in the WAC 310 implement components of the respective audio management system 100 instance. For example, a non- transitory processor-readable medium integrated therein has processor-executable instructions stored thereon, which, when executed, cause the set of processors to implement at least features of the respective ANC system 140 and / or AHS system 150 instances. As described herein, embodiments of the AHS system 150 include one or more types of artificial neural networks, corresponding trained network models, or the like. In some embodiment, such networks and / or models are implemented using specialized hardware, such as neuromorphic chips. In other embodiments, such networks and / or models are implemented by using processor-readable instructions to reconfigure general-purpose computing hardware (e.g., a central processing unit, CPU), specialized Al accelerators.
[0054] Turning specifically to FIG. 3 A, a first type of wearable audio environment 300a is shown in which one or both WACs 310 is in communication with a cloud computing environment (“cloud”) 340. For example, the cloud 340 includes a server, or several distributed servers, accessible via the Internet. Though the WAC 310 is shown as directly in communication with the cloud 340, such a connection can be facilitated by any suitable intermediary devices, such as routers, hubs, etc. As described further herein, automated attention handling features described herein rely on generation of an inference model that includes several neural networks and / or models. In the illustrated embodiments of FIG. 3 A, the inference model is generated by the local computation environment of the WAC 310 and / or based on information ported to the WAC 310 from the cloud 340.
[0055] As described herein, some embodiments involve enrollment of invocation names for use in name detection. In the embodiments of FIG. 3A, an enrollment application 330 is downloadedto the local computational environment of the WAC 310, and the enrollment application 330 is used for such enrollment. The enrollment application 330 can facilitate generation of the inference model, based on a foundation model 350. Different types of foundation model 350 can be developed for different name detection approaches. In some embodiments, the foundation model is a teacher model, such as a knowledge distillation model (KDM). The foundation model 350 can be stored in the cloud 340. In some embodiments, the foundation model 350 is accessible to the WAC 310 (e.g., to the enrollment application 330) via the cloud 340. Reference to the “inference model” can include some or all of a name embedding model, a deep image, a relation network, and a false rejection network. Each is described more fully below.
[0056] Turning to FIG. 3B, some embodiments of the wearable audio environment 300b further include a user computational device 320 separate from the WAC 310. For example, the user computational device 320 can be a smartphone, laptop computer, tablet computer, smart watch, portable audio player, or any other suitable device that is separate from the WAC 310 and includes its own one or more processors and its own one or more non-transitory storage media for storing processor-readable instructions. The user computational device 320 can be in communication with each WAC 310 via any suitable wired and / or wireless communication link, such as via an audio cable (e.g., via a 3.5 -millimeter or 1 / 4-inch analog audio jack), a universal wired connection (e.g., universal serial bus (USB)), a short-range universal wireless connection (e.g., Bluetooth, short- range radiofrequency, near field communication (NFC)), an optical connection (e.g., infrared), a proprietary connector, a multi-pin connector, an intermediary component or platform (e.g., a docking station or dongle), etc.
[0057] Similar to the embodiments of FIG. 3 A, the embodiments of FIG. 3B can include an enrollment application 330, communications with the cloud 340, use of a foundation model 350, etc. As illustrated in FIG. 3B, these features can be facilitated via the user computational device 320 (rather than directly by the WAC 310). For example (as described more fully below), when a user first registers the WAC 310 (e.g., first attempts to pair the earbuds with the user computational device 320), the user computational device 320 automatically accesses and / or downloads the enrollment application 330 or prompts the user to access and / or download the enrollment application 330. The enrollment application 330 can be downloaded from the cloud 340, or from any other suitable environment. The enrollment application 330 can then access the foundation model 350 via the cloud 340 and can use the foundation model 350 to generate an inference model for automated attention handling. For example, to facilitate LNE-AHS, the inference model includes some or all of a name embedding model, a deep image, a relation network, and a false rejection network. The generated inference model can then be ported to theWAC 310. For example, the inference model can be ported from the user computational device 320 to both earbuds; ported from the user computational device 320 to a master earbud, and from the master earbud to a slave earbud; etc.
[0058] Embodiments generally build an inference model for storing at the WAC 310 to enable the WAC 310 to subsequently use the inference model to perform automated attention handling, as described herein. FIG. 4 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block 400 and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein. The AS trigger detection block 400 can be an implementation of the AS trigger detection block 210 of FIG. 2. It is assumed that the AS trigger detection block 400 is implemented in the context of a WAC 310.
[0059] Embodiments can include an enrollment stage 410, an identification stage 430, and a verification stage 440. In general, the enrollment stage 410 occurs outside of normal operation of the WAC 310, such as when a user first sets up and / or uses the WAC 310, when the user first registers the WAC 310, when the user first configures the WAC 310 for automated attention handling, etc. The identification stage 430 and the verification stage 440 are real-time blocks that occur during normal operation of the WAC 310 to facilitate real-time automated attention handling. Embodiments generally use the enrollment stage 410 to obtain any information needed to set up an inference model for storing at the WAC 310 to enable the WAC 310 to perform automated attention handling features during normal operation. As shown, the inference model includes at least an embedding model 415, a relation network 435, and a false rejection network 445.
[0060] In the enrollment stage 410, the embedding model 415 can be used to convert an enrollment audio stream 405 into a set of reference embeddings based on a foundation model 350. The foundation model 350 can be stored in and accessed via the cloud 340 (e.g., and / or any other suitable communication network). The reference embeddings can be stored as a deep image 420. The deep image 420 can also be considered as part of the inference model.
[0061] During normal operation, in the identification stage 430, a real-time (RT) audio stream 407 is received. The embedding model 415 (i.e., the same embedding model 415 generated in the enrollment stage 410) is used to generate a RT embedding from the received RT audio stream 407. The relation network 435 can then compare the RT embedding with each of the reference embeddings in the deep image 420 to determine if there is a match. For example, the relation network 435 is configured to compute a similarity score for each comparison (e.g., correspondingto a mathematical correlation, or the like). The relation network 435 can determine if any similarity scores meet or exceed a predetermined threshold; if so, the reference embedding with the highest similarity score can be selected as a candidate matching embedding.
[0062] In the verification stage, the false rejection network 445 can then confirm the match by transforming the RT embedding and the candidate matching embedding into different mathematical spaces (e.g., domains) and determining whether the embeddings can be reliably discriminated from each other in any of the mathematical spaces. For example, the deep image 420 is configured to compute a discrimination score for each mathematical space. The deep image 420 can determine if any discrimination scores meet or exceed a predetermined threshold; if so, the match determined by the relation network 435 is determined to be a false match and is ignored. If the embeddings cannot be sufficiently discriminated in any mathematical spaces, the false rejection network 445 can output a signal that attention seeking audio has been detected.
[0063] As described herein, embodiments of the attention seeker (AS) trigger detection block 400 are configured to detect an invocation name as part of an attention handling system. The embedding model 415 is a name embedding model (e.g., a processor-executable name embedding model) that generates an output embedding from an audio sample. As used herein, the terms “audio sample” or an “audio signal” are used interchangeably in the context of an input to a component of the inference model; such an “audio sample” or an “audio signal” can be represented in any suitable manner, such as by any suitable number of digital samples. For example, reference to an input as an “audio sample” means an audio signal of a duration, or a sampled duration of an audio signal, at a sampling rate resulting in a large number of digital samples (e.g., one second of audio sampled at 16 kHz to yield 16,000 samples). As described above, the audio sample can be from an enrollment audio stream 405 in an enrollment stage 410, and the audio sample can be from a RT audio stream 407 during normal operation. The name embedding model 415 is trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications. The deep image 420 (e.g., a processor-readable deep image) includes reference name embeddings generated by the name embedding model 415 based on a set of invocation names provided by a user during an enrollment procedure. The relation network 435 (e.g., a processor-executable relation network) is coupled with the deep image 420 and the name embedding model 415 to output a candidate name embedding responsive to determining that one of the reference name embeddings has a highest similarity with a RT embedding generated by the name embedding model 415. The RT embedding can be from a RT audio stream 407 received from a reference microphone associated with an ANC system of the WAC 310. The false rejection network 445 (e.g., a processor-executable false rejection network) is coupled with the relationnetwork 435 to output a name invoked signal 450 responsive to determining that the real-time embedding and the candidate name embedding cannot be reliably discriminated. As described herein, the name invoked signal 450 can direct an ANC system of the WAC 310 automatically to enter a conversation mode.
[0064] As noted above, some embodiments of the interference model used for automated attention handling are based on a teacher model. The teacher model can be used to train one or more student models by transfer learning. In some implementations, the teacher model is a KDM, which can be used to train other portions of the inference model using knowledge distillation techniques. Embodiments of the foundation model 350 are generated from a large speech-audio corpus of diversified words spoken by diversified speakers. “Diversified words” refers herein to the speech-audio corpus including a wide variety of at least phonemes and linguistic information. “Diversified speakers” refers to the speech-audio corpus representing a wide variety of at least accents and prosody. The speakers can be further diversified with respect to age, gender, geography, etc. For example, the speech-audio corpus includes tens of thousands of words (i.e., classifications) spoken multiple times (e.g., 10 - 15 times) by hundreds of speakers from around the world.
[0065] The term “suprasegmental” is used herein as an umbrella term to encompass properties of a speaker’s influence when speaking words, such as accent, prosody, intonation, rhythm, and other non-segmental aspects of speech. Such suprasegmental features can be contrasted with segmental features pertaining to individual speech sounds or segments, such as vowels and consonants, and can span multiple segments or an entire utterance. Examples of suprasegmental features of an utterance (e.g., a word, name, etc.) can include accent (including accent-influenced variations in pitch, loudness, and duration), prosody (including rhythm, intonation, and melody of speech), intonation (i.e., the rise and fall of pitch in speech), rhythm and / or rate (e.g., the temporal patterns of speech, such as duration and timing of sounds, syllables, and pauses), and stress (e.g., emphasis placed on a particular syllable). For example, a large speech-audio corpus of diversified words spoken by diversified speakers may include hundreds or thousands of samples of a particular word being spoken with wide suprasegmental variance over the samples.
[0066] Training of the foundation model 350 is described in detail below. In general, the training can begin with an encoder-decoder architecture, transformer network, conformer network, or the like, which are types of neural network architecture designed to learn compact representations of data, such as so-called “latent features,” audio tokens, or a combination thereof.In the context of embodiments described herein, the auto-encoder architecture is used to extract meaningful features from raw audio data to be used for automatic speech recognition (ASR).
[0067] In some embodiments, the goal of training is for the foundation model 350 to learn how to convert many different instances of input labels that all represent suprasegmentally varying samples of a same class into a common set of output labels to represent that class, and to learn how to do that for a large speech-audio corpus of diversified classes. The terms “class,” “word,” and “name” are used interchangeably herein and are intended to mean any type of word, name, or utterance that could reasonably be used to get someone’s attention, such as “John,” “mister,” “hey,” “excuse me,” etc. For example, the foundation model 350 is trained to automatically segment a spoken sample of a word into a same set of acoustical segments, regardless of suprasegmental influence on the sample by the speaker (e.g., the speaker’s accent, prosody, etc.).
[0068] In general, the foundation model 350 architecture includes three high-level stages: an encoder, a bottleneck layer, and a decoder. The encoder receives a high-dimensionality input (i.e., the raw audio data) and includes a multi-layer network to progressively reduce the dimensionality of the data using transformations. For example, the ASR information is represented as a sequence of audio frames, each including features that can be mathematically described as Mel-frequency cepstral coefficients (MFCCs), or in some other manner. Each layer of transformations (e.g., linear operations followed by non-linear operations, such as a rectified linear unit (ReLU) function) seeks to extract increasingly abstract and higher-level features from the input data.
[0069] The bottleneck layer is so called because it typically includes significantly lower dimensionality than the layers of the encoder before it and or the decoder after it. This reduced dimensionality effectively forces the network to learn a compressed and informative representation of the input data. Effective training results in the bottleneck layer producing a highly compact, but highly meaningful representation of the input data; effectively extracting the most salient features for the desired task. Some embodiments can directly extract outputs from the bottleneck layer for training and refinement of networks, and / or to support other features. In such cases, the outputs from the bottleneck layer are referred to herein as “embeddings” or “embedding frames.”
[0070] In some embodiments, the decoder operates in essentially a reverse manner to that of the encoder. It takes the output of the bottleneck layer as its input data (i.e., a lowest-dimensionality representation) and applies multiple layers of transformations to generate increasingly higherdimensional representations of the data. In effect, the decoder progressively increases the dimensionality with a goal of reconstructing the original input data from the highly compressed representation in the bottleneck layer. In such cases, although the decoder output ideally matchesthe encoder input (i.e., the raw audio data), there will practically be some difference, referred to as reconstruction loss. In such cases, training of the auto-encoder seeks to optimize the auto-encoder network (e.g., the bottleneck layer) to minimize reconstruction loss (e.g., to minimize mean squared error loss, or the like). The compact representation in the bottleneck layer can be referred to as the “latent space,” and the auto-encoder’s trained (learned) knowledge of how to encode raw input data into the latent space is embedded in a set of weights. The set of weights can be represented as a feature vector, a set of feature vectors, or in any suitable manner. For example, one second of audio can be represented by 16,000 samples (i.e., at a 16 kHz sampling rate), and the 16,000 samples can effectively be represented by a set of 256 weights (e.g., or 128 weights, or another suitable number). The set of weights can represent the embedding from the auto-encoder, the filter bank energies, the MFCCs, etc.
[0071] In other embodiments, the decoder is not the reverse of the encoder. Instead, the decoder seeks to convert the bottleneck features into a particular set of output labels, such as a posterior probability matrix (PPM) and / or an ordered acoustical segment vector (OASV), as described more fully below. The decoder takes the output of the bottleneck layer as its input data (i.e., a lowest- dimensionality representation) and applies multiple layers of transformations to reach a representation of the data matching the desired output labels.
[0072] Various approaches for initially training a foundation model are described in more detail below. In particular, the following description begins with reference to FIGS. 5 - 9B to provide more detail regarding acoustic segmentation approaches. The description continues with reference to FIGS. 10 - 14 to provide higher-level descriptions of enrollment, automated attention handling, conversation start and end detection, and other features, which can be applicable to any of the name-detection approaches described herein. The description then provides several illustrative systems and methods for name detection and / or automated attention handling using acoustic segmentation (FIGS. 15 - 16), The description proceeds to and follows with more detail regarding linguistic name embedding approaches (FIG. 17), unified sound conversion approaches (FIGS. 18A - 22), and hybrid approaches (FIG. 23). An illustrative computational system for implementing various embodiments is then described with reference to FIG. 24. Subsequently, novel embodiments are described herein with reference to FIGS. 25 - 29 for using progressive learning to further develop and refine the foundation models generated according to other approaches.
[0073] Beginning with acoustic segmentation approaches for name detection, FIGS. 5 - 9B illustrate training of a foundation model 350 for such contexts. FIG. 5 shows a trainingenvironment 500 for implementing a first phase of training a knowledge distillation model (KDM) 550. As noted above, the KDM 550 is an implementation of the foundation model 350 for use with such approaches. As illustrated, the training environment 500 includes a spoken word audio repository 510 and a training auto-supervisor 520. In the context of FIG. 5, the KDM is labeled as KDM 550’, representing that the KDM is in the first training phase. In the first training phase, the KDM 550’ is trained to convert class audio samples from the spoken word audio repository 510 into corresponding posterior probability matrices (PPMs) 530.
[0074] The spoken word audio repository 510 can include one or more large corpuses of spoken audio data. Preferably, the corpuses include many words (classes), and many diversified samples in each class, so that each class includes many versions of the same word spoken with wide suprasegmental variance. For example, a given word may be spoken 10,000 times by different speakers from around the world. Thus, for each class, the spoken word audio repository 510 can output a large number of diversified audio samples for the class, which can be referred to as the “class audio samples” for the class.
[0075] The class audio samples for each class can form the entire set of spoken audio samples for that class. For example, if the spoken word audio repository 510 includes S samples for a particular class (S is a positive integer), there are S class audio samples. In other embodiments, additional variation in the audio samples for each class is created by passing the class audio samples through an augmenter (part of the training auto-supervisor 520, not explicitly shown). The augmenter uses one or more augmentation models to generate an augmented set of class audio Samples with variations in features, such as speech speed (e.g., lengthening or shortening of the audio sample, lengthening or shortening of some or all vowel sounds, etc.), modeled suprasegmental variations, models of noise profiles and / or ambient noise features (e.g., traffic sounds, background conversation sounds, etc.), etc. Some embodiments of the augmenter are implemented in the same manner as the name augmenter 1020 of FIG. 10B (e.g., and the augmentation models used by the training auto-supervisor 520 can be the same as, or different from the augmentation models 1015 of FIG. 10B). Other embodiments of the augmenter introduce more and / or different types of variation using more and / or other augmentation models. In embodiments of the training auto-supervisor 520 that include the augmenter, the augmented set of audio samples for the class is used as the class audio samples. For example, if the augmenter produces A augmentations for each of the S class audio samples from the spoken word audio repository 510 (A is a positive integer), there will be A * S class audio samples used by the training auto-supervisor 520. In some cases, the spoken word audio repository 510 may includedifferent numbers of samples for different classes, and / or the augmenter may apply different types of augmentations for different classes, such that the values of A and / or S may be class dependent.
[0076] As noted above, in the first training phase, the training auto-supervisor 520 trains the KDM 550’ to generate PPMs 530. The training auto-supervisor 520 is automated and is implemented by a processor. Embodiments of the KDM 550’ can use any suitable neural network architecture tailored for capturing features in audio data, such as a convolutional neural network (CNN), a conformer network, a transformer network, a recurrent neural network (RNN), a convolutional recurrent neural network (CRNN), etc. Embodiments of the training auto-supervisor 520 can begin by pre-processing the class audio samples into suitable input labels for use by the encoder (e.g., the input layer) of the KDM 550’. For example, the class audio samples can be resampled and / or normalized, and certain features can be extracted, such as using spectrograms or Mel-frequency cepstral coefficients (MFCCs). The input layer(s) of the KDM 550’ can also be tailored to receiving of the pre-processed audio samples, such as by having a number of dimensions corresponding to the number of MFCCs, or the like.
[0077] As described above, the encoder portion of the KDM 550’ can include several layers, such as convolutional and / or recurrent layers, to progressively reduce the dimensionality of the input class audio samples into corresponding, highly compressed representations. The layers seek to identify the most salient acoustical features based on temporal dependencies, frequency patterns, and / or other relevant information patterns. Generation of the PPMs 530 can be considered as a classification task, such that the decoder portion of the KDM 550’ is a classifier that includes as many nodes as there are cells in the PPM 530. For example, the output layer(s) of the KDM 550’ can effectively implement an activation function that allows the input to be a member of multiple classes (e.g., the sigmoid activation function), such that the values at the output nodes of the KDM 550’ represent the likelihood of the input belonging to a corresponding cell in the PPM 530. In some implementations, as illustrated, each PPM 530 is a J x K matrix, such that the output of the KDM 550’ includes J*K classification nodes.
[0078] For example, FIG. 6 shows an example of a PPM 530 represented as a J x K matrix (i.e., with J*K cells). The illustrative PPM 530 represents sounds of the English language using 19 “bodies” (columns) and 11 “souls” (rows). Each body generally corresponds to a particular consonant sound or set of related consonant sounds. For example, the body labeled ‘s’ represents the sounds ‘s’ and ‘sh.’ Each soul generally corresponds to a vowel sound (or corresponding range thereof). Each cell (i.e., each column-row intersection) corresponds to an acoustical unit, and the value in that cell represents the posterior probability of an input audio sample including thatacoustical unit. Further, the PPM 530 includes an additional row to account for lone bodies (i.e., without a soul) and an additional column to account for lone souls (i.e., without a body). As such, the illustrated PPM 530 includes 20 columns (i.e., J = 20) and 7 rows (i.e., K = 7). For example, a value in cell 631 represents the posterior probability of the acoustical unit ‘t’ (labeled as “Pp(‘t’)”), and a value in cell 632 represents the posterior probability of the acoustical unit ‘ba’ (labeled as “Pp ba’ .
[0079] It can be seen that the illustrated PPM 530 does not include all the letters in the English language. For example, the PPM 530 does not include ‘h’, ‘v’, or ‘w’; as those consonant sounds can tend to be reliably represented in their spoken context by other acoustical units. Further, the cells of the illustrated PPM 530 do not map directly to all of the phonemes in the English language. For example, many linguists classify the English language into 44 phonemes, and the illustrated PPM 530 includes 140 cells (i.e., 20 x 7). Other implementations of the PPM 530 can include any suitable number of cells corresponding to any suitable set of acoustical units. For example, the PPM 530 can be tailored to different languages, dialects or regional variations, etc.
[0080] Returning briefly to FIG. 5, the training is an iterative and automated process. As illustrated, the training auto-supervisor 520 repeatedly directs the KDM 550’ to generate PPMs 530, receives the generated PPMs 530 as feedback, and adjusts the KDM 550’ until all class audio samples representing a same class (or at least a threshold number) yield a same PPM 530 for the class. For example, the training auto-supervisor 520 seeks to minimize a loss function (e.g., crossentropy) to find the most representative PPM 530 for each class. When the KDM 550’ is trained in accordance with the first phase, it can be considered as KDM 550” (i.e., moved to a second training phase).
[0081] FIG. 7 shows a training environment 700 for implementing a second phase of training the knowledge distillation model (KDM) 350. As in the training environment 500 of FIG. 5, the training environment 700 includes the spoken word audio repository 510 and the training autosupervisor 520. In the context of FIG. 7, the KDM is labeled as KDM 550”, representing that the KDM is in the second training phase. In the second training phase, the KDM 550” is trained to use the PPMs 530 as a guide for acoustically segmenting the class audio samples from the spoken word audio repository 510 into corresponding acoustical segment vector (OASVs) 735.
[0082] As shown, the second training phase can include two sub-phases. In a first sub-phase, embodiments of the training auto-supervisor 520 automatically segment class audio samples into candidate segmentations based on ortho-segmentation rules 725. As used herein, “orthosegmentation” refers to segmentation of a word into orthographic units that are based on theorthography (i.e., the written form) of the word. In some embodiments, the spoken word audio repository 510 includes a lexical entry for each of some or all of the classes, which can be used directly as “class text.” For example, the term “INDEPENDENCE” can have hundreds of diversified spoken audio samples for the word, all stored in association with a lexical entry (i.e., the text) for the word. In other embodiments, the spoken word audio repository 510 may not include lexical entries for classes, or may not include a lexical entry for one or more classes. In such embodiments, for any class that does not have an associated lexical entry, one or more of the repository spoken audio samples is fed to a speech-to-text (STT) engine 710, which generates the class text from the class audio sample(s) as received from the spoken word audio repository 510.
[0083] The class text (whether received from the spoken word audio repository 510 or the STT engine 710, is passed to an ortho-segmenter 720. The ortho-segmenter 720 is a parser that converts the class text to a candidate segmentation based on ortho- segmentation rules 725. The ortho-segmentation rules 725 is represented as storage in FIG. 7, indicating that, embodiment of the ortho-segmentation rules 725 are stored in a non-transitory, processor-readable storage medium. For example, the ortho-segmentation rules 725 can be stored as a set of functions, scripts, or the like, which can be executed by the ortho-segmenter 720 on the class text. An example set of ortho-segmentation rules 725 is as follows: a) Segment before or between constriction (i.e., where the lips touch together or the tongue touches the upper or lower palate), such as / p / , / ph / , / t / , / th / , / d / , / dh / , / ! / , / lh / , / b / , / bh / , / g / , / k / , / n / , / m / , / s / , / sh / , hd l zl z l . b) For trilling: (1) segment before trilling, if / r / is not followed by plosives such as / t / , / k / , / g / , or / d / ; (2) segment before plosives such as / t / , / k / , / g / , and / d / that succeed trilling / r / ; (3) segment the trilling / r / after the vowel, if succeeded by vowels, such as / a / , / e / , / i / , / o / , / u / ; and (4) segment the trilling / r / alone, if it is not proceeded or succeeded by the aforesaid plosives or vowels, respectively. c) Segment before a nasal phoneme (n, m), if it continues with carriers or vowels. Otherwise, segment after the nasal phoneme. d) Segment before and after fricatives, such as / sh / , / ch / , / f / , / x / , / z / , / zh / . e) Treat parallel vowels, such as / j / and / q / , separately by combining / j / and / q / with succeeding vowels. f) Segment consecutive carriers, if both carriers are succeeded and preceded by a body. g) Combine end-plosives, such as / k / , / d / , / t / , / b / , / g / , and / I / .
[0084] The candidate segmentation for each class automatically generated by the ortho- segm enter 720 can be fed into an audio segm enter 730, along with some or all of the class audio samples for the corresponding class. The output of the audio segmenter 730 is a sequence of audio chunks of each class audio sample, where each audio chunk corresponds to a respective unit of the candidate segmentation. The audio chunks can be fed into the KDM 550” as input labels for the second training phase. For example, feeding the audio chunks into the KDM 550” can involve preprocessing the audio chunks into MFCCs, or the like. As illustrated, the second training phase trains the KDM 550” to generate OASVs 735 from the sequences of audio chunks.
[0085] Each OASV 735 is a 1 x L vector, where L is a positive integer (e.g., 16) corresponding to a maximum number of acoustical units that can be used for acoustical segmentation by the KDM 550”. In the first training phase, the KDM 550’ is trained as a classifier, where the classification output nodes correspond to the J*K cells of the PPM 530, In the second training phase, the classification knowledge of the KDM 550” is used to classify each audio chunk sequentially as a corresponding one of the cells of the PPM 530. For example, the KDM 550” tries to use all of the first audio chunks from all of the class audio samples for a particular class (in accordance with the candidate segmentation) to figure out a best-matching cell from the PPM 530 to represent the audio chunk. Classifying the sequence of audio chunks results effectively in a sequence of PPM 530 cells determined to represent the sequence of acoustical segments that best correspond to the sequence of audio chunks, and that sequence of PPM 530 cells can be represented as the OASV 735. Embodiments of the KDM 550” can be implemented with an output layer having L output nodes corresponding to the L elements of the OASV 735. Where fewer than L acoustical segments are used, the remaining elements of the OASV 735 can include a default value (e.g., ‘-1’) that does not correspond to any of the cells of the PPM 530.
[0086] For example, FIG. 8 shows an illustrative candidate segmentation and an illustrative corresponding PPM 810 and OASV 830. The PPM 810 can be an example of PPM 530, and OASV 830 can be an example of OASV 735. In the example, the class (name) “MONISHA” has been classified to generate PPM 810. The PPM 810 shows cells having a value of ‘ 1’ where the corresponding acoustical segments is found by the classification to be present in the class audio samples for “MONISHA”. In the illustrated example, a ‘ 1’ is present in the cells corresponding to acoustical units ‘mo’, ‘ni’, ‘s[h]’, and ‘a’ (as mentioned above, the unit ‘s’ also represents the fricative ‘sh’).
[0087] FIG. 8 also shows an illustrative index matrix 820. The index matrix 820 is the same size as the PPM 810, and each cell of the index matrix 820 has a unique value that represents an indexto the corresponding cell of the PPM 810. For example, the acoustical segment ‘ni’ corresponds to cell index ‘65’. Each audio chunk can be classified as the one of the index values from the index matrix 820 corresponding to the PPM 810 cell that is the best-matching acoustical segment. In some implementations, the classification of each audio chunk yields a value, and the value is rounded to the nearest cell index value in the index matrix 820. In one implementation, rather than each index being separated from its neighbors by ‘ 1’ (as shown), each index can be separated from its neighbors by ‘ 100’. For example, instead of indexing the cells as ‘O’, ‘ 1’, ‘2’, etc., they can be indexed as ‘O’, ‘ 100’, ‘200’, etc. (i.e., each index shown in index matrix 820 can be multiplied by 100). In such an implementation, ‘mo’ corresponds to index ‘8900’, and any classification result between 8850 and 8949 can be classified as ‘mo’. The difference between neighboring index values can effectively operate as a quantization resolution, and different implementations can use any suitable quantization resolution.
[0088] In the example illustrated by FIG. 8, the class “MONISHA” has been segmented into a candidate segmentation: ‘MO’ / ‘NI’ / ‘S[H]’ / ‘A’. For example, the class has automatically been ortho-segmented by the ortho-segm enter 720 according to ortho- segmentation rules 725. The illustrated OASV 830 is a 16 x 1 vector. Because “MONISHA” was segmented into four segments, the first four elements of the OASV 830 point to a sequence of cells of the PPM 810, and the remaining 12 elements show a default entry of ‘-1’. The first four elements index the sequence of acoustical segments that best represent the sequence of audio chunks according to the candidate segmentation. It can be seen that, if the candidate segmentation produced the correct acoustical segmentation (i.e., an acoustical segmentation matching the spoken form of the class), the acoustical segments identified by the OASV 830 will match those predicted by the PPM 810 (as happens to be the case in the illustrated example).
[0089] Returning to FIG. 7, the training auto-supervisor 520 can include an evaluator 740 that automatically determines whether the candidate segmentation appears to produce a good acoustical segmentation. Embodiments of the evaluator 740 can evaluate the generated OASV 735 for a class based on the generated PPM 530 for the class to determine whether the set of acoustical segments represented by the OASV 735 matches those in the PPM 530. Embodiments of the PPM 530 indicate which acoustical segments are probabilistically present in the class audio samples, but it may not represent the order of those segments. If the OASV 735 represents an accurate acoustical segmentation of the class audio samples, it should indicate the same set of acoustical segments as indicated by the PPM 530 (and the order of those acoustical segments).
[0090] Words are frequently pronounced in a manner that does not match a relatively small and rigid set of rules based on the word’s orthography (i.e., ortho- segmentation rules 725). As such, it can be expected that automated segmentation by the ortho-segmenter 720 based on orthosegmentation rules 725 will yield some incorrect candidate segmentations. After the first subphase of the second training phase, there will be some percentage (e.g., X%) of candidate segmentations determined by the evaluator 740 to be “correct,” and some percentage (e.g., Y%) of candidate segmentations determined by the evaluator 740 to be “incorrect.”
[0091] As illustrated, the classes that were not correctly segmented by the ortho-segmenter 720 can be identified for performance of the second sub-phase of the second training phase: acoustical re-segmentation 750. In some embodiments, the evaluator 740 automatically generates and outputs a set (e.g., a list) of the classes for which automated ortho-segmentation resulted in an incorrect acoustical segmentation. The acoustical re-segmentation 750 can be performed on the identified set of incorrectly segmented classes. In some embodiments, the acoustical resegmentation 750 is a manual process (e.g., the only manual portion of the training) by which a human trainer or trainers can attempt to find a re-segmentation that better represents the acoustic segments. In other embodiments, the acoustical re-segmentation 750 is a fully automated, or partially automated process. For example, in each iteration of the second training phase, embodiments can use a different subset of ortho-segmentation rules (e.g., from the stored rules 725), can modify previously applied ortho-segmentation rules (e.g., in random or pre-defined ways), etc.
[0092] For example, FIGS. 9A and 9B show example OASVs 910 resulting from an illustrative automated ortho-segmentation and an illustrative re-segmentation, respectively. Turning first to FIG. 9A, the class “CHOCOLATE” is automatically segmented by the ortho-segmenter 720 in accordance with stored ortho-segmentation rules 725, resulting in a candidate segmentation: ‘C[H]’ / ‘O’ / ‘CO’ / ‘LA’ / ‘TE’. This candidate segmentation, after classification by the KDM 550”, results in an OASV 910a of [2, 80, 82, 33, 46] (the remaining elements in the vector are unused, as represented by the value ‘-1’). It can be assumed that this OASV 910a does not sufficiently correspond to the PPM for the class.
[0093] Turning to FIG. 9B, the class “CHOCOLATE” is now re-segmented (according to the acoustical re-segmentation 750 sub-phase, such as manually) into a different candidate segmentation: ‘C[H]A’ / ‘K’ / ‘LE’ / ‘T’. This candidate segmentation, after classification by the KDM 550”, results in a OASV 910b of [22, 1, 53, 6] (the remaining elements in the vector are unused, as represented by the value ‘-1’). It may be that this OASV 910b does sufficientlycorrespond to the PPM for the class. If not, the class may be passed back through the acoustical re-segmentation 750 sub-phase.
[0094] Returning to FIG. 7, the second sub-phase of the second training phase can be iterative. For example, after automated ortho-segmentation (i.e., the first sub-phase of the second training phase), X may be 80 and Y may be 20, such that 20% of the classes were incorrectly segmented by the ortho-segmenter 720. Those 20% are passed to the acoustical re-segmentation 750 sub-phase and are re-segmented. All of the classes are again passed through the KDM 550” to generate corresponding OASVs 735. Passing all of the classes back through the KDM 550” (i.e., as opposed to repeating the process only for those classes that were incorrectly segmented in the previous iteration) can help to correct any inherent error in the KDM 550” itself. For example, it is possible that a class that was correctly segmented in one iteration will be incorrectly segmented in a subsequent iteration because of changes to the KDM 550”, but it is assumed that this still represents an overall improvement to the KDM 550”. After this second iteration (i.e., after manual re-segmentation) X may now be 90 and Y may now be 10, such that only 10% of the classes are now being incorrectly segmented. Those 10% can be passed again to the acoustical resegmentation 750 sub-phase and can be re-segmented differently.
[0095] The second sub-phase process can repeat until a training satisfaction level is reached: either X is above a predetermined threshold, Y is below a predetermined threshold, or the segmentations of all classes result in correct acoustical segmentations. As illustrated, once the training satisfaction level is reached, the KDM 550” can be considered as the KDM 550 for use in training the inference model for name-detection-based attention handling, as described herein. With training of the KDM 550 complete, the KDM 550 is capable of automatically generating a correct acoustical segmentation from an input audio sample to at least a predetermined confidence level. Moreover, the training is such that even suprasegmentally varied versions of a same class will be converted by the KDM 550 into a same OASV 735.
[0096] Returning to FIG. 4, the trained KDM 550 can now be used to train the name embedding model 415 of the inference model. Training of the name embedding model 415 is performed by applying knowledge distillation from the KDM 550 based on a smaller corpus of real-world name data. For example, the KDM 550 can use a large number (e.g., 11,000) of classifications to generate the PPMs 530 and OASVs 735, and the name embedding model 415 (deep feature generation model, or DFGNet) can be trained on a smaller number (e.g., 500 - 1000) of classifications, each associated with a linguistically distinct name.
[0097] Training of the name embedding model 415 by knowledge distillation generally involves determining which and how many layers and connections of the KDM 550 can be removed without reducing the automated acoustical segmentation performance by too much. In general, the knowledge distillation involves copying the KDM 550 as a first (largest) iteration of the name embedding model 415, running a batch of input data to produce “correct” results (i.e., assuming that any results produced by the KDM 550 in its entirety are considered to be correct), and freezing the input and output data (e.g., the input and output labels). The name embedding model 415 can be iteratively distilled. In each iteration, the frozen input labels are provided to the distilled model, and the resulting output labels are compared to the frozen output labels to determine an amount of error that resulted from the distillation. If the error produced by the name embedding model 415 relative to the KDM 550 is within a predetermined tolerance, the name embedding model 415 can be further distilled in another iteration. If not, the previous distillation can be undone; and the name embedding model 415 can either be finalized as is (e.g., if it is sufficiently compact for the desired runtime environment), or a different type of distillation can be attempted.
[0098] In each iteration, the knowledge distillation can involve any suitable distillation task. One example of a distillation task is encoder simplification, in which the number of layers of the neural network can be reduced to make the model more lightweight. Another example of a distillation task is layer-wise distillation; rather than removing layers, knowledge can be selectively distilled from one or more layers of the teacher model to focus on only the most informative layers (e.g., and to help prevent information loss). Another example of a distillation task is reducing network connections. For example, the teacher model may have extensive interlayer connections (e.g., skip connections between encoder and decoder layers). In such cases, in addition to reducing the numbers and / or complexity of layers, complexity can be reduced by simplifying and / or removing some of these inter-layer connections in the student model. Another example of a distillation task is downsampling, or the like. For example, the teacher model may process input streams at certain sampling rates, temporal resolutions, etc.; and those resolutions can be reduced in the student model (e.g., by downsampling, using smaller temporal step sizes, reducing the number of recurrent layers in an RNN, etc.). Similarly, precision of weight parameters can be simplified in some cases (e.g., 32-bit floating-point weights can be reduced to 8- bit weights, or lower), which can appreciably reduce computational complexity. Other examples of distillation tasks can include cases where the KDM includes complex attention mechanisms (e.g., multi-head attention in transformers), and the attention mechanism can be simplified (e.g., by reducing the number of attention heads); or if the output layer of the teacher model includesmultiple output heads, and the student model may be able to operate reliably with fewer heads or a modified (simplified) structure.
[0099] Each of these or other types of distillation tasks (e.g., each distillation iteration) will potentially add some amount of error to the performance of the name embedding model 415. Such distillation error in each iteration can be evaluated in any suitable manner. In some embodiments, the name embedding model 415 is trained with a total error that is a weighted combination of the “original task error” (e.g., cross-entropy loss) and an additional “knowledge distillation error.” The knowledge distillation error measures the similarity between predictions of the KDM 550 and those of the name embedding model 415. For example, an objective function can be mathematically described as:where X is a hyperparameter controlling the importance (weight) of the distillation error and i is an index of a model layer.
[0100] Ultimately, the goal of training the name embedding model 415 is to distill the KDM 550 (as the teacher model) into the name embedding model 415 (as the student model) by transferring the knowledge of the KDM 550 to the name embedding model 415 in such a way that the name embedding model 415 can achieve comparable performance with appreciably reduced computational resources. It is generally assumed herein that the KDM 550 is too large and too complex to practically run in real-time within the resource confines of a WAC. For example, continuous real-time running of KDM 550 would require too many computational resources, too much memory, too much power, and / or too many other resources to be practical. As such, the goal of the knowledge distillation is to distill the knowledge of the KDM 550 into a name embedding model 415 with a size and complexity that can practically be run continuously and in real-time within the computational environment of a WAC.
[0101] As noted above, the name embedding model 415 is trained on a smaller corpus of name audio samples. Real audio samples used to train the name embedding model 415 can correspond to people’s names from various regions and languages, and sample names can be chosen to cover most phonetic usage in each region. Implementations of the name embedding model 415 can be trained to recognize any suitable number of name classifications. For example, it may be impractical impossible to train the model for all possible names and their variants everywhere in the world, and a practical number of more common names (e.g., 1,000) can be chosen instead. In some implementations, different versions of the name embedding model 415 can be generatedand / or trained differently for different user groupings (e.g., geographical regions, ethnicities, etc.) to capture the most popular names for the corresponding groupings. For example, grouping information can be entered by the user as part of enrollment, obtained for the user from account information, assumed for the user based on location or other demographic information, etc. Further, the name embedding model 415 can be designed with as much complex as needed to generate proper acoustical segmentations of the name classifications used for training with enough reliability to be suitable for name-detection-based attention handling. In other approaches (e.g., linguistic name embedding), the name embedding model 415 can be designed with as much complex as needed to discriminate between the name classifications used for training with enough reliability to be suitable for name-detection-based attention handling. For example, a higher- complexity model (e.g., where the number of layers and / or connections is larger) may be able to more reliably discriminate among a larger number of name classifications, but use of such a model in real time will involve more computational resources (e.g., which may correspond to more processing time, more battery usage, more heat generation, etc.).
[0102] Once the name embedding model 415 has been trained (i.e., sufficiently distilled), it can be used to generate a reference embedding for each invocation name. Some embodiments of the name embedding model 415 generate an OASV 735 for each invocation name and store the OASVs 735 in the deep image 420. Some embodiments strip the output layers from the name embedding model 415, leaving only the encoding (input) and bottleneck layers, so that the output of the name embedding model 415 can be the output labels (e.g., audio tokens, or other latent space representation) of the bottleneck layer, which can be an N-dimensional vector of weights. The weights can effectively represent a highly compressed version of the input audio sample that includes only those features determined to be most salient for automated acoustical segmentation. N can be any suitable integer number to provide sufficiently reliable classification. In one implementation, N is 128. In another implementation, N is 256. For example, the name embedding model 415 generates each reference embedding as the N-dimensional vector and stores the vectors in the deep image 420.
[0103] Some embodiments of the name embedding model 415 generate both types of reference embedding for each invocation name: both a corresponding OASV 735 from a classifier portion of the name embedding model 415 and a corresponding latent space representation from the bottleneck layer of the name embedding model 415. The deep image 420 stores both reference embeddings for each invocation name. In such embodiments, the name embedding model 415 is also configured to generate both types OASVs 735 and latent space representations for the realtime embeddings. In some implementations, the relation network 435 is trained to generate theinitial identification of candidate matches using the OASVs 735 of the reference and real-time embeddings, and the false rejection network 445 is trained to discriminate true and false matches using the latent space representations of the reference and real-time embeddings.
[0104] FIGS. 10A and 10B show block diagrams 1000 of illustrative uses of the name embedding model 415 to generate the deep image 420. Turning first to FIG. 10A, a user 1005 provides a set of M invocation names 1010 via an enrollment application 330 (M is a positive integer). For example, the user 1005 speaks each name one or more times, types each name using its proper spelling, types each name phonetically, etc. The M invocation names 1010 are passed to the name embedding model 415, which generates M corresponding reference embeddings. As noted above, each reference embedding is an N-dimensional vector corresponding to a set of N weights in the name embedding model 415 that represents the invocation name that yielded that reference embedding. The name embedding model 415 generates each reference embedding as the N-dimensional vector and stores the vectors in the deep image 420, such that the deep image 420 stores M N-dimensional vectors, a single M-by-N-dimensional matrix, or the like.
[0105] Turning to FIG. 10B, a user 1005 again provides a set of M invocation names 1010 via an enrollment application 330 (M is a positive integer). Unlike in FIG. 10A, the M invocation names 1010 are passed to a name augm enter 1020, which augments the user-provided set of invocation names 1010 to generate an augmented set of invocation names 1010’. The name augm enter 1020 can include, or be in communication with, an augmentation model 1015. Embodiments of the augmentation model 1015 include mathematical transformations to apply to each of some or all of the invocation names 1010. The name augmenter 1020 can generate G augmentations (G is a positive integer) for each of the M invocation names 1010, so that the augmented set of invocation names 1010’ includes M * G names. For example, a user 1005 enrolls four invocation names, nine augmentations are applied to each invocation name to generate ten total names for each invocation name, or forty total entries in the augmented set of invocation names 1010’.
[0106] In some implementations, the name augmenter 1020 adds time-based augmentations to each of some or all of the invocation names 1010, such as by time-stretching and / or timecompressing a user-provided audio sample of the invocation name. In some implementations, the name augmenter 1020 adds accent-based augmentations to each of some or all of the invocation names 1010, such as by mathematically applying different vowel changes, regional variations, pronunciations, etc. to the invocation name. In some implementations, the name augmenter 1020 adds suprasegmental augmentations to each of some or all of the invocation names 1010, such asby mathematically applying different syllable accenting, intonation, volume, pitch, etc. Other augmentations can account for differences across genders, ages, etc. Other augmentations can account for noise models, such as models of ambient background noise, television or music noise, traffic noise, road noise, engine noise, air conditioning noise, running water noise, etc. The M * G invocation names 1010’ are passed to the name embedding model 415, which generates M * G corresponding reference embeddings. For example, the name embedding model 415 generates each reference embedding as an N-dimensional vector for storage in the deep image 420 (e.g., as M * G N-dimensional vectors, as a (M * G)-by-N-dimensional matrix, or the like). In some implementations, the name augm enter 1020 applies different augmentations to different invocation names, and / or different numbers of augmentations to different invocation names. As one example, different augmentations can be applied based on whether the invocation name is characterized more by its vowel content, or more by its consonant content. As another example, a more common term enrolled as an invocation name (e.g., “boss,” “mom”), or a shorter name enrolled as an invocation name (e.g., “Max,” “Tim”) may be augmented differently than less common terms, longer names, etc.
[0107] Returning to FIG. 4, the relation network 435 is trained with the linear and non-linear features that characterize the name embedding model 415. Embodiments of the relation network 435 are trained to output a similarity score (e.g., a probability of a match) responsive to two inputs: one of the reference embeddings from the deep image 420, and a real-time embedding generated from a real-time audio sample received via the reference microphone. For example, the one of the reference embeddings from the deep image 420 is a N-dimensional vector previously generated by the name embedding model 415 during enrollment, and the real-time embedding is an N- dimensional vector generated by the name embedding model 415 in real-time. As described above, the reference embedding vector essentially represents salient linear and non-linear features for proper acoustical segmentation (e.g., or salient linear and non-linear features of associated classifications in linguistic name embedding and / or other approaches), and the real-time embedding vector essentially represents the same salient linear and non-linear features of the realtime audio sample. The relation network can map those same linear and non-linear features between real-time embeddings and reference embeddings to find candidate matches. For example, in a scenario where 40 classifications are generated (i.e., the deep image 420 is a 40-by-N matrix), the relation network 435 can compute a correspondence between the real-time embedding and each of the 40 reference embeddings. This can be performed as 40 serial computations (e.g., iterative), 40 parallel computations, or in any suitable manner. Some embodiments of the relation network 435 are implemented as a two-dimensional convolutional neural network (CNN). Someother embodiments of the relation network 435 are implemented as a one-dimensional CNN, a time-delay neural network (TDNN), or another suitable neural network. Some other embodiments of the relation network 435 are implemented using simple cosine similarity or equilidian distance estimation. For example, thresholding is performed based on the measured metric, and either a high value of the cosine similarity score represents a high relationship (for simple cosine similarity), or a minimum score represents a high relationship (for equlidian distance).
[0108] Embodiments compute a similarity score (e.g., a mathematical correlation) between a present real-time embedding (RTE) and each of the reference embeddings and determine whether the similarity score exceeds a predetermined matching threshold (e.g., 0.3) for any one or more of the reference embeddings. If none of the reference embeddings yields a similarity score exceeding the predetermined matching threshold, embodiments determine that there is no name match and ignore the analyzed portion of the real-time audio signal (i.e., discards the RTE). If one of the reference embeddings yields a similarity score exceeding the predetermined matching threshold, the class associated with that reference embodiment is selected as a candidate matching name (i.e., that reference embedding is selected as the candidate matching reference embedding, or CMRE). If multiple reference embeddings yield similarity scores exceeding the predetermined matching threshold, the reference embedding associated with the highest similarity score is selected as the CMRE.
[0109] Embodiments of the relation network 435 are trained to output a similarity score (e.g., a probability of a match) responsive to two inputs: one of the reference embeddings from the deep image, and a real-time embedding generated from a real-time audio sample received via the reference microphone. During training of the relation network 435, a training audio sample can be used as the real-time audio sample, which is fed into the name embedding model 415 to generate a training embedding (corresponding to the real-time embedding generated during normal operation). For example, the reference embedding for a particular invoked name classification is input to the relation network 435, and a training embedding is generated by the name embedding model 415 for a training audio sample: if the training audio sample is known to correspond to a particular invocation name, the relation network 435 is trained to output ‘ 1’, ‘ 100 percent’, etc. when fed the corresponding reference and training embeddings; if the training audio sample is known not to correspond to a particular invocation name, the relation network 435 is trained to output ‘O’, ‘0 percent’, etc. when fed the corresponding reference and training embeddings. In some embodiments, the relation network 435 is trained based on the same corpus of audio samples (or a portion thereof) used to train foundation model 350 (e.g., KDM 550). In other embodiments,the relation network 435 is trained based on the same corpus of audio samples (or a portion thereof) used to train the name embedding model 415.
[0110] The false rejection network (FRNet) 445 seeks to determine whether the CMRE and the RTE can be discriminated. In effect, the relation network 435 seeks to find a candidate match, and the false rejection network 445 seeks to determine whether the candidate match is a false match. Embodiments of the false rejection network 445 apply multiple mathematical transformations (e.g., rotations), each transformation designed to transform both the CMRE and the RTE into a corresponding domain and / or space to see whether the two datasets continue to match. For example, suppose a user has enrolled the invocation name, “Jonathan,” and the real-time audio signal includes the phrase “on a thin.” In such a scenario, the relation network 435 may find a candidate match (i.e., a similarity score exceeding the threshold), but the false rejection network 445 may determine that the candidate match is likely not a match and can be rejected. Some embodiments of the false rejection network 445 are implemented as a progressive layered extraction (PLE) neural network. Some other embodiments of the false rejection network 445 are implemented as a probabilistic linear discriminant analysis (PLDA) network.[OHl] Embodiments of the false rejection network 445 are trained to output a discrimination score (e.g., a likelihood ratio representing probability of a false match) responsive to two inputs: one of the reference embeddings from the deep image 420, and a real-time embedding generated from a real-time audio sample received via the reference microphone. The training of the false rejection network 445 can be similar to the training of the relation network 435. For example, during training of the false rejection network 445, a training audio sample can be used as the realtime audio sample, which is fed into the name embedding model 415 to generate a training embedding (corresponding to the real-time embedding generated during normal operation).Unlike the relation network 435, the false rejection network 445 is trained to apply transformations to the two inputs to look for a particular domain or space in which the two can be discriminated. For example, the training can use some training audio samples that are similar to a particular invoked name classification and other audio samples that are completely different (e.g., effectively linguistically orthogonal) to the invoked name classification. The false rejection network 445 is trained to find transformations that reliably discriminate involved name classifications from audio samples that sound like those invoked names but actually carry a different linguistic meaning. In some embodiments, the false rejection network 445 is trained based on the same corpus of audio samples (or a portion thereof) used to train the foundation model 350 (e.g., KDM 550). In other embodiments, the false rejection network 445 is trained based on the same corpus of audio samples (or a portion thereof) used to train the name embedding model 415.
[0112] The name embedding model 415, the relation network 435, and the false rejection network 445 can all be trained together (e.g., in parallel, or serially). In the acoustic segmentations embodiments described herein, the name embedding model 415 is trained by knowledge distillation from the KDM 550 using a corpus of real-world name data. The input is an audio sample, and the output (after removing the output layers) is an N-dimensional weighting vector. The specific invocation names (e.g., including augmentations) are used to generate reference embeddings for each of a set of invocation name classifications, which are stored as the deep image 420. Embodiments of the relation network 435 are trained to output a respective similarity score between a real-time embedding generated from a real-time audio sample received via the reference microphone and each of the reference embeddings from the deep image 420. Embodiments of the false rejection network 445 are trained to output a respective discrimination score between a real-time embedding generated from a real-time audio sample received via the reference microphone and each of the reference embeddings from the deep image 420. Because the relation network 435 only computes similarity scores and matches them to a threshold, the relation network 435 can be very lightweight (e.g., resource-efficient). For example, even in the context of a small processor and a small battery, such as in an earbud), the relation network 435 can run continuously without using excessively processor computation cycles, without draining excessive power, without generating excessive heat, etc. Embodiments of the false rejection network 445, which may use appreciably more resources to perform transformations, etc., only run when a candidate match has been identified. Alternative embodiments can combine the functionality of the relation network 435 and the false rejection network 445, such as in contexts where resources are not as limited (e.g., implemented in over-ear headphones that include wired power).
[0113] As described above, some or all of the name embedding model 415, the relation network 435, the false rejection network 445, and the deep image 420 can be treated as a single inference model (or a “name detection model”). For example, the foundation model 350 is a common model that is computed in and / or stored in the cloud 340. When invocation names are first enrolled, an enrollment application 330 is downloaded to a user device. For example, the application is downloaded to the user’s laptop computer, tablet computer, smartphone, smart watch, portable audio player, headset, etc. In some implementations, the WAC 310 is associated with a case, such as for storage and / or charging; and the enrollment application 330 can be downloaded to a computational environment stored in the case.
[0114] For the sake of illustration, FIG. 11 shows several example screenshots from an example enrollment application 330 running on a user device. At a first screen 1110, the user begins aname enrollment process. By clicking “NEXT” using a user interface of the user device (e.g., a touchscreen), the user can proceed to a second screen 1120. At the second screen 1120, the user is prompted to enroll an invocation name. For example, the second screen 1120 includes a button to activate a microphone of the user device by which to receive an audio sample from the user representing the invocation name being enrolled. Additionally or alternatively, the second screen 1120 (or another screen) can include interface elements for receiving text, etc. Proceeding to a third screen 1130 (e.g., by clicking “NEXT”), the user is presented with several options, such as an option to re-record the enrollment name, to enroll another name, or to end the enrollment process. Some implementations can present additional options, such as permitting the user to select any previously enrolled name to re-record, to delete, etc. In some cases, opting to re-record or to enroll another name can bring the user back to the second screen 1120, or another similar screen. Opting to end the enrollment can bring the user to a fourth screen 1140, which indicates to the user that the enrollment is complete.
[0115] In some implementations, conclusion of the user enrollment of invocation names automatically triggers the enrollment application 330 to compute (generate) some or all of the name detection model. In other implementations, subsequent to the user enrollment of invocation names, the user is prompted to continue with generation of some or all of the name detection model. In some implementations, some or all of the name detection model is generated separately from the enrollment application 330. After the name detection model is generated, the name detection model can be ported to the WAC 310 for local execution. Some embodiments of the enrollment application 330 permit the user, at any suitable time, to enroll additional invocation names, delete enrolled invocation names, etc.
[0116] Some embodiments described herein assume joint participation of a cloud-based computational platform, a local computational platform separate from the WAC 310 (e.g., a smartphone), and the computational platform integrated in the WAC 310. Different arrangements of features, components, etc. can be implemented depending on the computing, power, storage, and / or other resources of these computational platforms. In one implementation, the application is downloaded directly to the WAC 310 (or is previously loaded to the WAC 310), and the name detection model is computed directly by the WAC 310 (i.e., there is no need for a separate computational platform. In another implementation, enrollment information is exchanged with cloud-based processing resources to generate some or all of the name detection model. For example, audio samples corresponding to the invocation names (e.g., including augmentations thereof) are sent to the cloud, cloud-based resources are used to compute the name detection model, and the name detection model is ported (e.g., directly from the cloud, or via one or moreintermediary devices) to the WAC 310. In other implementations, the application is directly ported to the WAC 310, and it is then downloaded to, or installed on, the local computational platform separate from the WAC 310 (e.g., the smartphone, etc.), if the local computational platform does not already have it while pairing.
[0117] FIG. 12 shows a flow diagram of an illustrative method 1200 for audio management that includes automated attention handling in a wearable audio component (WAC), according to embodiments described herein. Embodiments of the method 1200 can be implemented using any of the system implementations described herein, or any other suitable variation thereof. Some embodiments begin at stage 1204 by receiving a real-time audio signal (e.g., from a reference microphone associated with an active noise control (ANC) system). The real-time audio signal is ambient audio received via the WAC while a user is listening to desired audio with ANC in an active (ambient sound suppression) mode.
[0118] At stage 1208, embodiments can detect whether the real-time audio signal includes attention seeking (AS) audio. For example, as described above with reference to FIG. 1, an AHS system 150 can be used to detect when a second-party attention seeker is trying to get the attention of a user while the user is wearing the WAC and is listening to desired audio 165 with the ANC system 140 in the active mode. In particular, the AHS system 150 listens for presence of attention seeking (AS) audio 157 within the ambient audio 155. For example, embodiments of the AHS system 150 are configured to listen for AS audio 157 corresponding to a previously enrolled invocation name (e.g., a name of the user). As described herein, the detection in stage 1208 can be performed using automated attention handling based on automated acoustic segmentation, linguistic name embedding, universal sound conversion, or any suitable hybrid thereof. As illustrated, the detection at stage 1208 can rely on prior training of a name detection (i.e., inference) model at stage 1220, generation and storage of reference embeddings based on a set of enrolled invocation names using the name detection model at stage 1222, generation of real-time embeddings from the real-time audio signal using the name detection model at stage 1224, and comparison of the real-time embeddings with the reference embeddings to determine whether the AS audio is present at stage 1226.
[0119] A determination block at stage 1212 represents the result of the determination at stage 1208. If no AS audio is detected, embodiments of the method 1200 return to stage 1204. For example, embodiments continue to listen to the real-time audio signal, and the ANC system remains in active mode. If AS audio is detected, embodiments proceed to stage 1216 by triggering the ANC system automatically to switch to a conversation mode. For example, referring back toFIG. 1, when the AS audio 157 is detected by the AHS system 150, the AHS system 150 automatically directs the ANC system 140 to switch from the active mode to the conversation mode. As described herein, the conversation mode can include disabling ANC, lowering the volume of desired audio, pausing the desired audio, enhancing conversationally relevant audio, suppressing feedback of the user’s own speech, etc.
[0120] As illustrated by off-page reference “A”, some embodiments of the method 1200 include an enrollment phase prior to stage 1204. FIG. 13 shows a flow diagram of an illustrative method 1300 for such an enrollment phase. In some embodiments, the method 1300 begins at stage 1304 when a new WAC is detected. Such a detection can occur when a WAC is first paired with a user device, first paired with a network, first set to perform name detection, etc. For example, the term “new” in this context can simply indicate that the WAC is new with respect to features of automated attention handling described herein. In response to the detection at stage 1304, embodiments can obtain an enrollment application (e.g., from the cloud) at stage 1308.
[0121] At stage 1312, embodiments can receiving a set of invocation names (e.g., see stage 1212 of FIG. 12) from the user. At stage 1316, embodiments can generate a reference name embedding for each of the invocation names by the processor-executable name embedding model. Stage 1316 can correspond to stage 1222 of FIG. 12. In some embodiments, generating the reference name embedding at stage 1316 includes applying a plurality of augmentation transformations to each of the set of invocation names to generate an augmented set of invocation names and generating a reference name embedding for each of the augmented set of invocation names by the processorexecutable name embedding model. At stage 1320, embodiments can storing the reference name embeddings in a non-transitory deep image.
[0122] Returning to FIG. 12, embodiments of the method 1300 can also include conversation end detection subsequent to stage 1216, as indicated by off-page reference “B.” FIG. 14 shows a flow diagram of an illustrative method 1400 for such conversation end detection. For example, it is assumed that the output of the name invoked signal at stage 1216 of FIG. 12 indicates the beginning of a conversation involving the user and a second party. At stage 1404, embodiments detect a conversation end trigger subsequent to stage 1216 (i.e., after the name invoked signal directed the ANC system automatically to enter the conversation mode). At stage 1408, in response to the detection at stage 1404, embodiments can output a conversation end signal responsive to detecting the conversation end trigger. The name invoked signal directs the ANC system automatically to switch from an ambient sound suppression mode to a conversation mode,and the conversation end signal directs the ANC system automatically to switch from the conversation mode to the ambient sound suppression mode.
[0123] Acoustic Segmentation
[0124] FIG. 15 shows a flow diagram of an illustrative method 1500 for training an automated acoustic segmentation (AAS) system for use with embodiments described herein. Embodiments of the method 1500 can be performed using an AAS system, such as the system of FIG. 7. Embodiments begin at stage 1504 by receiving an orthographic representation of each of a large number of words. As described above, the words can be received from a spoken word audio repository having stored thereon one or more speech-audio corpuses of suprasegmentally diversified speech-audio samples of phonetically diversified words. Each word of the phonetically diversified words is associated with multiple spoken audio samples including those of the suprasegmentally diversified speech-audio samples representing respective instances of the word. The orthographic representation is the written form of the word. There may be only one written form associated with each word (i.e., the class text). In some cases, the orthographic representation is received from the repository. In other cases, the orthographic representation is generated by a speech-to-text engine, or in any other suitable manner.
[0125] At stage 1508, embodiments can automatically ortho-segment the orthographic representations of each word based on pre-stored ortho-segmentation rules to generate a respective candidate segmentation for each word. At stage 1512, embodiments can automatically segment audio of each of the speech-audio samples for a word based on the candidate segmentation of the word, thereby generating a large number of candidate segmented audio samples for the word. At stage 1516, embodiments can update training of a knowledge distillation model (KDM) automatically to generate and output, for each word, a candidate ordered acoustical segmentation vector (OASV) based on automatically identifying salient features of the candidate segmented audio samples. As described herein, elements of the candidate OASVs map to an index matrix having cells corresponding to a predefined set of representative acoustical segments for a spoken language. At stage 1520, embodiments can automatically determine whether the candidate OASV output by the KDM for each word is consistent with a posterior probability matrix (PPM) for the word. The PPMs have cells corresponding to those of the index matrix. Based on the determination, at stage 1524, embodiments can output a set of X correctly segmented words for which the candidate OASV is determined to be consistent with the PPM for the word, and a set of Y incorrectly segmented words for which the candidate OASV is determined to be inconsistent with the PPM for the word, X and Y being positive integers.
[0126] A determination is made at stage 1528 as to whether Y is below a predetermined threshold (i.e., whether at least a threshold number of words can been correctly acoustically segmented). If not, at stage 1532, embodiments can re-segment at least the Y incorrectly segmented words to generate updated candidate segmentations. In some implementations, the resegmentation at stage 1532 is manual in some or all iterations. In other implementations, the resegmentation at stage 1532 is automatic, or partially automatic, in some or all iterations. Embodiments can then iterate back through stages 1512 - 1528 with the updated candidate segmentations. In some implementations, in each iteration, only the re-segmented words are run back through stages 1512 - 1528. In other implementations, all words are run back through stages 1512 - 1528. For example, any of the X correctly segmented words from a prior iteration are passed back through with the same segmentation used in that prior iteration. As described herein, the first pass through stages 1504 - 1528 can be referred to as a first training phase (or sub-phase), and subsequent passes through stages 1532 and 1512 - 1528 can be referred to as a second training phase (or sub-phase). After one or more iterations, Y will be determined at stage 1528 to fall below the threshold, and the method 1500 can end. For example, at that point, the KDM can be frozen and used for knowledge distillation-based training of the inference model (e.g., the name embedding model).
[0127] FIG. 16 shows a flow diagram of an illustrative method 1600 for automated acoustic segmentation-based attention handling in a wearable audio component (WAC), according to embodiments described herein. Embodiments can begin during runtime operation of an active noise control (ANC) system of the WAC, while a user is wearing the WAC and the ANC system is operating in an ambient sound suppression mode. Such embodiments can begin at stage 1604 by receiving a real-time audio signal. At stage 1608, embodiments can generate a real-time embedding from the real-time audio signal by a name embedding model trained automatically to acoustically segment a corpus of real-world name audio samples in accordance with a predefined set of representative acoustical segments for a spoken language. In some embodiments, the name embedding model is trained further by knowledge distillation from a knowledge distillation model (KDM). The KDM is an artificial neural network trained (e.g., according to the method 1500 of FIG. 15) automatically to acoustically segment a speech-audio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers in accordance with the predefined set of representative acoustical segments.
[0128] At stage 1612, embodiments can obtain a stored number of reference name embeddings previously generated by the name embedding model based on a set of invocation names provided by a user during an enrollment procedure. At stage 1616, embodiments can determine (e.g., by apre-trained relation network) whether any one of the reference name embeddings has a highest similarity with the real-time embedding and that the highest similarity exceeds a predetermined similarity threshold. If not, at stage 1632, embodiments can ignore the real-time audio signal and can return to stage 1604 to receive a next real-time audio signal.
[0129] At stage 1620, embodiments can output the one of the reference name embeddings as a candidate name embedding responsive to determining at stage 1616 that one of the reference name embeddings has the highest similarity with the real-time embedding and that the highest similarity exceeds the predetermined similarity threshold. At stage 1624, embodiments can determine (e.g., by a pre-trained false rejection network), responsive to the outputting at stage 1620, whether the real-time embedding and the candidate name embedding can be discriminated in excess of a predetermined discrimination threshold in any of several mathematical spaces. If not, at stage 1632, embodiments can ignore the real-time audio signal and can return to stage 1604 to receive a next real-time audio signal. If so (i.e., responsive to determining that the real-time embedding and the candidate name embedding cannot be discriminated in excess of the predetermined discrimination threshold in any of several mathematical spaces), at stage 1628, embodiments can output a name invoked signal, which directs the ANC system automatically switch from the ambient sound suppression mode to a conversation mode.
[0130] In some embodiments, generating the real-time embeddings at stage 1608 includes generating a real-time bottleneck feature embedding (BFE) by a bottleneck layer of the name embedding model and generating a real-time ordered acoustical segmentation vector (OASV) by one or more output layers of the name embedding model. In such embodiments, each of the stored plurality of reference name embeddings is also previously generated by the name embedding model to include a reference BFE and a reference OASV. In some such embodiments, the determining at stage 1616 includes determining whether one of the reference OASVs has a highest similarity with the real-time OASV, the one of the reference OASVs being the respective reference OASV of the candidate name embedding. In some such embodiments, the determining at stage 1624 includes determining whether the real-time BFE and the respective reference BFE of the candidate name embedding cannot be discriminated in excess of the predetermined discrimination threshold. In some implementations, each reference BFE and the real-time BFE is generated by the name embedding model as a latent space representation vector and / or as a set of audio tokens. In some implementations, each reference OASV and the real-time OASV are generated by the name embedding model as a 1-by-L vector of index values, each index value either indicating an unused element of the OASV, or pointing to a cell of a J-by-K index matrix, each cell of the J-by-K index matrix corresponding to a respective one of the predefined set of representative acoustical segments.
[0131] Referring back to the method 1200 of FIG. 12, the AAS-based approach described in FIG. 16 can be used for automated attention handling. For example, stage 1604 of FIG. 16 can be an implementation of stage 1204 of FIG. 12, in which a real-time audio signal is received by a reference microphone associated with an active noise control (ANC) system while the ANC system is in an ambient sound suppression mode. It can be assumed that the ANC system is integrated into a wearable audio component being worn by a first party (e.g., the user). Stages 1608 - 1624 of FIG. 16 can be an implementation of stages 1208 and 1212 of FIG. 12, in which a pre-trained inference model is used to detect whether the real-time audio signal includes attentionseeking audio spoken by a second party (i.e., the attention seeker, who is someone other than the user). The attention-seeking audio is predetermined to indicate that the second party is seeking attention of the first party. Stage 1628 of FIG. 16 can be an implementation of stage 1216 of FIG. 12, in which a name invoked signal is output to the ANC system automatically in response to determining that the real-time audio signal includes the attention-seeking audio, the name invoked signal directing the ANC system automatically to switch from the ambient sound suppression mode to a conversation mode. As illustrated in FIG. 16 and described in the context of FIG. 12, some embodiments can further include enrollment (e.g., according to the method 1300 of FIG. 13) and / or detection of a conversation end trigger (e.g., according to the method 1400 of FIG. 14).
[0132] Linguistic Name Embedding
[0133] FIG. 17 shows a flow diagram of an illustrative method 1700 for audio management that includes automated attention handling in a wearable audio component (WAC) using linguistic name embedding techniques, according to embodiments described herein. Embodiments of the method 1700 can be implemented using any of the system implementations described herein, or any other suitable variation thereof. Embodiments of the method 1700 can be a linguistic-name- embedding-specific implementation of the method 1200 of FIG. 12. Similar to stage 1204 of FIG. 12, some embodiments of the method 1700 begin at stage 1704 by receiving a real-time audio signal (e.g., from a reference microphone associated with an active noise control (ANC) system). The real-time audio signal is ambient audio received via the WAC while a user is listening to desired audio with ANC in an active (ambient sound suppression) mode.
[0134] At stage 1708, embodiments can generate a real-time embedding from the real-time audio signal by a processor-executable name embedding model trained to classify a corpus of real- world name audio samples into a linguistically differentiated set of name classifications. In someembodiments, prior to receiving the real-time audio signal at stage 1708, the name embedding model is trained at stage 1706. For example, the name embedding model is trained by transfer learning from a foundation model that is an artificial neural network trained to linguistically classify a speech-audio corpus of phonetically diversified words spoken by a plurality of accent- diversified speakers.
[0135] At stage 1712, embodiments can select a candidate name embedding based on determining, by a processor-executable relation network, that one or more of a stored plurality of reference name embeddings has at least a threshold similarity with the real-time embedding, the stored plurality of reference name embeddings previously generated by the name embedding model based on a set of invocation names provided by a user during an enrollment procedure. In some embodiments, the selecting at stage 1712 includes computing a similarity score between each of the reference name embeddings and the real-time embedding, the similarity score indicating a mathematical correlation. Such embodiments can determine whether the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold and can output the candidate name embedding as the one of the reference name embeddings yielding a highest similarity score when it is determined that the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold.
[0136] At stage 1716, embodiments can output a name invoked signal to the ANC system based on determining that the real-time embedding and the candidate name embedding cannot be discriminated by a processor-executable false rejection network in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces. The name invoked signal can direct the ANC system automatically to enter a conversation mode. In some embodiments, the outputting at stage 1716 includes transforming the real-time embedding and the candidate name embedding into a plurality of mathematical spaces. For each mathematical space, a discrimination score can be computed indicating a likelihood of discriminating between the real-time embedding and the candidate name embedding in the mathematical space. The name invoked signal can be output responsive to the discrimination score computed for at least one of the mathematical spaces falling below the predetermined discrimination threshold. Such a determination can also be performed by computing a likelihood ratio, such that the name invoked signal can be output responsive to determining that the likelihood ratio exceeds the discrimination threshold.
[0137] Embodiments of the method 1700 can include preceding and / or subsequent stages, such an indicated by off-page references. As in the method 1200 of FIG. 12, off-page reference “A”optionally references a preceding enrollment phase. An example of such an enrollment phase is described with reference to FIG. 13. Also as in the method 1200 of FIG. 12, off-page reference “B” optionally references a subsequent conversation end detection phase. An example of such a conversation end detection phase is described with reference to FIG. 14.
[0138] Unified Sound Conversion
[0139] Some of the embodiments described above automatically classify spoken audio into one of a number of linguistically distinct classifications (i.e., each classification corresponds to a unique word, regardless of suprasegmental variations). An embedding model essentially converts enrolled invocation names into an enrolled set of classifications and stores associated reference embeddings. The same embedding model can then try to classify real-time audio to produce a real-time embedding. Other models are then used to determine whether the real-time embedding matches any of the reference embeddings, which would indicate detection of one of the enrolled invocation names in the real-time audio and can be used automatically to switch ANC into a conversation mode.
[0140] Other embodiments convert spoken audio of a class (i.e., a word) into a unified sound that mimics what would be generated by a speech synthesizer for that class (i.e., the unified sound is stripped of the speaker’s suprasegmental influence on the audio). An embedding model essentially converts enrolled invocation names into an enrolled set of unified sound names and stores them as reference embeddings. The same embedding model can then convert real-time audio into a unified sound to produce a real-time embedding. Other models are then used to determine whether the real-time embedding matches any of the reference embeddings, which would indicate detection of one of the enrolled invocation names in the real-time audio and can be used automatically to switch ANC into a conversation mode.
[0141] FIGS. 18A and 18B show block diagrams of a training environment 1800 for training a foundation model 350 to support unified sound conversion (USC) for automated name-detectionbased attention handling, according to some embodiments described herein. As illustrated, the training environment 1800 can begin with a spoken word audio repository 1810. The repository 1810 can include one or more large corpuses of spoken audio data. Preferably, the corpuses include many words (classes), and many diversified samples in each class, so that each class includes many versions of the same word spoken with wide suprasegmental variance. For example, a given word may be spoken 10,000 times by different speakers from around the world. Thus, for each class, the repository 1810 can output a large number of diversified audio samples for the class, which can be referred to as the “repository spoken audio samples” for the class.
[0142] The repository spoken audio samples for each class form the entire set of spoken audio samples for that class, referred to as the “class spoken audio samples.” For example, if the repository 1810 includes S samples for a particular class (S is a positive integer), there are S class spoken audio samples. In other embodiments, additional variation in the audio samples for each class is created by passing the class spoken audio samples through an augmenter 1820. The augmenter 1820 uses one or more augmentation models 1815 to generate an augmented set of class spoken audio samples with variations in features, such as speech speed (e.g., lengthening or shortening of the audio sample, lengthening or shortening of some or all vowel sounds, etc.), modeled suprasegmental variations, models of noise profiles and / or ambient noise features (e.g., traffic sounds, background conversation sounds, etc.), etc. Some embodiments of the augmenter 1820 are implemented in the same manner as the name augmenter 1020 of FIG. 10B (e.g., and the augmentation models 1815 can be the same as, or different from the augmentation models 1015 of FIG. 10B). Other embodiments of the augmenter 1820 introduce more and / or different types of variation using more and / or other augmentation models 1815. In embodiments that include the augmenter 1820, the augmented set of audio samples for the class is used as the class spoken audio samples. For example, if the augmenter 1820 produces A augmentations for each of the S repository spoken audio samples (A is a positive integer), there are A * S class spoken audio samples. In some cases, the repository 1810 may include different numbers of samples for different classes, and / or the augmenter 1820 may apply different types of augmentations for different classes, such that the values of A and / or S may be class dependent.
[0143] Embodiments also generate a unified audio sample for each class, referred to as a “class unified audio sample.” As illustrated in FIG. 18 A, in some embodiments, the repository 1810 includes a lexical entry for each of some or all of the classes, which can be used directly as “class text.” For example, the term “INDEPENDENCE” can have hundreds of diversified spoken audio samples for the word, all stored in association with a lexical entry (i.e., the text) for the word. In other embodiments, the repository 1810 may not include lexical entries for classes, or may not include a lexical entry for one or more classes. In such embodiments, for any class that does not have an associated lexical entry, one or more of the repository spoken audio samples is fed to a speech-to-text (STT) engine 1830, which generates the class text from the repository spoken audio sample(s). The class text, whether derived from a lexical entry in the repository 1810 or generated by the STT engine 1830, is passed to a text-to-speech (TTS) synthesizer 1835 to produce the class unified audio sample. In some implementations, the class unified audio sample is stripped of all accent, prosody, etc. For example, the class unified audio sample is a purely phonetic representation of the class in a standardized set of phonemes. In other implementations, the TTSsynthesizer 1835 produces the class unified audio sample to have synthesized suprasegmental features (i.e., the class unified audio sample is stripped of the speaker’s suprasegmental influence, but it may still have synthesized suprasegmental features).
[0144] FIG. 18B shows alternative embodiments for generating the class unified audio sample. Rather than generating class unified audio sample from TTS synthesis, the class unified audio sample is generated using a selected audio sample for the class. As illustrated, embodiments can include a selector 1870. The repository 1810 includes spoken audio samples from many different speakers for each class (i.e., the class spoken audio samples), and the class unified audio sample for each class is generated by the selector 1870 by selecting one of the class spoken audio samples, and / or selecting one of the speakers of the class spoken audio samples. In one implementation, the selected spoken audio sample and / or speaker is the same for all classes. For example, the selector 1870 uses a particular identifier (e.g., index, etc.) to select a same speaker’s audio contribution as the unified sound for all classes. In other implementations, the selector 1870 randomly (or otherwise) selects a selected speaker from among the available speakers and / or samples for each class. For example, a particular speaker may not have provided audio contributions for all classes. In other implementations, the selection is based on quality features. For example, if the repository 1810 aggregates data from multiple corpuses, some corpuses may provide better quality audio samples than others; or within a particular one or more corpuses, some audio samples may be of better quality than others. In such cases, the selector 1870 is configured to select one of the class spoken audio samples as the class unified audio sample for each class based on the quality of the sample. In some cases, a particular corpus or portion of a corpus may provide better data for a particular geographic region in which the user resides. In such cases, the selector 1870 is configured to select one of the class spoken audio samples as the class unified audio sample for each class based on regional factors, such as suprasegmental features more common to a particular geographic region. For example, some implementations may ignore certain very strong regional accents for the purposes of generating the class unified audio sample, except in cases where the user resides in geographic proximity to speakers with that regional accent.
[0145] As illustrated in both FIGS. 18A and 18B, embodiments include a feature extractor 1840, which takes an audio waveform as its input and outputs a computed a set of features, referred to as a feature vector. In some implementations, the computed feature vectors are sets of filter bank energies, which are a set of values representing the respective energy content at each of a set of different frequency bands in the input audio signal. For example, the input audio signal is passed through a set of bandpass filters (e.g., each designed to respond to a specific range of frequencies that mimics the frequency response of the human auditory system). The filtered signals can thenbe rectified (e.g., to preserve magnitude information and remove phase information) and smoothed (e.g., averaged), resulting in the so-called filter bank energies. As illustrated, passing the class spoken audio samples through the feature extractor 1840 produces spoken feature vectors 1842, and passing the class unified audio sample through the feature extractor 1840 produces a unified feature vector 1844. In some implementations, the feature vectors can be generated with a representation other than filter bank energies. In some such implementations, a complex fast Fourier transform (FFT) is used to generate the feature vectors as an array of complex numbers. Each complex number in the array represents a frequency component in the audio signal, where the magnitude of the complex number represents an amplitude of a corresponding frequency component, and the phase of the complex number represents a phase shift of the corresponding component relative to a reference. In other such embodiments, MFCCs can be used (e.g., or the filter bank energies can be converted to MFCCs) for a more compact representation, but such representations may not yield sufficient accuracy for all applications (e.g., the filter resolution may effectively be too low for successful training of the foundation model 1850). In other such embodiments, sample-to-sample modeling can be used, but such implementations may be too bulky to be practical in all applications.
[0146] As illustrated, during training, a foundation model 1850 takes the spoken feature vectors 1842 as its input labels and the unified feature vector 1844 as its output label. The foundation model 1850 is iteratively trained based on the input and output labels, so that inputting any of the spoken feature vectors 1842 for a particular class into the foundation model 1850 will cause the foundation model 350 to output the same unified feature vector 1844 for that class. In some embodiments, the feature extractor 1840 is integrated with the foundation model 1850 so that the class spoken audio samples can be input to the foundation model 1850, and the foundation model 1850 will try to generate an output unified audio sample that mimics the class synthesized audio sample (or to output a feature vector that mimics the unified feature vector 1844).
[0147] As one example, each audio sample at the input to the feature extractor 1840 is one second of audio sampled at 16 kHz (i.e., 16,000 samples). The sample is split into 160 frames of ten milliseconds each. The feature extractor 1840 extracts 30 Mel frequency bins (e.g., using bandpass filters tuned to the Mel scale) from each frame, thereby compressing each audio input of 16,000 samples into 30 * 100, or 3,000, filter bank energies. In this example, the feature extractor 1840 has an input dimension of 16,000 and an output dimension of 3,000. Suppose the repository 1810 has 10,000 repository spoken audio samples for a particular word, and the augm enter 1820 generates ten augmentations for each repository spoken audio sample, resulting in 100,000 class spoken audio samples. The feature extractor 1840 converts each of the 100,000 class spoken audiosamples into a respective 3,000 filter bank energies (i.e., the spoken feature vectors 1842), and converts the single class unified audio sample into another respective 3,000 filter bank energies (i.e., the unified feature vector 1844); and the foundation model 1850 uses all those filter bank energies to figure out how to convert any of the spoken feature vectors 1842 into the same unified feature vector 1844.
[0148] Ultimately, the foundation model 1850 is trained so that inputting any input audio sample 1860 will cause the foundation model 1850 to generate an output unified feature vector 1865 that is stripped of any speaker influence on suprasegmental features. For example, the generated output unified feature vector 1865 mimics the unified feature vector 1844 that would be generated if the input audio sample were to represent a class from the repository 1810, class text corresponding to the class were passed through the TTS synthesizer 1835 to generate a class synthesized audio sample, and the class synthesized audio sample were passed through the feature extractor 1840 to generate the unified feature vector 1844. In other words, the foundation model 1850 is trained to convert any input audio sample 1860 into a unified sound representation.
[0149] Having trained the foundation model 1850 to generate a unified sound version of an input audio sample, use of the foundation model 1850 in automated attention handling can be similar to what is described with reference to LNE-AHS above. FIG. 19 shows a simplified block diagram of stages of an illustrative implementation of an attention seeker (AS) trigger detection block 1900 and the association of each stage with a corresponding portion of a generated inference model, according to embodiments described herein. The AS trigger detection block 1900 can be implemented in a similar manner to the AS trigger detection block 400 of FIG. 4 and consistent reference designators are used to indicate such similarities. For example, both include an enrollment stage 410 that serves the same overall purpose with respect to the system and are shown with the same reference designator, accordingly; but each enrollment stage 410 yields a different type of embedding model, and the embedding models are shown with different reference designators, accordingly. Like the AS trigger detection block 400 of FIG. 4, the AS trigger detection block 1900 of FIG. 19 can be implemented in the context of a WAC 310.
[0150] Embodiments can include an enrollment stage 410, an identification stage 430, and a verification stage 440. In general, the enrollment stage 410 occurs outside of normal operation of the WAC 310, such as when a user first sets up and / or uses the WAC 310, when the user first registers the WAC 310, when the user first configures the WAC 310 for automated attention handling, etc. The identification stage 430 and the verification stage 440 are real-time blocks that occur during normal operation of the WAC 310 to facilitate real-time automated attentionhandling. Embodiments generally use the enrollment stage 410 to obtain any information needed to set up an inference model for storing at the WAC 310 to enable the WAC 310 to perform automated attention handling features during normal operation. As shown, the inference model includes at least an embedding model 1915, a relation network 1935, and a false rejection network 1945.
[0151] In the enrollment stage 410, the embedding model 1915 can be used to convert an enrollment audio stream 405 into a set of reference embeddings based on a foundation model 1850. The foundation model 1850 can be stored in and accessed via the cloud 340 (e.g., and / or any other suitable communication network). As described with reference to FIG. 18, in the context of USC-based embodiments, the foundation model 1850 can be trained to convert an input audio stream into either a unified output audio sample (i.e., corresponding to the input audio stream stripped of any speaker influence on accent, prosody, etc.), or an output unified feature vector 1865 (e.g., a set of filter bank energies that describe the unified output audio sample). In this context, the embedding model 1915 is trained by transfer learning from the foundation model 1850 on a smaller corpus of name audio samples. For example, the encoder portion of the foundation model 1850 is taken by the embedding model 1915, and a new decoder is trained on the smaller name audio corpus.
[0152] As described above (e.g., with reference to FIGS. 4, 10A, 10B, 11, and 13), during the enrollment stage 410, the enrollment audio stream 405 can include several user-enrolled audio samples of invocation names and / or a set of augmentations to those invocation names. In this context, the “reference embeddings” refer to either the set of unified output audio samples generated from the set of invocation names (e.g., and their augmentations), or the set of output unified feature vectors 1865 generated from the set of invocation names (e.g., and their augmentations). The reference embeddings can be stored as a deep image 1920. The deep image 1920 can also be considered as part of the inference model.
[0153] During normal operation, in the identification stage 430, a real-time (RT) audio stream 407 is received. The embedding model 1915 (i.e., the same embedding model 1915 generated in the enrollment stage 410) is used to generate a RT embedding from the received RT audio stream 407. In this context, the “RT embedding” refers to either a unified output audio sample generated from the RT audio stream 407, or an output unified feature vector 1865 generated from the RT audio stream 407. The relation network 1935 can then compare the RT embedding with each of the reference embeddings in the deep image 1920 to determine if there is a match. For example, the relation network 1935 is configured to compute a similarity score for each comparison (e.g.,corresponding to a mathematical correlation, or the like). The relation network 1935 can determine if any similarity scores meet or exceed a predetermined threshold; if so, the reference embedding with the highest similarity score can be selected as a candidate matching embedding. Further options and / or features of the relation network 1935 can be understood with reference to descriptions of the relation network 435 above.
[0154] In the verification stage, the false rejection network 1945 can then confirm the match by transforming the RT embedding and the candidate matching embedding into different mathematical spaces (e.g., domains) and determining whether the embeddings can be reliably discriminated from each other in any of the mathematical spaces. For example, the deep image 1920 is configured to compute a discrimination score for each mathematical space. The deep image 1920 can determine if any discrimination scores meet or exceed a predetermined threshold; if so, the match determined by the relation network 1935 is determined to be a false match and is ignored. If the embeddings cannot be sufficiently discriminated in any mathematical spaces, the false rejection network 1945 can output a signal that attention seeking audio has been detected (i.e., a name invoked signal 450). As described herein, the name invoked signal 450 can direct an ANC system of the WAC 310 automatically to enter a conversation mode. As noted with reference to FIG. 4 above, embodiments of the embedding model 1915 are a processor-executable name embedding model, embodiments of the deep image 1920 are a processor-readable deep image, embodiments of the relation network 1935 are a processor-executable relation network, and embodiments of the false rejection network 1945 are a processor-executable false rejection network. Further options and / or features of the false rejection network 1945 can be understood with reference to descriptions of the false rejection network 445 above.
[0155] It can be seen that, although the stages used in unified sound conversion embodiments are similar to those used in LNE-AHS embodiments, the underlying modeling approach used for name detection is different. For example, in both categories of embodiments, input audio can be encoded using types of language modeling, linguistic embedding, etc. However, the LNE-AHS embodiments use a classifier architecture that essentially converts the encoded linguistically salient features into a classification, while the unified sound conversion embodiments use an encoder-decoder architecture that essentially converts the encoded linguistically salient features into a unified audio output. For example, as described above, the LNE-AHS embodiments can rely on architectures, such as two-dimensional convolution with long short-term memory (LSTM) networks, etc.; while unified sound conversion embodiments can rely on architectures, such as two-dimensional convolution with recurrent neural networks (e.g., a gated recurrent unit (GRU) network).
[0156] As described above, some USC-based embodiments implement the embedding model 1915 to include both an encoder and a decoder, so that the output of the embedding model 1915 is the output unified feature vector 1865 (e.g., synthesized filter bank energies). As illustrated in FIGS. 18A and 18B, in other embodiments, the embedding model 1915 is implemented only to include the encoder portion of the model (i.e., all layers of the model after the bottleneck layer are removed after training). As described above, encoder-decoder architectures (e.g., auto-encoder architectures) include a neural network architecture designed to learn compact representations of data, referred to as “latent features.” The encoder portion receives a high-dimensionality input (i.e., the raw audio data) and includes a multi-layer network to progressively reduce the dimensionality of the data using transformations, until a lowest dimensionality (i.e., most compressed) representation is reached at the bottleneck layer. The vector space that is represented at the bottleneck layer can be referred to as a latent space embedding of the audio sample. In some embodiments, the latent space embedding generated by the encoder (e.g., at the bottleneck layer) is used as a unified sound code word (USCW) 1867.
[0157] In such embodiments, the embedding model 1915 is trained to generate respective USCWs 1867 as the reference embeddings stored in the deep image 1920 and to generate a respective USCW 1867 as the RT embedding, the relation network 1935 is trained to determine candidate matches based on comparing the USCWs 1867, and the false rejection network 1945 is trained to reject false matches by discriminating between the USCWs 1867. With proper training and implementation, the performance of the AS trigger detection block 1900 can be almost the same for embodiments of the inference model that include a decoder and are based on output unified feature vectors 1865 as for embodiments of the inference model that do not include a decoder and are based on UCSWs 1867.
[0158] FIG. 20 shows a flow diagram of an illustrative method 2000 for automated audio management based on name detection, according to USC-based embodiments described herein. Embodiments of the method 2000 begin at stage 2004 by obtaining (e.g., by one or more processors) spoken audio samples for each word of a large number of phonetically diversified words from a spoken word audio repository. As described herein, the spoken word audio repository includes one or more speech-audio corpuses of suprasegmentally diversified speechaudio samples of the phonetically diversified words. As such, the spoken audio samples for each word of the phonetically diversified words includes those of the suprasegmentally diversified speech-audio samples representing respective instances of the word.
[0159] At stage 2008, embodiments can obtain (e.g., by the one or more processors) a respective text representation of each word of the phonetically diversified words. For some or all the phonetically diversified words, some embodiments can convert one or more of the class spoken audio samples associated with the word into the respective text representation of the word. For example, a class spoken audio sample is speech-to-text converted in stage 2006 to generate the textual representation. In some implementations, for each of at least some of the phonetically diversified words, the spoken word audio repository includes a lexical entry for the word (e.g., stored in association with at least one of the class spoken audio samples for the word). In such implementations, embodiments of the method 2000 can obtain the respective text representation of each of some or all the phonetically diversified words in stage 2008 by obtaining the lexical entry for the word from the spoken word audio repository.
[0160] At stage 2012, embodiments can synthesize (e.g., by the one or more processors) a respective class unified audio sample for each word from a respective text representation of the word. At stage 2016, embodiments can convert, for each word (e.g., by the one or more processors), the respective class unified audio sample for the word into respective output labels representing a unified feature vector of the respective class unified audio sample. At stage 2020, embodiments can convert, for each word (e.g., by the one or more processors), the class spoken audio samples associated with the word into respective input labels representing spoken feature vectors of the class spoken audio samples. In some embodiments, the spoken feature vectors represent filter bank energies of the plurality of class spoken audio samples, and the unified feature vector represents filter bank energies of the respective class unified audio sample.
[0161] Some embodiments, at stage 2018, prior to the converting in stage 2020, can add class augmented audio samples to the class spoken audio samples associated with the word by applying one or more augmenter models to one or more of the class spoken audio samples from the spoken word audio repository (i.e., the class spoken audio samples for the word include both the class spoken audio samples from the spoken word audio repository and the class augmented audio samples). For example, each of the augmenter models mathematically represents a respective one of several noise models and / or suprasegmental feature models (e.g., models of different accents, prosody, etc.).
[0162] At stage 2024, embodiments can train a foundation model. As described herein, the foundation model is an artificial neural network trained, based on the input labels and the output labels for each word, to convert any particular one of the class spoken audio samples for the word into the respective class unified audio sample for the word by automatically removingsuprasegmental differences between the particular class spoken audio sample and the class unified audio sample. For example, once properly trained, the foundation model is able to remove a speaker’s suprasegmental influence on spoken audio received from the speaker as the input stream, thereby generating a unified version of the audio that mimics what would be generated by text-to- speech synthesis.
[0163] FIG. 21 shows a flow diagram of an illustrative method 2100 for training a unified sound conversion (USC) inference model. As represented by off-page reference “C,” the method 2100 can proceed based on and subsequent to the training of the foundation model in stage 2024 of FIG. 20. At stage 2104, embodiments train the USC inference model (e.g., a processor-executable USC inference model) based on transfer learning from the foundation model. The training is such that the USC inference model can automatically: receive a real-time audio stream; generate a real-time embedding representing the real-time audio stream stripped of suprasegmental features; and output a name invoked signal based on matching, with at least a predetermined threshold confidence level, the real-time embedding to one of a stored plurality of reference name embeddings generated to represent unified audio representations of each of a set of invocation names stripped of the suprasegmental features, the name invoked signal to direct an ANC system automatically to enter a conversation mode.
[0164] In some embodiments, the training in stage 2104 can proceed according to stages 2108 - 2116. At stage 2108, embodiments can train a name embedding model, by transfer learning from the foundation model based on a corpus of real-world name audio samples, automatically to remove the suprasegmental features from an input audio sample to generate a corresponding unified audio output. In such embodiments, the real-time embedding is generated by the name embedding model from the real-time audio signal received during an operational time, and the reference name embeddings are generated by the name embedding model based on enrollment audio samples received during an enrollment time.
[0165] At stage 2112, embodiments can train a relation network to output a candidate name embedding responsive to determining that one of the reference name embeddings has a highest similarity with the real-time embedding and that the highest similarity exceeds a predetermined similarity threshold. For example, the relation network is trained to output the candidate name by: computing a similarity score between each of the reference name embeddings and the real-time embedding, the similarity score indicating a mathematical correlation; determining whether the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold; and outputting the candidate name embedding as the one of the referencename embeddings yielding a highest similarity score when it is determined that the similarity score computed for any of the reference name embeddings exceeds the predetermined similarity threshold.
[0166] At stage 2116, embodiments can train a false rejection network to output the name invoked signal responsive to determining that the real-time embedding and the candidate name embedding cannot be discriminated in excess of a predetermined discrimination threshold in any of a plurality of mathematical spaces. For example, the false rejection network is trained to output the name invoked signal by: transforming the real-time embedding and the candidate name embedding into a plurality of mathematical spaces; computing, for each mathematical space, a discrimination score indicating a likelihood of discriminating between the real-time embedding and the candidate name embedding in the mathematical space; and outputting the name invoked signal responsive to the discrimination score computed for at least one of the mathematical spaces falling below the predetermined discrimination threshold. In embodiments that perform stages 2112 and 2116, the predetermined confidence level in stage 2104 can be based on a combination of the predetermined similarity threshold and the predetermined discrimination threshold.
[0167] FIG. 22 shows a flow diagram of an illustrative method 2200 for using the trained unified sound conversion (USC) inference model during an operation time. As represented by off- page reference “D,” the method 2200 can proceed based on and subsequent to the training of the USC inference model in stage 2104 of FIG. 21. Embodiments of the method 2200 begin at stage 2204 by receiving a real-time audio signal by the USC inference model (e.g., from a reference microphone associated with the ANC system). At stage 2208, embodiments can generate a realtime embedding from the real-time audio signal by the USC inference model. At stage 2212, embodiments can output the name invoked signal in response to (e.g., only after determining) determining (e.g., only after determining), by the USC inference model, that the real-time embedding matches any of the stored plurality of reference name embeddings with at least the predetermined threshold confidence level.
[0168] Similar to the method 1700 of FIG. 17, embodiments of the method 2200 can include additional preceding and / or subsequent stages, such an indicated by off-page references “A” and “B”. As described above, off-page reference “A” optionally references a preceding enrollment phase. An example of such an enrollment phase is described with reference to FIG. 13. Also as described above, off-page reference “B” optionally references a subsequent conversation end detection phase. An example of such a conversation end detection phase is described with reference to FIG. 14.
[0169] Hybrid Unified Linguistic Name Embedding (ULNE)
[0170] Some embodiments described above perform automated name detection-based attention handing using linguistic name embedding (LNE). LNE-based embodiments rely on a classification model to convert names into linguistically distinct classifications based on linguistically salient features. Such a classification model tends to have a performance tradeoff between a true positive rate and a false acceptance rate for names and sounds that are linguistically similar. Additionally, to ensure good performance of such LNE-based approaches when there is discrepancy between the prosody or accent of enrolled names and those of microphone-captured names uttered by attention seekers, such LNE-based approaches tend to involve augmentation of audio samples with many types of time scaling, pitch shifting, etc., which tends to increase the computation load of the identification stage.
[0171] Other embodiments described above perform automated name detection-based attention handing using unified sound conversion (USC). USC-based embodiments rely on an encoderdecoder model to convert names into unified sound representations that stripped of any speaker influence on suprasegmental features of the audio. This suprasegmental unification (e.g., normalization) can help USC-based embodiments avoid some of the limitations of LNE-based embodiments. However, some USC-based embodiments can have their own limitations, such as challenges in classifying similar sounding names from the resulting types of embedding or normalized features.
[0172] Other embodiments perform automated name detection-based attention handing using a hybrid unified linguistic name embedding (ULNE) approach, which combines features of USC- based embodiments with features of LNE-based embodiments. FIG. 23 shows a simplified block diagram of an embedding environment 2300 incorporating an illustrative hybrid name embedding model 2310 for use in a unified linguistic name embedding attention handling system (ULNE- AHS). As illustrated, the hybrid name embedding model 2310 includes two encoding stages. In a first stage, a USC-trained embedding model includes a USC encoder 2320 and a USC decoder 2325, which can be components of embedding model 1915, as described with reference to FIGS. 18A and 18B. The USC-trained embedding model is trained to convert an input audio stream 2305 into a unified sound representation as a set of unified features. For example, the unified features are an output unified feature vector 1865. In a second stage of the hybrid name embedding model 2310, an LNE-trained embedding model includes an LNE encoder 2330 and an LNE classifier 2335, which can be components of embedding model 415, as described withreference to FIG. 4. The LNE-trained embedding model is trained to convert the unified features at the output of the USC-trained embedding model into an output classification 2340.
[0173] Ultimately, the output classification 2340 at the output of the hybrid name embedding model 2310 can be the same classification output as in the LNE-based embodiments described above (e.g., at the output of name embedding model 415). As such, ULNE-based embodiments can use the same relation network 435 and the same false rejection network 445 used in the LNE- based embodiments. For example, referring to FIG. 4, the embedding model 415 can be replaced with the hybrid name embedding model 2310 from FIG. 23. In that context, during the enrollment stage 410, the hybrid name embedding model 2310 converts the enrollment audio stream 405 into corresponding classifications for storage in the deep image 420 by converting each name audio sample into unified features and converting the unified features into an appropriate classification. As such, during the enrollment stage 410, the output classifications 2340 generated by the hybrid name embedding model 2310 are used as the reference embeddings. During real-time operation, the hybrid name embedding model 2310 converts the RT audio stream 407 into a corresponding classification by converting the received audio into unified features and converting the unified features into an appropriate classification. As such, during the enrollment stage 410, the output classification 2340 generated by the hybrid name embedding model 2310 is used as the RT embedding. The classifications represented by the reference and RT embeddings are then used by the relation network 435 and the false rejection network 445 to determine whether to output the name invoked signal 450.
[0174] In effect, unified sound conversion is used at the front-end to normalize prosody and accent prior to performing linguistic name embedding, so that the linguistic name embedding is performed on a unified version of the name stripped of any particular accent and / or prosody influence by the attention seeker. This can enable LNE-based classification to generate well- defined and linguistically discriminative name embedding with sustained and effective performance across a wide variety of accents and prosody. As such, performance of ULNE-based embodiments can be better than embodiments based on either LNE alone or USC alone.
[0175] Referring back to the method 1200 of FIG. 12, any of the acoustic segmentation-based, LNE-based, USC-based, or ULNE-based approaches described herein can be used for automated attention handling. For example, in stage 1204, a real-time audio signal is received by a reference microphone associated with an active noise control (ANC) system while the ANC system is in an ambient sound suppression mode. It can be assumed that the ANC system is integrated into a wearable audio component being worn by a first party (e.g., the user). At stage 1208, a pre-trainedinference model is used to detect whether the real-time audio signal includes attention-seeking audio spoken by a second party (i.e., the attention seeker, who is someone other than the user). The attention-seeking audio is predetermined to indicate that the second party is seeking attention of the first party. At stages 1212 and 1216, embodiments can output a name invoked signal to the ANC system automatically in response to determining that the real-time audio signal includes the attention-seeking audio, the name invoked signal directing the ANC system automatically to switch from the ambient sound suppression mode to a conversation mode. As described there, some embodiments can further include enrollment (e.g., according to the method 1300 of FIG. 13) and / or detection of a conversation end trigger (e.g., according to the method 1400 of FIG. 14).
[0176] Generally, the attention-seeking audio can be predetermined to indicate that the second party is seeking attention of the first party based on a set of invocation names previously enrolled by the first party that is previously converted by the pre-trained inference model into a set of reference embeddings. As such, the detecting can generally involve: converting the real-time audio signal to a real-time embedding by the pre-trained inference model; and determining whether the real-time embedding matches any of the reference embeddings with at least a threshold confidence level. The converting of the real-time audio signal to the real-time embedding can be performed by a processor-executable name embedding model.
[0177] For example, in LNE-based embodiments, the name embedding model is an LNE-based name embedding model trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications by transfer learning from a foundation model, and the foundation model is an artificial neural network trained to linguistically classify a speech-audio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers. In USC-based embodiments, the name embedding model is a USC-based name embedding model trained, by transfer learning from a foundation model based on a corpus of real- world name audio samples, automatically to remove suprasegmental features from an input audio sample to generate a corresponding unified audio output; and the foundation model is an artificial neural network trained, based on a speech-audio corpus of phonetically diversified words spoken by a plurality of accent-diversified speakers, to generate class unified audio samples from class spoken audio samples by removing the suprasegmental features.
[0178] In ULNE-based embodiments, the name embedding model is a ULNE-based name embedding model (i.e., hybrid name embedding model 2310) that can include: a first embedding stage trained automatically to remove suprasegmental features from a corpus of real-world name audio samples to generate corresponding unified audio outputs; and a second embedding stagecoupled with the first embedding stage and trained to classify a corpus of real-world name audio samples into a linguistically differentiated set of name classifications. In such embodiments, the converting of the real-time audio signal to the real-time embedding by the name embedding model can involve converting the real-time audio signal, by the first embedding stage, into a unified realtime audio signal stripped of suprasegmental influence by the second party; and converting the unified real-time audio signal, by the second embedding stage, into the real-time embedding, the real-time embedding representing a name classification.
[0179] Computational Systems
[0180] FIG. 24 provides a schematic illustration of an illustrative computational system 2400 that can implement various system components and / or perform various steps of methods provided by various embodiments. Embodiments of the computational system 2400 can be integrated in a WAC, such as an earbud, headset, etc. Embodiments of the computational system 2400 can implement some or all of the audio management system 100 of FIG. 1, including embodiments of the AHS 150, the ANC 140, and / or the audio processing system (APS) 160 described herein. Additionally or alternatively, embodiments can implement any of the training environments described herein, and / or can execute any of the machine learning models described herein. FIG. 24 is meant only to provide a generalized illustration of various components, any or all of which may be utilized as appropriate. FIG. 24, therefore, broadly illustrates how individual system elements may be implemented in a relatively separated or relatively more integrated manner.
[0181] The computational system 2400 is shown including hardware elements that can be electrically coupled via a bus 2405 (or may otherwise be in communication, as appropriate). The hardware elements may include one or more processors 2410, including, without limitation, one or more general -purpose processors and / or one or more special-purpose processors (such as digital signal processing chips, graphics acceleration processors, video decoders, and / or the like); one or more input devices 2415; and one or more output devices 2420. In the WAC context, the input devices 2415 can include wired and / or wireless ports, buttons, switches, microphones, touch interfaces, and / or any other suitable input device 2415; and the output devices 2420 can include indicator lights, displays, speakers, and / or any other suitable output devices 2420.
[0182] The computational system 2400 may further include (and / or be in communication with) one or more non-transitory storage devices 2425, which can comprise, without limitation, local and / or network accessible storage, and / or can include, without limitation, a disk drive, a drive array, an optical storage device, a solid-state storage device, such as a random access memory (“RAM”), and / or a read-only memory (“ROM”), which can be programmable, flash-updateableand / or the like. Such storage devices may be configured to implement any appropriate data stores, including, without limitation, various file systems, database structures, and / or the like. In some embodiments, the storage devices 2425 include the deep image 420 and / or an inference model 2427. As described herein, the inference model can include one or more types of name embedding models, relation networks, false rejection networks, etc. for implementing name detection-based attention handling.
[0183] The computational system 2400 can also include a communications subsystem 2430, which can include, without limitation, a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset (such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, cellular communication device, etc.), and / or the like. As described herein, the communications subsystem 2430 supports multiple communication technologies. Further, as described herein, the communications subsystem 2430 can provide communications with one or more networks 140, and / or other networks. For example, embodiments of the communications subsystem 2430 can communicate with aKDM 550 foundation model 350 via the cloud 350. Though not explicitly shown, some embodiments interface via the communications subsystem 2430, and / or via input devices 2415 and output devices 2420, with one or more user computational devices 320.
[0184] In many embodiments, the computational system 2400 will further include a working memory 2435, which can include a RAM or ROM device, as described herein. The computational system 2400 also can include software elements, shown as currently being located within the working memory 2435, including an operating system 2440, device drivers, executable libraries, and / or other code, such as one or more application programs 2445, which may include computer programs provided by various embodiments, and / or may be designed to implement methods, and / or configure systems, provided by other embodiments, as described herein. Merely by way of example, one or more procedures described with respect to the method(s) discussed herein can be implemented as code and / or instructions executable by a computer (and / or a processor within a computer); in an aspect, then, such code and / or instructions can be used to configure and / or adapt a general-purpose computer (or other device) to perform one or more operations in accordance with the described methods. In some embodiments, the operating system 2440 and the working memory 2435 are used in conjunction with the one or more processors 2410 to implement some or all of the audio management system 100 components, such as the ANC 140, AHS 150, and / or APS 160.
[0185] A set of these instructions and / or codes can be stored on a non-transitory computer- readable storage medium, such as the non-transitory storage device(s) 2425 described above. In some cases, the storage medium can be incorporated within a computer system, such as computer system 2400. In other embodiments, the storage medium can be separate from a computer system (e.g., a removable medium, such as a compact disc), and / or provided in an installation package, such that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the instructions / code stored thereon. These instructions can take the form of executable code, which is executable by the computational system 2400 and / or can take the form of source and / or installable code, which, upon compilation and / or installation on the computational system 2400 (e.g., using any of a variety of generally available compilers, installation programs, compression / decompression utilities, etc.), then takes the form of executable code.
[0186] It will be apparent to those skilled in the art that substantial variations may be made in accordance with specific requirements. For example, customized hardware can also be used, and / or particular elements can be implemented in hardware, software (including portable software, such as applets, etc.), or both. Further, connection to other computing devices, such as network input / output devices, may be employed.
[0187] As mentioned above, in one aspect, some embodiments may employ a computer system (such as the computer system 2400) to perform methods in accordance with various embodiments of the invention. According to a set of embodiments, some or all of the procedures of such methods are performed by the computational system 2400 in response to processor 2410 executing one or more sequences of one or more instructions (which can be incorporated into the operating system 2440 and / or other code, such as an application program 2445) contained in the working memory 2435. Such instructions may be read into the working memory 2435 from another computer-readable medium, such as one or more of the non-transitory storage device(s) 2425. Merely by way of example, execution of the sequences of instructions contained in the working memory 2435 can cause the processor(s) 2410 to perform one or more procedures of the methods described herein.
[0188] The terms “machine-readable medium,” “computer-readable storage medium” and “computer-readable medium,” as used herein, refer to any medium that participates in providing data that causes a machine to operate in a specific fashion. These mediums may be non-transitory. In an embodiment implemented using the computer system 2400, various computer-readable media can be involved in providing instructions / code to processor(s) 2410 for execution and / or can be used to store and / or carry such instructions / code. In many implementations, a computer-readable medium is a physical and / or tangible storage medium. Such a medium may take the form of a non-volatile media or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as the non-transitory storage device(s) 2425. Volatile media include, without limitation, dynamic memory, such as the working memory 2435. Common forms of physical and / or tangible computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, any other physical medium with patterns of marks, a RAM, a PROM, EPROM, a FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.
[0189] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to the processor(s) 2410 for execution. Merely by way of example, the instructions may initially be carried on a magnetic disk and / or optical disc of a remote computer. A remote computer can load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received and / or executed by the computer system 2400. The communications subsystem 2430 (and / or components thereof) generally will receive signals, and the bus 2405 then can carry the signals (and / or the data, instructions, etc., carried by the signals) to the working memory 2435, from which the processor(s) 2410 retrieves and executes the instructions. The instructions received by the working memory 2435 may optionally be stored on a non-transitory storage device 2425 either before or after execution by the processor(s) 2410.
[0190] Refinement by Progressive Learning
[0191] As described above, several approaches are available for name detection, including using acoustic segmentation approaches, linguistic name embedding (LNE) approaches, universal sound conversion (USC) approaches, and hybrid (e.g., LTLNE) approaches. Although such approaches represent significant improvements over prior approaches, they still tend to have certain limitations. For example, name detection becomes increasingly difficult for short (e.g., monosyllabic) names, in presence of large person-to-person variations in the manner of speaking a same name, in presence of noise and / or other ambient sounds, in presence of sources of distortion, etc. As one example, when seeking to detect a short name (or short predefined hot word or phrase), there may not be a long enough sample for reliable discrimination and detection, and monosyllabic utterances are more likely to be part of long irrelevant utterances. These and other factors can lead to higher false acceptance rates for background voices and / or other undesirable consequences.
[0192] The various name detection approaches described herein can all be generally described as having three blocks: a name embedding model, an identification model, and a verification model. In general, the name embedding model generates embeddings for the enrolled names. As described herein, the embeddings generated during enrollment are referred to as a deep image, and each embedding in the deep image is referred to herein as a reference embedding. The embedding model also generates embeddings in for the real-time (RT) audio stream as captured by the microphone while a wearer of a WAC is engaging with music or audio.
[0193] The embeddings represent a highly compact representation of the names that preserve only the most salient features for the particular type of name detection. For example, in LNE- based approaches, the name embedding model is trained so that the embeddings represent the most salient features for automatically classifying spoken audio samples into linguistically distinct name classifications. In USC-based approaches, the name embedding model is trained so that the embeddings represent the most salient features for automatically removing speaker influence from spoken audio samples. In acoustical segmentation-based approaches, the name embedding model is trained so that the embeddings represent the most salient features for automatically converting spoken audio samples into a sequence of acoustical segments. In hybrid approaches, USC can be used at the front-end of LNE and / or acoustical segmentation approaches to effectively performed those approaches on unified versions of a name stripped of any particular suprasegmental influence by the attention seeker.
[0194] For each type of name detection, the identification model generally determines similarity scores between real-time embeddings and the reference embeddings in the deep image to determine whether there is a match. For example, if any similarity scores computed by the identification model meet or exceed a predetermined threshold, the reference embedding having the highest similarity score is selected as a candidate matching embedding.
[0195] For each type of name detection, the verification model generally confirms whether the candidate match is a real match. For example, the verification model transforms the real-time embedding and the candidate matching embedding into different mathematical spaces or domains and determines whether the embeddings can be reliably discriminated from each other in any of the mathematical spaces. The verification model can compute a discrimination score for each mathematical space and can check if any discrimination scores meet or exceed a predetermined threshold. If the embeddings can be sufficiently discriminated in any mathematical spaces, the match determined by the identification model is determined to be a false match and is ignored.Otherwise, the verification model can output a signal that attention-seeking audio has been detected, and the detected name can be retrieved and used as an indication for the wearer.
[0196] Embeddings generated by the name embedding model can become damaged when the real-time audio stream contains background noises or competing speech influences. The enrollment can be performed in a quiet enough environment, as the enrollment tends to occur only once, at a time and in a location under user control. However, general usage of the WAC can be at any time and place, in which there may be little or no wearer control over the environment leaving any name detection susceptible to ambient noises or interfering voices. It can be desirable therefore to provide reliable name detection in such diverse environments.
[0197] Indeed, differences between the enrollment environment and subsequent usage environments can manifest differences between the types of audio used to generate reference embeddings and the types of audio used to generate real-time embeddings. For example, real-time embeddings generated by the name embedding model will have more influence from noise than the reference embeddings, and / or they may be generated from audio received from different locations or distances than what was used during enrollment (i.e., yielding different kinds of ambient impacts, distortions, etc.). One approach is to train the foundation model with noise augmentation. However, certain embeddings for names may not be sufficiently segregated in the vector space to be well discriminated or verified due to the overriding impact of noises and / or background voices. This can tend to reduce performance of name detection in noisy and / or other environments. Another approach is to apply a noise reduction model at the front end. Although such an approach can effectively reduce noise-related concerns, it tends to add significant computation and memory. For name detection to operate in a desirable manner, name detection may always be on (at least while ANC is active), such that any front-end noise reduction model would also always be on, thereby quickly draining battery resources, etc. Further, reducing noise tends not to address performance reductions resulting from competing speech in the environment.
[0198] Embodiments described in this section refine name detection approaches described above using a multi-stage progressive learning approach. The progressive training seeks to improve name detection performance in the presence of noise and other competing speech. Embodiments begin by building an accurate name embedding model using audio data recorded in a quiet ambient environment. A refined embedding model (REM) can be generated by transfer learning from the name embedding model to remove embedding variations across common linguistic information, content, or names due to speaker influences such as prosody, emotion, accent, etc. (e.g., suprasegmental factors). A noise-robust REM (NRREM) can be generated by expanding themodel (e.g., adding a few layers at the output of the bottleneck layer of the embedding model) to filter noises (e.g., kitchen noise, music, babble, cafe, traffic, driving, wind, etc.) and generate an embedding that is very close to one obtained when noise influences are not there for a particular name or class. Output layers can be reintroduced, as needed, based on the type of name detection being used. For example, in LNE-based approaches, classification layers can be added so that the name embedding model outputs a classification.
[0199] As described above, there are several approaches to generating a foundation model to act as a name embedding model. For example, LNE-based approaches can be used to train the foundation model to remove speaker influence (e.g., suprasegmental factors) from words to generate classifications, so that all audio samples representing the same linguistic information (with different suprasegmental factors) are in the same class. Similarly, acoustic segmentationbased approaches can be used to train the foundation model to remove speaker influence from words to generate acoustic segments, so that all audio samples representing the same linguistic information generate the same acoustical segments. For example, hundreds of speakers can provide audio samples of them saying the name “SHIVA” with different accents, prosody, etc., and the foundation model is trained to pull out all the salient features so that all those audio samples will either be classified into a single classification associated with the word “SHIVA,” or will generate a same sequence of acoustical samples representing ‘S(H)F / ‘VA’, or the like.
[0200] As noted above, the foundation models generally include input layers, a bottleneck layer, and output layers, and embeddings can be pulled from the bottleneck layer. Thus, refinement of the foundation models into refined embedding models (REMs) and noise-robust REMs (NRREMs) can be performed on embedding models that do not include the output (e.g., decoder) layers. For example, for LNE-based foundation models, the functionality of the decoder can be that of a simple classifier without any modeling or transformation, such that the embedding models are able to provide well discriminable embeddings across different classes. In some implementations, this is accomplished by using name audio samples captured in a quiet environment (i.e., having a relatively low noise floor and generally free of ambient noises, competing speech, etc.) to train the embedding model, and using very few layers with a basic dense network as decoders.
[0201] FIG. 25 shows a block diagram of a first refinement stage training environment 2500 for an embedding model, according to embodiments described herein. The training environment 2500 can be tailored to refinement of embedding models trained for LNE, USC, UNLE, acoustic segmentation, and / or other name embedding approaches. To avoid overcomplicating the description, refinement of the embedding model is described with reference to a foundation model2515 trained for LNE or ULNE. As described above, for any given class (e.g., word), there may be multiple (e.g., 10, 100, etc.) class audio samples representing that same class spoken in different ways (e.g., by different speakers, at different speeds, with different prosody, etc.). The goal of training the foundation model 2515 is to be able to classify with sufficient accuracy, so that all class audio samples representing a same class will be converted to embeddings that are sufficiently close to the same as to be accurately classified (i.e., spoken audio representing the same linguistic information will be placed into the same class).
[0202] For the first refinement stage, the foundation model 2515 can be converted to an embedding model by using the linguistic name embedding model with the output classifier layers removed. As such, embeddings of input audio samples can be captured from the bottleneck layer of the model. In some embodiments, the refinement is performed for all classes (i.e., for all names).
[0203] In other embodiments, the refinement is performed only on classes determined to generate embeddings that deviate more than a predetermined threshold. In such embodiments, for each class, the foundation model 2515 is analyzed for different audio segments representing that class. For each audio sample, an embedding is obtained. If the embedding deviation across the class of audio segments is greater than a predefined threshold, the foundation model 2515 is further optimized by tuning hyperparameters, by using an enhanced network topology, or in any other suitable manner. Once the embedding deviation for a given class is below a certain threshold, one of the audio segments for that class is selected (e.g., randomly) and processed by the foundation model 2515, and the corresponding embedding is generated and stored in memory as an initial reference embedding.
[0204] Thus, an initial reference embedding is obtained for each class and is stored in a latest reference embedding data store 2530. A refined embedding model (REM) 2520 can then be initialized using transfer learning from the foundation model 2515. Transfer learning in this context involves leveraging knowledge (features, weights, biases, etc.) from a pre-existing “teacher network” (i.e., the foundation model 2515) to accelerate and enhance the learning process of a new “student network” (i.e., the REM 2520) on a related but extended task. Here, the foundation model 2515 has already been successfully trained to generate embeddings by its bottleneck layer that are sufficiently discriminable for linguistic classification. The bottleneck layer acts as a compact representation or embedding of the learned features, encapsulating the distilled knowledge of the network about that problem. The trained bottleneck layer (e.g., and input layers) from the foundation model 2515 can be directly used or fine-tuned in the REM 2520.This allows the REM 2520 to not start from scratch, but instead to build upon the learned representations of the foundation model 2515, adapting and extending them to refine the embedding generation.
[0205] In particular, the first refinement stage seeks to train the transfer-learned REM 2520 until all embeddings generated for a particular class are effectively the same. In some embodiments, this is measured by training until the Euclidean distance across embeddings of each class of audio segments meets a predefined acceptance criterion. As illustrated, the training environment 2500 can use a repository of class audio samples for training, illustrated as a clean spoken word audio (SWA) repository 2505. The clean SWA repository 2505 includes multiple SWA samples for each of many classes collected in a substantially noise-free environment. In some embodiments, the clean SWA repository 2505 is the SWA repository 510 of FIG. 5 and 7, or the SWA repository 1810 of FIG. 18A and 18B.
[0206] For LNE purposes, each SWA sample in the clean SWA repository 2505 can be tagged with a class identifier. In each of a series of training frames, an input audio stream 2510 is generated by obtaining an audio sample for a class from the clean SWA repository 2505. The audio sample is loaded to the REM 2520. As illustrated, in some embodiments, the audio sample is passed through a trained USC block 2535 prior to being loaded to the REM 2520, and the USC block 2535 implements universal sound conversion as described herein. Referring to FIG. 23, the USC block 2535 can include a USC encoder 2320 and a USC decoder 2325 that convert the input audio stream 2510 into unified features, such as by stripping the audio sample of suprasegmental features and / or by normalizing the audio samples as if all spoken in a same way by a same speaker.
[0207] The REM 2520 generates a frame embedding corresponding to the audio sample received at its input based at least on transfer learning from the foundation model 2515. In the same training frame, the class identifier for the input audio stream 2510 is provided to the latest reference embedding data store 2530, so that the latest reference embedding data store 2530 outputs a latest reference embedding for the identified class. The frame embedding output from the REM 2520 is compared to a latest reference embedding from the latest reference embedding data store 2530, thereby yielding an error that is back-propagated for use in refining the embedding generation by the REM 2520. For example, a subtractor 2525 represents a loss function (e.g., mean squared error (MSE), cosine similarity, etc.) that results in a calculated error, and the calculated error is used to adjust the weights of the REM 2520 through back-propagation.
[0208] The output of the REM 2520 is also fed back to update the latest reference embedding data store 2530, such that the frame embedding generated for that class becomes the new latest reference embedding for that class. In some embodiments, a delay block 2545 is in the path between the REM 2520 output and the latest reference embedding data store 2530 to ensure that the updating of the latest reference embedding data store 2530 is delayed until the end of the refinement frame. In effect, each training frame seeks to refine the REM 2520 so that the embeddings for all audio samples in each class are as close as possible (i.e., at least in satisfaction of the predetermined acceptance criterion. By the end of the first refinement stage, the REM 2520 is trained to effectively generate a unified embedding for all audio samples in each class.
[0209] FIG. 26 shows a block diagram of a second refinement stage training environment 2600 for an embedding model, according to embodiments described herein. The training environment 2600 can be used after the training environment 2500 of FIG. 25. As noted with respect to FIG. 25, although embodiments can be tailored to LNE, USC, UNLE, acoustic segmentation, and / or other name embedding approaches, refinement of the embedding model is described with reference to LNE or ULNE approaches. Embodiments assume that, at the end of the first refinement stage, the REM 2520 layers are frozen and the latest reference embedding data store 2530 is frozen as REM embedding data 2630, so that the REM embedding data 2630 is the set of unified embeddings for the classes as generated by the fully trained (and now frozen) REM 2520.
[0210] In the second refinement training stage, a noise-robust refined embedding model (NR- REM) 2620 is produced that can provide accurate embedding (e.g., linguistic embedding) even for highly noisy data. In the first refinement training stage, the training dataset was based on a clean SWA repository 2505 (i.e., on relatively noise-free audio samples). In the second refinement training stage, the training dataset is a noisy SWA repository 2605. The noisy SWA repository 2605 can include audio samples for some or all of the same classes represented in the clean SWA repository 2505. The noisy SWA repository 2605 can be generated by collecting real-time SWA samples (e.g., name data) in a very noisy environment and / or augmenting clean SWA samples with different noises that do not have linguistic influences (such as kitchen noise, music, babble, cafe, traffic, driving, wind, etc.).
[0211] Each noisy SWA sample in the noisy SWA repository 2605 can be tagged with a class identifier. In each of a series of training samples, an input audio stream 2610 is generated by obtaining an audio sample for a class from the noisy SWA repository 2605. The audio sample is loaded to the NR-REM 2620. As illustrated, in some embodiments, the audio sample is passed through a trained USC block 2535 prior to being loaded to the NR-REM 2620, and the USC block2535 implements universal sound conversion as described herein (i.e., to convert the audio into unified features. The NR-REM 2620 generates a sample embedding corresponding to the audio sample received at its input. In the same training frame, the class identifier for the input audio stream 2610 is provided to the REM embedding data 2630, so that the REM embedding data 2630 outputs the unified embedding for the identified class. The sample embedding output from the NR-REM 2620 is compared to a universal embedding from the REM embedding data 2630, thereby yielding an error that is back-propagated for use in refining the embedding generation by the NR-REM 2620. For example, a subtractor 2525 represents a loss function (e.g., mean squared error (MSE), cosine similarity, etc.) that results in a calculated error, and the calculated error is used to adjust the weights of the NR-REM 2620 through back-propagation.
[0212] In some embodiments, the NR-REM 2620 is initialized using transfer learning from the REM 2520. The trained bottleneck layer (e.g., and input layers) from the REM 2520 can be directly used or fine-tuned in the NR-REM 2620. This allows the NR-REM 2620 to not start from scratch, but instead to build upon the learned representations of the REM 2520, adapting and extending them to further refine the embedding generation. In other embodiments, a noise-robust embedding generator is implemented using a new network topology based on the REM embedding data 2630 and the noisy SWA repository 2605.
[0213] In effect, each training frame of the second refinement training stage seeks to further refine the NR-REM 2620 to ignore the influence of noise in embedding generation. This can be measured by continuing to train the NR-REM 2620 until the loss function meets a predetermined acceptance criterion, such as a MSE that is below a predetermined threshold value. At the end of the second refinement training stage, the NR-REM 2620 is trained so that the same universal embedding is generated for any particular class, regardless of the presence of noise.
[0214] When a real-time audio stream contains competing talk, or another voice in addition to the attention seeker’s speech, the linguistic embedding generated by the NR-REM 2620 can tend to become distorted. In such cases, the embedding generated by the NR-REM 2620 may not sufficiently correlate with a corresponding enrolled name, if any. This is because the audio segments contain a combination of multiple linguistic information, including at least the name uttered by the attention seeker, and other linguistic information present in the background voices of the other speakers. Some embodiments further refine generation of name embeddings to operate in context of competing speech.
[0215] FIG. 27 shows a block diagram of a third refinement stage training environment 2700 for an embedding model, according to embodiments described herein. The training environment 2700can be used after the training environment 2600 of FIG. 26. As noted with respect to FIGS. 25 and 26, although embodiments can be tailored to LNE, USC, UNLE, acoustic segmentation, and / or other name embedding approaches, refinement of the embedding model is described with reference to LNE or ULNE approaches. Embodiments assume that, at the end of the second refinement stage, the NR-REM 2620 layers are frozen.
[0216] In the third refinement training stage, a distortion-robust REM (DR-REM) model 2720 is produced that can provide accurate embedding (e.g., linguistic embedding) even for audio signals collected in environments with noise and competing speech. Embodiments of the third refinement training stage can use the clean SWA repository 2505 (i.e., on relatively noise-free audio samples) as its training set. As illustrated, the training environment 2700 includes at least the NR-REM 2620, the DR-REM 2720, a segment selector block 2705, an augmenter block 2740, and a training controller 2730. To increase the model’s robustness in the context of competing speech, the training controller 2730 directs the training environment 2700 to operate over a series of training frames, each configured either as an acceptance (+ve) training frame or a rejection (— ve) training frame. The +ve training frames are those for which input data contains an SWA sample from a a training class spoken with different prosody or accent by different speakers, or by the same speaker in different instances. The — ve training frames are those for which input data does not contain any speech correlated with the training class. The training frames can follow a distribution as [X% + ve : (100 — X)% — ve]. Some embodiments build more negative cases to reject false class detection completely, even if that tends to result in slightly lower sensitivity for real class presence in very competitive scenarios.
[0217] Embodiments can begin by selecting a set of N class IDs (N is a positive integer) as training classes. In some embodiments, the training classes are N randomly selected classes from the clean SWA repository 2505. In other embodiments, the training classes are N classes determined to be of interest for name detection training, such as classes corresponding to a representative set of names, attention-seeking utterances, etc. In each training frame, an SWA sample (e.g., a one-second audio segment) is chosen by the segment selector block 2705 for one of the N training classes. The selected SWA sample for the frame is illustrated as training class sample 2710-1. The training class sample 2710-1 is loaded to the NR-REM 2620 (optionally via the USC block 2535-1), so that the NR-REM 2620 outputs the noise-robust universal embedding for the corresponding class as a reference embedding for the training frame. As further described below, this reference embedding can be used both as a masking signal 2622 and as an acceptance ground truth signal 2624-A in acceptance training frames.
[0218] Also in each training frame, the training controller 2730 determines whether the training frame is an acceptance training frame or a rejection training frame. If the current training frame is an acceptance training frame, a training class sample 2710-2 is output by the segment selector block 2705. The training class sample 2710-2 can be the same segment as the training class sample 2710-1, or it can be a different SWA sample for the same training class. If the current training frame is a rejection training frame, a non-training class sample 2715 is output by the segment selector block 2705. The non-training class sample 2715 can be a SWA sample from any class of the clean SWA repository 2505 other than a training class.
[0219] The “outputting” of the training class sample 2710-2 or the non-training class sample 2715 by the segment selector block 2705 can be implemented in several ways. In some implementations, the segment selector block 2705 receives a control signal from the training controller 2730 indicating whether the current training frame is an acceptance training frame or a rejection raining frame, and the segment selector block 2705 obtains either one of the training class sample 2710-2 or the non-training class sample 2715 from the clean SWA repository 2505, accordingly. In other implementations, the segment selector block 2705 obtains both the training class sample 2710-2 and the non-training class sample 2715 in each training frame and one of the two is selected for output based on the control signal from the training controller 2730. As illustrated, embodiments can include a switching network 2735-2 that is controlled by the control signal from the training controller 2730, and the switching network 2735-2 can effectively select between the training class sample 2710-2 and the non-training class sample 2715 for output. For example, the switching network 2735-2 is in the position indicated by solid lines for acceptance training frames, and the switching network 2735-2 is in the position indicated by dashed lines for rejection training frames. The switching network 2735-2 can be part of the segment selector block 2705 or separate from the segment selector block 2705.
[0220] The selected audio sample output from the segment selector block 2705 (i.e., the training class sample 2710-2 or the non-training class sample 2715) is passed to an augmenter block 2740. The augmenter block 2740 applies one or more augmentation models 2745 to generate an augmented audio sample for the training frame. Each of the augmentation models 2745 can effectively distort the audio sample with one or more added noises, background speech samples (e.g., from a conversational speech corpus), competing linguistic information, etc. The augmented audio sample is loaded to the DR-REM 2720 (optionally via the USC block 2535-2).
[0221] As illustrated another switching network 2735-1 can be controlled by the training controller 2730 (e.g., by the same control signal from the training controller 2730) to select whichof two signals to use as a ground truth signal 2624 for the frame. During acceptance training frames, the switching network 2735-1 is in the position indicated by solid lines, so that the ground truth signal 2624 is an acceptance ground truth signal 2624-A corresponding to the reference embedding output by the NR-REM 2620. During rejection training frames, the switching network 2735-1 is in the position indicated by dashed lines, so that the ground truth signal 2624 is a rejection ground truth signal 2624-R corresponding to a non-linguistic reference (NLR) signal 2717. In one implementation, the NLR signal 2717 is a zero vector.
[0222] In each training frame, the training of the DR-REM 2720 is based on using the DR-REM 2720 to generate a frame embedding from the augmented audio sample that is masked according to the masking signal 2622. The masking signal 2622 essentially passes, from the NR-REM 2620 to the DR-REM 2720, only the linguistic information relevant to embedding the selected class for the training frame. In one implementation, the masking signal 2622 operates to pass embedding tokens from the NR-REM 2620 that relate to the selected class and to mask all other tokens. As a result, the output of the DR-REM 2720 is effectively forced to be enhanced whenever the augmented audio sample includes linguistic information relevant to the selected class (even though it also includes augmentations) and is effectively forced to be filtered whenever the augmented audio sample does not include linguistic information relevant to the selected class.
[0223] As illustrated, a loss function (represented by subtractor 2525) is computed between the frame embedding output by the DR-REM 2720 and ground truth signal 2624. The calculated error resulting from the loss function is used to adjust the weights of the DR-REM 2720 through back- propagation. For any acceptance training frame, the loss function is essentially a comparison between a reference embedding and a frame embedding that both correspond to the same classrelevant linguistic information, such that the back-propagation tends to enhance the embedding output for those cases. For any rejection training frame, the loss function is essentially a comparison between the NLR signal 2717 and a frame embedding, neither having class-relevant linguistic information, such that the back-propagation tends to filter out the embedding output for those cases. In this way, the DR-REM 2720 learns to pay attention to class-relevant linguistic features.
[0224] In some embodiments, the DR-REM 2720 is initialized using transfer learning from the NR-REM 2620. The trained bottleneck layer (e.g., and input layers) from the NR-REM 2620 can be directly used or fine-tuned in the DR-REM 2720, and layers can be added on the output side of the DR-REM 2720 for purposes of this third training stage. This allows the DR-REM 2720 to notstart from scratch, but instead to build upon the learned representations of the NR-REM 2620, adapting and extending them to further refine the embedding generation.
[0225] FIG. 28 shows a block diagram of a final training environment 2800 for an embedding model, according to embodiments described herein. The training environment 2800 can be used after the training environment 2700 of FIG. 27, or it can be modified for use after the training environment 2600 of FIG. 26. Embodiments assume that, at the end of the third refinement stage, the DR-REM 2720 layers are frozen. A final robust name embedding (RNE) model 2820 can be transfer-learned from the DR-REM 2720 in this training environment 2800 by augmenting the frozen DR-REM 2720 with output layers in accordance with the desired name detection approach. For example, as illustrated, the final RNE model 2820 can be trained as embedding model 415 of FIG. 4, embedding model 1915 of FIG. 19, or the like; and the RNE model 2820 can be used in the manners described in those contexts.
[0226] The training environment 2800 is similar to the training environment 2700 of FIG. 27, except that it is tailored for training output layers for desired name detection. As in FIG. 27, a clean SWA repository 2505 can be used as the training set. The training controller 2730 directs the training environment 2700 to operate over a series of training frames, each configured either as an acceptance (+ve) training frame or a rejection (— ve) training frame. Embodiments can begin by selecting a set of N class IDs (N is a positive integer) as training classes. In some implementations, the training in FIG. 28 uses the same number (N) of training classes as in the training of FIG. 27. In some implementations, the training in FIG. 28 uses the same training classes as in the training of FIG. 27. In some implementations, the training in FIG. 28 uses a different number (N) of training classes and / or different training classes from the training of FIG. 27.
[0227] The selected SWA sample for each frame is illustrated as training class sample 2710-1. The training class sample 2710-1 is loaded to the NR-REM 2620 (optionally via the USC block 2535-1), so that the NR-REM 2620 outputs the noise-robust universal embedding for the corresponding class as a masking signal 2622 for the frame. Also in each training frame, the training controller 2730 determines whether the training frame is an acceptance training frame or a rejection training frame. If the current training frame is an acceptance training frame, a training class sample 2710-2 (i.e., the same or a different SWA sample for the same selected training class) is output by the segment selector block 2705. If the current training frame is a rejection training frame, a non-training class sample 2715 (i.e., any SWA sample from any class other than the selected training class) is output by the segment selector block 2705. As in FIG. 27, the outputtingcan be via a switching network 2735-2, or in any other suitable manner. The selected audio sample output from the segment selector block 2705 (i.e., the training class sample 2710-2 or the non-training class sample 2715) is passed to the augmenter block 2740, where one or more augmentation models 2745 are applied to generate an augmented audio sample for the training frame. The augmented audio sample is loaded to the RNE model 2820 (optionally via the USC block 2535-2).
[0228] In each training frame, the training of the RNE model 2820 is based on using the RNE model 2820 to generate a frame embedding from the augmented audio sample that is masked according to the masking signal 2622 and classified according to the added classification layers (added output layers 2810). A loss function (represented by subtractor 2525) is computed between the frame embedding output by the RNE model 2820 and a training target ground truth 2830.
[0229] As illustrated, the RNE model 2820 has added output layers 2810. As described herein, for LNE and UNLE approaches, the output layers 2810 implement a classifier, so that the output of the RNE model 2820 is a classification of the input audio into one of several classes based on its linguistic content. The training target ground truth 2830 is the class (e.g., the class identifier) corresponding to the selected training class for the training frame (i.e., the target classification). The calculated error resulting from the loss function is used to adjust the weights of the RNE model 2820 through back-propagation, thereby training the RNE model 2820 be able to classify SWA samples according to linguistic content.
[0230] As illustrated, the RNE model 2820 has added output layers 2810. As described herein, for acoustic segmentation approaches, the output layers 2810 seek to generate a posterior probability matrix (PPM) and / or an ordered acoustical segment vector (OASV) at the output of the RNE model 2820. The training target ground truth 2830 can be a target posterior probability matrix (PPM) and / or an ordered acoustical segment vector (OASV) associated with the selected training class for the training frame. For example, the training target ground truth 2830 can be generated using techniques, such as those described with reference to FIGS. 5 - 9B. The calculated error resulting from the loss function is used to adjust the weights of the RNE model 2820 through back-propagation, thereby training the RNE model 2820 be able to output an accurate acoustic segmentation of input SWA samples.
[0231] Referring back to FIG. 4 (or FIG. 19), the refined RNE model 2820 can be used as part of an enrollment stage 420 and an identification stage 430 to support name detection, automated attention handling, and / or other features described herein.
[0232] FIG. 29 shows a flow diagram of an illustrative method 2900 for refinement of a name embedding model, according to embodiments described herein. At stage 2904, embodiments can begin by generating a foundation model based on a corpus of clean (e.g., substantially noise-free) SWA data for a large number of classes. As described herein, the corpus can include a large number (e.g., tens or hundreds) of suprasegmentally diverse clean SWA samples for each of a large number (e.g., thousands or tens of thousands) of classes. The foundation model is trained to automatically generate output embeddings in accordance with a name detection paradigm. For example, the foundation model can generate output embeddings in accordance with LNE-based name detection, USC-based name detection, ULNE-based name detection, acoustic segmentationbased name detection, or any other feasible name detection paradigm.
[0233] At stage 2908, embodiments can train a refined embedding model (REM) based on the foundation model (e.g., by transfer learning) and the clean SWA data. As described above (e.g., with reference to FIG. 25), the training is to reduce within-class embedding variance, so that any SWA sample for a particular class will yield substantially the same unified embedding. The term “unified” is used in this context to mean that the variance across all embeddings generated for all SWA samples for a same class will be below a predetermined threshold. For example, the Euclidean distance between all embeddings generated for a same class will be smaller than a predetermined threshold.
[0234] In some embodiments, training the REM includes, for each of the classes, storing, in a data store, a respective latest embedding for the class as the output embedding generated for any of the clean SWA samples for the class by the foundation model. In such embodiments, the training can further include iterating several steps for each of multiple (e.g., hundreds or thousands of) training frames. In each iteration, a selected sample of the clean SWA samples and a class identifier for the selected sample is obtained from the corpus of clean SWA samples. The selected sample is loaded into the REM to generate a frame embedding. The respective latest embedding for the class is obtained from the data store based on the class identifier (i.e., as a reference embedding). An error is then computed based on comparing the frame embedding with the respective latest embedding for the class, and the error is back-propagated to the REM. The respective latest embedding for the class is updated in the data store to be the frame embedding. For example, the stored latest embedding for the class is replaced with the frame embedding. In some such embodiments, loading the selected sample into the REM to generate the frame embedding includes: loading the selected sample into a universal sound conversion (USC) block having a machine learning model trained to remove suprasegmental features from the selectedsample to output a unified feature representation of the selected sample; and loading the unified feature representation to into the REM.
[0235] At stage 2912, embodiments can train a noise-robust REM (NR-REM) based on the respective unified embeddings generated by the REM and based on noisy SWA data. Some embodiments can use transfer learning. The noisy SWA data can be a corpus of noisy SWA samples including multiple suprasegmentally diverse noisy SWA samples for each of the classes. In some embodiments, the NR-REM is initialized by transfer-learning from the REM. As described above (e.g., with reference to FIG. 26), the training is to increase the model’s robustness to noisy environments, so that any SWA sample for a particular class will still yield substantially the same unified embedding even in the presence of noise.
[0236] In some embodiments, training the NR-REM includes, for each of the classes, storing, in a data store, a respective REM embedding for the class as generated for any of the clean SWA samples for the class by the REM. Such embodiments can further include iterating through steps for each of multiple (e.g., hundreds or thousands of) training frames. In each iteration, a selected sample of the noisy SWA samples and a class identifier for the selected sample can be obtained from the corpus of noisy SWA samples. The selected sample can be loaded into the NR-REM to generate a frame embedding, and the respective REM embedding for the class can be obtained from the data store based on the class identifier (i.e., as a reference embedding). An error can be computed based on comparing the frame embedding with the respective REM embedding for the class, and the error can be back-propagated to the NR-REM. In some such embodiments, loading the selected sample into the NR-REM to generate the frame embedding includes: loading the selected sample into a universal sound conversion (USC) block having a machine learning model trained to remove suprasegmental features from the selected sample to output a unified feature representation of the selected sample; and loading the unified feature representation to into the NR-REM.
[0237] Some embodiments, at stage 2916, train a distortion-robust REM (DR-REM) based on the NR-REM, clean SWA data, and augmentation models. Some embodiments can use transfer learning. As described above (e.g., with reference to FIG. 26), the training is to increase the model’s robustness to environments with competing speech and other competing noise, so that any SWA sample for a particular class will still yield substantially the same unified embedding even in the presence of competing speech and other competing noise.
[0238] In some embodiments, the training iterates for each of a number (e.g., hundreds or thousands) of training frames. In each iteration, a training class is selected from the classes. Afirst of the clean SWA samples for the training class is selected. The NR-REM generates a reference embedding automatically in response to loading the first of the clean SWA samples to the NR-REM (the reference embedding corresponds to the respective unified embedding for the training class). A training controller determines (i.e., directs) whether the training frame is an acceptance training frame or a rejection training frame. An augmented audio signal by selecting and distorting either: a second of the clean SWA samples for the training class, if the training frame is one of the plurality of acceptance training frames; a clean SWA sample for any of the plurality of classes other than the training class, if the training frame is one of the plurality of rejection training frames. The distorting involves augmenting the audio sample with competing linguistic content, competing acoustical content, and / or other competing noise. The DR-REM generates a frame embedding automatically in response to loading the augmented audio signal to the DR-REM and based on masking according to the reference embedding (as described above). A loss function is computed based on comparing the frame embedding with a ground truth signal. The ground truth signal is either: the reference embedding, if the training frame is one of the plurality of acceptance training frames; or a non-linguistic representation vector (e.g., a zero vector), if the training frame is one of the plurality of rejection training frames. The DR-REM is then refined by back-propagating an error based on the loss function.
[0239] At stage 2920, embodiments train a robust name embedding (RNE) model based on the NR-REM or the DR-REM to produce target name detection outputs. In some embodiments, the training involves transfer learning an input portion and a bottleneck portion of the RNE model from the NR-REM or the DR-REM, adding an output portion to the RNE model, and training the output portion to automatically generate the output embeddings in accordance with the name detection paradigm. In LNE-based and UNLE-based contexts, the target name detection outputs can be output classifications. In acoustic segmentation-based contexts, the target name detection outputs can be posterior probability matrices (PPMs), ordered acoustical segment vectors (OASVs), and / or the like.
[0240] Embodiments of the method 2900, or any of the training environments described in FIGS. 25 - 28, can be implemented using a computational system, such as the system 2400 of FIG. 24. In some embodiments, storage devices 2425 of such a computational system store one or more of: the corpus of clean SWA samples including suprasegmentally diverse clean SWA samples for each of many classes; the corpus of noisy SWA samples including suprasegmentally diverse noisy SWA samples for each of the classes; machine learning networks, including the foundation model, the REM, the NR-REM, the DR-REM, and the RNE model. Further, as described with reference to FIG. 24, embodiments can store instructions which, when executed, cause one or moreprocessors 2410 to perform progressive training steps or features described herein. In some embodiments, the method 2900 is implemented by a computer program.
[0241] Having described several example configurations, various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the disclosure. For example, the above elements may be components of a larger system, wherein other rules may take precedence over or otherwise modify the application of the invention. Also, a number of steps may be undertaken before, during, or after the above elements are considered.
Claims
WHAT IS CLAIMED IS:
1. A method for progressive training of a machine learning model for refined name embedding, the method comprising: training a foundation model, based on a corpus of clean spoken word audio (SWA) samples including a plurality of suprasegmentally diverse clean SWA samples for each of a plurality of classes, to automatically generate output embeddings in accordance with a name detection paradigm; training a refined embedding model (REM), based on the corpus of clean SWA samples and on transfer-learning from the foundation model, to generate a respective unified embedding for each class of the plurality of classes responsive to receiving any of the clean SWA samples for the class; and training a noise-robust REM (NR-REM), based on a corpus of noisy SWA samples including a plurality of suprasegmentally diverse noisy SWA samples for each of the plurality of classes, and based on the respective unified embeddings generated the REM, to generate the respective unified embedding for each class of the plurality of classes responsive to receiving any of the noisy SWA samples for the class.
2. The method of claim 1, further comprising: training a robust name embedding (RNE) model by transfer learning an input portion and a bottleneck portion of the RNE model from the NR-REM, adding an output portion to the RNE model, and training the output portion to automatically generate the output embeddings in accordance with the name detection paradigm.
3. The method of claim 1, further comprising: training a distortion-robust REM (DR-REM), based on the corpus of clean SWA samples and the NR-REM by, for each of a plurality of training frames: selecting a training class from the plurality of classes; selecting a first of the clean SWA samples for the training class; generating, by the NR-REM, a reference embedding automatically in response to loading the first of the clean SWA samples to the NR-REM, the reference embedding corresponding to the respective unified embedding for the training class; determining by a training controller whether the training frame is one of a plurality of acceptance training frames or one of a plurality of rejection training frames; generating an augmented audio signal by selecting and distorting a second of the clean SWA samples for the training class if the training frame is one of the pluralityof acceptance training frames, or by selecting and distorting a clean SWA sample for any of the plurality of classes other than the training class if the training frame is one of the plurality of rejection training frames; generating, by the DR-REM, a frame embedding automatically in response to loading the augmented audio signal to the DR-REM and based on masking according to the reference embedding; computing a loss function based on comparing the frame embedding with a ground truth signal, the ground truth signal being the reference embedding if the training frame is one of the plurality of acceptance training frames or being a non-linguistic representation vector if the training frame is one of the plurality of rejection training frames; and refining the DR-REM by back-propagating an error based on the loss function.
4. The method of claim 3, further comprising: training a robust name embedding (RNE) model by transfer learning an input portion and a bottleneck portion of the RNE model from the DR-REM, adding an output portion to the RNE model, and training the output portion to automatically generate the output embeddings in accordance with the name detection paradigm.
5. The method of claim 1, wherein the training the REM comprises: for each of the plurality of classes, storing, in a data store, a respective latest embedding for the class as the output embedding generated for any of the clean SWA samples for the class by the foundation model; and iteratively, for each of a plurality of training frames: obtaining, from the corpus of clean SWA samples, a selected sample of the clean SWA samples and a class identifier for the selected sample; loading the selected sample into the REM to generate a frame embedding; obtaining the respective latest embedding for the class from the data store based on the class identifier; back-propagating an error to the REM computed based on comparing the frame embedding with the respective latest embedding for the class; and updating the respective latest embedding for the class in the data store to be the frame embedding.
6. The method of claim 5, wherein: loading the selected sample into the REM to generate the frame embedding comprises: loading the selected sample into a universal sound conversion block comprising a machine learning model trained to remove suprasegmental features from the selected sample to output a unified feature representation of the selected sample; and loading the unified feature representation to into the REM.
7. The method of claim 1, wherein the REM is trained to generate the respective unified embedding for each class, such that a Euclidean distance between all embeddings generated by the REM for all clean SWA samples corresponding to a same class is smaller than a predetermined threshold.
8. The method of claim 1, wherein the training the NR-REM comprises: for each of the plurality of classes, storing, in a data store, a respective REM embedding for the class as generated for any of the clean SWA samples for the class by the REM; and iteratively, for each of a plurality of training frames: obtaining, from the corpus of noisy SWA samples, a selected sample of the noisy SWA samples and a class identifier for the selected sample; loading the selected sample into the NR-REM to generate a frame embedding; obtaining the respective REM embedding for the class from the data store based on the class identifier; back-propagating an error to the NR-REM computed based on comparing the frame embedding with the respective REM embedding for the class.
9. The method of claim 8, wherein the training the NR-REM further comprises: initializing the NR-REM by transfer-learning from the REM.
10. The method of claim 8, wherein: loading the selected sample into the NR-REM to generate the frame embedding comprises:loading the selected sample into a universal sound conversion block comprising a machine learning model trained to remove suprasegmental features from the selected sample to output a unified feature representation of the selected sample; and loading the unified feature representation to into the NR-REM.
11. The method of claim 1, wherein: the name detection paradigm is a linguistic name embedding (LNE) name detection paradigm; and the training the foundation model is to automatically output a classification vector for any sample of the corpus of clean SWA samples based on linguistic content of the sample.
12. The method of claim 1, wherein: the name detection paradigm is an acoustic segmentation name detection paradigm; and the training the foundation model is to automatically output an acoustic segmentation vector for any sample of the corpus of clean SWA samples based on the acoustic content of the sample.
13. A computer program comprising instructions for implementing the method of claim 1.
14. A system for progressive training of a machine learning model for refined name embedding, the system comprising: one or more processors; and non-transitory memory having stored thereon: a corpus of clean spoken word audio (SWA) samples including a plurality of suprasegmentally diverse clean SWA samples for each of a plurality of classes; a corpus of noisy SWA samples including a plurality of suprasegmentally diverse noisy SWA samples for each of the plurality of classes; machine learning networks including a foundation model, a refined embedding model (REM), and a noise-robust REM (NR-REM); and instructions which, when executed, cause the one or more processors to perform steps comprising: training the foundation model, based on the corpus of clean SWA, to automatically generate output embeddings in accordance with a name detection paradigm;training the REM, based on the corpus of clean SWA samples and on transfer-learning from the foundation model, to generate a respective unified embedding for each class of the plurality of classes responsive to receiving any of the clean SWA samples for the class; and training the NR-REM, based on the corpus of noisy SWA samples and on the respective unified embeddings generated the REM, to generate the respective unified embedding for each class of the plurality of classes responsive to receiving any of the noisy SWA samples for the class.
15. The system of claim 14, wherein: the machine learning networks further include a robust name embedding (RNE) model; and the steps further comprise training the RNE model by transfer learning an input portion and a bottleneck portion of the RNE model from the NR-REM, adding an output portion to the RNE model, and training the output portion to automatically generate the output embeddings in accordance with the name detection paradigm.
16. The system of claim 14, wherein: the machine learning networks further include a distortion-robust REM (DR-REM) stored thereon; and the steps further comprise training the DR-REM, based on the corpus of clean SWA samples and the NR-REM by, for each of a plurality of training frames: selecting a training class from the plurality of classes; selecting a first of the clean SWA samples for the training class; generating, by the NR-REM, a reference embedding automatically in response to loading the first of the clean SWA samples to the NR-REM, the reference embedding corresponding to the respective unified embedding for the training class; determining by a training controller whether the training frame is one of a plurality of acceptance training frames or one of a plurality of rejection training frames; generating an augmented audio signal by selecting and distorting a second of the clean SWA samples for the training class if the training frame is one of the plurality of acceptance training frames, or by selecting and distorting a clean SWA sample for any of the plurality of classes other than the training class if the training frame is one of the plurality of rejection training frames;generating, by the DR-REM, a frame embedding automatically in response to loading the augmented audio signal to the DR-REM and based on masking according to the reference embedding; computing a loss function based on comparing the frame embedding with a ground truth signal, the ground truth signal being the reference embedding if the training frame is one of the plurality of acceptance training frames or being a non-linguistic representation vector if the training frame is one of the plurality of rejection training frames; and refining the DR-REM by back-propagating an error based on the loss function.
17. The system of claim 16, wherein: the machine learning networks further include a robust name embedding (RNE) model; and the steps further comprise training a robust name embedding (RNE) model by: transfer learning an input portion and a bottleneck portion of the RNE model from the DR-REM; adding an output portion to the RNE model; and training the output portion to automatically generate the output embeddings in accordance with the name detection paradigm.
18. The system of claim 14, wherein the training the REM comprises: for each of the plurality of classes, storing, in a data store, a respective latest embedding for the class as the output embedding generated for any of the clean SWA samples for the class by the foundation model; and iteratively, for each of a plurality of training frames: obtaining, from the corpus of clean SWA samples, a selected sample of the clean SWA samples and a class identifier for the selected sample; loading the selected sample into the REM to generate a frame embedding; obtaining the respective latest embedding for the class from the data store based on the class identifier; back-propagating an error to the REM computed based on comparing the frame embedding with the respective latest embedding for the class; and updating the respective latest embedding for the class in the data store to be the frame embedding.
19. The system of claim 14, wherein the training the NR-REM comprises: for each of the plurality of classes, storing, in a data store, a respective REM embedding for the class as generated for any of the clean SWA samples for the class by the REM; and iteratively, for each of a plurality of training frames: obtaining, from the corpus of noisy SWA samples, a selected sample of the noisy SWA samples and a class identifier for the selected sample; loading the selected sample into the NR-REM to generate a frame embedding; obtaining the respective REM embedding for the class from the data store based on the class identifier; back-propagating an error to the NR-REM computed based on comparing the frame embedding with the respective REM embedding for the class.
20. The system of claim 14, wherein: the name detection paradigm is one of a linguistic name embedding (LNE) name detection paradigm, a universal sound conversion (USC) name detection paradigm, a unified LNE (ULNE) name detection paradigm, or an acoustic segmentation name detection paradigm.