A cascade architecture for noise-robust keyword spotting.

The cascaded hotword detection architecture optimizes power consumption and accuracy by using a first-stage detector on a digital signal processor for coarse screening and a second-stage detector on a system-on-chip for accurate hotword detection in noisy environments, addressing the limitations of conventional methods in battery-powered devices.

JP7727783B2Active Publication Date: 2025-08-21GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024043843
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-08-21
Estimated Expiration
2040-04-08

AI Technical Summary

Technical Problem

Existing voice-enabled devices face challenges in accurately detecting hotwords in noisy environments while conserving battery life and computational resources, as conventional parallel hotword detection architectures consume excessive power and are unsuitable for battery-powered devices.

Method used

A cascaded hotword detection architecture is implemented, where a first-stage hotword detector on a digital signal processor coarsely screens multi-channel audio for candidates, triggering a second-stage detector on a system-on-chip to provide accurate detection, leveraging noise cleaning algorithms for improved accuracy and power efficiency.

Benefits of technology

The cascaded architecture optimizes power consumption, latency, and noise robustness, enabling efficient hotword detection in battery-powered devices by conserving computational resources and enhancing detection accuracy in noisy conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007727783000001
    Figure 0007727783000001
  • Figure 0007727783000002
    Figure 0007727783000002
  • Figure 0007727783000003
    Figure 0007727783000003
Patent Text Reader

Abstract

To provide a cascade architecture for noise-robust keyword spotting.SOLUTION: A method comprises: receiving a streaming multichannel audio captured by an array of microphones of a user device; and processing respective audio features by the use of a hot word detector of a first stage so as to determine whether a hot word is detected for each channel. When the hot word detector of the first stage detects a hot word, live audio data ticked using a first noise cleaning algorithm is provided so as to generate a clean monophonic audio chomp. The clean monophonic audio chomp is processed by the use of a hot word detector of a second stage so as to detect a hot word.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a cascaded architecture for noise-robust keyword spotting. [Background technology]

[0002] A voice-enabled environment (e.g., a home, a workplace, a school, an automobile, etc.) allows a user to speak queries or commands aloud to a computer-based system that processes and responds to the query and / or performs a function based on the command. The voice-enabled environment may be implemented using a network of connected microphone devices distributed throughout various rooms or areas of the environment. These devices may use hot words to help identify when a given utterance is directed to the system, as opposed to an utterance directed to another individual present in the environment. Thus, the device may operate in a sleep or hibernation state and wake up only if the detected utterance contains a hot word. These devices may include two or more microphones for recording multi-channel audio. Neural networks have recently emerged as an attractive solution for training models to detect hot words spoken by users in streaming audio. Typically, neural networks used to detect hot words in streaming audio receive a single channel of the streaming audio. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Patent Application PCT / US20 / 13705 Summary of the Invention [Means for solving the problem]

[0004] One aspect of the present disclosure provides a method for noise-robust keyword / hotword spotting in a cascaded hotword detection architecture. The method includes, in a first processor of a user device, receiving streaming multi-channel audio captured by an array of microphones in communication with the first processor, each channel of the streaming multi-channel audio including a respective audio feature captured by a separate dedicated microphone in the array of microphones. The method also includes, by the first processor, processing, using a first-stage hotword detector, each audio feature of at least one channel of the streaming multi-channel audio to determine whether a hotword is detected in the streaming multi-channel audio by the first-stage hotword detector. If the first-stage hotword detector detects a hotword in the streaming multi-channel audio, the method further includes, by the first processor, providing chopped multi-channel raw audio data to a second processor of the user device, each channel of the chopped multi-channel audio data corresponding to a respective channel of the streaming multi-channel audio and including respective raw audio data chopped from each channel of the streaming multi-channel audio. The method includes processing, by a second processor, each channel of the chopped multi-channel raw audio data using a first noise cleaning algorithm to generate a clean monophonic audio chomp, and processing the clean monophonic audio chomp using a second stage hot word detector to determine whether a hot word is detected by the second stage hot word detector in the clean monophonic audio chomp.In a clean monophonic audio chomp, if a hot word is detected by the second stage hot word detector, a second processor initiates a wake-up process for the user device to process the hot word and / or one or more other terms following the hot word in the streaming multi-channel audio.

[0005] Implementations of the present disclosure may include one or more of the following optional features: In some implementations, the respective raw audio data for each channel of the chopped multi-channel raw audio data includes an audio segment characterizing a hot word detected by the first-stage hot word detector in the streaming multi-channel audio. In these implementations, the respective raw audio data for each channel of the chopped multi-channel raw audio data further includes a prefix segment including a duration of audio immediately prior to the time at which the first-stage hot word detector detected the hot word in the streaming multi-channel audio.

[0006] In some examples, when streaming multi-channel audio is received at the first processor and audio features of each of at least one channel of the streaming multi-channel audio are processed by the first processor, the second processor operates in a sleep mode. In these examples, providing the chopped multi-channel audio raw data to the second processor wakes the second processor to transition from the sleep mode to a hot word detection mode. While in the hot word detection mode, the second processor may execute a first noise cleaning algorithm and a second stage hot word detector.

[0007] In some implementations, the method further includes, by a second processor, processing each raw audio data of one channel of the chopped multi-channel raw audio data using a second-stage hot word detector while processing the clean monophonic audio chomp in parallel to determine whether a hot word is detected by the second-stage hot word detector in the respective raw audio data. If a hot word is detected by the second-stage hot word detector in either the clean monophonic audio chomp or the respective raw audio data, the method includes, by the second processor, initiating a wake-up process for the user device to process the hot word and / or one or more other terms following the hot word in the streaming multi-channel audio. In these implementations, the method may further include, by the second processor, preventing initiation of a wake-up process for the user device if a hot word is not detected by the second-stage hot word detector in either the clean monophonic audio chomp or the respective raw audio data.

[0008] In some examples, processing each audio feature of at least one channel of the streaming multi-channel audio to determine whether a hot word is detected by the first-stage hot word detector in the streaming multi-channel audio includes processing each audio feature of the at least one channel of the streaming multi-channel audio without canceling noise from the respective audio features. In some implementations, the method includes processing each audio feature of each channel of the streaming multi-channel audio by a first processor to generate a multi-channel cross-correlation matrix. If the first-stage hot word detector detects a hot word in the streaming multi-channel audio, the implementation further includes, for each channel of the streaming multi-channel audio, using the multi-channel cross-correlation matrix to chop respective raw audio data from each audio feature of each channel of the streaming multi-channel audio, and providing the multi-channel cross-correlation matrix to a second processor by the first processor. In these implementations, processing each channel of the chopped multi-channel raw audio data to generate a clean monophonic audio chop includes calculating cleaner filter coefficients for a first noise cleaning algorithm using a multi-channel cross-correlation matrix provided by the first processor, and processing each channel of the chopped multi-channel raw audio data provided by the first processor with the first noise cleaning algorithm having the calculated cleaner filter coefficients to generate a clean monophonic audio chop.In these implementations, processing each audio feature of at least one channel of the streaming multi-channel audio to determine whether a hot word is detected by the first-stage hot word detector in the streaming multi-channel audio may include using the multi-channel cross-correlation matrix to calculate cleaner coefficients for a second noise cleaning algorithm executed in the first processor, while processing each channel of the streaming multi-channel audio through the second noise cleaning algorithm with the calculated filter coefficients to generate a monophonic clean audio stream. In these implementations, the method further includes processing the monophonic clean audio stream using the first-stage hot word detector to determine whether a hot word is detected by the first-stage hot word detector in the streaming multi-channel audio. The first noise cleaning algorithm may apply a first finite impulse response (FIR) including a first filter length to each channel of the chopped multi-channel raw audio data to generate chopped monophonic clean audio data, and the second noise cleaning algorithm may apply a second FIR including a second filter length to each channel of the streaming multi-channel audio to generate a monophonic clean audio stream, where the second filter length is shorter than the first filter length.

[0009] In some examples, the first processor includes a digital signal processor and the second processor includes a system-on-chip (SoC) processor. In additional examples, the user device includes a rechargeable limited power supply, the limited power supply providing power to the first processor and the second processor.

[0010] Another aspect of the present disclosure provides a system for noise-robust keyword spotting in a cascade architecture. The system includes data processing hardware including a first processor and a second processor, and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations in the first processor of a user device, including receiving streaming multi-channel audio captured by an array of microphones in communication with the first processor, where each channel of the streaming multi-channel audio includes a respective audio feature captured by a separate dedicated microphone in the array of microphones. The method also includes processing, by the first processor, audio features of each of at least one channel of the streaming multi-channel audio using a first-stage hotword detector to determine whether a hotword is detected in the streaming multi-channel audio by the first-stage hotword detector. The operations further include, if the first stage hot word detector detects a hot word in the streaming multi-channel audio, providing, by the first processor, the chopped multi-channel raw audio data to a second processor, each channel of the chopped multi-channel raw audio data corresponding to a respective channel of the streaming multi-channel audio and including respective raw audio data chopped from each channel of the streaming multi-channel audio, processing, by the second processor, each channel of the chopped multi-channel raw audio data using a first noise cleaning algorithm to generate a clean monophonic audio chomp, and processing the clean monophonic audio chomp with the second stage hot word detector to determine if a hot word was detected in the clean monophonic audio chomp by the second stage hot word detector.In a clean monophonic audio chomp, if a hot word is detected by the second stage hot word detector, a second processor initiates a wake-up process for the user device to process the hot word and / or one or more other terms following the hot word in the streaming multi-channel audio.

[0011] Implementations of the present disclosure may include one or more of the following optional features: In some implementations, the respective raw audio data for each channel of the chopped multi-channel raw audio data includes an audio segment characterizing a hot word detected by the first-stage hot word detector in the streaming multi-channel audio. In these implementations, the respective raw audio data for each channel of the chopped multi-channel raw audio data further includes a prefix segment including a duration of audio immediately prior to the time at which the first-stage hot word detector detected the hot word in the streaming multi-channel audio.

[0012] In some examples, when streaming multi-channel audio is received at the first processor and audio features of each of at least one channel of the streaming multi-channel audio are processed by the first processor, the second processor operates in a sleep mode. In these examples, providing the chopped multi-channel audio raw data to the second processor wakes the second processor to transition from the sleep mode to a hot word detection mode. While in the hot word detection mode, the second processor may execute a first noise cleaning algorithm and a second stage hot word detector.

[0013] In some implementations, the operations further include processing, by the second processor, each raw audio data of one channel of the chopped multi-channel raw audio data using a second-stage hot word detector while processing the clean monophonic audio chomp in parallel to determine whether a hot word is detected in the respective raw audio data by the second-stage hot word detector. If a hot word is detected in the clean monophonic audio chomp or the respective raw audio data by the second-stage hot word detector, the operations include initiating, by the second processor, a wake-up process for the user device to process the hot word and / or one or more other terms following the hot word in the streaming multi-channel audio. In these implementations, the operations may further include preventing, by the second processor, initiating a wake-up process for the user device if a hot word is not detected in either the clean monophonic audio chomp or the respective raw audio data by the second-stage hot word detector.

[0014] In some examples, the operation of processing each audio feature of at least one channel of the streaming multi-channel audio to determine if a hot word is detected in the streaming multi-channel audio by the first-stage hot word detector includes processing each audio feature of the at least one channel of the streaming multi-channel audio without canceling noise from the respective audio feature. In some implementations, the operation further includes processing, by the first processor, each audio feature of each channel of the streaming multi-channel audio to generate a multi-channel cross-correlation matrix. If the first-stage hot word detector detects a hot word in the streaming multi-channel audio, the operations further include, for each channel of the streaming multi-channel audio, chopping, by the first processor, respective raw audio data from respective audio features for each channel of the streaming multi-channel audio using a multi-channel cross-correlation matrix, and providing, by the first processor, the multi-channel cross-correlation matrix to the second processor. In these implementations, the operations of processing each channel of the chopped multi-channel raw audio data to generate a clean monophonic audio chomp include calculating cleaner filter coefficients for a first noise cleaning algorithm using the multi-channel cross-correlation matrix provided by the first processor, and processing each channel of the chopped multi-channel raw audio data provided by the first processor with the first noise cleaning algorithm having the calculated cleaner filter coefficients to generate a clean monophonic audio chomp. In these implementations, the operation of processing each audio feature of at least one channel of the streaming multi-channel audio to determine whether a hot word has been detected by the first-stage hot word detector in the streaming multi-channel audio includes calculating cleaner coefficients for a second noise cleaning algorithm executed in the first processor using the multi-channel cross-correlation matrix, while processing each channel of the streaming multi-channel audio through the second noise cleaning algorithm with the calculated filter coefficients to generate a monophonic clean audio stream. In these implementations, the operations further include processing the monophonic clean audio stream using the first-stage hot word detector to determine whether a hot word has been detected by the first-stage hot word detector in the streaming multi-channel audio.The first noise cleaning algorithm may apply a first finite impulse response (FIR) including a first filter length to each channel of the chopped multi-channel raw audio data to generate chopped monophonic clean audio data, and the second noise cleaning algorithm may apply a second FIR including a second filter length to each channel of the streaming multi-channel audio to generate a monophonic clean audio stream, where the second filter length is shorter than the first filter length.

[0015] In some examples, the first processor includes a digital signal processor and the second processor includes a system-on-chip (SoC) processor. In additional examples, the user device includes a rechargeable limited power supply, the limited power supply providing power to the first processor and the second processor.

[0016] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a schematic diagram of an exemplary system including a cascaded hot word detection architecture for noise-robust keyword spotting. [Figure 2A] FIG. 1 is a schematic diagram of an exemplary cascaded hotword detection architecture. [Figure 2B] FIG. 1 is a schematic diagram of an exemplary cascaded hotword detection architecture. [Figure 2C] FIG. 1 is a schematic diagram of an exemplary cascaded hotword detection architecture. [Figure 3] FIG. 2D is a schematic diagram of a cleaner task division for the cascade detection architecture of FIGS. 2B and 2C. [Figure 4]1 is a flowchart of an exemplary arrangement of operations for a method for detecting hot words in streaming multi-channel audio using a noise-robust cascaded hot word detection architecture. [Figure 5] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0018] Like reference symbols in the various drawings indicate like elements.

[0019] A voice-enabled device (e.g., a user device running a voice assistant) allows a user to speak queries or commands aloud, process and respond to the queries, and / or perform functions based on the commands. Through the use of “hot words” (also called “keywords,” “attention words,” “wake-up phrases / words,” “trigger phrases,” “invocation phrases,” or “voice action initiation commands”), which are predetermined terms / phrases reserved by consensus to be spoken to attract the attention of a voice-enabled device, a voice-enabled device can distinguish between utterances directed to the system (i.e., initiate a wake-up process to process one or more terms following the hot word in the utterance) and utterances directed to an individual in the environment. Typically, a voice-enabled device operates in a sleep state to conserve battery power and does not process input audio data unless the input audio data follows a spoken hot word. For example, while in a sleep state, the voice-enabled device captures streaming input audio via multiple microphones and uses a hot word detector trained to detect the presence of a hot word in the input audio. If a hot word is detected in the input audio, the voice-enabled device initiates a wake-up process to process the hot word and / or any other terms that follow the hot word in the input audio.

[0020] Hot word detection is akin to searching for a needle in a haystack, as it requires continuously listening to streaming audio and triggering accurately and instantly when the presence of a hot word is detected in the streaming audio. In other words, the hot word detector is tasked with ignoring the streaming audio unless the presence of a hot word is detected. To address the complexity of detecting the presence of a hot word in a continuous stream of audio, neural networks are commonly used by hot word detectors.

[0021] User devices (e.g., computing devices), and more specifically, mobile user devices such as smartphones, tablets, smartwatches, and headphones that are powered by a rechargeable, finite power source (e.g., a battery), are typically embedded systems with limited battery life and limited computational capabilities. That is, when a battery-powered device provides access to a voice-enabled application (e.g., a digital assistant), energy resources can be further limited when the device is tasked with constantly processing audio and / or other data to detect hotword signals for activating the voice-enabled application. In configurations where a battery-powered voice-enabled user device includes a device system-on-a-chip (SoC) (e.g., an application processor (AP)), when a user is interacting with the user device via voice, the device SoC can consume a significant proportion of energy compared to other subsystems (e.g., network processor, digital signal processor (DSP), etc.).

[0022] One of the design goals of voice-enabled user devices is to obtain noise robustness for accurate hot word detection. For user devices including two or more microphones, a statistical speech enhancement algorithm may operate on the multi-microphone noise signal to generate a monophonic audio stream with an improved signal-to-noise ratio (SNR). Therefore, user devices including two or more microphones may use a hot word cleaner algorithm that employs a statistical speech enhancement algorithm to improve the SNR and therefore improve hot word detection accuracy in noisy environments. Typically, the user device uses a hot word cleaner algorithm to obtain a clean monophonic audio stream and a parallel hot word detection architecture in two branches that use the same model but perform hot word detection independently on two different inputs: the raw microphone signal and the clean monophonic audio stream. The binary yes / no decisions made by the two branches, indicating whether a hot word is detected, are combined with a logical OR operation. Using a hotword cleaner algorithm in combination with a parallel hotword detection architecture results in uncompromised hotword detection accuracy in both clean and noisy acoustic environments, but parallel hotword detection architectures are typically not suitable for use in battery-powered devices (e.g., mobile devices) because parallel hotword detection uses a heavy computational load that requires increased power consumption that quickly depletes battery life.

[0023] Hotword detectors used by battery-powered user devices must implement hotword detection algorithms that not only detect hotwords with some accuracy but also achieve the conflicting objectives of low latency, small memory footprint, and light computational load. To achieve these objectives, the user device may use a cascaded hotword detection architecture that includes two hotword detectors: a first-stage hotword detector and a second-stage hotword detector. Here, the first-stage hotword detector resides on a dedicated DSP (e.g., the first processor), includes a small model size, and is computationally efficient for coarsely screening the input audio stream for hotword candidates. Detection of a hotword candidate in the input audio stream by the first-stage hotword detector triggers the DSP to pass / provide a small buffer / chomp of audio data of an appropriate duration to safely contain the hotword to the second-stage hotword detector, which resides / runs on the device SoC. A second-stage hot word detector on the device SoC (e.g., the main AP) then includes a larger model size and provides more computational power than the first-stage hot word detector to provide more accurate detection of hot words, and thus acts as the final arbiter for determining whether the input audio stream actually contains a hot word. This cascade architecture allows the more power-consuming device SoC to operate in sleep mode to preserve battery life until the first-stage hot word detector running / executing on the DSP detects a candidate hot word in the streaming input audio. Only when the candidate hot word is detected does the DSP trigger the device SoC to transition from sleep mode to hot word detection mode to run the second-stage hot word detector.These conventional hot word detection cascade architectures present on user devices with two or more microphones do not leverage streaming multi-channel audio input from two or more microphones to obtain noise robustness (e.g., adaptive noise cancellation) to improve hot word detection accuracy.

[0024] Implementations herein are directed to incorporating a hotword cleaning algorithm into a cascade architecture for hotword detection in a voice-enabled user device. In some examples, the voice-enabled user device is a battery-powered user device (e.g., a mobile device) constrained by limited battery life and limited computing power. As will become apparent, various architectures are disclosed for jointly optimizing power consumption, latency, and noise robustness by splitting and allocating the workload for hotword detection by the cleaner to a DSP (i.e., a first processor) of the user device and an application processor (AP) (i.e., a second processor) of the user device.

[0025] 1 , in some implementations, an exemplary system 100 includes one or more user devices 102 associated with a respective user 10. Each of the one or more devices 102 may correspond to a computing device such as a mobile phone, a computer, a wearable device, a smart appliance, an audio infotainment system, a smart speaker, etc., and comprises memory hardware 105 and data processing hardware collectively including a first processor 110 (e.g., a digital signal processor (DSP)) and a second processor 120 (e.g., an application processor (AP)). The first processor 110 consumes less power during operation than the second processor consumes during operation. As used herein, the first processor 110 may be interchangeably referred to as a DSP, and the second processor 120 may be interchangeably referred to as an “AP” or a “device SoC.” The first and second processors 110, 120 provide a cascaded hot word detection architecture 200 in which a first stage hot word detector 210 runs on the first processor 110 and a second stage hot word detector 220 runs on the second processor 120 to cooperatively detect the presence of hot words in the streaming multi-channel audio 118 in a manner that optimizes power consumption, latency, and noise robustness. The multi-channel streaming audio 118 includes two or more channels 119, 119a-n of audio.

[0026] Generally, the first-stage hotword detector 210 resides on the dedicated DSP 110, includes a smaller model size than the model associated with the second-stage hotword detector 220, and is computationally efficient for coarsely screening the input streaming multi-channel audio 118 for hotword candidates. Thus, the dedicated DSP 110 (e.g., the first processor) may be "always on" so that the first-stage hotword detector 210 is always running to coarsely screen the multi-channel audio 118 for hotword candidates, while all other components of the user device 102, including the main AP 120 (e.g., the second processor), are in a sleep state / mode to conserve battery life. Meanwhile, the second-stage hotword detector 220 resides on the main AP 120, includes a larger model size, and provides more computational power than the first-stage hotword detector 210 to provide more accurate detection of hotwords initially detected by the first-stage hotword detector 210. Thus, the second-stage hotword detector 220 may be more rigorous in determining whether a hotword is present in the audio 118. While the DSP 110 is "always on," the more power-hungry main AP 120 operates in sleep mode to preserve battery life until the first-stage hotword detector 210 in the DSP 110 detects a candidate hotword in the streaming multi-channel audio 118. Thus, only when a candidate hotword is detected does the DSP 110 trigger the main AP to transition from sleep mode to hotword detection mode to run the second-stage hotword detector 220.

[0027] In the illustrated example, when a user 10 speaks an utterance 104 including a hotword (e.g., "Hey Google"), the utterance 104 is captured by a user device 102 as multi-channel streaming audio 118. A cascaded hotword detection architecture 200 resident on the user device 102 is configured to detect the presence of the hotword in the utterance 104 to initiate / trigger a wake-up process in the user device 102 to process the hotword and / or one or more terms (e.g., a query or command) following the hotword in the utterance 104. For example, the wake-up process may include the user device 102 locally executing an automatic speech recognition (ASR) system to recognize (e.g., transcribe) the hotword and / or one or more terms following the hotword, or the wake-up process may include the user device 102 transmitting audio data including the hotword and / or one or more other terms to a remote computing device (e.g., a server or cloud computing environment) that includes an ASR system to perform speech recognition on the audio data.

[0028] One or more user devices 102 may include (or may communicate with) two or more microphones 107, 107a-n to capture speech 104 from a user 10. Each microphone 107 may individually record the speech 104 on a separate, dedicated channel 119 of multi-channel streaming audio 118. For example, a user device 102 may include two microphones 107, each recording the speech 104, and the recordings from the two microphones may be combined into two-channel streaming audio 118 (i.e., stereo audio or stereo). In some examples, a user device 102 includes three or more microphones. That is, three or more microphones are present on the user device 102. Additionally or alternatively, a user device 102 may communicate with two or more microphones separate / remote from the user device 102. For example, a user device may be a mobile device located in a vehicle and in wired or wireless communication (e.g., Bluetooth) with two or more microphones of the vehicle. In some configurations, the user device 102 communicates with at least one microphone 107 that resides on a separate device. In these configurations, the user device 102 may also communicate with one or more microphones 107 that reside on the user device 102.

[0029] When receiving the multi-channel streaming audio 118, the always-on DSP 110 executes / operates the first-stage hotword detector 210 to determine whether a hotword is detected in each audio feature of at least one channel 119 of the streaming multi-channel audio 118. In some examples, the first-stage hotword detector 210 calculates a probability score indicative of the presence of a hotword in each audio feature from a single channel 119 of the streaming multi-channel audio 118. In some examples, a determination that the probability score of the respective audio feature satisfies a hotword threshold (e.g., the probability score is greater than or equal to the hotword threshold) indicates that a hotword is present in the streaming multi-channel audio 118. In particular, the AP 120 may operate in a sleep mode while the multi-channel audio is received at the DSP 110 and the DSP 110 processes each audio feature of the at least one channel 119 of the streaming multi-channel audio 118. In some examples, the "processing" of each audio feature by the DSP 110 includes executing a cleaner 250 to process each audio feature of each channel 119 of the streaming multi-channel audio 118 to generate a monophonic clean audio stream 225, and then executing / operating a first-stage hot word detector 210 to determine whether a candidate hot word is detected in the monophonic clean audio stream 255. As described in more detail below, the cleaner 250 uses a noise cleaning algorithm to provide adaptive noise cancellation for the multi-channel noisy audio. In other examples, the "processing" of each audio feature by the DSP 110 includes omitting the use of the cleaner 250 and simply processing each audio feature of one channel 119 of the streaming multi-channel audio 118 without canceling noise from the respective audio feature. In these examples, the channel 119 for which each audio feature is processed may be selected arbitrarily.

[0030] If the first-stage hot word detector 210 detects a hot word in the streaming multi-channel audio 118, the DSP 110 provides the chopped multi-channel raw audio data 212, 212a-n to the AP 120. In some examples, the DSP 110 providing the chopped multi-channel raw audio data 212 to the AP 120 triggers / wakes up the AP 120 to transition from sleep mode to hot word detection mode. Optionally, the DSP 110 may provide another signal or instruction to trigger / wake up the AP 120 to transition from sleep mode to hot word detection mode. Each channel of the chopped multi-channel raw audio data 212a-n corresponds to a respective channel 119a-n of the streaming multi-channel audio 118 and includes chopped raw audio data from a respective audio feature of a respective channel 119 of the streaming multi-channel audio 118. In some implementations, each channel of the chopped multi-channel raw audio data 212 includes an audio segment characterizing a hotword detected by the first-stage hotword detector 210 in the streaming multi-channel audio 118. That is, the audio segment associated with each channel of the chopped multi-channel raw audio data 212 includes a duration sufficient to safely contain the detected hotword. In addition, each channel of the chopped multi-channel raw audio data 212 includes a prefix segment 214 that includes the duration of audio immediately prior to the point at which the first-stage hotword detector 210 detected the hotword in the streaming multi-channel audio 118. A portion of each channel of the chopped multi-channel raw audio data 212 may also include a suffix segment that includes the duration of audio following the audio segment 213 that includes the detected hotword.

[0031] When operating in hot word detection mode, the AP 120 executes / operates the cleaner 250 to leverage the streaming multi-channel audio 118 input from two or more microphones 107 to obtain noise robustness (e.g., adaptive noise cancellation) for improving hot word detection accuracy. Specifically, the cleaner 250 includes a first noise cleaning algorithm that the AP 120 uses to process each channel of the chopped multi-channel raw audio data 212 to generate a clean monophonic audio chop 260. Importantly, the cleaner 250 requires that each channel of the chopped multi-channel raw audio data 212 include a prefix segment 214 of buffered audio samples immediately preceding the detected hot word to fully apply adaptive noise cancellation. The length of the prefix segment 214 needs to be longer when the cleaner 250 is used than in a configuration where the architecture does not include the cleaner 250. For example, the length of the prefix segment 214 only requires about 2 seconds without the cleaner. Generally, a longer prefix segment 214 (e.g., a longer duration of buffered audio samples) improves the performance of the cleaner 250, but also increases latency as the second-stage hot word detector 220 eventually processes the prefix segment 214 to keep up with real-time detection of hot words. Therefore, the cascaded hot word detection architecture 200 may select a prefix segment 214 length that balances latency and cleaner performance. The AP 120 then executes the second-stage hot word detector 220 to process the clean monophonic audio chomp 260 to determine whether a hot word is present within the clean monophonic audio chomp 260.

[0032] If a hotword is detected by the second-stage hotword detector 220, the AP 120 initiates a wake-up process for the user device 102 to process the hotword and / or one or more other terms following the hotword in the streaming multi-channel audio 118. Similar to the first-stage hotword detector 210, the second-stage hotword detector 220 may detect the presence of a hotword if the probability score associated with each clean monophonic audio chorp 260 or each raw audio data 212 satisfies a probability score threshold. The value of the probability score threshold used by the second-stage hotword detector 220 may be the same as or different from the value of the probability score threshold used by the first-stage hotword detector 210.

[0033] As described above, the DSP 110 may use a cleaner 250 that executes a second noise cleaning algorithm before executing the first-stage hot word detector 210 to obtain noise robustness (e.g., adaptive noise cancellation) for improving the hot word detection accuracy of the first-stage hot word detector 210. The filter models for the first and second noise cleaning algorithms may be the same, but the second noise cleaning algorithm may include a shorter length (e.g., fewer filtering parameters) because the DSP 110 involves a lower computational burden than the computational capabilities of the AP 120. Thus, the cleaner 250 used by the DSP 110 sacrifices some performance (e.g., signal-to-noise ratio (SNR) performance) compared to the cleaner used by the AP 120, but still provides sufficient noise robustness to improve the accuracy of the first-stage hot word detector 210.

[0034] The AP 120 may process the clean monophonic audio chomp 260 in parallel with processing the respective raw audio data 212a of one channel of the chopped multi-channel raw audio data 212 to determine whether a hotword is detected by the second-stage hotword detector 220. If the second-stage hotword detector 220 detects a hotword in either the clean monophonic audio chomp 260 or the respective raw audio data 212a, the AP initiates / triggeres a wake-up process for the user device 102 to process the hotword and / or one or more other terms following the hotword in the streaming multi-channel audio 118. If the second-stage hotword detector 220 does not detect a hotword in either the clean monophonic audio chomp 260 or the respective raw audio data 212a, the AP 120 prevents the wake-up process for the user device 102. The wake-up process may include the user device 102 performing voice recognition locally on the hotword and / or one or more other terms, or the wake-up process may include the user device 102 sending audio data including the hotword and / or one or more other terms to a remote server for voice recognition on the hotword and / or one or more other terms. In some examples, the user device 102 may send audio data including the hotword detected by the AP 120 to a remote server to verify that the hotword is present, thus functioning as a third-stage hotword detector.

[0035] 2A-2C illustrate examples of cascaded hot word detection architectures 200, 200a-c that may reside on the user device 102 of FIG. 1 for detecting the presence of hot words in spoken utterances 104. Referring to FIG. 2A, the example cascaded architecture 200a includes only a single cleaner 250, 250a that resides on the AP 120 for use by the second-stage hot word detector 220, such that the first-stage hot word detector 210 does not benefit from any noise cleaning algorithm.

[0036] In the illustrated example, for simplicity, the streaming multi-channel audio 118 includes two channels 119a, 119b, each including a respective audio feature captured by a separate dedicated microphone 107a-b in the array of two microphones 107. However, the streaming multi-channel audio 118 may include more than two channels without departing from the scope of this disclosure.

[0037] 2A shows an always-on DSP 110 (e.g., a first processor) that includes / executes a first-stage hotword detector 210 and an audio chopper 215. The first-stage hotword detector 210 receives as input only audio features from a single channel 119 a of the streaming multi-channel audio 118, and the channel 119 a received by the first-stage hotword detector 210 may be arbitrary. Here, the DSP 110 uses / executes the first-stage hotword detector 210 to process the audio features of each channel 119 a to determine whether a hotword has been detected by the first-stage hotword detector 210 in the streaming multi-channel audio 118. The first-stage hotword detector 210 may calculate a probability score indicative of the presence of a hotword in each audio feature from the single channel 119 a of the streaming multi-channel audio 118. In some examples, a determination that the probability score of the respective audio feature satisfies the hotword threshold (eg, the probability score is greater than or equal to the hotword threshold) indicates that the hotword is present in the streaming multi-channel audio 118.

[0038] If the first-stage hotword detector 210 detects a hotword in the streaming multi-channel audio 118 (e.g., in the first channel 119a), the DSP 110 triggers / starts the audio chopper 215 to generate and provide chopped multi-channel raw audio data 212, 212a-b, to the AP 120, where each channel of the chopped multi-channel raw audio data 212 corresponds to a respective channel 119a-b of the streaming multi-channel audio 118 and includes respective chopped raw audio data from the respective channel 119a-b that includes the hotword detected by the first-stage hotword detector 210. The provision of the chopped multi-channel raw audio data 212 from the DSP 110 to the AP 120 wakes the AP 120 to transition from sleep mode to hotword detection mode, in which the AP 120 executes a first noise cleaning algorithm in the cleaner engine 250a and the second-stage hotword detector 220. 2A , the DSP 110 does not employ a noise cleaning algorithm on the streaming multi-channel audio 118, and therefore the first-stage hot word detector 210 does not benefit from adaptive noise cancellation when determining whether the presence of a hot word is detected in the noisy respective audio feature of a single channel 119 a. In other words, processing the respective audio of one channel 119 a of the streaming multi-channel audio 118 occurs without canceling noise from the respective audio feature.

[0039] In the cascaded hot word detection architecture 200a of FIG. 2A, the cleaner 250 runs entirely on the AP 220 as the cleaner engine 250a, which uses a first noise cleaning algorithm to process each channel of the chopped multi-channel raw audio data 212 to generate a clean monophonic audio chomp 260. As described in more detail below with reference to FIG. 3, the cleaning operation performed by the cleaner engine 250a is relatively computationally complex and requires that each channel of the chopped multi-channel raw audio data 212a, 212b include a prefix segment 214 of a longer duration than the duration of the prefix segment in the chomp if the cleaning algorithm is not applied. For example, if the cleaning algorithm is not applied at the AP 120 such that only the raw audio data chomp is passed to the second-stage hot word detector 220, the prefix segment 214 of each chomp may include a duration of approximately 2 seconds. However, when the cleaning engine 250a is used to process each channel of the chopped multi-channel raw audio data 212, the cleaning engine 250a needs to determine the duration of the audio preceding the hot word, which contains only noise, in order to estimate an effective noise cancellation filter for generating the clean monophonic audio chop 260. Therefore, the performance (e.g., SNR) of the cleaning engine 250a improves with the duration of the prefix segment 214, which contains noisy audio, preceding the audio segment 213 containing the detected hot word. However, a longer prefix segment 214 results in an increased latency cost because the cleaning engine 250a and the second-stage hot word detector 220 must process the longer prefix segment 214 in the multi-channel raw audio data 212.This increased latency is due to the second stage hot word detector 220 having to process longer prefix segments 214 to keep up with real-time detection of hot words contained within the audio segments 213. In one example, the audio chopper 215 in the DSP generates prefix segments 214 with a duration / length equal to approximately 3.5 seconds to balance clean performance and latency. A portion of each channel of the chopped multi-channel raw audio data 212 may also include a suffix segment containing the duration of audio following the audio segment 213 containing the detected hot word.

[0040] While in hotword detection mode, the second-stage hotword detector 220 executing on the AP 120 is configured to process the clean monophonic audio chomp 260 output from the cleaner engine 250a to determine whether a hotword is detected in the clean monophonic audio chomp 260. In some examples, the second-stage hotword detector 220 corresponds to a parallel hotword detection architecture that performs hotword detection independently on two branches 220a, 220b using the same model but on two different inputs: the raw audio data 212a for each channel of the chopped multi-channel raw audio data 212, and the clean monophonic audio chomp 260. The channel associated with each raw audio data 212a provided as input to the second branch 220b of the second-stage hotword detector 220 may be arbitrary. Thus, the AP 120 may process the clean monophonic audio chomp 260 in the first branch 220a of the second-stage hot word detector 220 in parallel with processing the respective raw audio data 212a of one channel of the chopped multi-channel raw audio data 212 in the second branch 220b of the second-stage hot word detector 220. In the illustrated example, if the logical OR 270 operation indicates that a hot word has been detected by the second-stage hot word detector 220 in either the clean monophonic audio chomp 260 (e.g., in the first branch 220a) or the respective raw audio data 212a (e.g., in the second branch 220b), the AP 120 initiates a wake-up process for the user device 102 to process the hot word and / or one or more other terms following the hot word in the streaming multi-channel audio 118.Similar to the first-stage hot word detector 210, the second-stage hot word detector 220 may detect the presence of a hot word if the probability score associated with each clean monophonic audio chorp 260 or each raw audio data 212 satisfies a probability score threshold. The probability score threshold used by the second-stage hot word detector 220 may be the same or different from the probability score threshold used by the first-stage hot word detector 210.

[0041] To minimize the length of the prefix segment 214 for each channel of the chopped multi-channel raw audio data 212 provided by the DSP 110, and therefore reduce the latency of the cleaner engine 250a in processing each channel of the chopped multi-channel raw audio data 212 to generate a clean monophonic audio chomp 260, the example cascaded hotword architecture 200b of FIG. 2B includes the DSP 110 executing a cleaner front end 252 tasked with updating and buffering a multi-microphone cross-correlation matrix 254. As will become apparent, the cleaner front end 252 is configured to extract acoustic features from the streaming multi-channel raw audio data 212 prior to chopping in the audio chopper 215, and thereby track these extracted acoustic features for use by the cleaner engine 250a, similar to a multi-channel full-band cross-correlation matrix or a multi-channel sub-band coherence matrix, up until the time the audio is chopped. Here, each multi-microphone cross-correlation matrix 254 corresponds to noise cancellation between respective audio features of the streaming multi-channel audio 118. In contrast to the cascaded hot word detection architecture 200a of FIG. 2A , the cleaner engine 250a in the AP 120 is no longer tasked with calculating / generating the multi-microphone cross-correlation matrix 254 and can therefore process chopped multi-channel raw audio data 212 with prefix segments 214 that include shorter durations of audio immediately prior to the time when the first-stage hot word detector detected the hot word in the streaming multi-channel audio 118. For example, the length of the prefix segments 114 may be less than 3.5 seconds. Latency is essentially improved by allowing the AP 120 to process raw audio data 212 with prefix segments 214 of shorter duration.In the illustrated example, for simplicity, the streaming multi-channel audio 118 includes two channels 119a, 119b, each including a respective audio feature captured by a separate dedicated microphone 107a-b in the array of two microphones 107. However, the streaming multi-channel audio 118 may include more than two channels without departing from the scope of this disclosure.

[0042] 2B shows an always-on DSP 110 (e.g., a first processor) that includes / executes a first-stage hotword detector 210, an audio chopper 215, and a cleaner front end 253. The first-stage hotword detector 210 receives as input only audio features from a single channel 119a of the streaming multi-channel audio 118, and the channel 119a of the audio features may be optional. Here, the DSP 110 uses / executes the first-stage hotword detector 210 to process the audio features to determine whether a hotword has been detected by the first-stage hotword detector 210 in the streaming multi-channel audio 118. Simultaneously, the audio chopper 215 and the cleaner front end 252 receive the respective audio features of each channel 119 a, 119 b of the streaming multi-channel audio 118 such that the cleaner front end 252 generates a multi-channel cross-correlation matrix 254 associated with calculating noise cancellation between the respective audio features of each channel 119 a, 119 b of the streaming multi-channel audio 118. More specifically, the cleaner front end 252 continuously calculates / updates and buffers the multi-channel cross-correlation matrix 254 as the streaming multi-channel audio 118 is received. The cleaner front end 252 may buffer the multi-channel cross-correlation matrix 254 within the memory hardware 105 ( FIG. 1 ) of the user device 102.

[0043] 3 shows a schematic diagram 300 illustrating example cleaning subtasks performed by the cleaner front end 252 executing in the always-on DSP 110 and the cleaner engine 250a executing in the AP 120 when the AP 120 is in hot word detection mode. The cleaner front end 252 may include short-time Fourier transform (STFT) modules 310, 310a-b, each configured to convert a respective audio feature of each channel of the streaming multi-channel audio 118 into an STFT spectrum, whereby the respective converted audio features are provided as inputs to a matrix computer 320 in the cleaner front end 252 and a cleaned STFT spectrum computer 330 in the cleaner engine 250a.

[0044] The matrix computer 320 in the cleaner front end 252 is configured to continuously calculate / update and buffer a multi-channel cross-correlation matrix 254 based on the respective transformed audio features in each channel 119a, 119b. The matrix computer 320 may buffer the matrix 254 in a matrix buffer 305. The matrix buffer 305 is in communication with the DSP 110 and may reside on the memory hardware 105 (FIG. 1) of the user device 102. When the first-stage hot word detector 210 detects a hot word that triggers / activates the AP 120 to transition from sleep mode to hot word detection mode, the DSP 110 may pass the buffered multi-channel cross-correlation matrix 254 to the cleaner engine 250a in the AP 120. More specifically, the cleaner engine 250a includes a cleaner filter coefficient computer 340 configured to calculate cleaner filter coefficients 342 for a first noise cleaning algorithm based on the multi-channel cross-correlation matrix 254 received from the DSP 110. Here, the cleaned STFT spectrum computer 330 corresponds to a noise cancellation filter that executes a first noise cleaning algorithm with calculated cleaner filter coefficients 342, whereby the STFT output 332 of the cleaned STFT spectrum computer 330 is transformed by an STFT inverse module 334 to generate the clean monophonic audio chord 260. In some examples, the cleaned STFT spectrum computer 330 includes a finite impulse response (FIR) filter.

[0045] 2B , the first-stage hotword detector 210 may calculate a probability score indicating the presence of a hotword in each audio feature from a single channel 119 a of the streaming multi-channel audio 118. In some examples, a determination that the probability score of the respective audio feature satisfies a hotword threshold (e.g., the probability score is greater than or equal to the hotword threshold) indicates the presence of a hotword in the streaming multi-channel audio 118. In some embodiments, if the first-stage hotword detector 210 detects a hotword in the streaming multi-channel audio 118, the DSP 110 triggers / starts the audio chopper 215 to use the multi-channel cross-correlation matrix 254 generated by the cleaner front end 252 and stored in the buffer 305 to chop the respective raw audio data 212 a, 212 b from each audio feature of each channel 119 a, 119 b of the streaming multi-channel audio 118. Accordingly, the audio chopper 215 provides the chopped multi-channel raw audio data 212 to the cleaner engine 250 a in the AP 120. Here, provision of chopped multi-channel raw audio data 212 from DSP 110 to AP 120 may wake / trigger AP 120 to transition from sleep mode to hot word detection mode, in which AP 120 executes first noise cleaning algorithm 250a and second stage hot word detector 220.

[0046] The detection of a hot word by the first stage hot word detector 210 also causes the DSP 110 to instruct the cleaner front end 252 to provide a multi-channel cross-correlation matrix 254 to the cleaner engine 250a of the AP 120. Here, the cleaner engine 250a uses the multi-channel cross-correlation matrix 254 to calculate cleaner filter coefficients 342 for a first noise cleaning algorithm. The cleaner engine 250a then executes the first noise cleaning algorithm with the calculated cleaner coefficients 342 to process each channel of the chopped multi-channel audio data 212 provided from the audio chopper 215 to generate a clean monophonic audio chopper 260.

[0047] While in hotword detection mode, the second-stage hotword detector 220 executing on the AP 120 is configured to process the clean monophonic audio chomp 260 to determine whether a hotword is detected in the clean monophonic audio chomp 260. In some examples, the second-stage hotword detector 220 corresponds to a parallel hotword detection architecture that performs hotword detection independently on two branches 220a, 220b using the same model but on two different inputs: the raw audio data 212a for each channel of the chopped multi-channel raw audio data 212, and the clean monophonic audio chomp 260. The channel associated with each raw audio data 212a provided as input to the second branch 220b of the second-stage hotword detector 220 may be arbitrary. Thus, the AP 120 may process the clean monophonic audio chomp 260 in the first branch 220a of the second-stage hot word detector 220 in parallel with processing the respective raw audio data 212a of one channel of the chopped multi-channel raw audio data 212 in the second branch 220b of the second-stage hot word detector 220. In the illustrated example, if the logical OR 270 operation indicates that a hot word has been detected by the second-stage hot word detector 220 in either the clean monophonic audio chomp 260 (e.g., in the first branch 220a) or the respective raw audio data 212a (e.g., in the second branch 220b), the AP 120 initiates a wake-up process for the user device 102 to process the hot word and / or one or more other terms following the hot word in the streaming multi-channel audio 118. Similar to the first stage hot word detector 210, the second stage hot word detector 220 may detect the presence of a hot word when the probability score associated with each clean monophonic audio chorp 260 or each raw audio data 212 meets a probability score threshold.The probability score threshold used by the second-stage hot word detector 220 may be the same or different from the probability score threshold used by the first-stage hot word detector 210 .

[0048] Generally, hotword detection performance is measured by two error rates: a false accept rate (FAR) (e.g., incorrectly detecting a hotword) and a false reject rate (FRR) (e.g., failing to detect the current hotword). Therefore, a hotword can be identified by either of the cascade hotword detection architectures 200 only if both the first-stage hotword detector 210 and the second-stage hotword detector 220 detect the hotword. Therefore, the overall FAR of the cascade hotword detection architectures 200a, 200b is lower than the FARs of both the first-stage hotword detector 210 and the second-stage hotword detector 220. Additionally, the overall FRR of the cascade hotword detection architectures 200a, 200b is higher than the FRRs of both the first-stage hotword detector 210 and the second-stage hotword detector 220. For example, if the FRR of the first-stage hotword detector 210 is kept low, the overall FRR will be approximately the same as the FRR of the second-stage hotword detector 220. In some examples, the FAR of the first-stage hotword detector 210 is set to a reasonable value so that the second-stage hotword detector 220 is not triggered frequently to reduce power consumption by the AP 120. However, in the cascaded hotword architectures 200a, 200b of Figures 2A and 2B, the first-stage hotword detectors 210 do not benefit from the cleaner 250, and therefore, their FRR in noisy environments will be high even if the FAR of the first-stage hotword detectors 210 is adjusted to a higher value. Therefore, the cascaded hot word architectures 200a, 200b of FIGS. 2A-2C experience lower performance than the cascaded hot word detection architecture 200c of FIG. 2C, which uses a light cleaner 250b (eg, cleaner light) in the DSP 110.

[0049] To achieve an optimal balance between small footprint, low latency, and maximized accuracy in both clean and noisy environments, the exemplary cascaded hot word detection architecture 200c of FIG. 2C includes the DSP 110 employing a first stage cleaner 250b (e.g., cleaner-lite) that processes the respective audio features of each channel 119 of the streaming multi-channel audio 118 and executes a second noise cleaning algorithm to generate a monophonic clean audio stream 255 before executing / operating the first stage hot word detector 210 to determine whether a candidate hot word is detected in the monophonic clean audio stream 255. In other words, the cleaner-lite 250b employed in the DSP 110 executes a second noise cleaning algorithm to provide adaptive noise cancellation to the streaming multi-channel audio 118 such that the resulting monophonic clean audio stream 255 input to the first-stage hot word detector 210 contains an improved SNR compared to the raw audio features of each of the single channels 119 input to the detector 210 in the architectures 200a, 200b of Figures 2A and 2B. Thus, the hot word detection accuracy of the first-stage hot word detector 210 is improved when the first-stage hot word detector 210 can coarsely screen the monophonic clean audio stream 255 for hot word candidates as opposed to coarsely screening the raw audio features of the single channels for hot word candidates.

[0050] The filter models for the first and second noise cleaning algorithms may be the same, or alternatively, substantially similar, but because the DSP 110 includes lower computing power than that of the AP 120, the second noise cleaning algorithm executed on the cleaner engine 250b in the DSP 110 may include a shorter length (e.g., fewer filtering parameters) than the first noise cleaning algorithm executed on the cleaner engine 250a in the AP 120. For example, the first noise cleaning algorithm may apply a first finite impulse response (FIR) to each channel of the chopped multi-channel raw audio data 212 to generate a clean monophonic audio chop 260, and the second noise cleaning algorithm may apply a second FIR to each channel 119 of the streaming multi-channel audio 118 to generate a monophonic clean audio stream 255. In this example, the first FIR in cleaner engine 250a may include a first filter length, and the second FIR in cleaner-lite 250b may include a second filter length that is shorter than the first filter length. Thus, cleaner-lite 250b used by DSP 110 sacrifices some performance (e.g., signal-to-noise ratio (SNR) performance) compared to cleaner engine 250a used by AP 120, but still provides sufficient noise robustness to improve the accuracy of first-stage hot word detector 210.

[0051] 2C shows an always-on DSP 110 (e.g., a first processor) that includes / executes a cleaner-lite 250 (e.g., a cleaner), a first-stage hot word detector 210, an audio chopper 215, and a cleaner front-end 252. The cleaner-lite 250b receives as input audio features of both channels 119a, 119b of the streaming multi-channel audio 118 and executes a second noise cleaning algorithm to generate a monophonic clean audio stream 255 from each channel 119a, 119b of the streaming multi-channel audio 118. When generating the monophonic clean audio stream 255, the DSP 110 uses / executes the first-stage hot word detector 210 to process the monophonic clean audio stream 255 to determine whether any hot words have been detected in the monophonic clean audio stream 255 by the first-stage hot word detector 210. Optionally, and provided that the DSP 110 is not constrained by computational limitations, the first-stage hotword detector 210 may support a parallel hotword detection architecture that performs hotword detection in two branches using the same model, but independently on two different inputs: the raw audio features of each of one channel 119a of the streaming multi-channel audio 118, and the monophonic clean audio stream 255. In this optional configuration, logic and / or operations may be used to determine that a hotword is present in the streaming multi-channel audio 118 if a hotword is detected in either of the two hotword detection branches of the first-stage hotword detector 210.

[0052] At the same time that cleaner-lite 250b is executing the second noise cancellation algorithm, audio chopper 215 and cleaner front-end 252 receive respective audio features of each channel 119a, 119b of streaming multi-channel audio 118, such that cleaner front-end 252 generates a multi-channel cross-correlation matrix 254 associated with calculating noise cancellation between respective audio features of each channel 119a, 119b of streaming multi-channel audio 118. More specifically, and as discussed above with reference to FIGS. 2B and 3, cleaner front-end 252 continuously calculates / updates and buffers multi-channel cross-correlation matrix 254 as streaming multi-channel audio 118 is received. Cleaner front-end 252 may buffer multi-channel cross-correlation matrix 254 within memory hardware 105 (FIG. 1) of user device 102, as discussed above with reference to FIG. 3.

[0053] The first-stage hotword detector 210 may calculate a probability score indicating the presence of a hotword in the monophonic clean audio stream 255 of the streaming multi-channel audio 118. In some examples, a determination that the probability score of the monophonic clean audio stream 255 satisfies a hotword threshold (e.g., the probability score is greater than or equal to the hotword threshold) indicates the presence of a hotword in the streaming multi-channel audio 118. In some embodiments, if the first-stage hotword detector 210 detects a hotword in the streaming multi-channel audio 118, the DSP 110 triggers / starts the audio chopper 215 to use the multi-channel cross-correlation matrix 254 generated by the cleaner front end 252 and stored in the buffer 305 (FIG. 3) to chop the respective raw audio data 212a, 212b from the respective audio features of the respective channels 119a, 119b of the streaming multi-channel audio 118. Thus, the audio chopper 215 provides the chopped multi-channel raw audio data 212 to the cleaner engine 250a in the AP 120. Here, providing the chopped multi-channel raw audio data 212 from the DSP 110 to the AP 120 may wake / trigger the AP 120 to transition from sleep mode to hot word detection mode, in which the AP 120 executes the first noise cleaning algorithm on the cleaner engine 250a and the second stage hot word detector 220.

[0054] The detection of a hot word by the first-stage hot word detector 210 also causes the DSP 110 to instruct the cleaner front end 252 to provide a multi-channel cross-correlation matrix 254 to the cleaner engine 250a of the AP 120. Here, the cleaner engine 250a uses the multi-channel cross-correlation matrix 254 to calculate cleaner filter coefficients for a first noise cleaning algorithm. The cleaner engine 250a then executes the first noise cleaning algorithm with the calculated cleaner coefficients to process each channel of the chopped multi-channel audio data 212 provided from the audio chopper 215 to generate a clean monophonic audio chopper 260.

[0055] While in hotword detection mode, the second-stage hotword detector 220 executing on the AP 120 is configured to process the clean monophonic audio chomp 260 to determine whether a hotword is detected in the clean monophonic audio chomp 260. In some examples, the second-stage hotword detector 220 corresponds to a parallel hotword detection architecture that performs hotword detection independently on two branches 220a, 220b using the same model but on two different inputs: the raw audio data 212a for each channel of the chopped multi-channel raw audio data 212, and the clean monophonic audio chomp 260. The channel associated with each raw audio data 212a provided as input to the second branch 220b of the second-stage hotword detector 220 may be arbitrary. Thus, the AP 120 may process the clean monophonic audio chomp 260 in the first branch 220a of the second-stage hot word detector 220 in parallel with processing the respective raw audio data 212a of one channel of the chopped multi-channel raw audio data 212 in the second branch 220b of the second-stage hot word detector 220. In the illustrated example, if the logical OR 270 operation indicates that a hot word has been detected by the second-stage hot word detector 220 in either the clean monophonic audio chomp 260 (e.g., in the first branch 220a) or the respective raw audio data 212a (e.g., in the second branch 220b), the AP 120 initiates a wake-up process for the user device 102 to process the hot word and / or one or more other terms following the hot word in the streaming multi-channel audio 118. Similar to the first stage hot word detector 210, the second stage hot word detector 220 may detect the presence of a hot word when the probability score associated with each clean monophonic audio chorp 260 or each raw audio data 212 meets a probability score threshold.The probability score threshold used by the second-stage hot word detector 220 may be the same or different from the probability score threshold used by the first-stage hot word detector 210 .

[0056] In some examples, the second-stage hotword detector 220 utilizes a multi-channel hotword model trained to detect hotwords in the multi-channel input. In these examples, the second-stage hotword detector 220b is configured to take all of the chopped multi-channel raw audio data 212 and make a determination of whether a hotword is detected in the chopped multi-channel raw audio data 212. Similarly, in these examples, the cleaner engine 250a may be adapted to replicate the clean monophonic audio chomp 260 to the multi-channel output such that the multi-channel hotword model in the first branch 220a of the second-stage hotword detector 220 captures the clean multi-channel audio chomp 260. Instead of generating a multi-channel output, the cleaner engine 250a may instead be adapted to take the entire chopped multi-channel raw audio data 212 to generate the clean multi-channel audio chomp 260. The multi-channel hot word model may include a memorized neural network with a three-dimensional (3D) single value decomposition filter (SVDF) input layer and sequentially stacked SVDF layers, as disclosed in International Patent Application PCT / US20 / 13705, filed January 15, 2020, the contents of which are incorporated by reference in its entirety. In another example, the second-stage hot word detector 220 utilizes a multi-channel hot word model trained to detect hot words in both the raw multi-channel audio data 212 and the clean multi-channel audio chomp 260.

[0057] 4 is a flowchart of an exemplary arrangement of operations for a method 400 of detecting hot words in streaming multi-channel audio 118 using the noise-robust cascaded hot word detection architecture 200. At operation 402, the method 400 includes receiving, at a first processor 110 of the user device 102, streaming multi-channel audio 118 captured by an array of microphones 107 in communication with the first processor 110, where the first processor 110 may include an always-on DSP. Each channel 119 of the streaming multi-channel audio includes a respective audio feature captured by a separate dedicated microphone in the array of microphones 107.

[0058] At operation 404, the method 400 includes processing, by the first processor 110, audio features of each of at least one channel of the streaming multi-channel audio 118 using the first-stage hotword detector 210 to determine whether a hotword is detected by the first-stage hotword detector 210. If the first-stage hotword detector 210 detects a hotword in the streaming multi-channel audio 118, the method 400 includes, at operation 406, providing, by the first processor 110, the chopped multi-channel raw audio data 212 to a second processor of the user device 102. Each channel of the chopped multi-channel raw audio data 212 corresponds to a respective channel 119 of the streaming multi-channel audio 118 and includes respective raw audio data chopped from each channel 119 of the streaming multi-channel audio 118. The second processor 120 may include a device SoC, such as an AP. Before detecting a hotword in the first-stage hotword detector 210, the second processor 120 may operate in a sleep mode to conserve power and computational resources. When detecting a hotword in the first-stage hotword detector 210, the first processor 110 triggers / wakes up the second processor 120 to transition from sleep mode to hotword detection mode. The passing of the chopped multi-channel raw audio data 212 from the first processor 110 to the second processor 120 may serve as a basis for waking / triggering the second processor 120 to transition to hotword detection mode. Thus, the first processor 110 is configured to transition the second processor from sleep mode to hotword detection mode when the first-stage hotword detector 210 detects a hotword in the streaming multi-channel audio 118. The hotword may be a given term / phrase of one or more words, for example, "Hey Google," and / or any other term / phrase that can be used to initialize an application.The hotwords may be custom hotwords in some configurations.

[0059] At operation 408, the method 400 also includes, by the second processor 120, processing each channel of the chopped multi-channel raw audio data 212 using the first noise cleaning algorithm 250 to generate a clean monophonic audio chomp 260. Each channel of the chopped multi-channel raw audio data 212 includes a respective audio segment 213 including a detected hot word and a respective prefix segment 214 including a duration of noisy audio preceding the detected hot word. The prefix segment 214 includes a duration sufficient for the first noise cleaning algorithm 250 to process enough noisy audio preceding the detected hot word to generate the clean monophonic audio chomp 260. A prefix segment 214 with a longer duration improves the performance of the first noise cleaning algorithm, but a longer prefix segment also increases latency. Thus, the respective prefix segment 214 of each channel of the multi-channel raw audio data 212 may include a duration that optimizes cleaning performance and latency.

[0060] At operation 410, the method 400 includes processing the clean monophonic audio chomp 260 using the second stage hot word detector 220, by the second processor 120, to determine whether a hot word is detected by the second stage hot word detector 220 in the clean monophonic audio chomp 260. At operation 412, if a hot word is detected by the second stage hot word detector 220 in the clean monophonic audio chomp 260, the method 400 also includes initiating, by the second processor 120, a wake-up process for the user device 102 to process the hot word and / or one or more other terms following the hot word in the streaming multi-channel audio 118.

[0061] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform tasks. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0062] Non-transitory memory may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by a computing device. Non-transitory memory may be volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.

[0063] 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems and methods described in this document. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functionality are intended to be exemplary only and are not intended to limit the implementation of the invention described and / or claimed in this document.

[0064] Computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 that connects to memory 520 and a high-speed expansion port 550, and a low-speed interface / controller 560 that connects to a low-speed bus 570 and storage device 530. Each of components 510, 520, 530, 540, 550, and 560 are interconnected using various buses and may be mounted on a common motherboard or otherwise as desired. Processor 510 can process instructions for execution within computing device 500, including instructions stored in memory 520 or on storage device 530, to display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 580 coupled to high-speed interface 540. In other implementations, multiple processors and / or multiple buses may be used, as desired, along with multiple memories and types of memory. Also, multiple computing devices 500 may be connected (eg, as a server bank, a group of blade servers, or a multiprocessor system), with each device providing a portion of the required operations.

[0065] The memory 520 stores information non-transiently within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 520 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.

[0066] The storage device 530 can provide mass storage for the computing device 500. In some implementations, the storage device 530 is a computer-readable medium. In various different implementations, the storage device 530 can be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or a device including a device in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 520, the storage device 530, or memory on the processor 510.

[0067] The high-speed controller 540 manages bandwidth-intensive operations for the computing device 500, and the low-speed controller 560 manages less bandwidth-intensive operations. Such role assignments are merely exemplary. In some implementations, the high-speed controller 540 is coupled to the memory 520, the display 580 (e.g., via a graphics processor or accelerator), and the high-speed expansion port 550, which may accept various expansion cards (not shown). In some implementations, the low-speed controller 560 is coupled to the storage device 530 and the low-speed expansion port 590. The low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device, such as a switch or router, for example, via a network adapter.

[0068] Computing device 500 may be implemented in several different forms, as shown in the figure. For example, computing device 500 may be implemented as a standard server 500a, or multiple times within a group of servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.

[0069] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage device, at least one input device, and at least one output device.

[0070] These computer programs (also known as programs, software, software applications, or code) include machine language for programmable processors and can be implemented in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0071] The processes and logic flows described herein may be performed by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by special-purpose logic circuitry, e.g., an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, e.g., magnetic, magneto-optical, or optical disks, for storing data, or is operably coupled to receive data from, transmit data to, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0072] To provide for user interaction, one or more aspects of the present disclosure may be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide input to the computer. Other types of devices may be used to provide for user interaction as well, and the feedback provided to the user may be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic input, voice input, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from devices used by the user, e.g., by sending web pages to the web browser in response to requests received from the web browser on the user's client device.

[0073] Although several implementations have been described, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]

[0074] 10 users 100 systems 102 User Devices 104 utterances 105 Memory Hardware 107, 107a~n microphones 110 first processor, dedicated DSP, DSP 114 prefix segments 118 Streaming Multi-Channel Audio, Audio, Multi-Channel Streaming Audio, Two-Channel Streaming Audio 119, 119a~n channels 120 second processor, main AP, AP 200, 200a-c Cascade hot word detection architecture, cascade architecture 200a Cascaded Hotword Detection Architecture, Cascaded Hotword Architecture 200b Cascaded Hotword Architecture, Cascaded Hotword Detection Architecture 200c Cascaded Hotword Detection Architecture 210 First Stage Hot Word Detector, Detector 212 Streaming Multi-Channel Raw Audio Data 212a, 212b Raw audio data 212, 212a-n Engraved multi-channel raw audio data 213 Audio Segments 214 prefix segment 215 Audio Chomp 220 Second Stage Hot Word Detector 220a Branch Branch 220b, second stage hot word detector 225 monophonic clean audio streams 250, 250a cleaner 250a cleaner engine 250b lightweight cleaner, Cleaner-Lite 252 Cleaner Front End 254 Multi-microphone cross-correlation matrix, matrix 255 monophonic clean audio streams 260 clean monophonic audio chops, clean multichannel audio chops 270 Logical OR 300 Schematic 305 Matrix Buffer 320 Matrix Computer 330 Cleaned STFT Spectrum Computer 332 STFT output 334 STFT Inverse Module 340 Cleaner filter coefficient computer 342 Cleaner filter coefficient 500 computing devices 500a Server 500b laptop computer 500c Rack Server System 510 processor 520 memory 530 Storage Devices 540 High-Speed ​​Interface / Controller 550 High-Speed ​​Expansion Port 560 Low-Speed ​​Interface / Controller 570 Slow Bus

Claims

1. 1. A computer-implemented method executed on data processing hardware of a user device to cause the data processing hardware to perform an operation, the operation comprising: receiving streaming multi-channel audio captured by an array of microphones, each channel of the streaming multi-channel audio including a respective audio feature captured by a separate dedicated microphone in the array of microphones; processing the respective audio features of each channel of the streaming multi-channel audio using a noise cleaning algorithm to generate a clean multi-channel audio output; processing the respective audio features of at least one channel of the streaming multi-channel audio using a first-stage hot word detector to determine that a hot word has been detected in the streaming multi-channel audio by the first-stage hot word detector; using a second-stage hot word detector based on determining that the hot word was detected in the streaming multi-channel audio by the first-stage hot word detector; processing the cleaned multi-channel audio output produced using the noise cleaning algorithm to determine if the hot word is detected in the cleaned multi-channel audio output by the second stage hot word detector; and processing the respective audio features of one of the channels selected from the streaming multi-channel audio to determine whether the hot word is detected by the second-stage hot word detector in the respective audio feature of the selected channel; and initiating a wake-up process for the user device when the hot word is detected by the second stage hot word detector in either the clean multi-channel audio output or the respective audio characteristics of the selected channels; A method comprising:

2. 10. The method of claim 1, wherein the operations further comprise: preventing initiation of a wake-up process for the user device when the hotword is not detected by the second stage hotword detector in either the clean multi-channel audio output or the respective audio features of the selected channels.

3. The data processing hardware includes: a first processor that executes the first stage hot word detector; a second processor that executes the second stage hot word detector; The method of claim 1 , comprising:

4. the first processor includes a digital signal processor; The method of claim 3 , wherein the second processor comprises a system-on-chip (SoC) processor.

5. when the streaming multi-channel audio is received by the first processor and the respective audio features of the at least one channel of the streaming multi-channel audio are processed by the first stage hot word detector, the second processor operates in a sleep mode; 4. The method of claim 3, wherein if the first stage hot word detector detects the hot word in the streaming multi-channel audio, the first processor wakes up the second processor to transition from the sleep mode to a hot word detection mode.

6. The method of claim 3 , wherein the user device includes a rechargeable limited power source, the limited power source providing power to the first processor and the second processor.

7. The method of claim 1 , wherein the user device comprises a smartphone.

8. The method of claim 1 , wherein the user device comprises a tablet.

9. The method of claim 1 , wherein the user device comprises smart headphones.

10. The method of claim 1 , wherein the user device comprises a smart speaker.

11. data processing hardware of the user device; and memory hardware in communication with the data processing hardware and storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations comprising: receiving streaming multi-channel audio captured by an array of microphones, each channel of the streaming multi-channel audio including a respective audio feature captured by a separate dedicated microphone in the array of microphones; processing the respective audio features of each channel of the streaming multi-channel audio using a noise cleaning algorithm to generate a clean multi-channel audio output; processing the respective audio features of at least one channel of the streaming multi-channel audio using a first-stage hot word detector to determine that a hot word has been detected in the streaming multi-channel audio by the first-stage hot word detector; using a second-stage hotword detector based on determining that the hotword was detected in the streaming multi-channel audio by the first-stage hotword detector; processing the cleaned multi-channel audio output produced using the noise cleaning algorithm to determine if the hot word is detected in the cleaned multi-channel audio output by the second stage hot word detector; and processing the respective audio features of one of the channels selected from the streaming multi-channel audio to determine whether the hot word is detected by the second-stage hot word detector in the respective audio feature of the selected channel; and initiating a wake-up process for the user device when the hot word is detected by the second stage hot word detector in either the clean multi-channel audio output or the respective audio characteristics of the selected channels; A system comprising:

12. 12. The system of claim 11, wherein the operations further comprise preventing initiation of a wake-up process for the user device when the hotword is not detected by the second stage hotword detector in either the clean multi-channel audio output or the respective audio features of the selected channels.

13. The data processing hardware includes: a first processor that executes the first stage hot word detector; a second processor that executes the second stage hot word detector; The system of claim 11 , comprising:

14. the first processor includes a digital signal processor; The system of claim 13 , wherein the second processor comprises a system-on-chip (SoC) processor.

15. when the streaming multi-channel audio is received by the first processor and the respective audio features of the at least one channel of the streaming multi-channel audio are processed by the first stage hot word detector, the second processor operates in a sleep mode; 14. The system of claim 13, wherein if the first stage hot word detector detects the hot word in the streaming multi-channel audio, the first processor wakes up the second processor to transition from the sleep mode to a hot word detection mode.

16. 14. The system of claim 13, wherein the user device includes a rechargeable limited power source, the limited power source providing power to the first processor and the second processor.

17. The system of claim 11 , wherein the user device comprises a smartphone.

18. The system of claim 11 , wherein the user device comprises a tablet.

19. The system of claim 11 , wherein the user device comprises smart headphones.

20. The system of claim 11 , wherein the user device comprises a smart speaker.

Citation Information

Patent Citations

  • PCT/US20/13705

  • Trigger word based beam selection

    US10304475B1

  • Computerized device with voice command input capability

    US20180330727A1

  • Voice command processing in low power devices

    US20190207777A1