History-based ASR Error Correction
By using a similarity metric between initial and subsequent queries, the method improves transcription accuracy in continuous conversation mode, addressing the need for hotwords and reducing misinterpretation in voice-responsive systems.
Patent Information
- Application Number
- JP2025501472
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-11
- Filing Date
- 2023-07-10
- Publication Date
- 2025-07-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing voice-responsive systems require users to repeatedly speak hotwords for subsequent queries, leading to an unnatural user experience and potential misinterpretation of unintended audio inputs as valid queries.
A method that utilizes an initial query to bias the selection of candidate hypotheses for subsequent queries based on similarity to the initial query, using a similarity metric between query-based salient term vectors to improve transcription accuracy in continuous conversation mode.
Enables accurate transcription of subsequent queries without requiring hotwords, reducing misinterpretation and maintaining user experience by leveraging the context of the initial query.
Smart Images

Figure 2025524643000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to history-based automatic speech recognition (ASR) error correction.
Background Art
[0002] In a voice-responsive environment (e.g., home, workplace, school, automobile, etc.), a user can speak a query or command to a computer-based system, and the computer-based system can answer the query and / or execute a function based on the command. The voice-responsive environment can be implemented using a network of connected microphone devices distributed across various rooms or areas of the environment. The devices operate in a sleep state and can initiate a wake-up process in response to detecting a hotword that the user speaks prior to speaking, in order to perform speech recognition on the utterance directed to the system. After receiving a response from the computer-based system, the user can speak a follow-on query or command. It is unnatural and burdensome for the user to require the user to speak the hotword each time a subsequent query or command occurs.
Summary of the Invention
[0003] One aspect of the present disclosure is a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations, the operations including receiving subsequent audio data captured by an assistant-enabled device, the subsequent audio data corresponding to a subsequent query spoken by a user to a digital assistant after the user of the assistant-enabled device has submitted a prior query to the digital assistant. The operations also include processing the subsequent audio data using a speech recognizer to generate a plurality of candidate hypotheses, each candidate hypothesis corresponding to a candidate transcription of the subsequent query and represented by a respective sequence of hypothesized terms. For each corresponding candidate hypothesis among the plurality of candidate hypotheses, the operations also include determining a corresponding similarity metric between the prior query and the corresponding candidate hypothesis, and determining a transcription of the subsequent query spoken by the user based on the similarity metrics determined for the plurality of candidate hypotheses. The similarity metric indicates a similarity between a topic associated with the corresponding candidate hypothesis and a topic associated with the prior query.
[0004] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, the operation also includes determining a previous query-based salient terms (QBST) vector (previous QBST vector) related to a previous query submitted by a user, and for each corresponding candidate hypothesis, determining a corresponding candidate QBST vector related to the corresponding candidate hypothesis. Here, determining the corresponding similarity metric between the previous query and the corresponding candidate hypothesis includes determining the corresponding similarity metric between the previous QBST vector and the corresponding candidate QBST vector based on the previous QBST vector and the corresponding candidate QBST vector. In these embodiments, the previous QBST vector may indicate each set of salient terms related to the previous query, and each corresponding candidate QBST vector may indicate each set of salient terms related to the corresponding candidate hypothesis. In some examples, the corresponding similarity metric indicates a topical drift between the previous QBST vector and the corresponding candidate QBST vector. In additional examples, the corresponding similarity metric includes a cosine score between the previous QBST vector and the corresponding candidate QBST vector.
[0005] In some examples, the operation further includes receiving initial audio data captured by the assistant-enabled device while the assistant-enabled device is in a sleep state, where the initial audio data includes a hotword and a previous query submitted by the user to the digital assistant, and where detection of the hotword by the assistant-enabled device causes the assistant-enabled device to resume from the sleep state and trigger a speech recognizer to perform speech recognition on at least a portion of the initial audio data that includes the previous query. In these examples, after the speech recognizer performs speech recognition on at least a portion of the initial audio data, the operation also includes instructing the assistant-enabled device to operate in a subsequent query mode. Here, receiving subsequent audio data by the assistant-enabled device includes receiving the subsequent audio data during operation of the assistant-enabled device in the subsequent query mode, and the presence of the hotword is not in the subsequent audio data.
[0006] In some embodiments, the operation further includes performing query interpretation on the transcription to identify an operation specified by a subsequent query, instructing the digital assistant to perform the operation specified by the subsequent query, and receiving from the digital assistant a subsequent response indicating performance of the operation specified by the subsequent query. In these embodiments, the operation may also include presenting the subsequent response for output from the assistant-enabled device.
[0007] For each corresponding candidate hypothesis among the plurality of candidate hypotheses generated by the speech recognizer, the operation may also include obtaining the corresponding likelihood score assigned by the speech recognizer to the corresponding candidate hypothesis, and ranking the plurality of candidate hypotheses based on the corresponding likelihood scores assigned by the speech recognizer to the plurality of candidate hypotheses and the corresponding similarity scores determined for each corresponding candidate hypothesis among the plurality of candidate hypotheses. Here, determining the transcription of the subsequent query may be based on ranking the plurality of candidate hypotheses.
[0008] In particular, the data processing hardware may execute the speech recognizer and may be present either within the assistant-responsive device or within a remote server. The remote server may communicate with the assistant-responsive device via a network.
[0009] In some examples, the speech recognizer includes an end-to-end speech recognition model. In other examples, the speech recognizer includes an acoustic model and a language model. Additionally, the assistant-responsive device may communicate with one or more microphones configured to capture subsequent audio data and initial audio data corresponding to a preceding query. The assistant-responsive device may optionally include a battery-powered device.
[0010] Other aspects of the disclosure provide a system that includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to receive subsequent audio data captured by an assistant-enabled device, where the subsequent audio data corresponds to a subsequent query spoken by a user to a digital assistant after the user of the assistant-enabled device has submitted a prior query to the digital assistant. The operations also include processing the subsequent audio data using a speech recognizer to generate a plurality of candidate hypotheses, where each candidate hypothesis corresponds to a candidate transcription of the subsequent query and is represented by a respective sequence of hypothesized terms. For each corresponding candidate hypothesis among the plurality of candidate hypotheses, the operations also include determining a corresponding similarity metric between the prior query and the corresponding candidate hypothesis and determining a transcription of the subsequent query spoken by the user based on the similarity metrics determined for the plurality of candidate hypotheses. The similarity metric indicates a similarity between a topic associated with the corresponding candidate hypothesis and a topic associated with the prior query.
[0011] This aspect may include one or more of the following optional features. In some embodiments, the operation also includes determining a query-based significant term (QBST) vector related to a previous query submitted by a user, and for each corresponding candidate hypothesis, determining a corresponding candidate QBST vector related to the corresponding candidate hypothesis. Here, determining the corresponding similarity metric between the previous query and the corresponding candidate hypothesis includes determining the corresponding similarity metric between the previous QBST vector and the corresponding candidate QBST vector based on the previous QBST vector and the corresponding candidate QBST vector. In these embodiments, the previous QBST vector may indicate each set of significant terms related to the previous query, and each corresponding candidate QBST vector may indicate each set of significant terms related to the corresponding candidate hypothesis. In some examples, the corresponding similarity metric indicates the topic drift between the previous QBST vector and the corresponding candidate QBST vector. In additional examples, the corresponding similarity metric includes a cosine score between the previous QBST vector and the corresponding candidate QBST vector.
[0012] In some examples, the operation further includes receiving initial audio data captured by the assistant-enabled device while the assistant-enabled device is in a sleep state, where the initial audio data includes a hot word and a previous query submitted by the user to the digital assistant, and where detection of the hot word by the assistant-enabled device causes the assistant-enabled device to resume from the sleep state and trigger a speech recognizer to perform speech recognition on at least a portion of the initial audio data that includes the previous query. In these examples, after the speech recognizer performs speech recognition on at least a portion of the initial audio data, the operation also includes instructing the assistant-enabled device to operate in a subsequent query mode. Here, receiving subsequent audio data by the assistant-enabled device includes receiving the subsequent audio data during operation of the assistant-enabled device in the subsequent query mode, and the subsequent audio data does not include the presence of a hot word.
[0013] In some embodiments, the operation further includes performing query interpretation on the transcription to identify an operation specified by a subsequent query, instructing the digital assistant to perform the operation specified by the subsequent query, and receiving, from the digital assistant, a subsequent response indicating performance of the operation specified by the subsequent query. In these embodiments, the operation may also include presenting the subsequent response for output from the assistant-enabled device.
[0014] For each corresponding candidate hypothesis among a plurality of candidate hypotheses generated by the speech recognizer, the operation may also include obtaining a corresponding likelihood score assigned by the speech recognizer to the corresponding candidate hypothesis, and ranking the plurality of candidate hypotheses based on the corresponding likelihood scores assigned by the speech recognizer to the plurality of candidate hypotheses and the corresponding similarity scores determined for each corresponding candidate hypothesis among the plurality of candidate hypotheses. Here, determining the transcription of the subsequent query may be based on ranking the plurality of candidate hypotheses.
[0015] In particular, the data processing hardware may execute the speech recognizer and may be present either within the assistant-responsive device or within a remote server. The remote server may communicate with the assistant-responsive device via a network.
[0016] In some examples, the speech recognizer includes an end-to-end speech recognition model. In other examples, the speech recognizer includes an acoustic model and a language model. Additionally, the assistant-responsive device may communicate with one or more microphones configured to capture subsequent audio data and initial audio data corresponding to a previous query. The assistant-responsive device may optionally include a battery-powered device.
[0017] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the description and drawings, as well as from the claims.
Brief Description of the Drawings
[0018]
Figure 1A
Figure 1B
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
[0019] Like reference symbols in the various drawings refer to like elements.
DETAILED DESCRIPTION OF THE INVENTION
[0020] Automatic Speech Recognition (ASR) systems are becoming increasingly popular on client devices as they continue to provide more accurate transcriptions of what a user says. Still, in some cases, an ASR system generates an inaccurate transcription that misrecognizes what the user actually said. This often occurs when words are acoustically similar or when the user speaks unique words that are unknown to the ASR system. For example, "I say" and "Ice age" have a very similar sound, and it can be difficult for an ASR system to clarify which phrase the user is trying to convey. In some instances, when a user is actually trying to convey "Ice age" and the client device incorrectly transcribes it as "I say," the user can correct the transcription using the client device (e.g., by entering the correct transcription via the client device's keyboard).
[0021] Voice-based interfaces such as digital assistants are becoming increasingly prevalent across a variety of devices including, but not limited to, mobile phones and smart speakers / displays that include microphones for capturing audio. A common way to initiate a voice conversation with an Assistant Enabled Device (AED) is to speak a fixed phrase, such as a hotword. During streaming audio, when this fixed phrase is detected by the AED, it triggers a wake-up process that causes the AED to start recording and processing subsequent audio to confirm the query spoken by the user. Thus, the hotword is an important component in the overall digital assistant interface stack as it allows the user to wake up the AED from a low-power state to a high-power state so that the AED can perform more expensive processing such as full ASR or server-based ASR.
[0022] Often, after receiving a response to an initial query, the user makes a follow-up query. For example, the user may speak an initial query to the AED such as "What's the temperature in Detroit?", and the AED returns a response to the user saying "It's 76 degrees Fahrenheit." Subsequently, the user may say a follow-up query to the AED such as "Is it sunny?" Here, in the continuous conversation mode or the follow-up query mode, after providing a response by keeping the microphone on and running ASR (on-device and / or server-side) to determine whether any of the subsequent voices are directed towards the AED, the AED is kept in a high-power state for a predefined time. After a predetermined time has elapsed, the AED returns to the low-power state.
[0023] However, there is a challenge in processing the audio input while the AED is in the high-power state to determine whether the subsequent voice is directed towards the AED without degrading the user experience. For example, the subsequently captured voice may be processed by the ASR system and treated as a follow-up query even if the subsequent voice is not intended for the AED (regardless of whether it is the user speaking, another user speaking, or background noise from a nearby device such as a TV or music player). In such a scenario, if the captured voice not intended for the AED is processed as a subsequent one even though it is not for the AED, the AED may return a response that the user did not request or perform some action. Furthermore, the response returned or the action performed in response to an unintended follow-up query is very likely to be unrelated to even the initial query spoken by the user, which starts with an explicit hotword spoken by the user to invoke the ASR in the first place. On the other hand, the initial query spoken by the user can be utilized to improve the accuracy and waiting time for transcribing the subsequent query spoken by the user based on the intuition that the subsequent query should be topically similar to the initial query leading to the follow-up query.
[0024] Accordingly, embodiments of the present specification are directed to a method of utilizing an initial query directed to an AED so as to bias the selection of one speech recognition candidate from among a plurality of speech recognition candidates for a subsequent query based on the similarity to the initial query. Specifically, a user submits an initial query including a hotword and a query to an AED having a digital signal processor (DSP) that operates in a hotword detection mode and is always on. When the DSP operating in the hotword detection mode detects the hotword in the initial query, the DSP starts a wake-up process on a second processor and performs more costly processing such as full ASR that provides a response to the initial query. The second processor may include an application processor (AP) or another type of system-on-chip (SoC) processor.
[0025] In response to receiving a response to the initial query, the AED instructs a second processor to operate in a subsequent query mode (also referred to as a "continuous conversation mode") while maintaining the wake state, in order to enable the ASR system to process subsequent queries that the user may speak without the user having to repeat the hotword. While in the continuous conversation mode, the ASR system (executed on the second processor and / or on a server communicating with the second processor) receives subsequent audio data corresponding to a subsequent query spoken by the user and processes that audio data to generate a plurality of candidate hypotheses. Each of the plurality of candidate hypotheses corresponds to a candidate transcription of the subsequent query and is represented by a respective sequence of hypothesized terms (e.g., words). Here, the ASR system generates a respective likelihood score for each of the candidate hypotheses indicating the probability that the candidate hypothesis is the correct transcription of the subsequent query spoken by the user. Embodiments herein are particularly directed to determining a similarity metric between the initial query and each candidate hypothesis (or among the N-best list of candidate hypotheses having the highest likelihood score among the plurality of candidate hypotheses) and then using these similarity metrics to bias the selection of the candidate hypothesis related to the topic most similar to the topic of the initial query. As will be described in more detail below, the similarity metric determined between the initial query and each candidate hypothesis may be based on respective query-based salient term (QBST) vectors generated for each of the initial query and each candidate hypothesis.
[0026] The user can explicitly opt-in to enable the functionality of the subsequent query mode and can opt-out to disable the functionality of the subsequent query mode. When the subsequent query mode is disabled, the second processor returns to the sleep state after the initial query is processed and wakes up only when called by the always-on DSP in response to the detection of a hot word to process subsequent speech. The embodiments described herein that use a similarity metric based on the QBST vector are equally applicable to scenarios where a subsequent query is spoken by the user when the subsequent query mode is disabled, or where a subsequent query is spoken after the subsequent query mode is disabled and the second processor has returned to the sleep state. Thus, a subsequent query can correspond to any query submitted after the initial query such that the user may speak a predetermined hot word or perform some other action (e.g., gesture, line of sight, raising the user device, etc.) to wake up and activate the second processor to process the subsequent query.
[0027] Referring to FIGS. 1A and 1B, in some embodiments, an exemplary system 100 is associated with one or more users 10 and includes an AED (i.e., user device) 102 that optionally communicates with a remote system 110 via a network 104. The AED 102 may correspond to a computing device such as a mobile phone, computer (laptop or desktop), tablet, smart speaker / display, smart appliance, smart headphones, wearable, vehicle infotainment system, etc., and includes data processing hardware 103 and memory hardware 105. The AED 102 includes or communicates with one or more microphones 106 for capturing streaming audio 118 in an environment of the AED 102 that may include utterances 119, 129 spoken by respective users 10. The remote system 110 may be a single computer, multiple computers, or a distributed system (e.g., cloud environment) having scalable / elastic computing resources 112 (e.g., data processing hardware) and / or storage resources 114 (e.g., memory hardware).
[0028] The data processing hardware 103 of the AED 102 includes a first processor 150 and a second processor 160. The first processor 150 may include an always-on DSP 150 (also referred to as DSP 150) configured to detect the presence of one or more hotwords 116 in the streaming audio 118 without performing semantic analysis or speech recognition processing on the streaming audio 118. The DSP 150 may receive the streaming audio 118 including the acoustic features extracted by the acoustic feature extractor from the utterance 119 spoken by the user 10. The DSP (i.e., the first processor) 150, and more generally the AED 102, may operate in a hotword detection mode (FIG. 1A) 152 and a subsequent query mode (i.e., subsequent mode) (FIG. 1B) 154. In some examples, due to storage / memory / processing constraints, the DSP 150 operates in only one of the hotword detection mode or the subsequent query mode at a time. That is, in these examples, either the hotword detection mode 152 is enabled and the subsequent mode 154 is disabled, or the hotword detection mode 152 is disabled and the subsequent mode 154 is enabled. However, in other examples, the operation of the DSP 150 in the hotword detection mode 152 remains enabled and the DSP 150 operates simultaneously in the subsequent query mode 154. As will become apparent, the always-on processor 150 may operate in the hotword detection mode 152 while the AED 102 is in a low power state and the second processor 160 is in a sleep state. While the always-on processor 150 is in the subsequent query mode 154, the second processor 160 may be in a wake / sleep state / active to perform speech recognition on the subsequent audio data 127 characterizing the subsequent query. Alternatively, the second processor 160 may be in a sleep state while the always-on processor 150 operates in the subsequent query mode 152, and as a result, the always-on processor 150 may perform voice activity detection (VAD) on the subsequent audio data 127 and resume the second processor 160 from the sleep state only when the VAD determines that the subsequent audio data 127 includes human speech, and perform speech recognition on the subsequent audio data 127.
[0029] Referring now to FIG. 1A, in some embodiments, the always-on DSP 150 operates in a hot-word detection mode 152 while the subsequent query mode 154 is disabled and the second processor 160 is asleep. During operation in the hot-word detection mode 152, the DSP 150 is configured to detect the presence of the hot-word 116 "OK Google" in the streaming audio 118 and initiate a wake-up process on the second processor 160 to process the hot-word 116 and / or an initial query 117 following the hot-word 116 in the streaming audio 118. In the illustrated example, the utterance 119 includes the hot-word 116, "OK Google," and the subsequent initial query 117, "What's the temperature in Detroit?" The AED 102 extracts acoustic features from the streaming audio 118 and stores the extracted acoustic features in a buffer of the memory hardware 105 for use in detecting whether the streaming audio 118 contains the presence of the hot-word 116. The DSP 150 executes a hot-word detection model configured to generate a probability score indicative of the presence of the hot-word 116 in the acoustic features of the streaming audio 118 captured by the AED 102, and can detect the hot-word 116 in the streaming audio 118 when the probability score meets a hot-word detection threshold. The DSP 150 may include a plurality of hot-word detection models, each trained to detect different hot-words associated with specific terms / phrases. These hot-words may be pre-defined hot-words and / or custom hot-words assigned by the user 10. In some embodiments, the hot-word detection mode includes a trained neural-network-based model received from a remote system 110 via the network 104.
[0030] In the illustrated example, when the DSP 150 detects an acoustic feature that is a feature of the hotword 116 within the streaming audio 118, the DSP 150 can determine that the utterance 119 "Okay Google, what's the temperature in Detroit?" includes the hotword 116 "Okay Google". For example, the DSP 150 generates MFCCs from the audio data and, based on classifying that the MFCCs include MFCCs similar to the MFCCs that are features of the hotword "Okay Google" stored in the hotword detection model, can detect that the utterance 119 "Okay Google, what's the temperature in Detroit?" includes the hotword 116 "Okay Google". As another example, the DSP 150 generates mel-scale filter bank energies from the audio data and, based on classifying that the mel-scale filter bank energies include mel-scale filter bank energies similar to the mel-scale filter bank energies that are features of the hotword "Okay Google" stored in the hotword detection model, can detect that the utterance 119 "Okay Google, what's the temperature in Detroit?" includes the hotword 116 "Okay Google".
[0031] In response to detecting the presence of the hot word 116 in the streaming audio 118 corresponding to the utterance 119 spoken by the user 10, the DSP 150 provides the audio data 120 characterizing the hot word event to start a wake-up process on the second processor 160 to confirm the presence of the hot word 116. The audio data 120 (also referred to as the hot word event) includes a first portion 121 characterizing the hot word 116 and a second portion 122 characterizing the initial query 117. Here, the second processor 160 can execute a more robust hot word detection model to confirm whether the audio data 120 contains the hot word 116. Additionally or alternatively, the second processor 160 can perform speech recognition on the audio data 120 via the ASR model 210 of the automatic speech recognition (ASR) system 200 to confirm whether the audio data 120 contains the hot word.
[0032] When the second processor 160 confirms that the audio data 120 includes the hot word 116, the second processor 160 executes the ASR model 210 to process the audio data 120, thereby generating an audio recognition result 215, and also executes the natural language understanding (NLU) module 220 of the ASR system 200 to perform semantic interpretation on the audio recognition result 215, thereby determining that the audio data 120 includes an initial query 117 for the digital assistant 109 to perform an operation. In this example, the ASR model 210 processes the audio data 120 to generate an audio recognition result 215 of "What is the temperature in Detroit?", and the NLU module 220 identifies this as the initial query 117 for the digital assistant 109 and can perform an operation to fetch a response 192 (i.e., a response to the initial query 117) indicating that "The temperature in Detroit is 76 degrees." The digital assistant 109 may provide the response 192 to the output from the AED 102. For example, the digital assistant 109 may audibly output the response 192 from the AED 102 as synthesized speech and / or display a text representation of the response 192 on the screen of the AED 102.
[0033] In some embodiments, an ASR system 200 including an ASR model 210 and an NLU module 220 is disposed on a remote system 110 in addition to, or instead of, the AED 102. When the DSP 150 triggers and wakes up the second processor 160 in response to the detection of the hot word 116 during the utterance 119, the second processor can transmit the audio data 120 corresponding to the utterance 119 to the remote system 110 via the network 104. The AED 102 can transmit the first portion 121 of the audio data 120 including the hot word 116 to the remote system 110 to confirm the presence of the hot word 116 for performing speech recognition via the ASR model 210. Alternatively, the AED 102 may transmit only the second portion 122 of the audio data 120 corresponding to the initial query 117 spoken in the utterance 119 after the hot word 116 to the remote system 110. The remote system 110 executes the ASR model 210 to generate the speech recognition result 215 of the audio data 120. The remote system 110 also executes the NLU module 220 to perform semantic interpretation on the speech recognition result 215 and identify the initial query 117 for the digital assistant 109 to perform an operation. Alternatively, the remote system 110 can transmit the speech recognition result 215 to the AED 102, and the AED 102 can execute the NLU module 220 to identify the initial query 117.
[0034] The digital assistant 109 may be disposed in the remote system 110 and / or the AED 102. The digital assistant 109 is configured to execute an operation specified by the initial query 117 from the second processor 160. In some examples, the digital assistant 109 accesses a search engine to fetch a response 192 associated with the initial query 117. In other embodiments, the digital assistant 109 accesses the memory hardware 105 of the AED 102 and / or the memory hardware 114 of the remote system to fetch a response 192 related to the initial query 117. Alternatively, the digital assistant 109 executes an operation related to the initial query 117 (i.e., "call mother").
[0035] In some embodiments, user 10 has a follow-up query based on response 192 to initial query 117. That is, after receiving response 192 that "the temperature in Detroit is 76 degrees", user 10 may have a follow-up query to inquire about additional weather conditions in Detroit. In current techniques, by keeping the second processor 160 powered on for a predetermined time before returning to the sleep state and processing all audio data following initial query 117, the user can provide a follow-up query without having to speak the hot word again. That is, the ASR system 200 executed on the second processor 160 and / or the remote system 100 can continuously perform speech recognition and / or semantic interpretation on all subsequent audio data captured by the user device 102 during the follow-up query model 154 to determine whether the user 10 has submitted a follow-up query for the digital assistant 109 to perform an operation. As will be described in more detail below with reference to FIGS. 1B and 2, the ASR system 200 uses the initial query 117 (and corresponding transcription 215) to determine a similarity metric between the initial query 117 and each of the plurality of candidate hypotheses 225 (FIGS. 1B and 2) output by the ASR model 210 in response to the processing of subsequent audio data 127 (FIGS. 1B and 2) captured by the user device 102 during the follow-up query mode 154. The candidate hypothesis 225 associated with the highest similarity metric indicates the candidate hypothesis associated with the topic most similar to the topic associated with the initial query, thereby enabling the ASR system 200 to select this candidate hypothesis 225 as the final transcription 175 of the follow-up query 129.
[0036] Referring now to FIG. 1B, in response to receiving response 192 to initial query 117 submitted by user 10, digital assistant 109 instructs DSP 150 (and optionally second processor 160) to operate in subsequent query mode 154. Here, instructing DSP 150 to operate in subsequent query mode 154 may disable hot word detection mode 152. First processor 150 and / or second processor 160 receive subsequent audio data 127 corresponding to subsequent query 129 spoken by user 10 and captured by AED 102 while operating in subsequent mode 154. In the example shown, user 10 speaks subsequent query 129 "Is it sunny?" in response to AED 102, and response 192 "The temperature in Detroit is 76 degrees" is provided to user 10. In particular, user 10 simply speaks subsequent query 129 without speaking hot word 116 again, as user 10 spoke when user 10 spoke initial utterance 119 of FIG. 1A. However, in some other examples, subsequent query 129 includes hot word 116.
[0037] In some examples, during operation in the subsequent query mode 154, the DSP 152 or the second processor 160 performs voice activity detection (VAD) trained to determine whether voice activity is present in the subsequent audio data 127. That is, the VAD model determines whether the subsequent audio data 127 contains voice activity such as voice spoken by a human or non-voice activity audio (i.e., stereo, speaker, background noise, etc.). The VAD model can be trained to output a voice activity score indicating the likelihood that the subsequent audio data 127 contains voice activity. Here, the first processor 150 or the second processor 160 may determine that the subsequent audio data 127 contains voice activity if the voice activity score meets a voice activity threshold. In some examples, the VAD model outputs a binary voice activity indication, where "1" indicates "yes" and "0" indicates "no", indicating whether the subsequent voice data 127 contains ( "yes") or does not contain ("no") voice activity. The VAD model can be trained to distinguish human voice from synthetic / synthesized voice. By detecting voice activity, the first processor 150 and / or the second processor 160 can call the ASR system 200 to perform speech recognition and semantic analysis on the subsequent audio data 127 to determine whether the subsequent audio data 127 contains a subsequent query 129 that specifies subsequent operations to be performed by the digital assistant 109.
[0038] Therefore, the ASR model 210 may process the subsequent audio data 127 to generate a plurality of candidate hypotheses 225 for the subsequent audio data 127. The ASR system 200 can also execute the NLU module 220 to perform semantic interpretation on one or more of the candidate hypotheses 225 to identify a subsequent query 129 for submission to the digital assistant 109. In the illustrated example, the ASR system 200 determines that the subsequent audio data 127 corresponds to a subsequent query 129 for "Is it sunny?", and the assistant 109 searches for a subsequent response 193 to the subsequent query 129. Here, the second processor 160 provides the subsequent response 193 from the assistant 109 as an output from the AED 102 in the form of synthesized voice and / or text, indicating that "It is currently cloudy and the clouds will clear in the afternoon."
[0039] Alternatively, when finally processing the subsequent audio data 127, the ASR system 200 can determine that the subsequent audio data 127 is not directed to the AED 102 and thus there is no subsequent query 129 for the digital assistant 109 to execute. In this scenario, the second processor 160 provides an indication that there is no subsequent query for the subsequent audio data 127 and can return to the sleep state or remain in the wake state for a predetermined time. In some examples, the second processor 160 prompts the user 10 to repeat the subsequent query 129 in response to an indication that there is no subsequent query 129 for the subsequent audio data 127.
[0040] As described above with reference to FIG. 1A, the ASR model 210 and the NLU module 220 can be disposed on the remote system 110 in addition to, or instead of, the AED 102. Here, the AED 102 can transmit subsequent audio data 127 to the remote system 110 via the network 104, whereby the remote system 110 can execute the ASR model 210 to generate candidate hypotheses 225 for the audio data 120. The remote system 110 can also execute the NLU module 220 to perform semantic interpretation on one or more of the candidate hypotheses 225 and identify subsequent queries 129 for the digital assistant 109 to execute operations. Alternatively, the remote system 110 can transmit one or more candidate hypotheses 225 to the AED 102, and the AED 102 can execute the NLU module 220 to identify subsequent queries 129.
[0041] FIG. 2 shows an exemplary ASR system 200 that biases the selection of a plurality of candidate hypotheses 225, 225a - 225n output by an ASR model 210 for subsequent audio data 127 based on the topic similarity of one or more preceding queries 117. In the examples herein, the preceding query 117 is referred to as the first query in a conversation with the digital assistant 109, but the topic similarity related to any one or more queries before the current subsequent query 127 can be utilized to select the candidate hypothesis 225 as the transcription of the subsequent query 127. The ASR system 200 includes an ASR model 210, a query - based salient term (QBST) module 230, a similarity scorer 240, and a reranker 250. FIG. 2 shows operations (A)-(E) that illustrate the data flow. As described herein, the AED 102 or the remote system 110 performs operations (A)-(E). In some examples, the AED 102 performs all operations (A)-(E) entirely on - device, while in other examples, the remote system 110 performs all operations (A)-(E). In some additional examples, the AED 102 performs the first part of the operation and the remote system performs the second part of the operation.
[0042] In the illustrated example, during operation (A), the ASR system 200 receives subsequent audio data 127 captured by the microphone(s) 106 of the AED 120, and the ASR model 210 processes the subsequent audio data 127 to determine / generate a plurality of candidate hypotheses 225 for a subsequent query 129. Here, each candidate hypothesis 225 corresponds to a candidate transcription of the subsequent query 129 and is represented by a respective sequence of hypothesized terms. For example, the data - processing hardware 105 of the AED 102 can execute the ASR model 210 that generates a word lattice 300 indicating a plurality of candidate hypothesis transcripts 225 possible for the subsequent query 129 based on the subsequent audio data 127. The ASR model 210 can evaluate the potential paths through the word lattice 300 to determine the plurality of candidate hypotheses 225.
[0043] FIG. 3A shows an example of word lattices 300, 300a that may be provided by the ASR model 210 of FIG. 2. The word lattice 300a represents a plurality of possible combinations of words that can form different candidate hypotheses 225 for a subsequent query 129 while the AED 102 operates in a continuous conversation model (i.e., subsequent query mode).
[0044] The word lattice 300a includes one or more nodes 202a-202g corresponding to possible boundaries between words. The word lattice 300a includes a plurality of edges 204a-204l of possible words of candidate hypotheses resulting from the word lattice 300a. Further, each of the edges 304a-304l can have one or more weights or probabilities that the edge is the correct edge from the node to which the edge corresponds. The weights are determined by the ASR model 210 and can be based, for example, on the reliability of the match between the subsequent audio data 127 and the word of that edge, and how well the word fits grammatically and / or lexically with other words of the word lattice 300a.
[0045] For example, initially, the most likely path through the word lattice 300a (e.g., the most likely candidate hypothesis 225) may include edges 204c, 204e, 204i, 204k and have the text "we’re coming about 11:30". The second most optimal path through the word lattice 300a (e.g., the second most optimal candidate hypothesis 225) may include edges 204d, 204h, 204j, 204l and have the text "deer hunting scouts 7:30".
[0046] Each pair of nodes may have one or more paths corresponding to alternative words within various candidate hypotheses 225. For example, the initial most likely path between a pair of nodes starting at node 202a and ending at node 202c is edge 204c “we’re (we are)”. This path has alternative paths including edges 204a, 204b “we are” and edge 204d “deer”.
[0047] FIG. 3B is an example of hierarchical word lattices 300, 300b that may be provided by the ASR model 210 of FIG. 2. Word lattice 300b includes nodes 252a - 252l representing words that make up various candidate hypotheses 225 for a subsequent query 129 while the AED 102 operates in a continuous conversation model (i.e., subsequent query mode). The edges between nodes 252a - 252l indicate that the possible candidate hypotheses 225 include: (1) nodes 252c, 252e, 252i, 252k “we’re coming about 11:30 (we are coming about 11:30)”, (2) nodes 252a, 252b, 252e, 252i, 252k “we are cominng about 11:30 (we are coming about 11:30)”, (3) nodes 252a, 252b, 252f, 252g, 252i, 252k “we are come at about 11:30 (we are coming at about 11:30)”, (4) nodes 252d, 252f, 252g, 252i, 252k “deer come at about 11:30 (deer is coming at about 11:30)”, (5) nodes 252d, 252h, 252j, 252k “deer hunting scouts 11:30”, (6) nodes 252d, 252h, 252j, 252l “deer hunting scouts 7:30”.
[0048] Here too, the edges between nodes 252a - 252l can have associated weights or probabilities based on the reliability of speech recognition (e.g., candidate hypotheses) and the resulting grammar / vocabulary analysis of the text. In this example, "we’re coming about 11:30" might be the best hypothesis at the moment, and "deer hunting scouts 7:30" might be the next best hypothesis. One or more segments 254a - 254d that group words and their alternatives together can be created within the word lattice 300b. For example, segment 254a includes the words "we’re" and the alternatives "weare" and "deer". Segment 252b includes the word "comming" and the alternatives "come at" and "hunting". Segment 254c includes the word "about" and the alternative "scouts", and segment 254d includes the word "11:30" and the alternative "7:30".
[0049] Referring back to FIG. 2, the ASR model 210 may generate a plurality of candidate hypotheses 225 from the word lattice 300. That is, for each candidate hypothesis 225 in the word lattice 300, the ASR model 210 generates a likelihood score 155. Each likelihood score 155 indicates the probability that the candidate hypothesis 225 is correct (e.g., matches the utterance of the subsequent query 129 spoken by the user 10). Each likelihood score 155 may include a combination of an acoustic modeling score and / or a prior likelihood score. In some embodiments, the ASR model 210 includes an end-to-end (E2E) speech recognition model configured to receive subsequent audio data 127 and generate the word lattice 300. In particular, the E2E speech recognition model processes the subsequent audio data 127 to generate, for each of the plurality of candidate hypotheses 225 from the word lattice 300, a corresponding likelihood score 155. In these embodiments, the ASR model 210 may utilize an integrated language model or an external language model to generate the likelihood scores 155. In some examples, the ASR model 210 includes a separate acoustic model, language model, and / or pronunciation model. The ASR model 210 may share the acoustic model and the language model as additional hypothesis scorers (e.g., the acoustic model and the language model), or may have separate acoustic and language models. As used herein, the ASR model 210 may be interchangeably referred to as a "speech recognizer" that includes a single E2E neural network or a combination of separate acoustic models, language models, and / or pronunciation models.
[0050] During operation (B), the ASR system 200 identifies a set of the highest-ranked candidate hypotheses 225 from among the plurality of candidate hypotheses 225 within the word lattice 300. For example, using the likelihood scores 155 from the ASR system 210, the ASR system 200 selects n candidate hypotheses 225 having the highest likelihood scores 155, where n is an integer. In some examples, the ASR system 200 selects candidate hypotheses 225 having likelihood scores 155 that meet a likelihood score threshold. Optionally, the ASR model 210 may rank the set of the highest-ranked candidate hypotheses 225 using the likelihood scores 155.
[0051] Continuing with the example of FIGS. 1A and 1B, the ASR model 210 generates candidate hypotheses 225 for the subsequent query "Is it sunny?" spoken by the user 10. In this example, the top two candidate transcriptions (e.g., the two most likely ones) are selected as the set of the highest-ranked candidate hypotheses 225. For simplicity, in step B, only the top two candidate hypotheses 225 are shown, and thus the set of the top candidate hypotheses 225 can include a set of two or more top candidate hypotheses 225 from the word lattice 300. The top candidate hypotheses 225 include a first candidate hypothesis 225 "Is it funny" with a likelihood score 155 of 0.55 and a second candidate hypothesis 135 "Is it sunny" with a likelihood score 155 of 0.45. Here, the higher the likelihood score 155, the higher the confidence that the candidate hypothesis 135 is correct. In particular, the first candidate hypothesis ("Is it funny") 225 with the highest likelihood score 155 does not include the utterance of the subsequent query 129 actually spoken by the user 10. Thus, if the ASR system 200 selects the first candidate hypothesis 135 as the final transcription 275, the NLU 220 (FIGS. 1A and 1B) performs semantic interpretation on the transcription 275 and misidentifies "Is it funny" as an utterance not directed to the AED 102 (since there is no context as to what / who the term "that" corresponds to), or as the subsequent query 129 for submission to the digital assistant 109. In the former case, the digital assistant 109 may either ignore the subsequent query or prompt the user to resubmit the subsequent query. In the latter case, if the incorrect transcription of "Is it funny" is identified as the subsequent query 129 submitted to the digital assistant 109, the digital assistant 109 searches for a subsequent response 193 not related to the subsequent query 129 actually spoken by the user 10. In either scenario, the digital assistant 109 cannot return a subsequent response 193 that answers the subsequent query 129 actually spoken by the user 10, resulting in a significant degradation of the user experience.
[0052] The ASR system 200 may perform operations (C) and (D) as additional processing steps. These processing steps use the initial / prior query 117 to bias the selection of candidate hypotheses 225 from a plurality of candidate hypotheses 225 based on the topic similarity to the initial / prior query 117. In operation (C), the QBST module 230 receives a set of top candidate hypotheses 225, 225a - 225n, and determines corresponding sets of candidate QBST vectors 235, 235a - 235n, each associated with a corresponding one of the set of top candidate hypotheses 225. The QBST module 230 similarly receives the prior query 117 (e.g., the initial query 117 such as "What is the temperature in Detroit?"), and determines the corresponding prior QBST vector 217. If the current subsequent query during processing corresponds to the second or subsequent subsequent query in an ongoing conversation between the user 10 and the digital assistant 109, the QBST 230 may optionally receive a plurality of prior queries 117 from the ongoing conversation and, for each of the plurality of prior queries 117, determine the corresponding prior QBST vector 217. For simplicity, in the present disclosure, only the QBST 230 that receives only a single prior query 117 is described.
[0053] Each QBST vector 217, 225 may represent a set of prominent terms, each having a corresponding weight. For example, the prior QBST vector 217 corresponding to the prior query 117 ("What is the temperature in Detroit?") can represent a corresponding set of prominent terms including, but not limited to, "temperature", "weather", "forecast", "warm", "cool", etc. On the other hand, the first candidate QBST vector 235a corresponding to the top-ranked but incorrect first candidate hypothesis 225 of "Is it interesting?" represents a corresponding set of prominent terms including "interesting", "joke", "comedy", etc., while the second candidate QBST vector 235b corresponding to the correct second candidate hypothesis of "Is it sunny?" can represent a corresponding set of prominent terms including "sun", "clear sky", "weather", "forecast", etc., which share more prominent terms with the prior query 117.
[0054] In operation (D), the similarity score calculator 240 receives the prior QBST vector and the candidate QBST vectors 217, 235, and for each candidate hypothesis 225, 225a - 225n, determines the corresponding similarity metric 245, 245a - 245n between the corresponding prior QBST vector 217 output from the QBST module 230 for the corresponding candidate hypothesis 225 and the corresponding candidate QBST vector 235. Put another way, the similarity score calculator 240 compares the prior QBST vector 217 with each of the candidate QBST vectors 235 to determine a similarity metric 245 that indicates the topic drift between the prior QBST vector 217 and the corresponding candidate QBST vector 235. Each similarity metric 245 may include a similarity score between the prior QBST vector 217 and the corresponding candidate QBST vector 235. In some examples, each similarity metric 245 includes a cosine score / distance between the prior QBST vector 217 and the corresponding candidate QBST vector 235 to account for the term weights of the pair of QBST vectors 217, 235. For example, continuing with the above example, the similarity metric 245 that includes the cosine score between the prior QBST vector 217 and each of the first and second candidate QBST vectors 235a, 235b may be represented as follows. Similarity (What is the temperature in Detroit, interesting?)=0.001 Similarity (What is the temperature in Detroit, sunny?)=0.793
[0055] Accordingly, for each of the top-ranked candidate hypotheses 225 identified during operation (B) to select the final transcription 275 to represent the subsequent query 129 that will ultimately be processed by the digital assistant 109, the ASR system 200 can use the similarity metric 245 (e.g., cosine score) output by the similarity scorer 240. For example, in the above example, the second candidate hypothesis 225b, "Is it sunny?", is associated with a higher similarity score 245, even though its likelihood score 155 is lower than that of the top-ranked second candidate hypothesis 225a, "Is it interesting?", and is thus more similar to the overall topic of the preceding query 117 in the conversation with the digital assistant 109. That is, by the intuition that the current subsequent query relates to the same overall topic as the preceding query, the QBST module 230 and the similarity scorer 240 enable the ASR system 200 to select one of the candidate hypotheses 225 that is most similar to the preceding query 117 as the correct final transcription 275 for the subsequent query 129 actually spoken by the user 10.
[0056] The similarity score calculator 240 may be adapted to consider multiple different similarities between the prior query 117 and the plurality of candidate hypotheses 225. In one consideration of similarity, as detailed above, the similarity score calculator 240 determines a similarity metric 245 between the prior query 117 and each of the N alternative candidate hypotheses 225 such that the candidate hypotheses can be sorted / ranked based on their similarity to the prior query. In an additional consideration of similarity, the similarity score calculator 240 determines a similarity metric 245 between the candidate hypotheses 225 to enable pruning of some of the candidate hypotheses 225 that have little meaning different from the prior query 117. For example, if the top candidate hypothesis 225 is "it is sunny" and one of the other alternative candidate hypotheses 225 is "is it sunny", that alternative is close enough to the top candidate hypothesis 225 that it would search for the same subsequent response 193, so it can be pruned and removed from consideration. In this scenario, two of the candidate hypotheses 225 can be selected for processing by the NLU model 220 and ultimately determine which one results in the optimal search for the subsequent response 193. In yet another additional consideration of similarity, the similarity score calculator 240 uses the following formula to determine how close an alternative candidate hypothesis 225 is in similarity to the prior query 118 compared to the original (e.g., top-ranked) candidate hypothesis 225. Similarity(prior query, alternative candidate hypothesis) - Similarity(prior query, top-ranked candidate hypothesis) Here, the formula estimates an improvement that enables the ASR system 200 to attempt to select an alternative candidate hypothesis 225 that has a different meaning from the highest-ranked candidate hypothesis but not a significantly different meaning from the previous query 118. For example, considering a previous query such as "temperature in New York" and candidate hypotheses for subsequent queries including "is there sun in Newark" and "is there fun in New York", by sharing prominent terms such as "sun", "weather", "temperature", the candidate hypothesis "is there sun in Newark" may be selected. However, in the alternative candidate, "New York" is replaced by "Newark", which reduces the likelihood of it being the correct transcription. Thus, the improvement judged by topic similarity is weighted accordingly, so that a less likely alternative candidate hypothesis is not selected even though it is on the same topic as the previous query 117.
[0057] During operation (E), the re-ranker 250 receives a similarity metric 245 for the candidate hypothesis 225 and a similarity score / metric 225 from the similarity score l240, and receives a likelihood score 155 for the candidate hypothesis 225 from the ASR model 210. The re-ranker 250 is configured to output a re-rank result 255 that includes ranking of multiple candidate hypotheses 225 based on the biased likelihood score 156. In the illustrated example, the re-ranked result 255 includes a candidate hypothesis 225 regarding "is it sunny" with a biased likelihood score 156 of 0.8 as the most likely correct transcription 275. The NLU model 220 can perform semantic analysis on the correct transcription 275 to identify an appropriate subsequent query 129 for the digital assistant 109 to execute.
[0058] Additionally or alternatively, the QBST module 230 may determine a QBST vector for the response 192 to the initial query 117 submitted by the user, and determine a corresponding similarity metric between each candidate QBST vector 235 and the QBST vector determined for the response 192 to the initial query 117. Thus, the selection of candidate hypotheses for the final transcription 275 of the subsequent query 129 may be biased based on the topic similarity between the response 192 to the initial query 117 and each of the candidate hypotheses.
[0059] FIG. 4 is a flowchart of an exemplary arrangement of operations for a method 400 of transcribing subsequent audio data during operation of the assistant-responsive device 102 in a continuous conversation mode. The method 400 may include a computer-implemented method executed on data processing hardware 510 (FIG. 5). The data processing hardware 510 may include the data processing hardware 103 of the assistant-responsive device 102 of FIG. 1, or the data processing hardware 112 of the remote system 110. In operation 402, the method 400 includes receiving subsequent audio data 127 captured by the AED 102. Here, the subsequent audio data 127 may correspond to a subsequent query 129 spoken by the user 10 of the AED 102 to the digital assistant 109 after the user 10 submitted an initial query 117 to the digital assistant 109. The AED 102 may capture the subsequent audio data 127 while operating in the continuous conversation mode after the user 10 submitted the initial query 117. During operation in the continuous conversation mode, the AED 102 is in a wake state and instructs the speech recognizer 210 to perform speech recognition on the audio data captured by the AED 102 without requiring the user to explicitly speak a predetermined hot word.
[0060] In operation 404, method 400 includes processing subsequent audio data 127 to generate a plurality of candidate hypotheses 225 using speech recognizer 210. Each candidate hypothesis 225 corresponds to a candidate transcription of subsequent query 129 and is represented by a respective sequence of hypothesized terms.
[0061] In operation 406, for each corresponding candidate hypothesis 225 among the plurality of candidate hypotheses, method 400 also includes determining a similarity metric 245 between the preceding query 117 and the corresponding candidate hypothesis 225. Here, the similarity metric indicates the similarity between the topic associated with the corresponding candidate hypothesis and the topic associated with the preceding query. In some examples, the similarity metric includes the cosine distance between the QBST vector 217 associated with the preceding query 117 and the corresponding candidate QBST vector 235 associated with the corresponding candidate hypothesis 225. In operation 408, method 400 includes determining a transcription of the subsequent query 129 spoken by the user based on the similarity metrics 245 determined for the plurality of candidate hypotheses 225.
[0062] A software application (i.e., a software resource) may refer to computer software that causes a computing device to execute a task. In some examples, a software application may be referred to as an "application", "app", or "program". Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.
[0063] A non-transitory memory may be a physical device used to store programs (e.g., sequences of instructions) or data (e.g., program state information) temporarily or permanently for use by a computing device. The non-transitory memory may be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., used for firmware such as a boot program typically). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0064] FIG. 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems and methods described in this document. Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended only as examples and are not intended to limit the embodiments of the invention described and / or claimed in this document.
[0065] The computing device 500 includes a processor 510, a memory 520, a storage device 530, a high-speed interface / controller 540 connected to the memory 520 and the high-speed expansion port 550, and a low-speed interface / controller 560 connected to the low-speed bus 570 and the storage device 530. Each component 510, 520, 530, 540, 550, and 560 is interconnected using various buses and can be mounted on a common motherboard or exist in other ways as needed. The processor (i.e., data processing hardware) 510 processes instructions for execution within the computing device 500, including instructions stored in the memory 520 or the storage device 530, and can display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 580 connected to the high-speed interface 540. The data processing hardware 510 can include the data processing hardware 103 of the assistant-responsive device 102 in FIG. 1 or the data processing hardware 112 of the remote system 110. In other embodiments, multiple processors and / or multiple buses may be used, along with multiple memories and memory types, as needed. Also, multiple computing devices 500 may be connected, and each device may perform a part of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).
[0066] The memory 520 stores information non - temporarily within the computing device 500. The memory 520 may be a computer - readable medium, a volatile memory unit(s), or a non - volatile memory unit(s). The non - temporary memory 520 may be a physical device used to store a program (e.g., an instruction sequence) or data (e.g., program state information) temporarily or permanently for use by the computing device 500. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., used for firmware such as a boot program usually). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and disks or tapes.
[0067] The storage device 530 can provide mass storage for the computing device 500. In some embodiments, the storage device 530 is a computer - readable medium. In various different embodiments, the storage device 530 may be an array of devices including a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid - state memory device, or a storage area network or other configured device. In additional embodiments, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer - readable medium or a machine - readable medium such as the memory 520, the storage device 530, or the memory on the processor 510.
[0068] The high-speed controller 540 manages the bandwidth-intensive operations of the computing device 500, and the low-speed controller 560 manages the low-bandwidth-intensive operations. Such role assignments are merely examples. In some embodiments, the high-speed controller 540 is coupled to a high-speed expansion port 550 that can accept the memory 520, the display 580 (e.g., via a graphics processor or accelerator), and various expansion cards (not shown). In some embodiments, the low-speed controller 560 is coupled to the storage device 530 and the low-speed expansion port 590. The low-speed expansion port 590, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be connected to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a network device such as a switch or router, for example via a network adapter.
[0069] As shown in the figure, the computing device 500 may be implemented in many different forms. For example, the computing device may be implemented as a standard server 500a, or multiple times within a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0070] Various embodiments of the systems and techniques described herein can be realized in digital and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be included in one or more computer programs executable and / or interpretable in a programmable system including at least one programmable processor, at least one input device, and at least one output device, which are coupled to receive data and instructions from and to transmit data and instructions to a storage system.
[0071] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0072] The processes and logical flows described in this specification can be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to act on input data and generate output. The processing and logical flows can also be implemented by special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose processors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read only memory, a random access memory, or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes, or is operatively coupled to receive data from, or transfer data to, or both, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks. The processor and memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0073] To provide interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, or a touch screen, and a keyboard and a pointing device (such as a mouse or a trackball) through which the user can input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input from the user can be received in any form, such as acoustic input, voice input, or tactile input. Further, the computer can interact with the user by sending documents to the devices used by the user and receiving documents from the devices used by the user, for example, by sending a web page to the web browser of the user's client device in response to a request received from a web browser.
[0074] A number of embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the following claims.
Claims
1. A computer-implemented method (400) that, when executed by data processing hardware (510), causes the data processing hardware (510) to perform operations, the operations comprising: Receiving subsequent audio data (127) captured by an assistant-enabled device (102), the subsequent audio data (127) corresponding to a subsequent query (119) spoken by the user to the digital assistant (109) after the user of the assistant-enabled device (102) submitted a prior query (117) to the digital assistant (109); Processing the subsequent audio data (127) using a speech recognizer (210) to generate a plurality of candidate hypotheses (225), each candidate hypothesis (225) corresponding to a candidate transcription of the subsequent query (119) and represented by a respective sequence of hypothesized terms; Determining, for each corresponding candidate hypothesis (225) among the plurality of candidate hypotheses (225), a corresponding similarity metric (245) between the prior query (117) and the corresponding candidate hypothesis (225), the similarity metric (245) indicating a similarity between a topic associated with the corresponding candidate hypothesis (225) and a topic associated with the prior query (117); Determining a transcription of the subsequent query (119) spoken by the user based on the similarity metrics (245) determined for the plurality of candidate hypotheses (225); A computer-implemented method (400) comprising the above.
2. The operations further comprise: Determining a prior query-based salient term (QBSST) vector (217) associated with the prior query (117) submitted by the user; Determining, for each corresponding candidate hypothesis (225), a corresponding candidate QBSST vector (235) associated with the corresponding candidate hypothesis (225); Including the above. Determining the corresponding similarity metric (245) between the previous query (117) and the corresponding candidate hypothesis (225) includes determining the corresponding similarity metric (245) between the previous QBST vector (217) and the corresponding candidate QBST vector (235) based on the previous QBST vector (217) and the corresponding candidate QBST vector (235). The method (400) implemented on a computer according to claim 1.
3. The previous QBST vector (217) indicates each set of prominent terms associated with the previous query (117), Each corresponding candidate QBST vector (235) indicates each set of prominent terms associated with the corresponding candidate hypothesis (225). The method (400) implemented on a computer according to claim 2.
4. The corresponding similarity metric (245) indicates a topic drift between the previous QBST vector (217) and the corresponding candidate QBST vector (235). The method (400) implemented on a computer according to claim 2 or 3.
5. The corresponding similarity metric (245) includes a cosine score between the previous QBST vector (217) and the corresponding candidate QBST vector (235). The method (400) implemented on a computer according to claim 2 or 3.
6. The operation is Receiving initial audio data (127) captured by the assistant-corresponding device (102) while the assistant-corresponding device (102) is in a sleep state, where the initial audio data (127) includes a hot word (116) and the previous query (117) submitted by the user to the digital assistant (109). When the hot word (116) is detected by the assistant-corresponding device (102), the assistant-corresponding device (102) is restored from the sleep state, triggering the speech recognizer (210) to perform speech recognition on at least a portion of the initial audio data (127) including the previous query (117). Receiving, After the voice recognizer (210) performs voice recognition on at least the portion of the initial audio data (127), instruct the assistant-responsive device (102) to operate in a subsequent query mode; further comprising Receiving the subsequent audio data (127) by the assistant-responsive device (102) includes receiving the subsequent audio data during operation of the assistant-responsive device (102) in the subsequent query mode, and the presence of the hot word (116) is not in the subsequent audio data (127). The computer-implemented method (400) according to any one of claims 1 to 5. **Claim 7** The operation is Performing query interpretation on the transcription to identify the operation specified by the subsequent query (119); Instructing the digital assistant (109) to perform the operation specified by the subsequent query (119); Receiving a subsequent response (193) from the digital assistant (109) indicating execution of the operation specified by the subsequent query (119); The computer-implemented method (400) according to any one of claims 1 to 6, further comprising. **Claim 8** The operation further includes presenting the subsequent response (193) for output from the assistant-responsive device (102). The computer-implemented method (400) according to claim 7. **Claim 9** If the operation is For each corresponding candidate hypothesis (225) among the plurality of candidate hypotheses (225) generated by the voice recognizer (210), obtaining the corresponding likelihood score (155) assigned by the voice recognizer (210) to the corresponding candidate hypothesis (225); Ranking the plurality of candidate hypotheses (225) based on the corresponding likelihood scores (155) assigned to the plurality of candidate hypotheses (225) by the voice recognizer (210) and the corresponding similarity scores (245) determined for each corresponding candidate hypothesis (225) among the plurality of candidate hypotheses (225); further comprising Determining the transcription of the subsequent query (119) is based on ranking the plurality of candidate hypotheses (225), and is a method (400) implemented on a computer according to any one of claims 1 to 8.
10. The data processing hardware (510) is present in the assistant-compatible device (102) and executes the speech recognizer (210), and is a method (400) implemented on a computer according to any one of claims 1 to 9.
11. The data processing hardware (510) is present in the remote server (110), executes the speech recognizer (210), and the remote server (110) communicates with the assistant-compatible device (102) via the network (104), and is a method (400) implemented on a computer according to any one of claims 1 to 9.
12. The speech recognizer (210) includes an end-to-end speech recognition model, and is a method (400) implemented on a computer according to any one of claims 1 to 11.
13. The speech recognizer (210) includes an acoustic model and a language model, and is a method (400) implemented on a computer according to any one of claims 1 to 11.
14. The assistant-compatible device (102) communicates with one or more microphones (106) configured to capture the subsequent audio data (127) and the audio data (118) corresponding to the preceding query (117), and is a method (400) implemented on a computer according to any one of claims 1 to 13.
15. The assistant-compatible device (102) includes a battery-powered device, and is a method (400) implemented on a computer according to any one of claims 1 to 14.
16. Data processing hardware (510), Memory hardware (520) that communicates with the data processing hardware (510) and stores instructions, and when the instructions are executed on the data processing hardware (510), causes the data processing hardware (510) to Receiving subsequent audio data (127) captured by an assistant-enabled device (102), where the subsequent audio data (127) corresponds to a subsequent query (119) spoken by the user to the digital assistant (109) after the user of the assistant-enabled device (102) submitted a preceding query (117) to the digital assistant (109). Processing the subsequent audio data (127) using a speech recognizer (210) to generate a plurality of candidate hypotheses (225), where each candidate hypothesis (225) corresponds to a candidate transcription of the subsequent query (119) and is represented by a respective sequence of hypothesized terms. For each corresponding candidate hypothesis (225) among the plurality of candidate hypotheses (225), determining a corresponding similarity metric (245) between the preceding query (117) and the corresponding candidate hypothesis (225), where the similarity metric (245) indicates the similarity between the topic associated with the corresponding candidate hypothesis (225) and the topic associated with the preceding query (117). Determining a transcription of the subsequent query (119) spoken by the user based on the similarity metrics (245) determined for the plurality of candidate hypotheses (225). Memory hardware (520) for causing the performance of operations including the above. A system (100) comprising the above.
17. The operations further include Determining a preceding query-based salient term (QBST) vector (217) associated with the preceding query (117) submitted by the user. For each corresponding candidate hypothesis (225), determining a corresponding candidate QBST vector (235) associated with the corresponding candidate hypothesis (225). And Determining the corresponding similarity metric (245) between the preceding query (117) and the corresponding candidate hypothesis (225) includes determining the corresponding similarity metric (245) between the preceding QBST vector (217) and the corresponding candidate QBST vector (235) based on the preceding QBST vector (217) and the corresponding candidate QBST vector (235). The system (100) according to claim 16.
18. The prior QBST vector (217) represents each set of prominent terms associated with the prior query (117), and each corresponding candidate QBST vector (235) represents each set of prominent terms associated with the corresponding candidate hypothesis (225), for the system (100) according to claim 17.
19. The corresponding similarity metric (245) indicates a topic drift between the prior QBST vector (217) and the corresponding candidate QBST vector (235), for the system (100) according to claim 17 or 18.
20. The corresponding similarity metric (245) includes a cosine score between the prior QBST vector (217) and the corresponding candidate QBST vector (235), for the system (100) according to claim 17 or 18.
21. The operation is receiving initial audio data (127) captured by the assistant - corresponding device while the assistant - corresponding device is in a sleep state, the initial audio data (127) including a hotword (116) and the prior query (117) submitted by the user to the digital assistant (109), and when the hotword (116) is detected by the assistant - corresponding device, causing the assistant - corresponding device to resume from the sleep state, triggering the speech recognizer (210) to perform speech recognition on at least a portion of the initial audio data (127) including the prior query (117); after the speech recognizer (210) performs speech recognition on at least the portion of the initial audio data (127), instructing the assistant - corresponding device (102) to operate in a subsequent query mode; and further includes receiving the subsequent audio data (127) by the assistant - corresponding device includes receiving the subsequent audio data (127) during the operation of the assistant - corresponding device in the subsequent query mode, and the hotword (116) is not present in the subsequent audio data (127), for the system (100) according to any one of claims 16 - 20.
22. The operation is Performing query interpretation on the transcription to identify the operation specified by the subsequent query; Instructing the digital assistant (109) to perform the operation specified by the subsequent query; Receiving, from the digital assistant (109), a subsequent response indicating execution of the operation specified by the subsequent query; The system (100) according to any one of claims 16 to 21, further comprising: **Claim 23** The system (100) according to claim 22, wherein the operation further comprises presenting the subsequent response for output from the assistant-compatible device (102). **Claim 24** When the operation is For each corresponding candidate hypothesis (225) among the plurality of candidate hypotheses (225) generated by the speech recognizer (210), obtaining the corresponding likelihood score (155) assigned by the speech recognizer (210) to the corresponding candidate hypothesis (225); Ranking the plurality of candidate hypotheses (225) based on the corresponding likelihood scores (155) assigned by the speech recognizer (210) to the plurality of candidate hypotheses (225) and the corresponding similarity scores (245) determined for each corresponding candidate hypothesis (225) among the plurality of candidate hypotheses (225); Further comprising Determining the transcription of the subsequent query is based on the ranking of the plurality of candidate hypotheses (225), the system (100) according to any one of claims 16 to 23. **Claim 25** The system (100) according to any one of claims 16 to 24, wherein the data processing hardware (510) is present within the assistant-compatible device and executes the speech recognizer (210). **Claim 26** The system (100) according to any one of claims 16 to 24, wherein the data processing hardware (510) is present within a remote server (110), the remote server (110) executes the speech recognizer (210), and the remote server (110) communicates with the assistant-compatible device (102) via a network (104). **Claim 27** The system (100) according to any one of claims 16 to 26, wherein the speech recognizer (210) includes an end-to-end speech recognition model. **Claim 28** The speech recognizer (210) is the system (100) according to any one of claims 16 to 26, including an acoustic model and a language model.
29. The assistant-responsive device communicates with one or more microphones (106) configured to capture the subsequent audio data (127) and the audio data (128) corresponding to the previous query (117), the system (100) according to any one of claims 16 to 28.
30. The system (100) according to any one of claims 16 to 28, wherein the assistant-responsive device includes a battery-powered device.
Citation Information
Patent Citations
Sentence voice recognition equipment
JP1987019899A
Information sharing system and information sharing method
JP2013073408A
Method, apparatus, device and computer readable storage media for voice interaction
JP2021076818A
Information processor, information processing method, and program
JP2021149909A