Enabling natural dialog for automated assistants
By combining streaming speech recognition and natural language understanding models to fulfill rules, the next interaction state of a dialogue session is dynamically determined, solving the problems of low efficiency and resource waste in turn-based dialogue sessions, and realizing natural dialogue and efficient user interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2021-12-01
- Publication Date
- 2026-04-24
AI Technical Summary
Existing turn-based dialogue sessions are inefficient when processing multiple spoken words from users, resulting in significant waste of computing resources and an unnatural user experience. Automated assistants struggle to respond accurately to incomplete speech.
The system employs streaming automatic speech recognition and natural language understanding models to process audio data streams. By combining performance rules and audio characteristics, it dynamically determines the next interaction state of a dialogue session and continuously iterates to promote natural dialogue.
It improves the efficiency and user experience of conversations, reduces the waste of computing resources, reduces the failure of automated assistants and the repetitive input of users, and enables faster conversation termination.
Smart Images

Figure CN116368562B_ABST
Abstract
Description
Background Technology
[0001] Humans can engage in human-computer dialogue through interactive software applications referred to herein as “automated assistants” (also known as “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “personal voice assistants,” “conversational agents,” etc.). Automated assistants typically rely on a pipeline of components in interpreting and responding to spoken utterances. For example, an Automatic Speech Recognition (ASR) engine can process audio data corresponding to a user’s spoken utterances to generate ASR outputs, such as phonological assumptions about the spoken utterances (i.e., sequences of terms and / or other tokens). Furthermore, a Natural Language Understanding (NLU) engine can process ASR outputs (or touch / typed input) to generate NLU outputs, such as the user’s intent when providing spoken utterances (or touch / typed input) and, optionally, slot values of parameters associated with that intent. Additionally, an execution engine can be used to process NLU outputs and generate execution outputs, such as structured requests to obtain response content to spoken utterances and / or actions performed in response to spoken utterances.
[0002] Typically, conversations with automated assistants are initiated by the user providing spoken words, and the automated assistant uses the aforementioned component pipeline to respond to these spoken words and generate responses. The user can continue the conversation by providing additional spoken words, and the automated assistant can use the same component pipeline to respond to these additional spoken words and generate additional responses. In other words, these conversations are usually turn-based, as the user takes turns in the conversation to provide spoken words, and the automated assistant takes turns in the conversation to respond to these spoken words. However, from the user's perspective, these turn-based conversations may feel unnatural because they do not reflect how humans actually converse with each other.
[0003] For example, a first person may provide multiple different utterances to convey a single idea to a second person, and the second person may be able to consider each of the multiple different utterances when formulating a response to the first person. In some cases, the first person may pause for different amounts of time between these multiple different utterances. It is worth noting that the second person may not be able to formulate a response to the first person simply based on the first utterance among the multiple different utterances or on each of the multiple different utterances in isolation.
[0004] Similarly, in these turn-based conversational sessions, without considering the context of a given utterance relative to multiple different utterances, the automated assistant may not be able to fully formulate a response to a user's given utterance. As a result, these turn-based conversational sessions can be prolonged because the user attempts to convey his / her thoughts to the automated assistant with a single utterance during a single turn of these conversational sessions, thus wasting computational resources. Furthermore, if the user attempts to convey his / her thoughts to the automated assistant with multiple utterances during a single turn of these turn-based conversational sessions, the automated assistant may simply fail, also wasting computational resources. For example, when the user provides a long pause while attempting to formulate a utterance, the automated assistant may prematurely infer that the user has finished speaking, process incomplete utterances, and fail because it is determined (from processing) that the incomplete utterance does not convey a meaningful intent. Additionally, turn-based conversational sessions can prevent user utterances provided during the rendering of the assistant's response from being meaningfully processed. This can require the user to wait for the rendering of the assistant's response to complete before providing utterances, thus prolonging the conversational session. Summary of the Invention
[0005] The implementations described herein relate to enabling an automated assistant to perform natural dialogue with a user during a conversational session. Some implementations can use a streaming automatic speech recognition (ASR) model to process an audio data stream generated by multiple microphones of a user's client device to generate an ASR output stream. The audio data stream can capture one or more spoken utterances from the user, addressed to an automated assistant that is at least partially implemented at the client device. Furthermore, a natural language understanding (NLU) model can be used to process the ASR output to generate an NLU output stream. Additionally, one or more performance rules and / or one or more performance models can be used to process the NLU output to generate a performance data stream. Additionally, audio-based characteristics associated with one or more spoken utterances can be determined based on the processing of the audio data stream. Based on the current state of the NLU output stream, the performance data stream, and / or the audio-based characteristics, the next interaction state to be implemented during the conversational session can be determined. The next interaction state to be implemented during the conversational session can be one of the following: (i) enabling the performance output generated based on the performance data stream to be implemented, (ii) enabling the natural dialogue output to be audibly rendered for presentation to the user, or (iii) avoiding enabling any interaction to be implemented. This allows for the determination of the next interaction state to facilitate the conversation. Furthermore, the determination of the next interaction state can occur iteratively during the conversation (e.g., continuously at 10Hz, 20Hz, or other frequencies), where each iteration is based on the corresponding current state of the NLU output, fulfillment data, and audio-based characteristics—and does not wait for the completion of a user or automated assistant response. Therefore, by using the techniques described herein to determine the next interaction state during the conversation, the automated assistant can determine whether and how to implement the next interaction state to facilitate the conversation, rather than simply responding to the user after they have provided a verbal utterance, as is often the case in turn-based conversations.
[0006] For example, suppose a user engages in a conversation with an automated assistant and provides the spoken utterance “turn on the…theuhmmm…”. When the user provides the spoken utterance, an ASR output stream, an NLU output stream, and a performance data stream can be generated based on the audio data stream that captures the spoken utterance. It is worth noting that in this example, the NLU output stream may indicate that the user intends to control some software application (e.g., a music application, a video application, etc.) or some device (e.g., a client device, a smart appliance, a smart TV, a smart speaker, etc.), but the user has not yet identified the exact intent of “turn on”. Nevertheless, the performance data stream can be processed to generate a set of performance outputs. Furthermore, audio-based features associated with the spoken utterance can be generated based on the processed audio data stream and can include, for example, tone and rhythm of speech indicating that the user is unsure of the exact intent of “turn on,” the duration between “turn on” and “theuhmmm,” the duration since “the uhmmm,” and / or other audio-based features. Additionally, the NLU output stream and audio-based features can be processed to generate a set of natural dialogue outputs.
[0007] In this example, the current state of the NLU output stream, the performance data stream, and the audio-based features associated with the spoken utterance can be determined. For example, a classification machine learning (ML) model can be used to process the NLU output stream, performance data stream, and audio-based features to generate predictive metrics (e.g., binary values, probabilities, log-likelihoods, etc.). Each next interaction state can be associated with a corresponding predictive metric, allowing the automation assistant to determine whether (i) the performance output is performed, (ii) the natural dialogue output is audibly rendered for presentation to the user, or (iii) any interaction is avoided. In this case, if the automation assistant determines that (i) the performance output is performed, it can select and perform the performance output from multiple performance outputs (e.g., assistant commands associated with turning on the TV, turning on the lights, turning on music, etc.). Furthermore, if the automation assistant determines (ii) that the natural dialogue output should be audibly rendered for presentation to the user, the automation assistant can select a natural dialogue output from multiple natural dialogue outputs and render it audibly (e.g., “What would you like me to turn on?”, “Are you still there?”, etc.). Additionally, if the automation assistant determines (iii) that it should avoid any interaction being implemented, the automation assistant can continue processing the audio data stream.
[0008] In this example, we further assume that the current moment of the conversation corresponds to two seconds after the user finishes delivering the spoken phrase "turn on the...the uhmmm". For the current moment, the corresponding predictive metric associated with the next interaction state can instruct the automated assistant to avoid enabling any interaction. In other words, even if it appears the user has temporarily finished speaking, the automated assistant has not determined with sufficient confidence that any fulfilled output has been achieved and should provide the user with additional time to gather his / her thoughts and identify exactly what he / she wants to turn on. However, we further assume an additional five seconds pass. At this subsequent moment, audio-based features can indicate that seven seconds have elapsed since the user last spoke and can update the current state based at least on audio-based features. Therefore, for this subsequent moment, the corresponding predictive metric associated with the next interaction state can instruct the automated assistant to make the natural conversational output "What do you want me to turn on?" audibly rendered to present to the user to re-engage the conversation. In other words, even if the user finishes speaking and the automated assistant still hasn't determined that any fulfillment output has been achieved with sufficient confidence, the automated assistant can prompt the user to continue the conversation and loop back to what the user previously instructed him / her to do (e.g., turn something on). Furthermore, let's further assume the user provides the additional spoken phrase "oh, the television." At this further subsequent moment, the data stream can be updated to indicate the user's intention to turn on the television, and the current state can be updated based on the updated data stream. Therefore, for this other subsequent moment, the corresponding predictive metric associated with the next interaction state can indicate that the automated assistant should turn on the television (and optionally, the synthesized speech can be audibly rendered, indicating that the automated assistant will turn on the television). In other words, at this other subsequent moment, the automated assistant has determined that a fulfillment output with sufficient confidence has been achieved.
[0009] In some implementations, the NLU data stream can be processed by multiple agents to generate a performance data stream. A set of performance outputs can be generated based on the performance data stream, and performance outputs can be selected from the set of performance outputs based on a predicted NLU metric associated with the NLU data stream and / or a predicted performance metric associated with the performance data stream. In some implementations, only one performance output can be implemented as the next interaction state, while in other implementations, multiple performance outputs can be implemented as the next interaction state. As used herein, a “first-party” (1P) agent, device, and / or system refers to an agent, device, and / or system controlled by the same party that controls the automation assistant referenced herein. Conversely, a “third-party” (3P) agent, device, and / or system refers to an agent, device, and / or system controlled by a different party that controls the automation assistant referenced herein, but communicatively coupled to one or more 1P agents, devices, and / or systems.
[0010] In some versions of those implementations, the multiple agents include one or more 1P agents. Continuing the example above, at the current moment of the conversation (e.g., two seconds after the user completes the "the uhmmm" part of the spoken utterance "turn on the...the uhmmm..."), the execution data stream can be processed by the 1P music agent to generate execution output associated with an assistant command that, upon implementation, causes music to be turned on; the 1P video streaming service agent generates execution output associated with an assistant command that, upon implementation, causes the video streaming service to be turned on; the 1P smart device agent generates execution output associated with an assistant command that, upon implementation, causes one or more 1P smart devices to be controlled, and so on. In this example, the execution data stream can be transmitted to the 1P agents via an application programming interface (API). In additional or alternative versions of those implementations, the multiple agents include one or more 3P agents. The execution output generated by the 3P agents can be similar to the execution output generated by the 1P agents, but generated by the 3P agents. In this example, the fulfillment data stream can be transmitted to a 3P agent via an application programming interface (API) and over one or more networks, and the 3P agent can transmit the fulfillment output back to the client device. Each fulfillment output generated by multiple agents (e.g., a 1P agent and / or a 3P agent) can be aggregated into a fulfillment output set.
[0011] While the above examples are described in relation to the set of fulfillment outputs for auxiliary commands, it should be understood that this is for illustrative purposes and not restrictive. In some implementations, the set of fulfillment outputs may additionally or alternatively include instances of synthesized speech audio data, which include corresponding synthesized speech and can be audibly rendered to be presented to a user via one or more speakers of a client device. For example, suppose a user provides spoken utterances instead of actual speech during a conversational session where the user says "set a timer for 15 minutes." Further suppose that prediction metrics associated with the NLU data stream and / or prediction metrics associated with the fulfillment data stream indicate that the user spoke for 15 minutes or 50 minutes. In this example, the fulfillment data stream can be processed by the 1P timer agent to generate a first fulfillment output associated with a first assistant command to set the timer to 15 minutes, a second fulfillment output associated with a second assistant command to set the timer to 50 minutes, a third fulfillment output associated with a first instance of synthesized speech audio data confirming that the assistant command is to be executed (e.g., "The timer is set for 15 minutes"), and a fourth fulfillment output associated with a first instance of synthesized speech audio data requesting the user to clear up any ambiguity in the assistant command to be executed (e.g., "Is that timer for 15 minutes or 50 minutes?").
[0012] In this example, the automation assistant is able to select one or more performance outputs from a set of performance outputs to be implemented based on a predicted NLU metric associated with the NLU data stream and / or a predicted performance metric associated with the performance data stream. For example, suppose the predicted metrics satisfy both a first threshold metric and a second threshold metric, indicating a very high confidence level for the user to say "set a timer for 15 minutes." In this example, the automation assistant is able to implement a first performance output associated with the first assistant command to set the timer for 15 minutes, without causing any instance of synthesized speech to be audibly rendered to be presented to the user. Conversely, suppose the predicted metrics satisfy the first threshold metric but not the second threshold metric, indicating a low confidence level for the user to say "set a timer for 15 minutes." In this example, the automation assistant enables the execution of a first fulfillment output associated with a first assistant command to set a timer for 15 minutes, and also enables the execution of a third fulfillment output associated with a first instance of synthesized speech audio data confirming that the assistant command should be executed (e.g., "The timer is set for 15 minutes"). This provides the user with an opportunity to correct the automation assistant in case of inaccuracies. However, suppose the prediction metric fails to meet both the first and second threshold metrics, indicating a low confidence level for the user to say "set a timer for 15 minutes." In this example, the automation assistant enables the execution of a fourth fulfillment output associated with a first instance of synthesized speech audio data requesting the user to dispel ambiguity in the assistant command (e.g., "Is that timer for 15 minutes or 50 minutes?"), and the automation assistant may avoid setting any timers due to the low confidence level.
[0013] In some implementations, the set of performance outputs may additionally or alternatively include instances of graphical content that can be visually rendered to be presented to a user via a display of a client device or an additional client device communicating with the client device. Continuing with the timer example above, the performance data stream can be processed by an 1P timer agent to additionally or alternatively generate a fifth performance output associated with a first graphical content depicting a timer set to 15 minutes, and a sixth performance output associated with a second graphical content depicting a timer set to 50 minutes. Similarly, the automation assistant can select one or more performance outputs from the set of performance outputs to be implemented based on a predicted NLU metric associated with the NLU data stream and / or a predicted performance metric associated with the performance data stream. For example, suppose the predicted metric satisfies a first threshold metric and / or a second threshold metric, and this indicates a very high and / or low confidence level for the user to say "set a timer for 15 minutes". In this example, the automation assistant can additionally or alternatively cause a fifth performance output associated with a first graphical content depicting a timer set to 15 minutes to be implemented, such that the graphical depiction of the timer set to 15 minutes is visually rendered and presented to the user. Conversely, suppose the prediction metric fails to satisfy both the first and second threshold metrics, indicating a low confidence level for the user to say "set a timer for 15 minutes." In this example, the automation assistant can additionally or alternatively cause a fifth performance output associated with the first graphical content and a sixth performance output associated with the second graphical content, where the first graphical content depicts a timer set to 15 minutes and the second graphical content depicts a timer set to 50 minutes, such that both the graphical depiction of the timer set to 15 minutes and the other graphical depiction of the timer set to 50 minutes are visually rendered and presented to the user. Therefore, even when the next interaction state indicates that a performance output should be implemented, the performance output can be dynamically determined based on the current state.
[0014] In some implementations, the automated assistant can partially perform a set of performance outputs before determining the next interaction state that would cause the performance output to be implemented. As noted above in the initial example where the user provides the verbal utterance “turn on the…the uhmmm…”, the set of performance outputs generated by multiple agents can include performance outputs associated with an assistant command that, when implemented, turns on music; performance outputs associated with an assistant command that, when implemented, turns on a video streaming service; performance outputs associated with an assistant command that, when implemented, controls one or more 1P smart devices, and so on. In this example, when the user anticipates requesting the automated assistant to perform some action relative to the software application and / or smart device, the automated assistant can preemptively establish connections with the software application (e.g., a music application, a video streaming application) and / or the smart device (e.g., a smart TV, a smart appliance, etc.). As a result, the latency for causing the performance output to be implemented as the next interaction state can be reduced.
[0015] In some implementations, a set of natural dialogue outputs can be generated based on NLU metrics associated with the NLU data stream and / or audio-based characteristics, and natural dialogue outputs can be selected from the set. In some versions of those implementations, a superset of natural dialogue outputs can be stored in one or more databases accessible to the client device, and the set of natural dialogue outputs can be generated from the superset based on NLU metrics associated with the NLU data stream and / or audio-based characteristics. These natural dialogue outputs can be implemented as the next interaction state to facilitate a dialogue session, but are not necessarily implemented as an action. For example, natural dialogue outputs can include instructions requesting user confirmation of continued interaction with the automation assistant (e.g., "Are you still there?", etc.), requests for additional user input to facilitate a dialogue session between the user and the automation assistant (e.g., "What did you want to turn on?", etc.), and filler speech (e.g., "Sure," "Alright," etc.).
[0016] In some implementations, even if the next interaction state to be achieved is to avoid the interaction from being realized, the ASR model can still be used to process the audio data stream to update the ASR output stream, NLU output stream, and performance data stream. Therefore, the current state can be iteratively updated, allowing the determination of the next interaction state to also occur iteratively during the dialogue session. In some versions of those implementations, the automation assistant can additionally or alternatively use a voice activity detection (VAD) model to process the audio data stream to monitor for the occurrence of voice activity (e.g., after the user has been silent for a few seconds). In these implementations, the automation assistant can determine whether the detected voice activity is directed at the automation assistant. For example, the updated NLU data stream can indicate whether the detected voice activity is directed at the automation assistant. If so, the automation assistant can continue to update the performance data stream and performance output set.
[0017] In some implementations, the current state of the NLU output stream, performance data stream, and audio-based features includes the most recent instance of NLU output generated based on the most recent spoken utterance in one or more spoken utterances, the most recent instance of performance data generated based on the most recent NLU output, and the most recent instance of audio-based features generated based on the most recent spoken utterance. Continuing the example above, at the current moment of the conversation session (e.g., two seconds after the user finishes providing the “the uhmmm” portion of the spoken utterance “turn on the...theuhmmm…”), the current state may correspond only to the NLU data, performance data, and audio-based features generated based on the “theuhmmm” portion of the spoken utterance. In additional or alternative implementations, the current state of the NLU output stream, performance data stream, and audio-based features further includes one or more historical instances of NLU output generated based on one or more historical spoken utterances preceding the most recent spoken utterance, one or more historical instances of performance data generated based on one or more historical instances of NLU output, and one or more historical instances of audio-based features generated based on one or more historical spoken utterances. Continuing with the example above, at the current moment of the dialogue session (e.g., two seconds after the user finishes providing the “the uhmmm” portion of the spoken utterance “turn on the…the uhmmm…”), the current state may correspond only to NLU data, performance data, and audio-based features generated based on the “the uhmmm” portion of the spoken utterance, the “turn on the” portion of the spoken utterance, and optionally any spoken utterances that occurred before the “turn on the” portion of the spoken utterance.
[0018] By using the techniques described herein, one or more technical advantages can be achieved. As a non-limiting example, the techniques described herein enable automated assistants to participate in natural conversations with users during a dialogue session. For example, the automated assistant can determine the next interaction state of the dialogue session based on the current state of the session, making it not limited to turn-based dialogue sessions or depending on determining that the user has finished speaking before responding. Therefore, while the user participates in these natural conversations, the automated assistant can determine when and how to respond to the user. This results in various technical advantages of saving computational resources at the client device and enables dialogue sessions to end more quickly and efficiently. For example, the number of automated assistant failures can be reduced because the automated assistant can wait for more information from the user before attempting to perform any action on behalf of the user. Furthermore, for example, the amount of user input received at the client device can be reduced because the number of times the user must repeat themselves or re-invoke the automated assistant can be reduced.
[0019] As used herein, a “conversational session” can include logically self-contained exchanges between a user and an automated assistant (and in some cases, other human participants). An automated assistant can distinguish between multiple conversational sessions with a user based on various signals, such as the elapsed time between sessions, changes in the user’s context between sessions (e.g., location, before / during / after a scheduled meeting, etc.), detection of one or more intermediate interactions between the user and the client device other than the conversation between the user and the automated assistant (e.g., the user switching applications for a period of time, the user leaving and then later returning to a standalone voice-activated product), the locking / sleep of the client device between sessions, changes in the client device used to interface with the automated assistant, and so on.
[0020] The above description is provided as an overview of only a few embodiments disclosed herein. These and other embodiments are described in more detail herein.
[0021] It should be understood that the techniques disclosed herein can be implemented locally on a client device, remotely by one or more servers connected to the client device via one or more networks, and / or both. Attached Figure Description
[0022] Figure 1 A block diagram of an example environment is depicted, which demonstrates various aspects of this disclosure and enables the implementation of the methods disclosed herein.
[0023] Figure 2 The use according to various implementation methods is described. Figure 1 The various components demonstrate example process flows for various aspects of this disclosure.
[0024] Figure 3A and Figure 3B The flowchart illustrates example methods for determining, according to various implementations, whether the next interaction state to be achieved during a dialogue session (i) enables the fulfillment of output, (ii) enables the natural dialogue output to be audibly rendered for presentation to the user, or (iii) avoids enabling any interaction.
[0025] Figure 4 Non-limiting examples are described regarding determining, according to various implementations, whether the next interactive state to be achieved during a dialogue session is (i) to enable the fulfillment of output, (ii) to enable the natural dialogue output to be audibly rendered for presentation to the user, or (iii) to avoid enabling any interaction.
[0026] Figure 5A , Figure 5B and Figure 5C Various non-limiting examples of implementations that enable the fulfillment of outputs during a dialogue session are described.
[0027] Figure 6 Exemplary architectures of computing devices according to various implementations are depicted. Detailed Implementation
[0028] Now go to Figure 1 A block diagram of an example environment is depicted, illustrating various aspects of this disclosure and enabling the implementation of the embodiments disclosed herein. The example environment includes a client device 110 and a natural dialogue system 180. In some embodiments, the natural dialogue system 180 can be implemented locally at the client device 110. In additional or alternative embodiments, the natural dialogue system 180 can be implemented remotely from the client device 110, such as... Figure 1 As depicted in the document. In these embodiments, the client device 110 and the natural dialogue system 180 may be communicatively coupled to each other via one or more networks 199, such as one or more wired or wireless local area networks (“LANs”, including Wi-Fi LANs, mesh networks, Bluetooth, near field communications, etc.) or wide area networks (“WANs”, including the Internet) .
[0029] Client device 110 may be one or more of the following: desktop computer, laptop computer, tablet computer, mobile phone, vehicle computing device (e.g., in-vehicle communication system, in-vehicle entertainment system, in-vehicle navigation system), independent interactive speaker (optionally with a display), smart appliance such as a smart TV, and / or wearable device of a user including a computing device (e.g., watch of a user with a computing device, glasses of a user with a computing device, virtual or augmented reality computing device). Additional and / or alternative client devices may be provided.
[0030] Client device 110 can execute automation assistant client 114. An instance of automation assistant client 114 can be an application separate from the operating system of client device 110 (e.g., installed "on top" of the operating system) or can alternatively be implemented directly by the operating system of client device 110. Automation assistant client 114 can interact with natural dialogue system 180 implemented locally at client device 110 or via, for example... Figure 1 The automation assistant client 114 (and optionally through interaction with other remote systems (e.g., servers)) can form content that appears from the user's perspective as a logical instance of the automation assistant 115, which the user can utilize to engage in human-computer dialogue. The instance of the automation assistant 115 is in... Figure 1 The diagram is depicted and surrounded by dashed lines representing the Automation Assistant Client 114 (client device 110) and the Natural Dialogue System 180. Therefore, it should be understood that a user participating in the Automation Assistant Client 114 executing on client device 110 can actually participate in his or her own logical instance of the Automation Assistant 115 (or a logical instance of the Automation Assistant 115 shared among family members or other user groups). For the sake of brevity and simplicity, as used herein, Automation Assistant 115 will refer to the Automation Assistant Client 114 executing locally on client device 110 and / or on one or more servers where the Natural Dialogue System 180 can be implemented.
[0031] In various embodiments, client device 110 may include a user input engine 111 configured to detect user input provided by a user of client device 110 using one or more user interface input devices. For example, client device 110 may be equipped with one or more microphones that capture audio data, such as audio data corresponding to the user's spoken words or other sounds in the environment of client device 110. Additionally or alternatively, client device 110 may be equipped with one or more vision components configured to capture visual data corresponding to images and / or movements (e.g., gestures) detected in the field of view of one or more vision components. Additionally or alternatively, client device 110 may be equipped with one or more touch-sensitive components (e.g., keyboard and mouse, stylus, touchscreen, touch panel, one or more hardware buttons, etc.) configured to capture signals corresponding to touch input to client device 110.
[0032] In various implementations, client device 110 may include rendering engine 112 configured to provide content for audible and / or visual presentation to a user of client device 110 using one or more user interface output devices. For example, client device 110 may be equipped with one or more speakers, enabling content to be audibly presented to the user via client device 110. Additionally or alternatively, client device 110 may be equipped with a display or projector, enabling content to be visually presented to the user via client device 110.
[0033] In various embodiments, client device 110 may include one or more presence sensors 113 configured to provide signals indicating the detected presence (particularly human presence) upon consent from the corresponding user(s). In some of these embodiments, the automation assistant 115 is capable of recognizing client device 110 (or another computing device associated with the user of client device 110) at least in part based on the presence of the user at client device 110 (or at another computing device associated with the user of client device 110) to satisfy a verbal utterance. The verbal utterance can be satisfied by rendering response content at client device 110 and / or the other computing device associated with the user of client device 110 (e.g., via rendering engine 112), by controlling client device 110 and / or the other computing device associated with the user of client device 110, and / or by performing any other action at client device 110 and / or the other computing device associated with the user of client device 110 to satisfy the verbal utterance. As described herein, the automation assistant 115 can use data determined by the presence sensor 113 to determine the location of the client device 110 (or other computing device) based on where the user is or has recently been located, and can provide the corresponding commands only to the client device 110 (or those other computing devices). In some additional or alternative embodiments, the automation assistant 115 can use data determined by the presence sensor 113 to determine whether any user (any user or a specific user) is currently near the client device 110 (or other computing device), and can optionally suppress the provision of data to and / or from the client device 110 (or other computing device) based on the number of users(s) near the client device 110 (or other computing device).
[0034] The presence sensor 113 can take various forms. For example, the client device 110 can detect the presence of a user using one or more of the user interface input components described above relative to the user input engine 111. Additionally or alternatively, the client device 110 may be equipped with other types of light-based presence sensors 113, such as passive infrared (“PIR”) sensors that measure infrared (“IR”) light radiated from objects within its field of view.
[0035] Additionally or alternatively, in some embodiments, the presence sensor 113 may be configured to detect other phenomena associated with human presence or device presence. For example, in some embodiments, the client device 110 may be equipped with the presence sensor 113, which detects various types of wireless signals (e.g., waves such as radio waves, ultrasonic waves, electromagnetic waves, etc.) emitted by other computing devices (e.g., mobile devices, wearable computing devices, etc.) and / or other computing devices carried / operated by the user. For example, the client device 110 may be configured to emit waves imperceptible to humans, such as ultrasonic waves or infrared waves, which can be detected by other computing devices (e.g., via an ultrasonic / infrared receiver, such as a microphone with ultrasonic capabilities).
[0036] Additionally or alternatively, client device 110 may emit other types of waves imperceptible to humans, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.) that can be detected by other computing devices carried / operated by the user (e.g., mobile devices, wearable computing devices, etc.) and used to determine the user's specific location. In some embodiments, GPS and / or Wi-Fi triangulation may be used, for example, to detect a person's location based on GPS and / or Wi-Fi signals to / from client device 110. In other embodiments, client device 110 may use other wireless signal characteristics, such as time-of-flight, signal strength, etc., individually or collectively, to determine the location of a specific person based on signals emitted by other computing devices carried / operated by the user.
[0037] Additionally or alternatively, in some embodiments, client device 110 may perform speaker identification (SID) to identify the user from the user's voice. In some embodiments, the speaker's movement can then be determined, for example, by the presence sensor 113 of client device 110 (and optionally, a GPS sensor, Soli chip, and / or accelerometer of client device 110). In some embodiments, based on such detected movement, the user's location can be predicted, and this location can be assumed to be the user's location when the proximity of client device 110 and / or other computing devices to the user's location makes it possible to render any content at client device 110 and / or other computing devices. In some embodiments, it can be simply assumed that the user is in the last position he or she engaged with automation assistant 115, especially if not much time has passed since the last engagement.
[0038] Furthermore, client device 110 and / or natural dialogue system 180 may include one or more memories for storing data and / or software applications 198, one or more processors for accessing data and executing software applications 198, and / or other components facilitating communication via one or more networks 199. In some embodiments, one or more of the software applications 198 can be locally installed at client device 110, while in other embodiments, one or more of the software applications 198 can be remotely hosted (e.g., by one or more servers) and can be accessed by client device 110 via one or more networks 199. Operations performed by client device 110, other computing devices, and / or by automation assistant 115 can be distributed across multiple computer systems. Automation assistant 115 can be implemented, for example, at client device 110 and / or via a network (e.g., Figure 1 A network (199) is a computer program running on one or more computers in one or more locations that are coupled to each other.
[0039] In some implementations, the operations performed by the automation assistant 115 can be implemented locally at the client device 110 via the automation assistant client 114. For example... Figure 1 As shown, the automation assistant client 114 may include an automatic speech recognition (ASR) engine 120A1, a natural language understanding (NLU) engine 130A1, an execution engine 140A1, and a text-to-speech (TTS) engine 150A1. In some implementations, the operations performed by the automation assistant 115 can be distributed across multiple computer systems, such as when... Figure 1 When the natural dialogue system 180 is implemented remotely from the client device 110 as depicted. In these embodiments, the automation assistant 115 may additionally or alternatively utilize the ASR engine 120A2, NLU engine 130A2, fulfillment engine 140A2 and TTS engine 150A2 of the natural dialogue system 180.
[0040] Each of these engines can be configured to perform one or more functions. For example, ASR engines 120A1 and / or 120A2 can use multiple streaming ASR models (e.g., recurrent neural network (RNN) models, transformer models, and / or any other type of ML model capable of performing ASR) stored in a database of multiple machine learning (ML) models 115A to process the audio data stream captured from spoken utterances and generated by the microphone of client device 110 to generate an ASR output stream. Notably, when generating the audio data stream, the streaming ASR model can be utilized to generate the ASR output stream. Furthermore, NLU engines 130A1 and / or 130A2 can use multiple NLU models (e.g., long short-term memory (LSTM), gated recurrent unit (GRU), and / or any other type of RNN or other ML model capable of performing NLU) and / or multiple rules based on syntax to process the ASR output stream to generate an NLU output stream. Furthermore, fulfillment engines 140A1 and / or 140A2 are capable of generating a set of fulfillment outputs based on the fulfillment data stream generated according to the NLU output stream. They can utilize, for example, one or more first-party (1P) agents 171 and / or one or more third-party (3P) agents 171 (e.g., as referenced). Figure 2 The TTS engine 150A1 and / or 150A2 can use multiple TTS models stored in the ML model database 115A to process text data (e.g., text defined by the automation assistant 115) to generate synthetic speech audio data including computer-generated synthesized speech. It is worth noting that the multiple ML models stored in the ML model database 115A can be on-device ML models locally stored at the client device 110 or shared ML models accessible to the client device 110 and / or remote systems (e.g., multiple servers).
[0041] In various implementations, the ASR output stream can include, for example, a stream of speech hypotheses (e.g., terminology hypotheses and / or transcription hypotheses) predicted to correspond to multiple spoken utterances of a user captured in the audio data stream, one or more corresponding prediction values (e.g., probability, log-likelihood, and / or other values) for each speech hypothesis, multiple phonemes predicted to correspond to multiple spoken utterances of a user captured in the audio data stream, and / or other ASR outputs. In some versions of those implementations, ASR engines 120A1 and / or 120A2 can (e.g., based on the corresponding prediction values) select one or more speech hypotheses as the recognized text corresponding to the spoken utterances.
[0042] In various implementations, the NLU output stream can include, for example, an annotated stream of recognized text, which includes one or more annotations for one or more (e.g., all) terms of the recognized text. For example, NLU engines 130A1 and / or 130A2 may include part-of-speech taggers (not shown) configured to annotate terms with the grammatical roles of words. Additionally or alternatively, NLU engines 130A1 and / or 130A2 may include entity taggers (not depicted) configured to annotate entity references in one or more segments of the recognized text, such as references to people (including, for example, literary figures, celebrities, public figures, etc.), organizations, locations (real and hypothetical), etc. In some implementations, data about entities may be stored in one or more databases, such as in a knowledge graph (not depicted). In some implementations, the knowledge graph may include nodes representing known entities (and in some cases, entity attributes), and edges connecting nodes and representing relationships between entities. Entity taggers can annotate entity references at a high-granularity level (e.g., to enable the identification of all references to entity classes such as people) and / or a low-granularity level (e.g., to enable the identification of all references to a specific entity such as a particular person). Entity taggers can rely on the content of the natural language input to parse specific entities and / or can optionally communicate with a knowledge graph or other entity database to parse specific entities. Additionally or alternatively, NLU engines 130A1 and / or 130A2 may include a coreference parser (not depicted) configured to group or “cluster” references to the same entity based on one or more contextual cues. For example, a coreference parser can be used to parse the term “them” in the natural language input “buy them” as “buy theatre tickets” based on “theatre tickets” mentioned in a client device notification rendered immediately before the input “buy them” is received. In some implementations, one or more components of NLU engines 130A1 and / or 130A2 may depend on annotations from one or more other components of NLU engines 130A1 and / or 130A2. For example, in some implementations, the entity tagger may depend on annotations from the coreference resolver when annotating all references to a particular entity. Furthermore, for example, in some implementations, the coreference resolver may depend on annotations from the entity tagger when clustering references to the same entity.
[0043] Although described relative to a single client device with a single user Figure 1However, it should be understood that this is for illustrative purposes and not intended to be limiting. For example, one or more additional client devices of the user may also implement the techniques described herein. For example, client device 110, one or more additional client devices, and / or any other computing devices of the user may form an ecosystem of devices capable of employing the techniques described herein. These additional client devices and / or computing devices may communicate with client device 110 (e.g., via network 199). As another example, a given client device may be used by multiple users in a shared setup (e.g., a group of users, a family).
[0044] As described herein, the automation assistant 115 is able to determine, based on the NLU output stream of the spoken utterance(s) captured in the audio data stream on which the NLU output stream and performance data stream are generated, the performance data stream, and the current state based on audio characteristics, whether the next interaction state to be implemented during the dialogue session between the user and the automation assistant 115 is (i) to enable the performance output, (ii) to enable the natural dialogue output to be audibly rendered for presentation to the user, or (iii) to avoid enabling any interaction. In making this determination, the automation assistant is able to utilize the natural dialogue engine 160. In various embodiments, and as described herein... Figure 1 As shown, the natural dialogue engine 160 may include an acoustic engine 161, a state engine 162, an execution output engine 163, a natural dialogue output engine 164, and a partial execution engine 165.
[0045] In some implementations, acoustic engine 161 is capable of determining audio-based characteristics based on processing an audio data stream. In some implementations, acoustic engine 161 is capable of processing the audio data stream using acoustic ML models stored in the ML model database(s) 115A to determine audio-based characteristics. In some implementations, acoustic engine 161 is capable of processing the audio data stream using one or more rules to determine audio-based characteristics. Audio-based characteristics can include, for example, prosodic attributes associated with spoken utterances(s) captured in the audio data stream, duration elapsed since the most recent spoken utterance was provided, and / or other audio-based characteristics. Prosodic attributes can include, for example, one or more attributes of syllables and larger speech units, including language functions such as intonation, pitch, stress, rhythm, beat, pitch, and rests. Furthermore, prosodic attributes can provide indications of, for example, emotional state; form (e.g., statement, question, or command); irony; sarcasm; speech rhythm; and / or emphasis. In other words, prosodic properties are characteristics of speech that are independent of the individual speech characteristics of a given user and can be dynamically determined during a conversation based on individual spoken utterances and / or combinations of multiple spoken utterances.
[0046] In some implementations, the state engine 162 is capable of determining the current state of the dialogue session based on the NLU output stream, the fulfillment data stream, and audio-based characteristics. Furthermore, the state engine 162 is capable of determining the next interaction state to be implemented to facilitate the dialogue session based on the current state of the dialogue session. The next interaction state can include, for example, (i) enabling the fulfillment output to be implemented, (ii) enabling natural dialogue output to be audibly rendered for presentation to the user, or (iii) avoiding enabling any interaction to be implemented. In other words, the state engine 162 is capable of analyzing signals generated by the various components described herein to determine the current state of the dialogue and, based on the current state, determining how the automation assistant 115 should continue to facilitate the dialogue session. Notably, the state engine 162 is capable of iteratively (e.g., continuously at 10Hz, 20Hz, or other frequencies) updating the current state throughout the dialogue session based on updates to the NLU output stream, the fulfillment data stream, and audio-based characteristics, without waiting for the completion of a response from the user or the automation assistant. As a result, the next interaction state to be implemented can be iteratively updated, allowing for continuous determination of how the automation assistant 115 should proceed to facilitate the dialogue session.
[0047] In some implementations, the current state of the NLU output stream, performance data stream, and audio-based features includes the most recent instance of NLU output generated based on the most recent spoken utterance in one or more spoken utterances, the most recent instance of performance data generated based on the most recent NLU output, and the most recent instance of audio-based features generated based on the most recent spoken utterance. In additional or alternative implementations, the current state of the NLU output stream, performance data stream, and audio-based features further includes one or more historical instances of NLU output generated based on one or more historical spoken utterances preceding the most recent spoken utterance, one or more historical instances of performance data generated based on one or more historical instances of NLU output, and one or more historical instances of audio-based features generated based on one or more historical spoken utterances. Therefore, the current state of a dialogue session can be determined based on the most recent spoken utterance and / or one or more previous spoken utterances.
[0048] In various implementations, the state engine 162 is capable of determining the next interaction state based on the current state of processing the NLU output stream, performance data stream, and audio-based features using a classification ML model stored in a database of multiple ML models 115A. The classification ML model is capable of generating a corresponding predictive metric associated with each next interaction state based on the current state of processing the NLU output stream, performance data stream, and audio-based features. The classification ML model can be trained based on multiple training instances. Each training instance can include training instance inputs and training instance outputs. Training instance inputs can include, for example, the training state of the training stream of NLU outputs, the training stream of performance data, and training audio-based features, and training instance outputs can include ground truth outputs associated with whether the automation assistant 115 should, based on the training instance, (i) enable the performance output, (ii) enable the natural dialogue output to be audibly rendered for presentation to the user, or (iii) avoid enabling any interaction. The training state can be determined based on historical dialogue sessions between the user and the automation assistant 115 (and / or other users and their respective automation assistants), or can be heuristically defined. When training a classification ML model, training instances can be used as input to generate predicted outputs (e.g., corresponding prediction metrics), and the predicted outputs can be compared with the true outputs to generate one or more losses. The classification ML model can be updated based on one or more losses (e.g., via backpropagation). Once trained, the classification ML model can be deployed for use by a state engine 162 as described in this paper.
[0049] In some implementations, the fulfillment output engine 163 is capable of selecting one or more fulfillment outputs to be implemented from the set of fulfillment outputs in response to determining that the next interaction state to be implemented is (i) causing the fulfillment output to be implemented. As described above, fulfillment engines 140A1 and / or 140A2 are capable of generating a set of fulfillment outputs based on a fulfillment data stream, and are capable of generating the fulfillment data stream using, for example, one or more of 1P agents 171 and / or one or more of 3P agents 171. Although one or more of the 1P agents 171 are depicted as in Figure 1The implementation may take place on one or more of the networks 199, but it should be understood that this is for illustrative purposes and not intended to be limiting. For example, one or more 1P agents 171 may be implemented locally at client device 110, and the NLU output stream may be transmitted to one or more 1P agents 171 via an application programming interface (API), and fulfillment data from one or more 1P agents 171 may be obtained by the fulfillment output engine 163 via the API and incorporated into the fulfillment data stream. One or more 3P agents 172 may be implemented via a corresponding 3P system (e.g., (multiple) 3P servers). The NLU output stream may additionally or alternatively be transmitted over one or more networks 199 on one or more 3P agents 172 (and one or more 1P agents 171 not implemented locally at client device 110) to enable one or more 3P agents 172 to generate fulfillment data. The fulfillment data may be transmitted back to client device 110 via one or more of the networks 199 and incorporated into the fulfillment data stream.
[0050] Furthermore, the performance output engine 163 can select one or more performance outputs from the performance output set based on NLU metrics associated with the NLU data stream and / or performance metrics associated with the performance data stream. NLU metrics can be, for example, probabilities, log-likelihoods, binary values, etc., indicating the confidence that the predicted intent of NLU engines 130A1 and / or 130A2(multiple) corresponds to the actual intent of the user providing the spoken utterance(s) captured in the audio data stream, and / or the confidence that the predicted slot values of(multiple) parameters(multiple) associated with(multiple) predicted intents correspond to the actual slot values of(multiple) parameters(multiple) associated with(multiple) predicted intents. NLU metrics can be generated when NLU engines 130A1 and / or 130A2 generate the NLU output stream and can be included in the NLU output stream. Performance metrics can be, for example, probabilities, log-likelihoods, binary values, etc., indicating the confidence that the predicted performance outputs of performance engines 140A1 and / or 140A2(multiple) correspond to the user's desired performance. Performance metrics can be generated when performance data is generated in one or more of 1P agents 171 and / or one or more of 3P agents 172, and can be incorporated into the performance data stream, and / or can be generated when performance engines 140A1 and / or 140A2 process performance data received from one or more of 1P agents 171 and / or one or more of 3P agents 172, and can be incorporated into the performance data stream.
[0051] In some implementations, the natural dialogue output engine 164 is capable of generating a set of natural dialogue outputs and is capable of selecting one or more natural dialogue outputs to be implemented from the set in response to determining that the next interaction state to be implemented is (ii) causing the natural dialogue outputs to be audibly rendered for presentation to the user. For example, the set of natural dialogue outputs can be generated based on NLU metrics associated with the NLU data stream and / or audio-based characteristics. In some versions of those implementations, a superset of natural dialogue outputs can be stored in one or more databases (not shown) accessible to the client device 110, and the set of natural dialogue outputs can be generated from the superset of natural dialogue outputs based on NLU metrics associated with the NLU data stream and / or audio-based characteristics. These natural dialogue outputs can be implemented as the next interaction state to facilitate a dialogue session, but are not necessarily implemented as an execution. For example, natural dialogue output can include instructions requesting user confirmation of continued interaction with the automation assistant 115 (e.g., "Are you still there?"), requests for additional user input to facilitate the conversation between the user and the automation assistant 115 (e.g., "What did you want to turn on?"), and fill-in-the-blank speech (e.g., "sure," "Alright," etc.). In various implementations, the natural dialogue engine 164 can utilize one or more language models stored in the ML model database(s) 115A when generating the set of natural dialogue outputs.
[0052] In some implementations, the partial fulfillment engine 165 can partially fulfill the fulfillment outputs in the fulfillment output set before determining the next interaction state in which the fulfillment outputs are implemented. For example, the partial fulfillment engine 165 can establish connections with one or more software applications 198, one or more 1P agents 171, one or more 3P agents 172, additional client devices communicating with client device 110, and / or one or more smart devices communicating with client device 110, which is associated with the fulfillment outputs included in the fulfillment output set. This enables the generation (but not audible rendering) of synthetic speech audio data including synthesized speech, the generation (but not visual rendering) of graphical content, and / or the execution of other partial fulfillments of one or more fulfillment outputs. As a result, the latency for implementing the fulfillment outputs into the next interaction state can be reduced.
[0053] Now go to Figure 2 It describes the use Figure 1Various components demonstrate example process flows for various aspects of this disclosure. ASR engines 120A1 and / or 120A2 are capable of processing audio data stream 201A using streaming ASR models stored in the ML model database(s) 115A to generate ASR output stream 220. NLU engines 130A1 and / or 130A2 are capable of processing ASR output stream 220 using NLU models stored in the ML model database(s) 115A to generate NLU output stream 230. In some embodiments, NLU engines 130A1 and / or 130A2 are additionally or alternatively capable of processing non-audio data stream 201B when generating NLU output stream 230. Non-audio data stream 201B can include visual data streams generated by the visual components(s) of client device 110, touch input streams provided by a user via the display of client device 110, typing input streams provided by a user via the display of client device 110 or peripheral devices (e.g., mouse and keyboard), and / or any other non-audio data. In some implementations, (multiple) 1P agents 171 are capable of processing the NLU output stream to generate 1P fulfillment data 240A. In additional or alternative implementations, (multiple) 3P agents 172 are capable of processing the NLU output stream 230 to generate 3P fulfillment data 240B. Fulfillment engines 140A1 and / or 140A2 are capable of generating fulfillment data stream 240 based on 1P fulfillment data 240A and / or 3P fulfillment data 240B. Furthermore, acoustic engine 161 is capable of processing audio data stream 201A to generate audio-based features 261 associated with audio data stream 201A.
[0054] State engine 162 is capable of processing NLU output stream 230, performance data stream 240, and / or audio-based features to determine the current state 262, and, as shown in box 299, is capable of determining whether the next interaction state is: (i) to enable performance output, (ii) to enable natural dialogue output to be audibly rendered for presentation to the user, or (iii) to prevent any interaction from being enabled. For example, a classification ML model can be used to process NLU output stream 230, performance data stream 240, and audio-based features 241 to generate predictive metrics (e.g., binary values, probabilities, log-likelihoods, etc.). Each next interaction state can be associated with a corresponding predictive metric in the predictive metrics, such that state engine 162 is capable of determining whether (i) a first corresponding predictive metric based on the predictive metric enables performance output, (ii) a second corresponding predictive metric based on the predictive metric enables natural dialogue output to be audibly rendered for presentation to the user, or (iii) a third corresponding predictive metric based on the predictive metric prevents any interaction from being enabled.
[0055] In this scenario, if the state engine 162 determines (i) that a performance output is to be implemented based on a first corresponding prediction metric, the performance output engine 163 can select one or more performance outputs 263 from the set of performance outputs and implement one or more performance outputs 263 as shown in 280. Furthermore, in this scenario, if the state engine 162 determines (ii) that a natural dialogue output is to be audibly rendered and presented to the user based on a second corresponding prediction metric, the natural dialogue output engine 164 can select one or more natural dialogue outputs 264 from the set of natural dialogue outputs and render one or more natural dialogue outputs 264 audibly and present them to the user as shown in 280. Furthermore, in this scenario, if the state engine 162 determines (iii) that any interaction should be avoided based on a third corresponding prediction metric, the automation assistant 115 can avoid implementing any interaction, as indicated in 280. Although relative to... Figure 2 The process flow describes a specific implementation method, but it should be understood that this is for illustrative purposes and does not imply limitation.
[0056] By using the techniques described herein, one or more technical advantages can be achieved. As a non-limiting example, the techniques described herein enable automated assistants to participate in natural conversations with users during a dialogue session. For example, the automated assistant can determine the next interaction state of the dialogue session based on the current state of the session, making it not limited to turn-based dialogue sessions or depending on determining that the user has finished speaking before responding. Therefore, while the user participates in these natural conversations, the automated assistant can determine when and how to respond to the user. This results in various technical advantages, such as saving computational resources on the client device and enabling dialogue sessions to end more quickly and efficiently. For example, the number of automated assistant failures can be reduced because the automated assistant can wait for more information from the user before attempting to perform any action on behalf of the user. Furthermore, for example, the amount of user input received on the client device can be reduced because the number of times the user must repeat themselves or re-invoke the automated assistant can be reduced.
[0057] Now go to Figure 3A The diagram depicts a flowchart of an exemplary method 300 for determining whether the next interaction state to be achieved during a dialogue session is (i) to enable the fulfillment of an output, (ii) to enable the natural dialogue output to be audibly rendered for presentation to the user, or (iii) to prevent any interaction from being implemented. For convenience, the operation of method 300 is described with reference to a system performing the operation. The system of method 300 includes (e.g., multiple) computing devices (e.g., Figure 1 Client device 110 Figure 4 Client device 110A, Figures 5A-5CClient device 110B, and / or Figure 6 The computing device 610, one or more servers, and / or other computing devices may include one or more processors, memories, and / or other components. Furthermore, although the operations of method 300 are shown in a specific order, this does not imply limitation. One or more operations may be reordered, omitted, and / or added.
[0058] At box 352, the system uses a streaming ASR model to process the audio data stream to generate an ASR output stream. The audio data stream can be generated by the microphone(s) of a user's client device participating in a conversational session with an automated assistant implemented at least partially at the client device. Furthermore, the audio data stream can capture one or more spoken utterances from the user directed to the automated assistant. In some implementations, the system may process the audio data stream in response to determining that the user has invoked the automated assistant via one or more specific words and / or phrases (e.g., hot words such as "Hey Assistant," "Assistant," etc.), actuating one or more buttons (e.g., software and / or hardware buttons), one or more gestures captured by the client device's (multiple) visual components (which invoke the automated assistant upon detection), and / or by any other means. At box 354, the system uses an NLU model to process the ASR output stream to generate an NLU output stream. At box 356, the system generates a performance data stream based on the NLU output stream. At box 358, the system determines audio-based characteristics associated with one or more spoken utterances captured in the audio data based on the processed audio data stream. Audio-based features can include, for example, one or more prosodic attributes (e.g., intonation, pitch, stress, rhythm, beat, pitch, rest, and / or other prosodic attributes) associated with each of one or more spoken utterances, the duration elapsed since the most recent spoken utterance in one or more spoken utterances provided by the user, and / or other audio-based features that can be determined based on processing the audio data stream.
[0059] At box 360, the system determines the next interaction state to be implemented based on the current state of the NLU output stream, performance data stream, and / or audio-based characteristics associated with one or more spoken utterances captured in the audio data. In some implementations, the current state can be based on the NLU output, performance data, and / or audio-based characteristics at the current moment in the dialogue session. In additional or alternative implementations, the current state can be based on the NLU output, performance data, and / or audio-based characteristics at one or more previous moments in the dialogue session. In other words, the current state can correspond to the state of the dialogue session within the context of the most recent spoken utterance provided by the user and / or the entire dialogue session as a whole. The system can use, for example... Figure 3BMethod 360A determines the next interactive state to be implemented based on the current state.
[0060] refer to Figure 3B Furthermore, at box 382, the system uses a classification ML model to process the current state of the NLU output stream, the performance data stream, and / or audio-based features to generate corresponding predictive metrics. For example, the system can generate a first predictive metric associated with (i) enabling the performance output to be implemented, a second predictive metric associated with (ii) enabling the natural dialogue output to be audibly rendered and presented to the user, and a third predictive metric associated with (iii) avoiding enabling any interaction. The corresponding predictive metrics can be, for example, binary values, probabilities, log-likelihoods, etc., indicating the likelihood that one of the following should be implemented as the next interaction state: (i) enabling the performance output to be implemented, (ii) enabling the natural dialogue output to be audibly rendered and presented to the user, and (iii) avoiding enabling any interaction.
[0061] At box 384, the system can determine, based on the corresponding predictive metric, whether the next interaction state is: (i) to enable the fulfillment of the output, (ii) to enable the natural dialogue output to be audibly rendered to be presented to the user, or (iii) to avoid enabling any interaction.
[0062] If, at the iteration of box 384, the system determines, based on the corresponding prediction metric, that the next interaction state (i) results in the fulfillment output being realized, then the system can proceed to box 386A. At box 386A, the system selects a fulfillment output from the set of fulfillment outputs based on the NLU prediction metric associated with the NLU output stream and / or the fulfillment metric associated with the fulfillment data stream (e.g., as referenced). Figure 1 and Figure 2 The execution output engine 163 describes this. The execution output set can be generated by multiple agents (e.g., as described in reference 163). Figure 1 and Figure 2 (As described by (multiple) 1P agents 171 and / or (multiple) 3P agents 172). The system is able to return to Figure 3A The system can execute the selected output at box 386A, and enable the next interactive state to be implemented. In other words, the system can enable the execution output to be implemented.
[0063] If, at the iteration in box 384, the system determines, based on the corresponding prediction metric, that the next interaction state is (ii) such that the natural dialogue output is audibly rendered to be presented to the user, then the system can proceed to box 386B. At box 386B, the system selects natural dialogue outputs from the set of natural dialogue outputs based on the NLU prediction metric associated with the NLU output stream determined at box 360 and / or on audio-based characteristics. The set of natural dialogue outputs can be generated based on one or more databases that include a superset of natural dialogue outputs (e.g., as referenced). Figure 1 and Figure 2 (As described in the Natural Dialogue Output Engine 164). The system is able to return to... Figure 3A The system selects box 362 and enables the next interactive state to be realized. In other words, the system enables the natural dialogue output selected at box 386B to be audibly rendered and presented to the user.
[0064] If, at the iteration in box 384, the system determines the next interaction state based on the corresponding prediction metric to be (iii) avoiding any interaction from being implemented, then the system can proceed to box 384C. At box 384BA, the system avoids the interaction from being implemented. The system can then return to... Figure 3A The system can trigger the next interactive state by displaying box 362. In other words, the system enables the automated assistant to continue processing the audio data stream without performing any action at that given moment in the conversation.
[0065] It is worth noting, and further referencing Figure 3A After the system enables the next interactive state to be achieved at box 362, the system returns to box 352. In subsequent iterations of box 352, the system continues to process the audio data stream to continue generating the ASR output stream, NLU data stream, and execution data stream. Furthermore, one or more additional audio-based features can be determined based on the processing of the audio data stream at this subsequent iteration. As a result, the current state can be updated for this subsequent iteration, and a further next interactive state can be determined based on the updated current state. In this way, the current state can be continuously updated to continuously determine the next interactive state of the dialogue session.
[0066] Now go to Figure 4 The text describes non-limiting examples of determining whether the next interaction state to be implemented during a dialogue session is (i) to enable the fulfillment of an output, (ii) to enable the natural dialogue output to be audibly rendered for presentation to the user, or (iii) to avoid enabling any interaction. The automated assistant can be implemented at least partially at the client device 110A (e.g., relative to...). Figure 1 The described automated assistant 115). The automated assistant is able to utilize natural dialogue systems (e.g., relative to...). Figure 1The described natural dialogue system (180) determines the next interaction state. Figure 4 The client device 110A depicted may include various user interface components, including, for example, multiple microphones for generating audio data based on spoken words and / or other audible input, multiple speakers for audibly rendering synthesized speech and / or other audible output, and a display 190A for receiving touch input and / or visually rendering transcription and / or other visual output. Although Figure 4 The client device 110A depicted is a mobile device, but it should be understood that this is for illustrative purposes and not intended to be limiting.
[0067] Furthermore, and as Figure 4 As shown, the display 190A of the client device 110A includes various system interface elements 191, 192, and 193 (e.g., hardware and / or software interface elements) that can be interacted with by a user of the client device 110A to cause the client device 110A to perform one or more actions. The display 190A of the client device 110A enables the user to interact with the content rendered on the display 190A via touch input (e.g., by directing user input to the display 190A or a portion thereof (e.g., to a text input box 194 or to another portion of the display 190A)) and / or via verbal input (e.g., by selecting a microphone interface element 195—or simply by speaking without having to select a microphone interface element 195) (i.e., the automation assistant can monitor one or more specific terms or phrases, gestures, gazes, mouth movements, lip movements, and / or other conditions activating verbal input at the client device 110A).
[0068] For example, suppose a user of client device 110A provides the spoken utterance 452, “Assistant, call the office,” to initiate a conversation. It is noteworthy that, upon initiating the conversation, the user is requesting an automated assistant to make a phone call on their behalf. The automated assistant can determine the next interactive state to facilitate the conversation based on the NLU data stream generated from processing the audio data stream capturing the spoken utterance 452, the fulfillment data stream, and / or the current state based on audio characteristics. For example, suppose the automated assistant determines, based on (e.g., the NLU data stream used in determining the current state) the spoken utterance 452 is associated with the intent to make a phone call on behalf of the user, and the slot value of the phone number parameter corresponds to “office.” In this case, the NLU data stream can be processed by multiple agents (e.g., Figure 1The (multiple) 1P agents 171 and / or (multiple) 3P agents 172) process to generate (e.g., utilized when determining the current state) a fulfillment data stream. Further assume that the fulfillment data stream indicates that there is no available contact entry for the "office" parameter, such that the slot value for the telephone number parameter is not parsed.
[0069] In this example, and based at least on the NLU data stream and the fulfillment data stream, the determined current state can indicate that one or more fulfillment outputs should be implemented as the next interaction state to facilitate the conversational session. One or more fulfillment outputs to be implemented can be selected from the set of fulfillment outputs. The set of fulfillment outputs can include, for example, an assistant command initiating a phone call on behalf of a user using the contact entry "Office," synthetic speech audio data including a synthesized voice requesting more information about the contact entry "Office," and / or other fulfillment outputs. Furthermore, in this example, since the fulfillment data stream indicates that there is no available contact entry for "Office," the automation assistant can select the next fulfillment output to be implemented as synthetic speech audio data, which includes synthetic speech requesting more information about the contact entry "Office." Therefore, the automation assistant can cause the synthetic speech 454 (which requests more information about the contact entry "Office") included in the synthetic speech audio data to be audibly rendered and presented to the user via the speaker(s) of the client device 110A. Figure 4 As shown, the synthesized speech 454 can include "Sure, what's the phone number?".
[0070] Further assuming that, in response to the synthesized speech 454 being audibly rendered to be presented to the user, the user provides spoken utterance 456 “It's uhhh…”, followed by a few seconds of pause, and then spoken utterance 458, followed by another few seconds of pause. While the user provides spoken utterances 456 and 458, the automation assistant is able to continue processing the audio data stream generated by the microphone(s) of the client device 110A to iteratively update the current state to determine the next interaction state to be implemented. For example, at a given moment in the conversational session after the user provides spoken utterance 456 and before the user provides spoken utterance 458, the tone or rhythm of voice, including the audio-based features utilized in determining the current state, could indicate that the user is unsure of the “office” phone number but is trying to find it. In this example, the automation assistant could determine to avoid enabling any interaction, thus giving the user time to complete his / her thought by locating the phone number. Therefore, even if the automation assistant has determined that the user has completed spoken utterance 456, the automation assistant can still wait for further information from the user. Furthermore, at a given moment in the conversation after the user provides utterance 458, the tone or rhythm of the voice, including the audio-based features utilized in determining the current state, can still indicate to the user that they are unsure of the “office” phone number. As a result, the automated assistant can still determine to avoid any interaction being implemented, thus giving the user time to complete their thought by locating the phone number. However, at subsequent moments in the conversation (e.g., a few seconds after the user provides utterance 458), the audio-based features can include the duration since the user provided utterance 458.
[0071] In this example, and based on the NLU data flow and fulfillment data flow (e.g., still indicating the need for an "office" phone number to initiate a call on behalf of the user), and the audio-based characteristic indicating the user is unsure of the phone number and has been silent for several seconds, the determined current state can indicate that one or more natural dialogue outputs should be implemented as the next interaction state to facilitate the conversation. Therefore, even if the automation assistant has determined that the user has completed their spoken utterance 458 at a given moment, the automation assistant can still wait for further information from the user, such that one or more natural dialogue outputs should not be implemented as the next interaction state until that subsequent moment (e.g., after the given moment).
[0072] The system can select one or more natural dialogue outputs to be implemented from a set of natural dialogue outputs. The set of natural dialogue outputs may include, for example, synthesized speech audio data including synthesized speech asking the user whether they want more information or their thoughts, synthesized speech audio data including synthesized speech asking the user whether they still wish to interact with the automation assistant, and / or other natural dialogue outputs. Furthermore, in this example, since the data stream indicates that there is no available contact entry for "Office," and the characteristics of the audio indicate that the user is unsure of the information associated with the contact entry for "Office," the automation assistant can select that the next natural dialogue output to be audibly rendered and presented to the user should be the synthesized speech audio data asking the user whether they want more information or their thoughts. Therefore, the automation assistant can cause the synthesized speech 460, including the synthesized speech audio data asking the user whether they want more information or their thoughts, to be audibly rendered and presented to the user via the speaker(s) of the client device 110A. Figure 4 As shown, the synthesized speech 460 is capable of including the question, "Do you need more time?"
[0073] Further assuming that, in response to the synthesized speech 460 being audibly rendered to be presented to the user, the user provides the spoken words 462 "Yes," and the automation assistant causes the synthesized speech 464 "Okay" to be audibly rendered to be presented to the user, confirming the user's expectation of a longer delay. In some implementations, even if the user invokes a delay (e.g., by providing the spoken words "Can you give me some time to find it?"), the automation assistant may continue processing the audio data stream to update the current state, but determining the next interactive state to be implemented corresponds to preventing any interaction from being implemented for a threshold duration (e.g., 10 seconds, 15 seconds, 30 seconds, etc.) unless the user directs further input (e.g., spoken input, touch input, and / or typed input) to the automation assistant. In additional or alternative implementations, even if the user invokes a delay, the automation assistant may continue processing the audio data stream to update the current state and determine the next interactive state to be implemented regardless of the delay.
[0074] For example, and as shown in 466, further assuming that 15 seconds have passed since the user provided the spoken words 462, and further assuming that voice activity has been detected, as shown in 468. In this example, the automation assistant is able to determine that voice activity has been detected by processing the audio data stream using a VAD model. Furthermore, the automation assistant is able to determine whether the detected voice activity is directed at the automation assistant based on an updated NLU data stream generated from the audio data stream that captured the voice activity. Assuming that the voice activity is indeed directed at the automation assistant, the automation assistant is able to determine the next interaction state to be implemented based on the current state, and cause the next interaction state to be implemented. However, assuming that the voice activity is not directed at the automation assistant, and as relative to... Figure 4 The described audio-based feature determines, based on processing the audio data stream, that voice activity not directed at the automated assistant has been detected, and that the duration since the user last provided spoken utterances to the automated assistant is 15 seconds (or some other duration that meets the threshold duration). Therefore, at that moment, the automated assistant can determine that one or more natural dialogue outputs should be implemented as the next interaction state to prompt the user. For example, the automated assistant can cause the synthesized speech 470 included in the synthesized speech audio data to request the user's indication of whether to continue interacting with the automated assistant. Figure 4 As shown, the synthesized speech 470 can include "Still there?" as a prompt to the user to indicate whether they wish to continue the conversation with the automation assistant.
[0075] Further assuming the user provides spoken utterance 472 in response to synthesized speech 470, “Yes, the number is 123-456-7890”, the automated assistant can update the current state based on the audio data stream of the captured spoken utterance 472. For example, the NLU output stream can be updated to indicate the slot value of the phone number parameter (e.g., “123-456-7890”). As a result, the fulfillment data stream can be updated to indicate that a phone call to the “office” can be initiated. Therefore, at this moment, and based on the updated current state, the automated assistant can cause the fulfillment output to be implemented. In this embodiment, the fulfillment output to be implemented can include an assistant command on behalf of the user to initiate a phone call to the “office” (e.g., the phone number “123-456-7890”), and can optionally include synthesized speech 474, “Alright, calling the office now,” to provide the user with an indication that a phone call is being initiated. Therefore, in Figure 4In the example, the automated assistant is able to iteratively determine the next interaction state to be achieved to facilitate the conversation, rather than simply responding to each spoken word as in a turn-based conversation.
[0076] It is worth noting that although the phone call is not initiated until the moment of the conversation associated with the synthesized speech 474, the automated assistant is able to partially execute the assistant command to initiate the phone call before that moment. For example, the automated assistant can establish a connection with the phone application running in the background of the client device in response to receiving spoken words 452 (or at another moment in the conversation session) to reduce the latency of initiating the phone call when the user finally provides the slot value for the phone number parameters (e.g., as relative to...). Figure 1 (As described in part of the execution engine 165). Furthermore, although the dialogue session is depicted as a transcription via an automated assistant application, it should be understood that this is for illustrative purposes and not intended to be limiting.
[0077] Now go to Figure 5A , 5B Sections 5C and 5C describe various non-limiting examples that enable the fulfillment of outputs during a dialogue session. The automated assistant can be implemented at least partially at the client device 110B (e.g., relative to...). Figure 1 The described automated assistant 115). Assuming the automated assistant determines the action to be performed as the next interaction state, the automated assistant can utilize a natural dialogue system (e.g., relative to...). Figure 1 The described natural dialogue system (180) determines the performance output to be implemented. Figure 5A , 5B The client device 110B depicted in 5C may include various user interface components, including, for example, multiple microphones for generating audio data based on spoken words and / or other audible input, multiple speakers for audibly rendering synthesized speech and / or other audible output, and a display 190B for receiving touch input and / or visually rendering transcription and / or other visual output. Although Figure 5A , 5B The client device 110B depicted in 5C is a stand-alone interactive speaker with a display 190B, but it should be understood that this is for illustrative purposes and not intended to be limiting.
[0078] In order to run through Figure 5A , 5BFollowing the 5C example, suppose user 101 provides verbal utterance 552, “Assistant, set a timer for 15 minutes.” Further suppose the automation assistant has determined, based on the current state generated from the NLU output stream, performance data stream, and / or audio-based characteristics generated from processing the audio data stream capturing verbal utterance 552, that the next interactive state to be implemented is to cause one or more performance outputs to be implemented. However, in these examples, the one or more performance outputs implemented by the automation assistant may vary based on the NLU metrics associated with the NLU data stream used to determine the current state and / or the performance metrics associated with the performance data stream used to determine the current state.
[0079] For details, please refer to the following: Figure 5A Furthermore, it is assumed that the NLU metric, the performance metric, and / or a combination of the NLU metric and the performance metric (e.g., the average of the metric, the minimum of the metric, etc.) indicate that the automation assistant is highly confident that the user has indicated a desire to set the timer to 15 minutes. For example, it is assumed that the NLU metric and / or the performance metric satisfies a first threshold metric and a second threshold metric (or, if considered separately, satisfies a corresponding first threshold metric and a corresponding second threshold metric). Thus, the NLU metric can indicate this high confidence and the slot value of the 15-minute duration parameter in the intent associated with setting the timer, and the performance metric can indicate this high confidence to set the timer to 15 minutes in the assistant command. Therefore, the automation assistant can select the following performance outputs: 1) implement the assistant command to set the timer to 15 minutes, 2) implement the provision of graphical content for visual presentation to the user as indicated by the 15-minute timer shown at display 190B; and / or 3) implement the synthesized speech 554A “Okay, I set a timer for 15 – one five – minutes”.
[0080] For details, please refer to the following: Figure 5BFurthermore, it is assumed that NLU metrics, performance metrics, and / or some combination of NLU metrics and performance metrics (e.g., the average of the metrics, the minimum of the metrics, etc.) indicate to the automation assistant a slight degree of confidence that the user has indicated a desire to set the timer to 15 minutes. For example, suppose that the NLU metrics and / or performance metrics meet a first threshold metric but not a second threshold metric (or, if considered individually, meet the corresponding first threshold metric but not the corresponding second threshold metric). Therefore, the NLU metrics can indicate this slight degree of confidence in the intent associated with setting the timer, as well as a slot value with a duration parameter of 15 minutes (but also consider a slot value with a duration parameter of 50 minutes), and the performance metrics can indicate this slight degree of confidence in setting the timer to 15 minutes in the assistant command (but also consider an assistant command to set the timer to 50 minutes). Therefore, the automation assistant can choose to fulfill the following outputs: 1) implement the graphical content indicated by the 15-minute timer shown on display 190B for visual presentation to the user; and / or 2) implement the synthesized voice 554B “Okay, you want a timer for 15 minutes, not 50 minutes, right?”. In this example, the automation assistant may not implement the assistant command to set the timer to 15 minutes until the user has confirmed that the timer will be set to 15 minutes instead of 50 minutes, but the assistant command may have been partially fulfilled (as indicated by implementing the graphical content provided for visual presentation to the user).
[0081] For details, please refer to the following: Figure 5CFurthermore, it is assumed that the NLU metric, the performance metric, and / or a combination of the NLU metric and the performance metric (e.g., the average of the metric, the minimum of the metric, etc.) indicate to the automation assistant that it is not certain the user intends to set the timer to 15 minutes. For example, it is assumed that both the NLU metric and / or the performance metric cannot satisfy both the first threshold metric and the second threshold metric (or, if considered individually, neither the corresponding first threshold metric nor the corresponding second threshold metric can be satisfied). Therefore, the NLU metric can indicate this low confidence level and the slot value of the duration parameter of 15 minutes or 50 minutes in the intent associated with setting the timer, and the performance metric can indicate this low confidence level in the assistant command to set the timer to 15 minutes or 50 minutes. Therefore, the automation assistant can choose to fulfill the following outputs: 1) implement the graphical content to be visually presented to the user as indicated by both the 15-minute and 50-minute timers shown on display 190B; and / or 2) implement the synthesized voice 554C “A timer, sure, but was that 15 minutes or 50 minutes?”. In this example, the automation assistant may not implement the assistant command to set the timer to 15 minutes or 50 minutes until the user has confirmed that the timer will be set to 15 minutes instead of 50 minutes, but the assistant command may have been partially fulfilled (as indicated by implementing the graphical content provided for visual presentation to the user).
[0082] Therefore, even when the automated assistant determines that the next interaction state to be implemented is one that leads to the fulfillment output being achieved (rather than making the natural dialogue output audibly rendered or avoiding the interaction from being achieved), the automated assistant is still able to determine what to include as the fulfillment output. Although this describes a specific fulfillment output achieved based on a specific threshold... Figure 5A , 5B And 5C, but it should be understood that this is for illustrative purposes and not intended to be limiting. For example, multiple additional or alternative thresholds can be used to determine the performance output to be achieved. Furthermore, for example, the performance output may differ from those depicted. For example, and refer back to the reference. Figure 4 A. An automation assistant can simply set the timer to 15 minutes for the user without providing synthesized speech 554A, because the automation assistant is highly confident that the user has indicated they want the timer set to 15 minutes.
[0083] Now go to Figure 6The diagram depicts a block diagram of an example computing device 610 that can be optionally used to perform one or more aspects of the techniques described herein. In some implementations, one or more of a client device, a cloud-based automation assistant(s), and / or other components may include one or more components of the example computing device 610.
[0084] Computing device 610 typically includes at least one processor 614 that communicates with a plurality of peripheral devices via a bus subsystem 612. These peripheral devices may include a storage subsystem 624 (including, for example, a memory subsystem 625 and a file storage subsystem 626), a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow users to interact with computing device 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0085] User interface input device 622 may include a keyboard, pointing devices (such as a mouse, trackball, touchpad, or graphics tablet), scanner, touchscreen integrated into a display, audio input devices (such as a voice recognition system, microphone), and / or other types of input devices. Generally, the term "input device" is intended to encompass all possible types of devices and methods for inputting information into computing device 610 or a communication network.
[0086] User interface output device 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual displays, such as via an audio output device. Generally, the term "output device" is intended to encompass all possible types of devices and methods for outputting information from computing device 610 to a user or another machine or computing device.
[0087] Storage subsystem 624 stores the functional programming and data structures provided by some or all of the modules described herein. For example, storage subsystem 624 may include selected aspects for performing the methods disclosed herein, as well as implementations. Figure 1 and Figure 2 The logic of the various components described in the text.
[0088] These software modules are typically executed by processor 614 alone or in combination with other processors. The memory 625 used in storage subsystem 624 can include multiple memories, including main random access memory (RAM) 630 for storing instructions and data during program execution and read-only memory (ROM) 632 for storing fixed instructions. File storage subsystem 626 provides persistent storage for program and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical disk drives, or removable media cartridges. Modules implementing the functionality of certain embodiments may be stored by file storage subsystem 626 in storage subsystem 624 or in other machines accessible by processor(s) 614.
[0089] Bus subsystem 612 provides a mechanism for enabling various components and subsystems of computing device 610 to communicate with each other as intended. Although bus subsystem 612 is schematically shown as a single bus, alternative implementations of bus subsystem 612 may use multiple buses.
[0090] The computing device 610 can be of various types, including workstations, servers, computing clusters, blade servers, server groups, or any other data processing system or computing device. Due to the constantly evolving nature of computers and networks, Figure 6 The description of the computing device 610 depicted herein is intended only as a specific example for illustrating some embodiments. Many other configurations of the computing device 610 may have... Figure 6 The computing device depicted in the text has more or fewer components.
[0091] In cases where the systems described herein collect or otherwise monitor personal information about users, or may utilize personal and / or monitored information, users may be given the opportunity to control whether programs or functionalities collect user information (e.g., information about the user's social networks, social actions or activities, occupation, user preferences, or the user's current geographic location) or to control whether and / or how content that may be more relevant to the user is received from content servers. Furthermore, certain data can be processed in one or more ways before it is stored or used, resulting in the removal of personally identifiable information. For example, a user's identity may be processed to the point that the user's personally identifiable information cannot be determined, or the user's geographic location may be generalized (e.g., to the city, zip code, or state level) if geographic location information is available, making it impossible to determine the user's specific geographic location. Therefore, users may have control over how information about themselves is collected and / or used.
[0092] In some implementations, a method implemented by one or more processors is provided, comprising: processing an audio data stream using an Automatic Speech Recognition (ASR) model to generate an ASR output stream, the audio data stream being generated by one or more microphones of a client device, and the audio data stream capturing one or more spoken utterances of a user, the one or more spoken utterances being directed to an automated assistant at least partially implemented at the client device; processing the ASR output stream using a Natural Language Understanding (NLU) model to generate an NLU output stream; generating a performance data stream based on the NLU output stream; determining, based on processing the audio data stream, an audio-based feature associated with one or more spoken utterances; determining, based on a current state of the NLU output stream, the performance data stream, and the audio-based feature associated with one or more spoken utterances, whether a next interaction state to be implemented is: (i) enabling a performance output generated based on the performance data stream to be implemented, (ii) enabling a natural dialogue output to be audibly rendered to be presented to the user, or (iii) avoiding enabling any interaction; and enabling the next interaction state to be implemented.
[0093] These and other embodiments of the technology disclosed herein may optionally include one or more of the following features.
[0094] In some implementations, determining whether the next interaction state to be implemented, based on the current state of the NLU output stream, the performance data stream, and the audio-based features associated with one or more of the spoken utterances, is (i) to enable the performance output generated based on the performance data stream, (ii) to enable the natural dialogue output to be audibly rendered to be presented to the user, or (iii) to prevent any interaction from being implemented, may include: using a classification machine learning (ML) model to process the current state of the NLU output stream, the performance data stream, and the audio-based features to generate a corresponding predictive response metric associated with each of the following: (i) enabling the performance output generated based on the performance data stream, (ii) enabling the natural dialogue output to be audibly rendered to be presented to the user, and (iii) preventing any interaction from being implemented; and determining whether the next interaction state to be implemented is to (i) enable the performance output generated based on the performance data stream, (ii) enable the natural dialogue output to be audibly rendered to be presented to the user, or (iii) prevent any interaction from being implemented, based on the corresponding predictive response metric.
[0095] In some implementations, the method may further include: determining a set of performance outputs based on the performance data stream; and selecting the performance outputs from the set of performance outputs based on a predicted NLU metric associated with the NLU data stream and / or a predicted performance metric associated with the performance data stream.
[0096] In some versions of those implementations, selecting the performance output from the set of performance outputs may be in response to determining that the next interaction state to be implemented is (i) such that the performance output generated based on the performance data stream is implemented. Selecting the performance output from the set of performance outputs based on the predicted NLU metric associated with the NLU data stream and / or the predicted performance metric associated with the performance data stream may include: selecting a first performance output as the performance output from the set of performance outputs in response to determining that the predicted NLU metric associated with the NLU data stream and / or the predicted performance metric associated with the performance data stream satisfies both a first threshold metric and a second threshold metric.
[0097] In some versions of those implementations, selecting the performance output from the set of performance outputs based on the predicted NLU metric associated with the NLU data stream and / or the predicted performance metric associated with the performance data stream may further include selecting a second performance output as the performance output from the set of performance outputs in response to determining that both the predicted NLU metric associated with the NLU data stream and / or the predicted performance metric associated with the performance data stream satisfy the first threshold metric but not the second threshold metric. The second performance output may be different from the first performance output.
[0098] In some additional or alternative versions of those implementations, selecting the performance output from the set of performance outputs based on the predicted NLU metric associated with the NLU data stream and / or the predicted performance metric associated with the performance data stream may further include selecting a third performance output as the performance output from the set of performance outputs in response to determining that the predicted NLU metric associated with the NLU data stream and / or the predicted performance metric associated with the performance data stream fails to satisfy both the first threshold metric and the second threshold metric. The third performance output may be different from both the first performance output and the second performance output.
[0099] In some versions of those implementations, determining the set of performance outputs based on the performance data stream may include processing the performance data using multiple first-party agents to generate corresponding first-party performance outputs; and incorporating the corresponding first-party performance outputs into the set of performance outputs. In some other versions of those implementations, determining the set of performance outputs based on the performance data stream may further include transmitting the performance data from the client device to one or more third-party systems via one or more networks; receiving the corresponding third-party performance outputs at the client device and from the one or more third-party systems via one or more of the networks; and incorporating the corresponding third-party performance outputs generated using multiple first-party verticals into the set of performance outputs. Transmitting the performance data to the one or more third-party systems causes each of the third-party systems to process the performance data using a corresponding third-party agent to generate a corresponding third-party performance output.
[0100] In some versions of those implementations, the method may further include: before any synthesized speech corresponding to the performance output generated based on the performance data stream is audibly rendered via one or more speakers of the client device: initiating partial performance for one or more of the performance outputs in the set of performance outputs, wherein the partial performance is specific to each of the performance outputs in the set of performance outputs.
[0101] In some implementations, the performance output may include one or more of the following: synthesized speech audio data, including synthesized speech corresponding to the performance output, to be audibly rendered to be presented to the user via one or more of the speakers of the client device; graphical content corresponding to the performance output, to be visually rendered to be presented to the user via the display of the client device or an additional client device communicating with the client device; or an assistant command corresponding to the performance output, which, when executed, causes the automation assistant to control the client device or an additional client device communicating with the client device.
[0102] In some implementations, the method may further include: maintaining a set of natural dialogue outputs in one or more databases accessible by the client device; and selecting the natural dialogue output from the set of natural dialogue outputs based at least on audio-based features associated with one or more of the spoken utterances. In some versions of those implementations, selecting the natural dialogue output from the set of natural dialogue outputs may be in response to determining that the next interaction state to be implemented is (ii) causing the natural dialogue output to be audibly rendered for presentation to the user. Selecting the natural dialogue output from the set of natural dialogue outputs may include selecting a first natural dialogue output as the natural dialogue output based on a prediction metric associated with the NLU output and the audio-based features associated with one or more of the spoken utterances. In some other versions of those implementations, selecting the natural dialogue output from the set of natural dialogue outputs may further include selecting a second natural dialogue output as the natural dialogue output based on the prediction metric associated with the NLU output and the audio-based features associated with one or more of the spoken utterances. The second natural dialogue output may be different from the first natural dialogue output.
[0103] In some embodiments, the method may further include, in response to determining that the next interaction state to be implemented is (iii) avoiding any interaction from being implemented: processing the audio data stream using a voice activity detection model to monitor the occurrence of voice activity; and in response to detecting the occurrence of the voice activity: determining whether the occurrence of the voice activity is directed at the automation assistant. In some versions of those embodiments, determining whether the occurrence of the voice activity is directed at the automation assistant may include: processing the audio data stream using the ASR model to continue generating the ASR output stream; processing the ASR output stream using the NLU model to continue generating the NLU output stream; and determining whether the occurrence of the voice activity is directed at the automation assistant based on the NLU output stream. In some other versions of those embodiments, the method may further include, in response to determining that the occurrence of the voice activity is directed at the automation assistant: continuing to generate the performance data stream based on the NLU output stream; and updating the performance output set from which the performance output is selected based on the performance data stream. In additional or alternative implementations of those embodiments, the method may further include, in response to determining that the occurrence of the voice activity is not directed at the automation assistant: determining whether a threshold duration has elapsed since the user provided the one or more spoken utterances; and in response to determining that the threshold duration has elapsed since the user provided the one or more spoken utterances: determining that the next interaction state to be implemented is (ii) such that natural dialogue output is audibly rendered to be presented to the user. In yet another version of those embodiments, the natural dialogue output to be audibly rendered to be presented to the user via one or more speakers of the client device may include one or more of the following: an indication of whether to continue interacting with the automation assistant; or a request from the user to provide additional user input to facilitate a conversational session between the user and the automation assistant.
[0104] In some implementations, the audio-based characteristics associated with one or more of the spoken utterances may include one or more of the following: one or more prosodic attributes associated with each of the one or more spoken utterances, wherein the one or more prosodic attributes include one or more of the following: intonation, pitch, stress, rhythm, beat, pitch, and rest; or the duration that has elapsed since the user provided the most recent spoken utterance among the one or more spoken utterances.
[0105] In some implementations, the current state of the NLU output stream, the performance data stream, and the audio-based features associated with one or more of the spoken utterances may include a recent instance of the NLU output generated based on the most recent spoken utterance among the one or more spoken utterances, a recent instance of the performance data generated based on the most recent NLU output, and a recent instance of the audio-based features generated based on the most recent spoken utterance. In some versions of those implementations, the current state of the NLU output stream, the performance data stream, and the audio-based features associated with one or more of the spoken utterances may further include one or more historical instances of the NLU output generated based on one or more historical spoken utterances preceding the most recent spoken utterance, one or more historical instances of the performance data generated based on the one or more historical instances of the NLU output, and one or more historical instances of the audio-based features generated based on the one or more historical spoken utterances.
[0106] In some implementations, a method implemented by one or more processors is provided, comprising: processing an ASR output stream using an Automatic Speech Recognition (ASR) model, the audio data stream being generated by one or more microphones of a client device, and the audio data stream capturing one or more spoken utterances of a user, the one or more spoken utterances being directed to an automated assistant implemented at least partially at the client device; processing the ASR output stream using a Natural Language Understanding (NLU) model to generate an NLU output stream; generating a performance data stream based on the NLU output stream; determining, based on processing the audio data stream, audio-based features associated with one or more spoken utterances; and, based on the NLU output stream, the... The system describes the current state of the data stream and the audio-based features associated with one or more spoken utterances, and determines (i) when to allow synthesized speech audio data, including synthesized speech, to be audibly rendered to be presented to the user, and (ii) what is included in the synthesized speech; and in response to determining that the synthesized speech audio data should be audibly rendered to be presented to the user: the synthesized speech audio data should be audibly rendered via one or more speakers of the client device; and in response to determining that the synthesized speech audio data should not be audibly rendered to be presented to the user: the synthesized speech audio data should be prevented from being audibly rendered via one or more speakers of the client device.
[0107] Additionally, some implementations include one or more processors (e.g., multiple central processing units (CPUs), multiple graphics processing units (GPUs), and / or multiple tensor processing units (TPUs)) of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in associated memory, and wherein the instructions are configured to cause any of the methods described above to be performed. Some implementations also include one or more non-transitory computers capable of reading storage media containing computer instructions executable by the one or more processors to implement any of the methods described above. Some implementations also include a computer program product comprising instructions executable by the one or more processors to implement any of the methods described above.
Claims
1. A method implemented by one or more processors, the method comprising: An audio data stream is processed using an automatic speech recognition (ASR) model to generate an ASR output stream, the audio data stream being generated by one or more microphones of a client device, and the audio data stream capturing one or more spoken words of a user, the one or more spoken words being directed to an automated assistant implemented at least in part at the client device; The ASR output stream is processed using a Natural Language Understanding (NLU) model to generate an NLU output stream; The execution data stream is generated based on the NLU output stream; Based on processing the audio data stream, determine audio-based features associated with one or more spoken utterances in the spoken utterances; Based on the current state of the NLU output stream, the performance data stream, and the audio-based characteristics associated with one or more spoken utterances, determine whether the next interaction state to be implemented is: (i) to enable the performance output generated based on the performance data stream to be implemented. (ii) to render the natural dialogue output audibly for presentation to the user, or (iii) Avoid implementing any interactive state; and Before determining the next interaction state to be implemented: Based on the performance data stream, multiple candidate given performance outputs are determined that are predicted to satisfy the user's one or more verbal utterances; as well as As the user continues to provide the one or more verbal statements: For each of the plurality of candidate given performance outputs, a corresponding partial performance is initiated, wherein the corresponding partial performance of the plurality of candidate given performance outputs includes at least: For a first candidate given execution output from the plurality of candidate given execution outputs, establish a corresponding connection with one of the following: a given software application accessible to the client device, a given first-party agent, a given third-party agent, or a given additional client device added to the client device; and For a second candidate given execution output from the plurality of candidate given execution outputs, establish a corresponding connection with another of the following: the given software application accessible to the client device, the given first-party agent, the given third-party agent, or the given additional client device added to the client device; An update stream for execution data is generated based on the NLU output stream; and Based on the update stream of the performance data, a given performance output to be implemented from the plurality of candidate given performance outputs is determined; and In response to determining the next interaction state to be implemented: This enables the next interactive state to be realized.
2. The method according to claim 1, wherein, Based on the NLU output stream, the performance data stream, and the current state of the audio-based characteristics associated with one or more of the spoken utterances, determining whether the next interaction state to be implemented is (i) to enable performance output generated based on the performance data stream, (ii) to enable natural dialogue output to be audibly rendered for presentation to the user, or (iii) to prevent any interaction from being implemented includes: A classification machine learning (ML) model is used to process the NLU output stream, the performance data stream, and the current state based on the audio-based features to generate a corresponding predictive response metric associated with each of the following: (i) enabling the performance output generated based on the performance data stream to be implemented, (ii) enabling the natural dialogue output to be audibly rendered for presentation to the user, and (iii) avoiding enabling any interaction; and Based on the corresponding predicted response metric, determine whether the next interaction state to be implemented should (i) enable the performance output generated based on the performance data stream to be implemented, (ii) enable the natural dialogue output to be audibly rendered to be presented to the user, or (iii) avoid enabling any interaction to be implemented.
3. The method according to claim 1, wherein, Based on the performance data stream, multiple candidate given performance outputs predicted to satisfy the user's one or more verbal utterances are determined, including: Determine the performance output set based on the performance data stream; and Based on the predicted NLU metric associated with the NLU output stream and / or the predicted performance metric associated with the performance data stream, the plurality of candidate given performance outputs are selected from the set of performance outputs.
4. The method according to claim 3, wherein, The plurality of candidate given performance outputs are selected from the set of performance outputs based on the predicted NLU metric associated with the NLU output stream and / or the predicted performance metric associated with the performance data stream, including: In response to determining that the predicted NLU metric associated with the NLU output stream and / or the predicted performance metric associated with the performance data stream satisfies both a first threshold metric and a second threshold metric, a first candidate given performance output is selected from the set of performance outputs to initiate the corresponding partial performance.
5. The method according to claim 4, wherein, Selecting the plurality of candidate given performance outputs from the set of performance outputs based on the predicted NLU metric associated with the NLU output stream and / or the predicted performance metric associated with the performance data stream further includes: In response to determining that the predicted NLU metric associated with the NLU output stream and / or the predicted fulfillment metric associated with the fulfillment data stream satisfies the first threshold metric but not the second threshold metric, a second candidate given fulfillment output is selected from the set of fulfillment outputs to initiate the corresponding partial fulfillment. The second candidate given execution output is different from the first candidate given execution output.
6. The method according to claim 3, wherein, Determining the set of performance outputs based on the performance data stream includes: Multiple first-party agents are used to process the performance data to generate corresponding first-party performance outputs; and The corresponding first-party performance output is incorporated into the performance output set.
7. The method according to claim 6, wherein, Determining the performance output set based on the performance data stream further includes: The performance data is transmitted from the client device to one or more third-party systems via one or more networks, wherein transmitting the performance data to the one or more third-party systems causes each of the third-party systems to use a corresponding third-party agent to process the performance data to generate a corresponding third-party performance output; Through one or more of the networks, the client device receives the corresponding third-party performance output from the one or more third-party systems; and The corresponding third-party fulfillment outputs generated using multiple first-party agents are incorporated into the fulfillment output set.
8. The method according to claim 1, wherein, The given performance output to be achieved includes one or more of the following: Synthetic speech audio data, including synthesized speech corresponding to the given performance output, will be audibly rendered and presented to the user via one or more speakers of the client device. The graphical content corresponding to the given performance output will be visually rendered to be presented to the user via the display of the client device or an additional client device communicating with the client device, or An assistant command corresponding to the given performance output, when executed, causes the automation assistant to control the client device or an additional client device communicating with the client device.
9. The method of claim 1, further comprising: Maintain a set of natural dialogue outputs in one or more databases accessible by the client device; as well as Natural dialogue outputs are selected from the set of natural dialogue outputs, based at least on the audio-based characteristics associated with one or more of the spoken utterances.
10. The method according to claim 9, wherein, Selecting the natural dialogue output from the set of natural dialogue outputs is in response to determining that the next interaction state to be implemented is (ii) causing the natural dialogue output to be audibly rendered to be presented to the user, and wherein selecting the natural dialogue output from the set of natural dialogue outputs includes: Based on the prediction metric associated with the NLU output and the audio-based features associated with one or more spoken utterances, a first natural dialogue output is selected from the set of natural dialogue outputs as the natural dialogue output.
11. The method according to claim 10, wherein, Selecting the natural dialogue output from the set of natural dialogue outputs further includes: Based on the prediction metric associated with the NLU output and the audio-based features associated with one or more spoken utterances, a second natural dialogue output is selected from the set of natural dialogue outputs as the natural dialogue output. The second natural dialogue output is different from the first natural dialogue output.
12. The method of claim 1, further comprising: In response to determining the next interaction state to be implemented, (iii) avoid making any interaction a reality: The audio data stream is processed using a voice activity detection model to monitor the occurrence of voice activity; and In response to the detection of the voice activity: Determine whether the occurrence of the voice activity is directed at the automated assistant.
13. The method according to claim 12, wherein, Determining whether the occurrence of the voice activity is directed at the automated assistant includes: The ASR model is used to process the audio data stream to continue generating the ASR output stream; The NLU model is used to process the ASR output stream to continue generating the NLU output stream; and The NLU output stream is used to determine whether the voice activity is directed at the automation assistant.
14. The method of claim 13, further comprising: In response to determining that the voice activity has occurred, the automated assistant: The execution data stream is then generated based on the NLU output stream. as well as The set of performance outputs from which the given performance output to be implemented is updated based on the performance data stream.
15. The method of claim 13, further comprising: In response to determining that the voice activity did not occur in relation to the automated assistant: Determine whether a threshold duration has elapsed since the user delivered the one or more spoken utterances; as well as In response to determining that the threshold duration has elapsed since the user delivered the one or more spoken utterances: The next interaction state to be achieved is (ii) such that the natural dialogue output is audibly rendered to be presented to the user.
16. The method according to claim 15, wherein, The natural dialogue output to be audibly rendered and presented to the user via one or more speakers of the client device includes one or more of the following: An indication of whether to continue interacting with the automated assistant; or The user provides additional user input to facilitate a conversational session between the user and the automated assistant.
17. A system comprising: At least one processor; as well as A memory storing instructions that, when executed, cause the at least one processor to: An automatic speech recognition (ASR) model is used to process an audio data stream to generate an ASR output stream, the audio data stream being generated by one or more microphones of a client device, and the audio data stream capturing one or more spoken words of a user, the one or more spoken words being directed to an automated assistant implemented at least in part at the client device; The ASR output stream is processed using a Natural Language Understanding (NLU) model to generate an NLU output stream; The execution data stream is generated based on the NLU output stream; Based on processing the audio data stream, determine audio-based features associated with one or more spoken utterances in the spoken utterances; Based on the current state of the NLU output stream, the performance data stream, and the audio-based characteristics associated with one or more spoken utterances, determine whether the next interaction state to be implemented is: (i) to enable the performance output generated based on the performance data stream to be implemented. (ii) to render the natural dialogue output audibly for presentation to the user, or (iii) Avoid implementing any interactive state; and Before determining the next interaction state to be implemented: Based on the performance data stream, multiple candidate given performance outputs are determined that are predicted to satisfy one or more of the user's verbal utterances; and As the user continues to provide the one or more spoken statements: For each of the plurality of candidate given performance outputs, a corresponding partial performance is initiated, wherein the corresponding partial performance of the plurality of candidate given performance outputs includes at least: For a first candidate given execution output from the plurality of candidate given execution outputs, establish a corresponding connection with one of the following: a given software application accessible to the client device, a given first-party agent, a given third-party agent, or a given additional client device added to the client device; and For a second candidate given execution output from the plurality of candidate given execution outputs, establish a corresponding connection with another of the following: the given software application accessible to the client device, the given first-party agent, the given third-party agent, or the given additional client device added to the client device; An update stream for execution data is generated based on the NLU output stream; and Based on the update stream of the performance data, a given performance output to be implemented from the plurality of candidate given performance outputs is determined; and In response to determining the next interaction state to be implemented: This enables the next interactive state to be realized.
18. A system comprising: At least one processor; as well as A memory for storing instructions, which, when executed, cause the at least one processor to perform an operation corresponding to any one of claims 1 to 16.
19. A non-transitory computer-readable storage medium storing instructions, which, when executed, cause at least one processor to perform an operation corresponding to any one of claims 1 to 16.
Citation Information
Patent Citations
Parsing to determine interruptible state in an utterance by detecting pause duration and complete sentences
US10832005B1
Natural assistant interaction
US20190295544A1
Intelligent digital assistant in a multi-tasking environment
US20200118568A1
Utilizing pre-event and post-event input streams to engage an automated assistant
WO2020171809A1