Enabling natural conversations using soft endpointing for an automated assistant
A natural conversation system for automated assistants processes audio data to determine utterance completion, addressing inefficiencies in turn-based dialog by pausing and completing actions only when user input is complete, enhancing efficiency and accuracy.
Patent Information
- Application Number
- JP2023569701
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-22
- Filing Date
- 2021-11-29
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2041-11-29
AI Technical Summary
Turn-based dialog sessions with automated assistants are often unnatural and inefficient, as they fail to consider the context of multiple utterances and may waste computational resources due to premature responses or incomplete utterances, leading to lengthy interactions.
Implement a natural conversation system that processes a stream of audio data using streaming ASR and NLU models to determine audio-based features, allowing the assistant to pause and provide natural conversation outputs or fulfill user intents based on the completion of utterances, reducing computational waste and enhancing interaction efficiency.
Enables more natural and efficient dialog sessions by allowing the assistant to wait for user completion, reducing computational resource waste and shortening interaction time, while improving response accuracy and user satisfaction.
Smart Images

Figure 0007705481000001 
Figure 0007705481000002 
Figure 0007705481000003
Abstract
Description
Background Art
[0001] Humans can participate in human-computer dialogues with interactive software applications, which are referred to in this specification as "automatic assistants" (also called "chatbots", "conversational personal assistants", "intelligent personal assistants", "personal voice assistants", "conversational agents", etc.). Automatic assistants typically rely on a pipeline of components in the interpretation and response of utterances (or touch / type inputs). For example, an automatic speech recognition (ASR) engine can process audio data corresponding to a user's utterance to generate an ASR output such as an utterance or a speech hypothesis of phonemes expected to correspond to the utterance (i.e., a sequence of terms and / or other tokens). Further, a natural language understanding (NLU) engine can process the ASR output (or touch / type input) to generate an NLU output such as the user's intent in providing the utterance (or touch / type input) and the slot values of parameters optionally associated with the intent. Further, a fulfillment engine can be used to process the NLU output and generate a fulfillment output such as a structured request to obtain the response content for the utterance and / or execute an action in response to the utterance, and a stream of fulfillment data can be generated based on the fulfillment output.
[0002] Generally, a dialog session with an automatic assistant is initiated by the user providing an utterance, and the automatic assistant can respond to the utterance using the pipeline of the aforementioned components to generate a response. The user can continue the dialog session by providing additional utterances, and the automatic assistant can respond to the additional utterances using the pipeline of the aforementioned components to generate additional responses. In other words, these dialog sessions are generally turn-based in that the user accepts the dialog session to provide an utterance and the automatic assistant accepts the dialog session to respond to the utterance when the user stops speaking. However, these turn-based dialog sessions may not be natural as they do not reflect how humans actually converse with each other from the user's perspective.
[0003] For example, a first human may provide a plurality of different utterances to convey a single idea to a second human, and the second human may consider each of the plurality of different utterances in order to devise a response to the first human. In some cases, the first human may pause for various amounts of time between these plurality of different utterances (or between various amounts of time providing a single utterance). In particular, the second human may not be able to fully devise a response to the first human based solely on the first utterance (or a part thereof) of the plurality of different utterances or based on each of the plurality of different separate utterances.
[0004] Similarly, in these turn-based dialog sessions, the automated assistant may not be able to fully devise a response to a given user utterance (or a part thereof) without considering the context of the given utterance with respect to multiple different utterances, or without waiting for the user to complete the provision of the given utterance. As a result, these turn-based dialog sessions can be lengthy as the user attempts to convey their thoughts to the automated assistant in a single utterance within a single turn of these turn-based dialog sessions, thereby wasting computational resources. Further, if the user attempts to convey their thoughts to the automated assistant in multiple utterances within a single turn of these turn-based dialog sessions, the automated assistant may simply fail, thereby also wasting computational resources. For example, if the automated assistant causes a long pause when the user attempts to formulate an utterance, the automated assistant may fail as a result of prematurely concluding that the user has finished speaking, processing an incomplete utterance, and determining (from the processing) that the incomplete utterance does not convey a meaningful intent, or determining (from the processing) an incorrect intent conveyed by the incomplete utterance. In addition, turn-based dialog sessions may prevent the user utterance provided during the rendering of the assistant response from being processed in a meaningful way. This may require the user to wait for the completion of the rendering of the assistant response before providing an utterance, thereby lengthening the dialog session. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0005] The implementations described in this specification are directed to enabling an automated assistant to conduct a natural conversation with a user during a dialog session. Some implementations can process a stream of audio data generated by a user's client microphone to generate a stream of ASR output, for example, using a streaming automatic speech recognition (ASR) model. The stream of audio data can capture portions of the user's utterances directed to an automated assistant implemented at least partially on the client device. Further, the ASR output can be processed using an NLU model to generate a stream of natural language understanding (NLU) output. Further, the NLU output can be processed using one or more fulfillment rules and / or one or more fulfillment models to generate a stream of fulfillment data. Additionally, based on processing the stream of audio data, it is possible to determine audio-based features associated with one or more of the utterances. Audio-based features associated with a portion of an utterance can include, for example, intonation, tone, stress, rhythm, tempo, pitch, elongated syllables, pauses, grammar associated with the pauses, and / or other audio-based features that can be derived from processing the stream of audio data. Based on the stream of NLU output and / or the audio-based features, the automated assistant can determine whether the user has paused or completed providing an utterance (e.g., soft endpointing).
[0006] In some implementations, in response to determining that the user has paused providing speech, the automatic assistant can provide a natural conversation output for presentation to the user (even if the automatic assistant determines that speech fulfillment can be performed in various implementations) to indicate that the automatic assistant is waiting for the user to complete providing speech. In some implementations, in response to determining that the user has completed providing speech, the automatic assistant can be provided to present a fulfillment output to the user. Thus, by determining whether the user has paused or completed providing speech, the automatic assistant does not simply respond to the user after the user has paused providing speech as in a turn-based dialog session, but rather, based on what the user said and how the user said it, the user can naturally wait for the user to complete their thoughts.
[0007] For example, assume that a user is participating in a dialog session with an automated assistant and provides an utterance such as "call Arnolllld's". When the user provides the utterance, it is possible for the streams of ASR output, NLU output, and full fulfillment data to be generated based on processing a stream of audio data that captures the utterance. In particular, in this example, at the time the utterance is received, the stream of ASR output may include the recognized text corresponding to the utterance (e.g., "call Arnold's"), the stream of NLU output may include a predicted "call" or "phone call" intent having a slot value for "Arnold" regarding the incoming call parameters associated with the predicted "call" or "phone call" intent, and the stream of full fulfillment data, when executed as a full fulfillment output, may include an assistant command to initiate a call with the user's contact entry associated with the entity reference "Arnold" to a client device or an additional client device communicating with the client device. Further, audio-based features associated with the utterance may be generated based on processing the stream of audio data and may include, for example, elongated syllables (such as those indicated by "llll" in "call Arnolllld's") indicating that it is unclear exactly what the user intends regarding the incoming call parameters. Thus, in this example, although the automated assistant may be able to fulfill the utterance by initiating a call (e.g., to the contact entry "Arnold" on a client device or an additional client device) based on the stream of NLU data, the automated assistant may refrain from fulfilling the utterance based on the audio-based features to determine that the user paused and provide additional time for the user to complete the utterance.
[0008] Rather, in this example, the automatic assistant can determine to provide a natural conversation output for presentation to the user. For example, in response to determining that the user has paused the provision of utterances (and optionally, after the user has paused for a threshold duration), the automatic assistant can cause a natural conversation output such as "Mmhmm" or "Uh huhh" (or other vocal cues) to be provided to the user via the speaker of the client device to indicate that the automatic assistant is waiting for the user to complete the provision of utterances. In some cases, the volume of the natural conversation output provided for audible presentation to the user can be made lower than other audible outputs provided for presentation to the user. Additionally or alternatively, in an implementation where the client device includes a display, the client device can render one or more graphical elements, such as a streaming transcription of the utterance, along with an ellipse that bounds to indicate that the automatic assistant is waiting for the user to complete the provision of utterances. Additionally or alternatively, in an implementation where the client device includes one or more light emitting diodes (LEDs), the client device can turn on one or more of the LEDs to indicate that the automatic assistant is waiting for the user to complete the provision of utterances. In particular, while the natural conversation output is being provided for audible presentation to the user of the client device, one or more automatic assistant components (e.g., ASR, NLU, fulfillment, and / or other components) can remain active to continue processing the stream of audio data.
[0009] In this example, while the natural conversation output is being provided for audible presentation, or after the natural conversation output has been provided for audible presentation, it is further assumed that the user provides an utterance of "Arnold's Trattoria" to complete the offering of the previous utterance, resulting in an utterance of "call Arnold's Trattoria", where "Arnold's Trattoria" is a fictional Italian restaurant. Thus, the ASR output stream, the NLU output stream, and the fulfillment data stream can be updated based on the user completing the utterance. In particular, the NLU output stream may still include the predicted intent of "call" or "phone call", but has a slot value for "Arnold's Trattoria" for the callee parameter associated with the predicted intent of "call" or "phone call" (e.g., not the contact entry "Arnold"), and the fulfillment data stream, when executed as a fulfillment output, can include an assistant command to initiate a call with the restaurant associated with the entity reference "Arnold's Trattoria" to the client device or an additional client device communicating with the client device. Further, the automated assistant can initiate a call to the client device or an additional client device communicating with the client device in response to determining that the utterance has been completed.
[0010] In contrast, it is further assumed that after natural conversation output has been provided for audible presentation (and optionally, during a threshold duration after natural conversation output has been provided for audible presentation), the user did not provide any utterance to complete the provision of the previous utterance. In this example, the automated assistant can determine additional natural conversation output to be provided for audible presentation to the user. However, the additional natural conversation can be explicitly requested by the user of the client device to complete an utterance (e.g., "You were saying?", "Did I miss something?", etc.), or the user of the client device can be explicitly requested to provide a specific slot value regarding a predicted intent (e.g., "Who did you want to call?", etc.). In some implementations, then, assuming the user provides an utterance of "Arnold's Trattoria" to complete the provision of the previous utterance, the streams of ASR output, NLU output, and fulfillment output can be updated, and the automated assistant can fulfill the utterance (e.g., by causing the client device to initiate a call with a restaurant associated with the entity reference "Arnold's Trattoria") as described above.
[0011] In an additional or alternative implementation, assuming the client device includes a display, the automatic assistant can provide a plurality of selectable graphical elements for visual presentation to the user, and each of the selectable graphical elements is associated with a different interpretation of one or more portions of the utterance. In this example, the automatic assistant, when selected, can provide a first selectable graphical element that causes the automatic assistant to initiate a call to the restaurant "Arnold's Trattoria" and a second selectable graphical element that causes the automatic assistant to initiate a call to the contact entry "Arnold" when selected. The automatic assistant can then initiate a call based on receipt of a user selection of a given one of the selectable graphical elements, or if the user does not select one of the selectable graphical elements within a threshold duration during which one or more selectable graphical elements are presented, can initiate a call based on the NLU measurement associated with the interpretation. For example, in this example, the automatic assistant can initiate a call to the restaurant "Arnold's Trattoria" if the user does not provide a selection of one or more of the selectable graphical elements within 5 seconds, 7 seconds, or any other threshold duration after one or more selectable graphical elements are provided for presentation to the user.
[0012] As another example, assume that a user participates in a dialog session with an automated assistant and provides an utterance of "Is there anything on my calendar forrrr (what's on my calendar forrrr)". When the user provides the utterance, the streams of ASR output, NLU output, and fulfillment data can be generated based on processing the stream of audio data that captures the utterance. In particular, in this example, at the time the utterance is received, the stream of ASR output may include the recognized text corresponding to the utterance (e.g., "Is there anything on my calendar for (what's on my calendar for)"), the stream of NLU output may include a predicted "calendar" or "calendar search" intent having an unknown slot value for the date parameter associated with the predicted intent, and the stream of fulfillment data may include an assistant command that, when executed as a fulfillment output, causes the client device to search for the user's calendar information. Similarly, the audio-based features associated with the utterance can be generated based on processing the stream of audio data and may include, for example, elongated syllables (such as those indicated by "rrrr" in "Is there anything on my calendar forrrr (what's on my calendar forrrr)") that indicate the user's lack of confidence about the date parameter. Thus, in this example, the automated assistant may be unable to fulfill the utterance based on the stream of NLU data (e.g., based on the unknown slot value) and / or the audio-based features of the utterance, and the automated assistant may refrain from fulfilling the utterance based on the audio-based features to determine that the user has paused and to provide additional time for the user to complete the utterance.
[0013] Similarly, in this example, the automatic assistant can determine to provide a natural conversation output for presentation to the user. For example, in response to determining that the user has paused providing an utterance (and optionally, after the user has paused for a threshold duration), the automatic assistant can provide an audible prompt to the user via the speaker of the client device, such as "Mmhmm" or "Uh huhh", to indicate that the automatic assistant is waiting for the user to complete providing the utterance and / or to indicate other instructions that the automatic assistant is waiting for the user to complete providing the utterance. However, further assume that after the natural conversation output has been provided for audible prompting (and optionally, for a threshold duration after the natural conversation output has been provided for audible prompting), the user has provided nothing to complete the previous utterance. In this example, the automatic assistant simply infers the current date slot value for the unknown date parameter associated with the predicted "Calendar" or "Calendar Search" intent, and can fulfill the utterance by providing (e.g., audibly and / or visually) calendar information regarding the current date to the user, even though the user has not completed the utterance. In additional or alternative implementations, the automatic assistant can utilize one or more additional or alternative automatic assistant components to resolve any ambiguity in an utterance, to confirm full fulfillment of any utterance, and / or to perform any other action before fulfilling any assistant command.
[0014] In various implementation forms such as the latter example where the user first provides the utterance "What's on my calendar forrrr", in contrast to the former example where the user first provides the utterance "Call Arnolllld's", the automatic assistant can determine one or more computational costs associated with fulfilling the utterance to be fulfilled and / or canceling the full fulfillment of the utterance if the utterance is erroneously fulfilled. For example, in the former example, the computational cost associated with fulfilling the utterance may at least include initiating a call with the contact entry "Arnold", and the computational cost associated with canceling the full fulfillment of the utterance may at least include ending the call with the contact entry associated with "Arnold", restarting the dialog session with the user, processing additional utterances, and initiating another call with the restaurant "Arnold's Trattoria". Further, in the former example, one or more user costs associated with initiating an unintended call may be relatively high. Also, for example, in the latter example, the computational cost associated with fulfilling the utterance may at least include providing calendar information regarding the current date for presentation to the user, and the computational cost associated with canceling the full fulfillment of the utterance may include providing calendar information regarding another date specified by the user for presentation to the user. Further, in the latter example, one or more user costs associated with providing incorrect calendar information to the user may be relatively low. In other words, the computational costs associated with fulfillment (and canceling fulfillment) in the former example are relatively higher than the computational costs associated with fulfillment (and canceling fulfillment) in the latter example.Thus, in the latter example, the automatic assistant may determine to fulfill an utterance using a date parameter inferred based on the latter's computational cost in an attempt to end the dialog session more quickly and efficiently, but not so in the former example due to the former's computational cost.
[0015] By using the techniques described herein, one or more technical advantages can be achieved. As one non-limiting example, the techniques described herein enable an automatic assistant to participate in a natural conversation with a user during a dialog session. For example, the automatic assistant can determine whether the user has paused or completed providing an utterance so that the automatic assistant is not limited to turn-based dialog sessions or does not depend on determining that the user has finished speaking before responding to the user, and can adapt the output provided for presentation to the user accordingly. Thus, the automatic assistant can determine when to respond to the user and how to respond to the user when the user participates in these natural conversations. This results in various technical advantages, such as saving computational resources on the client device and being able to end the dialog session more quickly and efficiently. For example, the automatic assistant can wait for more information from the user before attempting to perform any fulfillment on behalf of the user (even if the automatic assistant predicts that the fulfillment should be performed), so that the number of times the automatic assistant fails can be reduced. Also, for example, the amount of user input received on the client device can be reduced because the number of times the user has to repeat the same thing or the number of times the automatic assistant has to be recalled can be reduced.
[0016] As used herein, a "dialog session" may include a logically self - contained exchange between a user and an automated assistant (optionally with other human participants). The automated assistant may distinguish between multiple dialog sessions with a user based on various signals such as the passage of time between sessions, changes in the user context between sessions (e.g., location, before / during / after a scheduled meeting, etc.), detection of one or more intervening interactions between the user and the client device other than the dialog between the user and the automated assistant (e.g., the user switches applications for a while, the user leaves and then returns to a stand - alone voice - activated product), locking / sleeping of the client device between sessions, changes in the client device used to interact with the automated assistant, etc.
[0017] The above description is provided as an overview of only some of the implementations disclosed herein. Those implementations and others are described in more detail herein.
[0018] It should be understood that the techniques disclosed herein can be implemented locally on a client device, remotely by a server connected to the client device via one or more networks, and / or both.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 5E
Figure 6
Embodiments for Carrying Out the Invention
[0020] Proceeding now to FIG. 1, a block diagram of an exemplary environment is shown that demonstrates various aspects of the present disclosure and in which the implementation forms disclosed herein can be implemented. The exemplary environment includes a client device 110 and a natural conversation system 180. In some implementation forms, the natural conversation system 180 can be implemented locally on the client device 110. In additional or alternative implementation forms, the natural conversation system 180 can be implemented remotely from the client device 110 (e.g., on a remote server), as shown in FIG. 1. In these implementation forms, the client device 110 and the natural conversation system 180 can be communicatively coupled to each other via one or more networks 199, such as one or more wired or wireless local area networks (a “LAN” including a Wi-Fi LAN, a mesh network, Bluetooth, near-field communication, etc.) or a wide area network (a “WAN” including the Internet).
[0021] The client device 110 can be, for example, one or more of a desktop computer, a laptop computer, a tablet, a mobile phone, a vehicle computing device (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a stand-alone interactive speaker (optionally having a display), a smart home appliance such as a smart TV, and / or a user's wearable device including a computing device (e.g., a user's wristwatch having a computing device, a user's glasses having a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices can be provided.
[0022] Client device 110 can execute an automatic assistant client 114. An instance of the automatic assistant client 114 can be an application separate from the operating system of client 110 (e.g., installed "on top of" the operating system), or alternatively can be directly implemented by the operating system of client device 110. The automatic assistant client 114 can interact with a natural conversation system 180 that is implemented locally on client device 110 or remotely from client device 110 via one or more of the networks 199 as shown in FIG. 1 (e.g., at a remote server). The automatic assistant client 114 can (and optionally via interaction with a remote server) form what appears to the user to be a logical instance of an automatic assistant 115 through which the user can participate in a human-computer dialog from the user's perspective. An instance of the automatic assistant 115 is shown in FIG. 1 and is surrounded by a dashed line including the automatic assistant client 114 and the natural conversation system 180 of client device 110. Thus, it should be understood that a user interacting with the automatic assistant client 114 running on client device 110 is actually interacting with a logical instance of the user's own automatic assistant 115 (or a logical instance of the automatic assistant 115 shared among a household or other user group). For simplicity and brevity, the automatic assistant 115 as used herein refers to the automatic assistant client 114 that runs locally on client device 110 and / or remotely from client device 110 (e.g., at a remote server that may additionally or alternatively implement an instance of the natural conversation system 180).
[0023] In various implementations, the client device 110 may include a user input engine 111 configured to detect user input provided by a user of the client device 110 using one or more user interface input devices. For example, the client device 110 may include one or more microphones configured to generate audio data, such as audio data that captures speech of a user of the client device 110 or other sounds in the environment of the client device 110. Additionally or alternatively, the client device 110 may include one or more visual components configured to generate visual data that captures images and / or movement (e.g., gestures) detected within one or more of the fields of view of the one or more visual components. Additionally or alternatively, the client device 110 may include one or more touch sensing components (e.g., keyboard and mouse, stylus, touch screen, touch panel, one or more hardware buttons, etc.) configured to generate one or more signals that capture touch input directed to the client device 110.
[0024] In various implementations, client device 110 may include a rendering engine 112 configured to provide content for audible and / or visual presentation to a user of client device 110 using one or more user interface output devices. For example, client device 110 may include one or more speakers that enable content to be provided for audible presentation to a user of client device 110 via one or more speakers of client device 110. Additionally or alternatively, client device 110 may include a display or projector that enables content to be provided for visual presentation to a user of the client device via a display or projector of client device 110. In other implementations, client device 110 may communicate with one or more other computing devices (e.g., via one or more of networks 199), and one or more user interface input devices and / or one or more user interface output devices of the one or more other computing devices may be utilized to detect user input provided by a user of client device 110 and / or to provide content for audible and / or visual presentation to a user of client device 110. Additionally or alternatively, client device 110 may include one or more light emitting diodes (LEDs) that may be illuminated in one or more colors to provide an indication that the automatic assistant 115 is processing user input from a user of client device 110 and waiting for the user of client device 110 to continue providing user input, and / or to provide an indication that the automatic assistant 115 is performing any other function.
[0025] In various implementations, the client device 110 may include one or more presence sensors 113 configured to obtain approval from a corresponding user and provide a signal indicating a detected presence, particularly a human presence. In some of those implementations, the automatic assistant 115 can identify the client device 110 (or another computing device associated with the user of the client device 110) that satisfies the utterance, at least in part based on the presence of the user in the client device 110 (or in another computing device associated with the user of the client device 110). The utterance can be satisfied by rendering response content in the client device 110 and / or in another computing device associated with the user of the client device 110 (e.g., via the rendering engine 112), by controlling the client device 110 and / or another computing device associated with the user of the client device 110, and / or by causing any other action to be performed in the client device 110 and / or in another computing device associated with the user of the client device 110 to satisfy the utterance. As described herein, the automatic assistant 115 can utilize the data determined based on the presence sensor 113 when determining the client device 110 (or other computing device) based on whether the user is nearby or was recently nearby, and provide the corresponding command to only the client device 110 (or to another computing device).In some additional or alternative implementations, the automatic assistant 115 can utilize data determined based on the presence sensor 113 when determining whether any user (any user or a specific user) is currently in proximity to the client device 110 (or other computing device), and can optionally suppress the provision of data to and / or from the client device 110 (or other computing device) based on the user in proximity to the client device 110 (or other computing device).
[0026] The presence sensor 113 can be in various forms. For example, the client device 110 can utilize one or more of the user interface input components described above with respect to the user input engine 111 (e.g., the microphone, visual components, and / or touch sensing components described above) to detect the presence of a user. Additionally or alternatively, the client device 110 can include other types of light-based presence sensors 113, such as a passive infrared ("PIR") sensor that measures infrared ("IR") light emitted from objects within the field of view.
[0027] Additionally or alternatively, in some implementations, the presence sensor 113 may be configured to detect a human presence or other phenomena associated with the presence of a device. For example, in some embodiments, the client device 110 may include a presence sensor 113 configured to detect other computing devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by a user and / or various types of wireless signals (e.g., waves such as radio, ultrasonic, electromagnetic, etc.) emitted by other computing devices. For example, the client device 110 may be configured to emit waves that are imperceptible to humans, such as ultrasonic or infrared waves, that can be detected by other computing devices (e.g., via an ultrasonic / infrared receiver such as an ultrasonic-compatible microphone).
[0028] Additionally or alternatively, the client device 110 may emit other types of waves that are imperceptible to humans, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.), that can be detected by other computing devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by a user and used to determine the specific location of the user. In some implementations, for example, GPS and / or Wi-Fi triangulation may be used to detect a person's location based on GPS and / or Wi-Fi signals to / from the client device 110. In other implementations, other wireless signal characteristics, such as time-of-flight, signal strength, etc., may be used alone or collectively by the client device 110 to determine the location of a specific person based on signals emitted by other computing devices carried / operated by the user.
[0029] Additionally or alternatively, in some implementations, the client device 110 may perform speaker identification (SID) to recognize the user from the user's voice. In some implementations, then, for example, the movement of the speaker may be determined by the presence sensor 113 of the client device 110 (and optionally the GPS sensor, Soli chip, and / or accelerometer of the client device 110). In some implementations, based on such detected movement, the user's position may be predicted, which position may be assumed to be the user's position when any content is rendered in the client device 110 and / or other computing devices, based at least in part on the proximity of the client device 110 and / or other computing devices to the user's position. In some implementations, the user may simply be assumed to be at the last position where the user interacted with the assistant 115, particularly if not much time has elapsed since the last interaction.
[0030] Furthermore, client device 110 and / or natural conversation system 180 may include one or more memories for storing data (such as software applications, one or more first-party (1P) agents 171, one or more third-party agents (3P) 172, etc.), one or more processors for accessing and executing the data, and / or other components that facilitate communication via one or more of network 199, such as one or more network interfaces. In some implementations, one or more of the software applications, 1P agents 171, and / or 3P agents 172 may be installable locally on client device 110, but in other implementations, one or more of the software applications, 1P agents 171, and / or 3P agents 172 may be hosted remotely (e.g., by one or more servers) and accessible by client device 110 via one or more of network 199. Operations performed by client device 110, other computing devices, and / or automated assistant 115 may be distributed across multiple computing devices. Automated assistant 115 may be implemented, for example, as a computer program executed on client device 110 and / or one or more computers at one or more locations coupled to each other via a network (such as one or more of network 199 of FIG. 1).
[0031] In some implementations, the operations performed by the automatic assistant 115 may be implemented locally at the client 110 via the automatic assistant client 114. As shown in FIG. 1, the automatic assistant client 114 may include an automatic speech recognition (ASR) engine 120A1, a natural language understanding (NLU) engine 130A1, a fulfillment engine 140A1, and a text-to-speech (TTS) engine 150A1. In some implementations, the operations performed by the automatic assistant 115 may be distributed across multiple computer systems, such as when the natural conversation system 180 is implemented remotely from the client device 110 as shown in FIG. 1. In these implementations, in the implementation where the natural conversation system 180 is implemented remotely from the client device 110 (e.g., on a remote server), the automatic assistant 115 may additionally or alternatively utilize the ASR engine 120A2, the NLU engine 130A2, the fulfillment engine 140A2, and the TTS engine 150A2 of the natural conversation system 180.
[0032] As will be described in more detail with reference to FIG. 2, each of these engines can be configured to perform one or more functions. For example, the ASR engines 120A1 and / or 120A2 can capture at least a portion of an utterance and process a stream of audio data generated by a microphone of the client device 110 to generate a stream of ASR output using a streaming ASR model (e.g., a recurrent neural network (RNN) model, a transformer model, and / or any other type of ML model capable of performing ASR) stored in the machine learning (ML) model database 115A. In particular, the streaming ASR model can be utilized to generate a stream of ASR output as the stream of audio data is being generated. Further, the NLU engines 130A1 and / or 130A2 can process a stream of ASR output to generate a stream of NLU output using an NLU model (e.g., long short-term memory (LSTM), gated recurrent unit (GRU), and / or any other type of RNN or other ML model capable of performing NLU) stored in the ML model database 115A and / or grammar-based NLU rules. Further, the fulfillment engines 140A1 and / or 140A2 can generate a set of fulfillment outputs based on a stream of fulfillment data generated based on a stream of NLU output. The stream of fulfillment data can be generated using, for example, one or more of a software application, the 1P agent 171, and / or the 3P agent 172. Finally, the TTS engines 150A1 and / or 150A2 can process text data (e.g., text devised by the virtual assistant 115) to generate synthetic audio data including synthetic speech generated by a computer corresponding to the text data using a TTS model stored in the ML model database 115A.In particular, the ML models stored in the ML model database 115A can be on-device ML models stored locally in the client device 110, or shared ML models accessible to both the client device 110 and / or (e.g., in an implementation where the natural conversation system is implemented by a remote server) other systems.
[0033] In various implementations, the stream of ASR outputs can include, for example, a stream of speech hypotheses (e.g., term hypotheses and / or transcription hypotheses) predicted to correspond to the utterance (or one or more portions thereof) of the user of the client device 110 captured within the stream of audio data, one or more corresponding predicted values (e.g., probabilities, log-likelihoods, and / or other values) for each of the speech hypotheses, a plurality of phonemes predicted to correspond to the utterance of the user of the client device 110 captured within the stream of audio data, and / or other ASR outputs. In some variations of those implementations, the ASR engines 120A1 and / or 120A2 can select one or more of the speech hypotheses as the recognized text corresponding to the utterance (e.g., based on the corresponding predicted values).
[0034] In various implementations, the stream of NLU outputs can include a stream of annotated recognized text that includes one or more annotations of recognized text for one or more (e.g., all) of the terms included in the stream of ASR outputs, one or more predicted intents determined based on the recognized text for one or more (e.g., all) of the terms included in the stream of ASR outputs, predicted and / or estimated slot values for corresponding parameters associated with each of the one or more predicted intents determined based on the recognized text for one or more (e.g., all) of the terms included in the stream of ASR outputs, and / or other NLU outputs. For example, NLU engines 130A1 and / or 130A2 can include a part-of-speech tagger (not shown) configured to annotate terms with their grammatical roles. Additionally or alternatively, NLU engines 130A1 and / or 130A2 can include an entity tagger (not shown) configured to annotate entity references within one or more segments of the recognized text, such as references to people (including, e.g., literary characters, celebrities, public figures, etc.), organizations, places (real and fictional). In some implementations, data regarding entities can be stored in one or more databases, such as a knowledge graph (not shown). In some implementations, the knowledge graph can include nodes representing known entities (and optionally entity attributes), as well as edges connecting the nodes to represent relationships between the entities. The entity tagger can annotate references to entities at a high level of granularity (e.g., to enable identification of all references to an entity class, such as a person) and / or at a low level of granularity (e.g., to enable identification of all references to a specific entity, such as a particular person). The entity tagger can depend on the content of the natural language input to resolve a particular entity and / or can optionally communicate with the knowledge graph or other entity databases to resolve a particular entity.
[0035] Additionally or alternatively, the NLU engines 130A1 and / or 130A2 may include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more context cues. For example, the coreference resolver may be used to resolve the term "them" in the natural language input "buy them" to "buy theatre tickets" based on the mention of "theatre tickets" in a notification of the client device rendered immediately prior to receiving the input "buy them". In some implementations, one or more components of the NLU engines 130A1 and / or 130A2 may depend on annotations from one or more other components of the NLU engines 130A1 and / or 130A2. For example, in some implementations, the entity tagger may depend on annotations from the coreference resolver when annotating all mentions of a particular entity. Also, for example, in some implementations, the coreference resolver may depend on annotations from the entity tagger when clustering references to the same entity.
[0036] In various implementations, a stream of fulfillment data can include one or more fulfillment outputs generated by, for example, a software application, a 1P agent 171, and / or a 3P agent 172. One or more structured requests generated based on the NLU output stream can be sent to one or more of a software application, a 1P agent 171, and / or a 3P agent 172, and one or more of the software application, a 1P agent 171, and / or a 3P agent 172 can send a fulfillment output expected to satisfy the utterance in response to receiving one or more of the structured requests. The fulfillment engines 140A1 and / or 140A2 can include the fulfillment output received at the client device 110 within a set of fulfillment outputs corresponding to the stream of fulfillment data. In particular, the stream of fulfillment data can be generated when a user of the client device 110 provides an utterance. Further, the fulfillment output engine 164 can select one or more fulfillment outputs from the stream of fulfillment outputs, and the selected one or more of the fulfillment outputs can be provided for presentation to the user of the client device 110 to satisfy the utterance. The one or more fulfillment outputs can include, for example, audible content that is predicted to respond to the utterance and can be aurally rendered for presentation to the user of the client device 110 via a speaker, visual content that is predicted to respond to the utterance and can be visually rendered for presentation to the user of the client device 110 via a display, and / or an assistant command that, when executed, causes the client device 110 and / or another computing device communicating with the client device 110 (e.g., via one or more of the networks 199) to be controlled in response to the utterance.
[0037] Although FIG. 1 is described with respect to a single client device having a single user, it should be understood that this is for purposes of example and is not meant to be limiting. For example, one or more additional client devices of the user can also implement the techniques described herein. For example, client device 110, one or more additional client devices, and / or any other computer device of the user can form an ecosystem of devices that can use the techniques described herein. These additional client devices and / or computing devices can communicate with client device 110 (e.g., via one or more of network 199). As another example, a given client device can be utilized by multiple users in a shared setting (e.g., a group of users, a household, a shared living space, etc.).
[0038] As described herein, the automatic assistant 115 can determine whether to provide a natural conversation output for presentation to the user in response to determining that the user has paused providing an utterance and / or deciding when to fulfill the utterance. In making this determination, the automatic assistant can utilize the natural conversation engine 160. In various implementations, as shown in FIG. 1, the natural conversation engine 160 can include an acoustic engine 161, a pause engine 162, a time engine 163, a natural conversation output engine 164, and a fulfillment output engine 165.
[0039] In some implementations, the acoustic engine 161 can determine audio-based features based on processing a stream of audio data. In some variations of those implementations, the acoustic engine 161 can process a stream of audio data to determine audio-based features using an audio-based ML model stored in the ML model database 115A. In additional or alternative implementations, the acoustic engine 161 can process a stream of audio data to determine audio-based features using one or more rules. Audio-based features can include, for example, prosodic characteristics associated with utterances captured within the stream of audio data and / or other audio-based features. Prosodic characteristics can include, for example, one or more characteristics of syllables and larger speech units, including intonation, tone, stress, rhythm, tempo, pitch, elongated syllables, pauses, grammar associated with pauses, and / or other audio-based features derivable from processing the stream of audio data. Additionally, prosodic characteristics can provide, for example, an indication of an emotional state, form (e.g., statement, question, or command), irony, sarcasm, speaking tone, and / or emphasis. Put another way, prosodic characteristics are voice features that can be determined dynamically during a dialog session based on individual utterances and / or combinations of multiple utterances, without relying on the individual voice characteristics of a given user.
[0040] In some implementations, the pause engine 162 can determine whether the user of the client device 110 has paused the delivery of speech captured within the audio data stream or has completed the delivery of the speech. In some variations of those implementations, the pause engine 162 can determine that the user of the client device 110 has paused the delivery of speech based on the processing of audio-based features determined using the audio engine 161. For example, the pause engine 162 can process the audio-based features to generate an output using an audio-based classification ML model stored in the ML model database 115A and, based on the output generated using the audio-based classification ML model, determine whether the user of the client device 110 has paused the delivery of speech or has completed the delivery of the speech. The output can include, for example, one or more predicted measurements (e.g., binary values, log-likelihoods, probabilities, etc.) indicating whether the user of the client device 110 has paused the delivery of speech or has completed the delivery of the speech. For example, assume that the user of the client device 110 provides the utterance "call Arnolllld's", where "llll" indicates a drawn-out syllable included within the utterance. In this example, the audio-based features may include an indication that the utterance includes a drawn-out syllable, and as a result, the output generated using the audio-based classification ML model may indicate that the user has not completed the delivery of the speech.
[0041] In additional or alternative variations of those implementations, the pause engine 162 can determine that the user of the client device 110 has paused providing utterances based on a stream of NLU data generated using the NLU engine 130A1 and / or 130A2. For example, the pause engine 162 can process a stream of audio data based on a predicted intent, and / or a prediction and / or estimated slot value regarding a corresponding parameter associated with the predicted intent and / or a predicted slot value regarding a predicted slot value, regardless of whether the user of the client device 110 has paused providing an utterance or has completed providing an utterance. For example, assume that the user of the client device 110 provides an utterance of "call Arnolllld's", where "llll" indicates a prolonged syllable included in the utterance. In this example, the stream of NLU data can include the predicted intent of "call" and the slot value regarding the entity parameter of "Arnold". However, in this example, even if the automatic assistant 115 can access the contact entry associated with the entity "Arnold" (so that the utterance can be fulfilled), the automatic assistant 115 may not initiate a call to the entity "Arnold" based on the prolonged syllable included in the audio-based features determined based on processing the utterance. In contrast, in this example, if the user did not provide "Arnolllld" with the prolonged syllable and / or the user provided an explicit command (e.g., "call Arnold now", "call Arnold immediately", etc.) for the automatic assistant 115 to start full fulfillment of the utterance, the pause engine 162 can determine that the user of the client device 110 has completed providing the utterance.
[0042] In some implementations, the natural conversation output engine 163 can determine the natural conversation output to be provided for presentation to the user of the client device in response to determining that the user has paused providing utterances. In some variations of those implementations, the natural conversation output engine 163 can determine a set of natural conversation outputs and, based on NLU metrics and / or audio-based features associated with the stream of NLU data, select one or more (e.g., randomly or cyclically through the set of natural conversation outputs) of the natural conversation outputs to be provided for presentation to the user (e.g., audible presentation via one or more speakers of the client device 110) from among the set of natural conversation outputs. In some further variations of those implementations, a superset of natural conversation outputs can be stored (e.g., as text data converted to synthetic voice audio data and / or as synthetic voice audio data) in one or more databases (not shown) accessible by the client device 110, and the set of natural conversation outputs can be generated from the superset of natural conversation outputs based on NLU metrics associated with the stream of NLU data and / or audio-based features.
[0043] These natural conversation outputs can be implemented to facilitate the ongoing dialog session during speech, but do not necessarily have to be implemented as a full fulfillment of the speech. For example, natural conversation outputs can include requests for the user to confirm an instruction to continue the conversation with the automated assistant 115 (e.g., "Are you still there?"), requests for the user to provide additional user input to facilitate the dialog session between the user and the automated assistant 115 (e.g., "Who did you want me to call?"), and filler speech (e.g., "Mmmhmm", "Uh huhh", "Alright"). In various implementations, the natural conversation engine 163 can utilize one or more language models stored in the ML model database 115A when generating a set of natural conversation outputs. In other implementations, the natural conversation engine 163 can obtain a set of natural conversation outputs from a remote system (e.g., a remote server) and store the set of natural conversation outputs in the on-device memory of the client device 110.
[0044] In some implementations, the fullfillment output engine 164 can select one or more fullfillment outputs to be provided for presentation to the user of the client device from the stream of fullfillment outputs in response to determining that the user has completed providing an utterance or (e.g., as described with respect to FIG. 5C) in response to determining that although the user has not completed providing an utterance, the utterance should nevertheless be fulfilled. Although the 1P agent 171 and the 3P agent 172 are shown as being implemented via one or more of the networks 199 in FIG. 1, it should be understood that this is for example only and is not meant to be limiting. For example, one or more of the 1P agent 171 and / or the 3P agent 172 can be implemented locally on the client device 110, the stream of NLU outputs can be sent to one or more of the 1P agent 171 and / or the 3P agent 172 via an application programming interface (API), and the fullfillment output from one or more of the 1P agent 171 and / or the 3P agent 172 can be obtained via the API and incorporated into the stream of fullfillment data. Additionally or alternatively, one or more of the 1P agent 171 and / or the 3P agent 172 can be implemented remotely from the client device 110 (e.g., each in a 1P server and / or a 3P server), the stream of NLU outputs can be sent to one or more of the 1P agent 171 and / or the 3P agent 172 via one or more of the networks 199, and the fullfillment output from one or more of the 1P agent 171 and / or the 3P agent 172 can be obtained via one or more of the networks 199 and incorporated into the stream of fullfillment data.
[0045] For example, the full fulfillment output engine 164 can select one or more full fulfillment outputs from a stream of full fulfillment data based on NLU measurements associated with a stream of NLU data and / or full fulfillment measurements associated with a stream of full fulfillment data. The NLU measurements can be, for example, probabilities, log-likelihoods, binary values, etc., indicating to what extent the NLU engines 130A1 and / or 130A2 are confident that the predicted intent corresponds to the actual intent of the user who provided the utterance captured within the stream of audio data, and / or to what extent the estimation of the parameters associated with the predicted intent and / or the predicted slot values correspond to the actual slot values of the parameters associated with the predicted intent. The NLU measurements can be generated when the NLU engines 130A1 and / or 130A2 generate a stream of NLU outputs and can be included within the stream of NLU outputs. The full fulfillment measurements can be, for example, probabilities, log-likelihoods, binary values, etc., indicating to what extent the full fulfillment engines 140A1 and / or 140A2 are confident that the predicted full fulfillment output corresponds to the desired full fulfillment of the user. The full fulfillment measurements can be generated when one or more of the software application, the 1P agent 171, and / or the 3P agent 172 generate a full fulfillment output and can be incorporated into the stream of full fulfillment data, and / or can be generated when the full fulfillment engines 140A1 and / or 140A2 process the full fulfillment data received from one or more of the software application, the 1P agent 171, and / or the 3P agent 172 and can be incorporated into the stream of full fulfillment data.
[0046] In some implementations, in response to determining that the user has paused the provision of speech, the time engine 165 can determine the duration of the pause in the provision of speech and / or the duration of any subsequent pause. The automatic assistant 115 can utilize one or more of these pause durations in the natural conversation output engine 163 when selecting the natural conversation output to be provided for presentation to the user of the client device 110. For example, assume that the user of the client device 110 provides the utterance "call Arnolllld's", where "llll" indicates a prolonged syllable included in the utterance. Further assume that the user is determined to have paused the provision of speech. In some implementations, in response to determining that the user of the client device 110 has paused the provision of speech, a natural conversation output can be provided for presentation to the user (e.g., by auditorily rendering "Mmmhmm", etc.). However, in other implementations, the natural conversation output can be provided for presentation to the user in response to the time engine 165 determining that a threshold duration has elapsed since the user first paused. Further assume that the user of the client device 110 does not continue to provide speech in response to the natural conversation output being provided for presentation. In this example, additional natural conversation output can be provided for presentation to the user in response to the time engine 165 determining that an additional threshold duration has elapsed since the user first paused (or since the natural conversation output was provided for presentation to the user).Thus, when providing additional natural conversation output for presentation to the user, the natural conversation output engine 163 can select different natural conversation outputs that request the user of the client device 110 to complete speaking (e.g., "You were saying?", "Did I miss something?"), or that request the user of the client device 110 to provide a specific slot value regarding the predicted intent (e.g., "Who did you want to call?", "And how many people was the reservation for?").
[0047] In various implementations, while the automated assistant 115 is waiting for the user of the client device 110 to finish speaking, the automated assistant 115 can optionally cause fulfillment output within a set of fulfillment outputs to be partially fulfilled. For example, the automated assistant 115 can establish a connection with one or more of additional computing devices that communicate with the client device 110 (e.g., via one or more of the networks 199), such as software applications, 1P agents 171, 3P agents 172, and / or other client devices associated with the user of the client device 110, smart network devices, etc., based on one or more fulfillment outputs included within the set of fulfillment outputs, can cause synthetic audio data including synthetic speech to be generated (but not aurally rendered), can cause graphical content to be generated (but not visually rendered), and / or can perform any other partial fulfillment of one or more of the fulfillment outputs. As a result, it is possible to reduce the waiting time when providing the fulfillment output for presentation to the user of the client device 110.
[0048] Proceeding now to FIG. 2, an exemplary process flow is shown demonstrating various aspects of the present disclosure using the various components of FIG. 1. The ASR engines 120A1 and / or 120A2 can process the audio data stream 201A using the streaming ASR model stored in the ML model database 115A to generate the ASR output stream 220. The NLU engines 130A1 and / or 130A2 can process the ASR output stream 220 using the NLU model stored in the ML model database 115A to generate the NLU output stream 230. In some implementations, the NLU engines 130A1 and / or 130A2 can additionally or alternatively process the non-audio data stream 201B when generating the NLU output stream 230. The non-audio data stream 201B can include visual data generated by the visual components of the client device 110, a stream of touch inputs provided by the user via the touch-sensitive components of the client device 110, a stream of type inputs provided by the user via the touch-sensitive components of the client device 110 or peripheral devices (e.g., mouse and keyboard), and / or any other non-audio data generated by any other user interface input device of the client device 110. In some implementations, the 1P agent 171 can process the NLU output stream to generate the 1P fulfillment data 240A. In additional or alternative implementations, the 3P agent 172 can process the NLU output stream 230 to generate the 3P fulfillment data 240B.The fulfillment engines 140A1 and / or 140A2 can generate a stream 240 of fulfillment data based on 1P fulfillment data 240A and / or 3P fulfillment data 240B (and optionally, other fulfillment data generated based on one or more software applications accessible in the client device 110 that processes the stream 230 of NLU output). Further, the acoustic engine 161 can process the stream 201A of audio data to generate audio-based features 261 associated with the stream 201A of audio data, such as audio-based features 261 of one or more utterances (or portions thereof) included within the stream 201A of audio data.
[0049] As shown in block 262, the pause engine 162 can process the NLU output stream 230 and / or the audio-based feature 261 to determine whether the user of the client device has paused the provision of the utterance captured within the audio data stream 201A or has completed the provision of the utterance captured within the audio data stream 201A. The automatic assistant 115 can determine whether to provide a natural conversation output or a fulfillment output based on whether block 262 indicates that the user has paused the provision of the utterance or has completed the provision of the utterance. For example, assume that the automatic assistant 115 determines based on the indication in block 262 that the user has paused the provision of the utterance. In this example, the automatic assistant 115 can cause the natural conversation output engine 163 to select a natural conversation output 263, and the automatic assistant 115 can cause the natural conversation output 263 to be provided for presentation to the user of the client device 110. In contrast, assume that the automatic assistant 115 determines based on the indication in block 262 that the user has completed the provision of the utterance. In this example, the automatic assistant 115 can cause the fulfillment output engine 164 to select one or more fulfillment outputs 264, and the automatic assistant 115 can cause the one or more fulfillment outputs 264 to be provided for presentation to the user of the client device 110. In some implementations, the automatic assistant 115 can consider one or more pause durations 265 determined by the time engine 165 when determining whether to cause the natural conversation output 263 to be provided for presentation to the user of the client device 110 or whether to cause one or more fulfillment outputs 264 to be provided for presentation to the user of the client device 110. In these implementations, the natural conversation output 263 and / or the one or more fulfillment outputs 264 can be adapted based on the one or more pause durations.Regarding specific features and embodiments, reference has been made to FIGS. 1 and 2, which are for illustrative purposes only and are not meant to be limiting. For example, additional features and embodiments will be described below with reference to FIGS. 3, 4, 5A-5E, and 6.
[0050] Turning now to FIG. 3, a flowchart illustrating an exemplary method 300 is shown for determining whether to provide a natural conversation output for presentation to a user in response to determining that the user has paused the provision of speech and / or determining when to fulfill the speech. For convenience, the operations of method 300 will be described with reference to a system that performs the operations. This system of method 300 includes one or more processors, memory, and / or other components of a computing device (e.g., client device 110 of FIGS. 1 and 5A-5E, computing device 610 of FIG. 6, one or more servers, and / or other computing devices). Further, although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be rearranged, omitted, and / or added.
[0051] In block 352, the system uses a streaming ASR model to process a stream of audio data including a portion of the user's utterance to generate a stream of ASR output for an auto-assistant. The stream of audio data can be generated during a dialog session with the auto-assistant at least partially implemented on the client device by a microphone of the user's client device. In some implementations, in response to the system determining that the user has invoked the auto-assistant via one or more specific words and / or phrases (e.g., hot words such as "Hey Assistant", "Assistant"), one or more buttons (e.g., software buttons and / or hardware buttons), one or more gestures captured by a visual component of the client device that activates the auto-assistant, and / or any other means, the system can process the stream of audio data. In block 354, the system uses an NLU model to process the stream of ASR output to generate a stream of NLU output. In block 356, the system causes a stream of fulfillment data to be generated based on the stream of NLU output. In block 358, the system determines audio-based features associated with the portion of the utterance captured in the audio data based on processing the stream of audio data. The audio-based features can include, for example, one or more prosodic features (e.g., intonation, tone, stress, rhythm, tempo, pitch, pause, and / or other prosodic features) associated with the portion of the utterance, and / or other audio-based features that can be determined based on processing the stream of audio data. The operations of blocks 352-358 are described in more detail herein (e.g., with respect to FIGS. 1 and 2).
[0052] In block 360, the system determines whether the user has paused or completed providing speech based on the audio-based features associated with the portion of the speech captured within the NLU output stream and / or the audio data. In some implementations, the system can process the audio-based features to generate an output using an audio-based classification ML model, and the system can determine whether the user has paused or completed providing speech based on the output generated using the audio-based classification ML model. The output generated using the audio-based classification ML model can include one or more predictive measurements (e.g., binary values, probabilities, log-likelihoods, and / or other measurements) indicating whether the user has paused or completed providing speech. For example, assume the output includes a first probability of 0.8 associated with the prediction that the user has paused providing speech and a second probability of 0.6 associated with the prediction that the user has completed providing speech. In this example, the system can determine that the user has paused providing speech based on the predictive measurements. In additional or alternative implementations, the system can process or analyze the NLU output stream to determine whether the user has paused or completed providing speech. For example, if the system determines that the NLU measurements associated with the predicted intent, and / or the estimates and / or predicted slot values for the corresponding parameters associated with the predicted intent do not meet the NLU measurement threshold, or if the system determines that the slot value for the corresponding parameter associated with the predicted intent is unknown, the automatic assistant can determine that the user has paused providing speech. In particular, in various implementations, the system can determine whether the user has paused or completed providing speech based on both the audio-based features and the stream of NLU data.For example, if the system determines, based on a stream of NLU data, that an utterance can be fulfilled, but an audio-based feature indicates that the user has paused the provision of the utterance, any additional portion of the utterance that can be provided by the user may change how the user desires the utterance to be fulfilled, so the system may determine that the user has paused the provision of the utterance.
[0053] In the iteration of block 360, if the system determines that the user has completed the provision of the utterance, the system can proceed to block 362. In block 362, the system causes the automatic assistant to begin full fulfillment of the utterance. For example, the system can select one or more fulfillment outputs predicted to satisfy the utterance from a stream of fulfillment data and cause the one or more fulfillment outputs to be provided for presentation to the user via a client device or an additional computing device communicating with the client device. As described above with respect to FIG. 1, the one or more fulfillment outputs can include, for example, audible content that is predicted to respond to the utterance and can be aurally rendered for presentation to the user of the client device via a speaker, visual content that is predicted to respond to the utterance and can be visually rendered for presentation to the user of the client device via a display, and / or assistant commands that, when executed, cause the client device and / or other computing devices communicating with the client device to be controlled in response to the utterance. The system can return to block 352 and execute additional iterations of method 300 of FIG. 3.
[0054] In the iteration of block 360, if the system determines that the user has paused the provision of speech, the system can proceed to block 364. In block 364, the system determines the natural conversation output to be provided for audible presentation to the user. Further, in block 366, the system can cause the natural conversation output to be provided for audible presentation to the user in block 366. The natural conversation output can be selected from a set of natural conversation outputs stored in the on-device memory of the client device based on NLU measurements associated with the NLU data stream and / or audio-based features. In some implementations, one or more of the natural conversation outputs included in the set of natural conversation outputs can correspond to text data. In these implementations, the text data associated with the selected natural conversation output can be processed using a TTS model to generate synthetic audio data including a synthetic voice corresponding to the selected natural conversation output, and the synthetic audio data can be aurally rendered for presentation to the user via a speaker of the client device or an additional computing device.
[0055] In additional or alternative implementations, one or more of the natural conversation outputs included within a set of natural conversation outputs can correspond to synthetic audio data that includes a synthetic voice corresponding to the selected natural conversation output, and the synthetic audio data can be aurally rendered for presentation to the user via a speaker of the client device or an additional computing device. In particular, in various implementations, when providing a natural conversation output for audible presentation to the user, the volume at which the natural conversation output is played to the user can be lower than the volume of other outputs that are aurally rendered for presentation to the user. Further, in various implementations, when providing a natural conversation output for audible presentation to the user, one or more of the automatic assistant components can remain active while the natural conversation output is being provided for audible presentation to the user in order to enable the automatic assistant to continue processing the stream of audio data (e.g., ASR engines 120A1 and / or 120A2, NLU engines 130A1 and / or 130A2, and / or fulfillment engines 140A1 and / or 140A2).
[0056] In block 368, the system determines whether to fulfill the utterance following providing the natural conversation output for audible presentation to the user. In some implementations, the system can determine to fulfill the utterance in response to determining that the user has completed providing the utterance following providing the natural conversation output for audible presentation to the user. In these implementations, the streams of ASR output, NLU output, and full fulfillment data can be updated based on the user having completed providing the utterance. In additional or alternative implementations, the system can determine to fulfill the utterance in response to determining that, even if the user has not completed providing the utterance, it is possible to fulfill the utterance based on a portion of the utterance, based on one or more costs associated with starting the full fulfillment of the utterance by the automated assistant (as described in more detail with respect to, for example, FIG. 5C).
[0057] In the iteration of block 368, if the system decides to fulfill the utterance after providing a natural conversation output for audible prompting to the user, the system proceeds to block 362 to start the full fulfillment of the utterance by the automatic assistant as described above. In the iteration of block 368, if the system decides not to fulfill the utterance after providing a natural conversation output for audible prompting to the user, the system returns to block 364. In this subsequent iteration of block 364, the system can determine additional natural conversation output to be provided for audible prompting to the user. In particular, the additional conversation output to be provided for audible prompting to the user selected in this subsequent iteration of block 364 may be different from the natural conversation output to be provided for audible prompting to the user selected in the previous iteration of block 364. For example, the natural conversation output to be provided for audible prompting to the user selected in the previous iteration of block 364 may be provided to the user as an indication that the automatic assistant is still listening and waiting for the user to complete the utterance (e.g., "Mmhmm", "Okay", "Uh huhhh", etc.). However, the natural conversation output to be provided for audible prompting to the user selected in this subsequent iteration of block 364 may also be an indication that the automatic assistant is still listening and waiting for the user to complete the utterance, but can also be provided to the user as an indication that more explicitly prompts the user to complete the utterance or provide a specific input (e.g., "Are you still there?", "And how many people was the reservation for?"). The system can continue to execute the iterations of blocks 364 - 368 until the system decides to fulfill the utterance in the iteration of block 368 and proceeds to block 362 to start the full fulfillment of the utterance by the automatic assistant as described above.
[0058] In various implementations, one or more predictive measurements indicating whether a user has paused or completed providing speech can be used to determine whether and / or when to provide a natural conversation output for audible presentation to the user. For example, assume that the output generated using an audio-based classification ML model includes a first probability of 0.8 associated with the prediction that the user has paused providing speech and a second probability of 0.6 associated with the prediction that the user has completed providing speech. Further assume that the first probability of 0.8 meets a pause threshold indicating that the system is strongly confident that the user has paused providing speech. Thus, in the first iteration of block 364, the system can cause the audio backchannel to be used as a natural conversation output (e.g., "uh huh"). Further, in the second iteration of block 364, since the system is strongly confident that the user has paused providing speech, the system can cause another audio backchannel to be used as a natural conversation output (e.g., "Mmmhmm" or "I'm here"). In contrast, assume that the output generated using an audio-based classification ML model includes a first probability of 0.5 associated with the prediction that the user has paused providing speech and a second probability of 0.4 associated with the prediction that the user has completed providing speech. Further assume that the first probability of 0.5 does not meet a pause threshold indicating that the system is strongly confident that the user has paused providing speech. Thus, in the first iteration of block 364, the system can cause the audio backchannel to be used as a natural conversation output (e.g., "uh huh"). However, in the second iteration of block 364, rather than causing the utterance of another audio backchannel to be used as a natural conversation output, the system may require the user to confirm the predicted intent predicted based on the speech processing (e.g., "Did you want to call someone?").In particular, when determining the natural conversation output to be provided for audible presentation to the user, the system can randomly select a given natural conversation output to be provided for audible presentation to the user from a set of natural conversation outputs, can cycle through the set of natural conversation outputs when selecting a given natural conversation output to be provided for audible presentation to the user, or can determine the natural conversation output to be provided for audible presentation to the user in any other way.
[0059] Figure 3 is described herein without considering any temporal aspects in providing natural conversation output for audible presentation to the user, but it should be understood that it is for illustrative purposes only. In various implementations, as described below with respect to Figure 4, the system can cause only instances of natural conversation output to be provided for audible presentation to the user based on various thresholds of time. For example, in method 300 of Figure 3, the system can cause an initial instance of natural conversation output to be provided for audible presentation to the user in response to determining that a first threshold duration has elapsed since the user paused providing utterances. Further, in method 300 of Figure 3, the system can cause subsequent instances of natural conversation output to be provided for audible presentation to the user in response to determining that a second threshold duration has elapsed since the initial instance of natural conversation output was provided for audible presentation to the user. In this example, the first threshold duration and the second threshold duration can be the same or different and can correspond to any positive integer and / or fraction thereof (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.).
[0060] Here, turning to FIG. 4, a flowchart is shown that depicts another exemplary method 400 for determining whether to provide a natural conversation output for presentation to the user in response to determining that the user has paused the provision of speech, and / or for determining when to execute the speech. For convenience, the operations of method 400 are described with reference to the system that performs the operations. This system of method 400 includes one or more processors, memories, and / or other components of a computing device (e.g., the client device 110 of FIGS. 1 and 5A - 5E, the computing device 610 of FIG. 6, one or more servers, and / or other computing devices). Further, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.
[0061] In block 452, the system receives a stream of audio data that includes a portion of the user's speech and is directed to the automatic assistant. The stream of audio data can be generated by the microphone of the user's client device during a dialog session with the automatic assistant that is at least partially implemented on the client device. In block 454, the system processes the stream of audio data. The system can process the stream of audio data in the same or a similar manner as described above with respect to operation blocks 352 - 358 of method 300 of FIG. 3.
[0062] In block 456, the system determines whether the user has paused or completed providing an utterance based on the stream of NLU outputs and / or audio-based features associated with the portion of the utterance captured within the audio data determined in block 454 based on processing the utterance. The system can make this determination in the same or a similar manner as described with respect to the operation of block 360 of method 300 in FIG. 3. In an iteration of block 456, if the system determines that the user has completed providing an utterance, the system can proceed to block 458. In block 458, in the same or a similar manner as described with respect to the operation of block 360 of method 300 in FIG. 3, cause the automatic assistant to start fulfilling the utterance. The system returns to block 452 and performs additional iterations of method 400 in FIG. 4. In an iteration of block 456, if the system determines that the user has paused providing an utterance, the system can proceed to block 460.
[0063] In block 460, the system determines whether a user pause in providing speech satisfies an N threshold, where N is any positive integer and / or fraction thereof (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.). In an iteration of block 460, if the system determines that a user pause in providing speech does not satisfy the N threshold, the system returns to block 454 and continues to process the audio data stream. In an iteration of block 460, if the system determines that a user pause in providing speech satisfies the N threshold, the system proceeds to block 460. In block 462, the system determines the natural language conversation output to be provided for audible presentation to the user. In block 464, the system causes the natural conversation output to be provided for audible presentation to the user. The system can perform the operations of blocks 462 and 464 in the same or similar manner as described above with respect to the operations of blocks 364 and 366 of method 300 of FIG. 3. In other words, in an implementation that utilizes one or more aspects of method 400 of FIG. 4, in contrast to method 300 of FIG. 3, the system can wait for N seconds after the user first pauses in providing speech and before causing the natural conversation output to be provided for audible presentation to the user.
[0064] In block 466, the system determines whether a pause of the user when providing an utterance satisfies an M threshold after providing natural conversation output for audible presentation to the user, where M is any positive integer and / or its decimal (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.). In the iteration of block 466, if the system determines that the pause of the user when providing an utterance satisfies the M threshold, the system returns to block 462. Similar to the above description regarding FIG. 3, in this subsequent iteration of block 462, the system can determine additional natural conversation output to be provided for audible presentation to the user, and the additional conversation output to be provided for audible presentation to the user selected in this subsequent iteration of block 462 may be different from the additional natural conversation output to be provided for audible presentation to the user selected in the previous iteration of block 462. In other words, the system can determine the natural conversation output to be provided for audible presentation to the user selected in the previous iteration of block 462 to prompt the user to complete providing the utterance, while the system can determine additional natural conversation output to be presented for audible presentation to the user selected in the subsequent iteration of block 462 to explicitly request the user to complete providing the utterance. In the iteration of block 466, if the system determines that the pause of the user when providing an utterance does not satisfy the M threshold, the system proceeds to block 468.
[0065] In block 468, after the system provides the natural conversation output for audible presentation to the user, it determines whether to fulfill the utterance. In some implementations, after the system provides the natural conversation output (and / or any additional natural conversation output) for audible presentation to the user, in response to determining that the user has completed providing the utterance, the system can decide to fulfill the utterance. In these implementations, the streams of ASR output, NLU output, and fulfillment data can be updated based on the user having completed providing the utterance. In additional or alternative implementations, the system can decide to fulfill the utterance in response to determining that, based on one or more costs associated with starting the fulfillment of the utterance by the automated assistant (as described in more detail with respect to FIG. 5C, for example), the utterance can be fulfilled based on a portion of the utterance even if the user has not completed providing the utterance.
[0066] In an iteration of block 468, if the system decides to fulfill the utterance after providing the natural conversation output for audible presentation to the user, the system proceeds to block 458 to start the fulfillment of the utterance by the automated assistant as described above. In an iteration of block 468, if the system decides not to fulfill the utterance after providing the natural conversation output (and / or any additional natural conversation output) for audible presentation to the user, the system returns to block 462. The subsequent iteration of block 462 was described above. The system can continue to execute the iteration of blocks 462 - 468 until the system decides to fulfill the utterance in an iteration of block 468 and proceeds to block 458 to start the fulfillment of the utterance by the automated assistant as described above.
[0067] Proceeding now to FIGS. 5A-5E, various non-limiting examples are shown of determining whether to provide natural conversation output for presentation to the user in response to determining that the user has paused providing utterances and / or determining when to fulfill an utterance. The automated assistant can be implemented at least in part on the client device 110 (e.g., the automated assistant 115 described with respect to FIG. 1). The automated assistant can utilize a natural conversation system (e.g., the natural conversation system 180 described with respect to FIG. 1) to determine natural conversation output and / or fulfillment output to facilitate a dialog session between the automated assistant and the user 101 of the client device 110. The client device 110 shown in FIGS. 5A-5E can include various user interface components, such as a microphone for generating audio data based on utterances and / or other audible inputs, a speaker for aurally rendering synthesized speech and / or other audible outputs, and a display 190 for receiving touch inputs and / or visually rendering transcriptions and / or other visual outputs. The client device 110 shown in FIGS. 5A-5E is a stand-alone interactive speaker having a display 190, but it is to be understood that this is for example only and is not meant to be limiting.
[0068] For example, referring specifically to FIG. 5A, assume that user 101 of client device 110 provides utterance 552A1 of "Assistant, call Arnolllld's", and then pauses for N seconds as shown by 552A2, where N is any positive integer and / or its fraction (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.). In this example, the automatic assistant can use the ASR model to process a stream of audio data that captures utterance 552A1 and the pause shown by 552A2 to generate a stream of ASR output. Further, the automatic assistant can use the NLU model to process the stream of ASR output to generate a stream of NLU output. Further, the automatic assistant can use a software application accessible on client device 110, a 1P agent accessible on client device 110, and / or a 3P agent accessible on client device 110 to generate a stream of fulfillment data based on the stream of NLU output. In this example, based on processing utterance 552A1, the ASR output includes the recognized text corresponding to utterance 552A1 captured within the stream of audio data (e.g., the recognized text corresponding to "call Arnold's"), the stream of NLU data includes the predicted intention of "call" or "phone call" having a slot value of "Arnold" for the incoming person entity parameter related to the predicted intention of "call" or "phone call", and assume that the stream of fulfillment data includes an assistant command that causes client device 110, when executed, to initiate a call with a contact entry associated with a friend of user 101 named "Arnold".Accordingly, based on processing utterance 552A1 and without processing any additional utterances, the automatic assistant may determine that utterance 552A1 can be satisfied by executing an assistant command. However, even if the automatic assistant may determine that utterance 552A1 can be fulfilled, the automatic assistant may refrain from starting full fulfillment of the utterance.
[0069] In some implementations, the automatic assistant can process a stream of audio data using an audio-based ML model to determine audio-based features associated with utterance 552A1. Further, the automatic assistant can process the audio-based features using an audio-based classification ML model to generate an output indicating whether the user paused or completed providing utterance 552A1. In the example of FIG. 5A, assume that the output generated using the audio-based classification ML model indicates that user 101 paused providing utterance 552A1 (as indicated, for example, by the user providing a drawn-out phrase within “Arnolllld(Arnolllld's)”). Thus, in this example, the automatic assistant may refrain from starting fulfillment of utterance 552A1, at least partially based on the audio-based features of utterance 552A1.
[0070] In additional or alternative implementations, the automatic assistant can determine one or more computational costs associated with the fulfillment of utterance 552A1. The one or more computational costs can include, for example, the computational cost associated with executing the full fulfillment of utterance 552A1, the computational cost associated with canceling the executed fulfillment of utterance 552A1, and / or other computational costs. In the example of FIG. 5A, the computational cost associated with executing the full fulfillment of utterance 552A1 can include at least initiating a call with the contact entry associated with "Arnold" and / or other costs. Further, the computational cost associated with canceling the executed full fulfillment of utterance 552A1 can include at least ending the call with the contact entry associated with "Arnold", restarting the dialog session with user 101, processing additional utterances, and / or other costs. Thus, in this example, the automatic assistant can refrain from initiating the full fulfillment of utterance 552A1, based at least on the relatively high computational cost associated with quickly fulfilling utterance 552A1.
[0071] As a result, the automatic assistant may determine to provide a natural conversation output 554A, such as "Mmhmm", as shown in FIG. 5A, for audible presentation to the user 101 via the speaker of the client device 110 (and, optionally, in response to determining that the user 101 has paused for N seconds after providing the utterance 552A1, as shown by 552A2). The natural conversation output 554A can be provided for audible presentation to the user 101 to provide an indication that the automatic assistant is still listening and waiting for the user 101 to complete the provision of the utterance 552A1. In particular, in various implementations, while the automatic assistant provides the natural conversation output 554A for presentation to the user 101, the automatic assistant components (e.g., the ASR engines 120A1 and / or 120A2, the NLU engines 130A1 and / or 130A2, the fulfillment engines 140A1 and / or 140A2, and / or other automatic assistant components of FIG. 1 such as the acoustic engine 161 of FIG. 1) utilized in processing the audio data stream can remain active on the client device 110. Further, in various implementations, the natural conversation output 554A can be provided for audible presentation to the user 101 at a lower volume than other audible outputs to avoid interfering with the user 101's completion of the utterance 552A1 and to reflect a more natural conversation between actual humans.
[0072] In the example of FIG. 5A, it is further assumed that the user 101 completed the utterance 552A1 by providing the utterance 556A of "Call Arnold's Trattoria", where "Arnold's Trattoria" is a fictional Italian restaurant. Based on the fact that the user 101 completed the utterance 552A1 by providing the utterance 556A, the automatic assistant can update the streams of ASR output, NLU output, and fulfillment data. In particular, the automatic assistant can determine that the updated stream of NLU data still includes the predicted intent of "call" or "phone call", but with respect to the callee entity parameter associated with the previously predicted "call" or "phone call", it has a slot value of "Arnold's Trattoria" instead of "Arnold". Thus, in response to the user 101 completing the utterance 552A1 by providing the utterance 556A, the automatic assistant can cause the client device 101 (or an additional client device communicating with the client device 101, e.g., a mobile device associated with the user 101) to initiate a call to "Arnold's Trattoria" and, optionally, provide the synthetic voice 558A of "Okay, calling Arnold's Trattoria" for audible presentation to the user 101.In these and other ways, the automated assistant can refrain from prematurely and incorrectly acting on the predicted intent of user 101 as determined based on utterance 552A1 (e.g., by calling the contact entry “Arnold”), and user 101 can wait for their thought process to complete in order to correctly act on the predicted intent of user 101 as determined based on user 101 completing utterance 552A1 via utterance 556A (e.g., by calling the fictional restaurant “Arnold's Trattoria”).
[0073] As another example, particularly referring to FIG. 5B, assume that user 101 of client device 110 provides utterance 552B1 of "Assistant, call Arnolllld's", and then pauses for N seconds as shown by 552B2 and then resumes, where N is any positive integer and / or its fraction (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.). Similar to FIG. 5A, even if the automatic assistant may determine that utterance 552B1 can be fulfilled, the automatic assistant may refrain from starting to fulfill utterance 552B1 based on the audio-based features associated with utterance 552B1 and / or based on one or more computational costs associated with performing and / or canceling the fulfillment of utterance 552B1. Assume that the automatic assistant determines to provide a natural conversation output 554B1 such as "Mmhmm" as shown in FIG. 5B and causes the natural conversation output 554B1 to be provided for audible presentation to user 101 of client device 110. However, in the example of FIG. 5B, in contrast to the example of FIG. 5A, assume that user 101 of client device 110 did not complete utterance 554B1 within M seconds after causing the natural conversation output 554B1 to be provided for audible presentation to user 101 as shown by 554B2, where M is any positive integer and / or its fraction (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.) that may be the same as or different from N seconds as shown by 552B2.
[0074] As a result, in the example of FIG. 5B, the automatic assistant can determine additional natural conversation output 556B to be provided for an audible prompt to user 101 of client device 110. In particular, rather than causing the voice backchannel to be provided for an audible prompt to user 101 of client device 110, similar to natural conversation output 554B1 indicating that the automatic assistant is waiting for user 101 to complete utterance 552B1, additional natural conversation output 556B can more explicitly indicate that the automatic assistant is waiting for user 101 to complete utterance 552B1 and / or (as will be described below with respect to FIG. 5C) can request that user 101 provide a specific input for the facilitation of the dialog session. In the example of FIG. 5B, in response to additional natural conversation 556B being provided for an audible prompt to user 101, further assume that user 101 of client device 110 provides utterance 558B of "Sorry, call Arnold's Trattoria" to complete the provision of utterance 552B1. Thus, in response to user 101 completing the provision of utterance 552B1 by providing utterance 558B, the automatic assistant causes a call to "Arnold's Trattoria" to be initiated on client device 110 (or an additional client device communicating with client device 110, e.g., user 101's mobile device), and optionally, can cause synthetic voice 560B of "Okay, calling Arnold's Trattoria" to be provided for an audible prompt to user 101.Similar to FIG. 5B, the automatic assistant can refrain from prematurely performing the predicted intention of user 101 determined based on utterance 552B1 (e.g., by calling the contact entry "Arnold"), and even if user 101 can pause for a longer duration as in the example of FIG. 5B, user 101 can wait for his or her thoughts to be completed in order to correctly perform the predicted intention of user 101 when completing utterance 552B1 via utterance 558B (e.g., by calling the fictional restaurant "Arnold's Trattoria").
[0075] As yet another example, and with particular reference to FIG. 5C, assume that user 101 of client device 110 provides utterance 552C1 of "Assistant, make a reservation tonight at Arnold's Trattoria for six people", and then pauses for N seconds as shown by 552C2, where N is any positive integer and / or fraction thereof (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.). In this example, based on processing utterance 552C1, the ASR output includes the recognized text corresponding to utterance 552C1 captured within the audio data stream (e.g., the recognized text corresponding to "make a reservation tonight at Arnold's Trattoria for six people"), and the NLU data stream includes the predicted slot value of "Arnold's Trattoria" for the restaurant entity parameter associated with the predicted intention of "reservation" or "restaurant reservation", the slot value of "[today's date]" for the reservation date parameter associated with the predicted intention of "reservation" or "restaurant reservation", and the slot value of "6" for the number of people parameter associated with the predicted intention of "reservation" or "restaurant reservation", including the predicted intention of "reservation" or "restaurant reservation". In particular, when providing utterance 552C1, user 101 of client device 110 did not provide a slot value for the time parameter associated with the intention of "reservation" or "restaurant reservation".As a result, based on the stream of NLU data, the automatic assistant may determine that user 101 has paused providing utterance 552C1.
[0076] Furthermore, assume that when the stream of fulfillment data is executed, it includes an assistant command to cause client device 110 to make a restaurant reservation using a restaurant reservation software application accessible on client device 110 and / or a restaurant reservation agent accessible on client device 110 (e.g., one of the 1P agent 171 and / or 3P agents of FIG. 1). In the example of FIG. 5C, in contrast to the examples of FIGS. 5A and 5B, based on processing utterance 552C1 and without processing any additional utterances, the automatic assistant may determine that utterance 552C1 can be fulfilled by executing the assistant command. In this example, the automatic assistant may start fulfillment of utterance 552C1 based on NLU measurements associated with a stream of NLU data indicating that user 101 intends to make a restaurant reservation but did not provide a slot value for time parameters associated with the intent of simply "reservation" or "restaurant reservation". Thus, the automatic assistant can establish a connection with a restaurant reservation software application accessible on client device 110 and / or a restaurant reservation agent accessible on client device 110 (e.g., one of the 1P agent 171 and / or 3P agents of FIG. 1) and start providing a slot value to begin making a reservation, even though full fulfillment of utterance 552C1 may not be possible.
[0077] Specifically, when the automatic assistant starts fulfilling the utterance 552C1, since the automatic assistant determines that the user 101 has paused providing the utterance 552C1 based at least on the stream of NLU data, the automatic assistant still decides to provide a natural conversation output 554C1 such as "Uh huhh" as shown in FIG. 5C, and the natural conversation output 554C1 can be provided for audible presentation to the user 110 of the client device 110. However, in the example of FIG. 5C, similar to FIG. 5B, assume that the user 101 of the client device 110 does not complete the utterance 552C1 within M seconds after providing the natural conversation output 554C1 for audible presentation to the user 101 as indicated by 554C2, where M is any positive integer and / or its decimal (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.) that may be the same as or different from N seconds as indicated by 552C2.
[0078] As a result, in the example of FIG. 5C, the automatic assistant can determine additional natural conversation output 556C to be provided for an audible prompt to user 101 of client device 110. In particular, rather than causing an audio backchannel to be provided for an audible prompt to user 101 of client device 110, similar to natural conversation output 554C1 indicating that the automatic assistant is waiting for user 101 to complete utterance 552C1, the additional natural conversation output 556C can request that user 101 provide a specific input, such as "For what time?", to facilitate the dialog session, based on the fact that user 101 did not provide a slot value regarding a time parameter associated with the intent of "reservation" or "restaurant reservation". Further, in the example of FIG. 5C, in response to the additional natural conversation 556C being provided for an audible prompt to the user, assume that user 101 of client device 110 provides utterance 558C of "7:00PM" to complete the provision of utterance 552C1. Thus, the automatic assistant can complete fullfilment of the assistant command using a previously unknown slot value in response to user 101 completing utterance 552C1 by providing utterance 558C, and make a restaurant reservation on behalf of user 101. In these and other ways, the automatic assistant can wait for user 101 to complete their thought by providing natural conversation output 554C1, and then prompt user 101 to complete their thought by providing natural conversation output 556C if user 101 does not complete their thought in response to the provision of natural conversation output 554C1.
[0079] As yet another example, referring particularly to FIG. 5D, assume that user 101 of client device 110 provides utterance 552D1 of "Assistant, what's on my calendar forrrr", and then pauses for N seconds as shown by 552D2, where N is any positive integer and / or its fraction (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.). In this example, based on processing utterance 552C1, the ASR output includes recognized text corresponding to utterance 552D1 captured within the audio data stream (e.g., recognized text corresponding to "what's on my calendar for"), and assume that the NLU data stream includes a predicted "calendar" or "calendar search" intent having an unknown slot value regarding data parameters. In this example, the automatic assistant may determine that user 101 paused the provision of utterance 552D1 because the user did not provide a slot value regarding data parameters based on the NLU data stream. Additionally or alternatively, in this example, the automatic assistant may determine that user 101 paused the provision of utterance 552D1 as indicated by the elongated syllable (e.g., "rrrr" when providing "forrrr" within utterance 552D1) included within utterance 552D1 based on the audio-based features of utterance 552D1.
[0080] Furthermore, assume that when the fulfillment data stream is executed, it includes an assistant command that causes the client device 110 to search for the calendar information of the user 101 using the calendar software application accessible on the client device 110 and / or the calendar agent accessible on the client device 110 (e.g., one of the 1P agent 171 and / or 3P agents in FIG. 1). In the example of FIG. 5D, in contrast to the examples of FIGS. 5A - 5C, based on processing the utterance 552D1 and without processing any additional utterances, the automatic assistant may determine that the utterance 552D1 can be fulfilled by executing the assistant command. In this example, based on the NLU measurement associated with the NLU data stream indicating that the user 101 intends to search for one or more calendar entries but did not provide a slot value for the data parameter associated with the intention of simply "calendar" or "calendar search", the automatic assistant may initiate the fulfillment of the utterance 552D1. Thus, the automatic assistant can establish a connection with the calendar software application accessible on the client device 110 and / or the calendar agent accessible on the client device 110 (e.g., one of the 1P agent 171 and / or 3P agents in FIG. 1).
[0081] When the automated assistant begins fulfillment of utterance 552D1, since the automated assistant determines, based on a stream of NLU data, that user 101 has paused provision of utterance 552D1, the automated assistant decides to still provide a natural conversation output 554D1, such as "Uh huhh" as shown in FIG. 5D, and can cause the natural conversation output 554D1 to be provided for audible presentation to user 110 of client device 110. However, in the example of FIG. 5D, similar to FIGS. 5B and 5C, assume that user 101 of client device 110 did not complete utterance 552D1 within M seconds after providing the natural conversation output 554D1 for audible presentation to user 101, as indicated by 554D2, where M is any positive integer and / or fraction thereof (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.) that may be the same as or different from N seconds as indicated by 552D2.
[0082] However, in the example of FIG. 5D, the automatic assistant can determine to execute the full fulfillment of utterance 552D1 even if user 101 was unable to complete utterance 552D1. The automatic assistant can make this determination based on one or more of the computational costs associated with executing the full fulfillment and / or canceling any executed full fulfillment. In this example, one or more of the computational costs can include providing the synthetic voice 556D1 of "You have two calendar entries for today..." for audible presentation to user 101 in the full fulfillment of utterance 552D1, and providing other synthetic voice speech for audible presentation to user 101 if user 101 desires calendar information for another day. Thus, in an attempt to conduct the dialog session more quickly because the computational cost of doing so is relatively low, the automatic assistant can proceed and determine to execute the full fulfillment of utterance 552D1 using the estimated slot value for the current day for data parameters associated with the intent of "calendar" or "calendar search".
[0083] In particular, in various implementation forms, while the automatic assistant provides the synthesized voice 556D1 for presentation to the user 101, the automatic assistant components (e.g., ASR engines 120A1 and / or 120A2, NLU engines 130A1 and / or 130A2, fulfillment engines 140A1 and / or 140A2, and / or other automatic assistant components in FIG. 1 such as the acoustic engine 161 in FIG. 1) utilized when processing the audio data stream can remain active on the client device 110. Thus, in these implementation forms, if the user 101 interrupts the automatic assistant during the audible presentation of the synthesized voice 556D1 by providing another utterance requesting a different date than the estimated current date, the automatic assistant can quickly and efficiently adapt the fulfillment of the utterance 552D1 based on the different date provided by the user 101. In additional or alternative implementation forms, after providing the synthesized voice 556D1 for audible presentation to the user 101, the automatic assistant can auditorily render an additional synthesized voice 556D2, such as "Wait, did I cut you off a second ago?", to actively provide the user 101 with an opportunity to correct the fulfillment of the utterance 552D1. In these and other ways, the automatic assistant can balance waiting for the user 101 to complete their thought by providing a natural conversation output 554D1 and more quickly and efficiently ending the dialog session by fulfilling the utterance 552D1 when the computational cost of doing so is relatively low.
[0084] Regarding the example of FIGS. 5A - 5D, it has been described that natural conversation output is provided for audible presentation to user 101, but it should be understood that this is for example purposes and is not meant to be limiting. For example, referring briefly to FIG. 5E, assume that user 101 of client device 110 provides the utterance "Assistant, call Arnolllld's" and then pauses for N seconds and then resumes, where N is any positive integer and / or its fraction (e.g., 2 seconds, 2.5 seconds, 3 seconds, etc.). In the example of FIG. 5E, a streaming transcript 552E of the utterance can be provided for visual presentation to the user via the display 190 of client device 110. In some implementations, the display 190 of client device 110 can additionally or alternatively provide one or more graphical elements 191, such as an ellipse added to the streaming transcript 552E that can move on the display 190, indicating that the auto - assistant is waiting for user 101 to complete the utterance. The graphical element 191 shown in FIG. 5E is an ellipse added to the streaming transcript, but this is for example purposes and is not meant to be limiting, and it should be understood that any other graphical element can be provided for visual presentation to user 101 to indicate that the auto - assistant is waiting for user 101 to complete the utterance. In additional or alternative implementations, one or more LEDs can be lit (e.g., as indicated by dashed line 192) to indicate that the auto - assistant is waiting for user 101 to complete the utterance, which can be particularly advantageous when client device 110 does not have a display 190. Further, the examples of FIGS. 5A - 5E are provided for example purposes and are not meant to be limiting.
[0085] Furthermore, in an implementation where the client device 110 of the user 101 includes a display 190, when the user provides an utterance, one or more selectable graphical elements associated with various interpretations of the utterance can be provided for visual presentation to the user. The automatic assistant can start fulfilling the utterance based on receiving a user selection from the user 101 of a given one of the one or more selectable graphical elements and / or in response to no user selection from the user 101 being received within a threshold duration, based on the NLU measurement value associated with a given one of the one or more selectable graphical elements. For example, in the example of FIG. 5A, after receiving the utterance 552A1 of "Assistant, call Arnolllld's", a first selectable graphical element can be provided for presentation to the user 101 via the display, and when the first selectable graphical element is selected, it causes the automatic assistant to call the contact entry associated with "Arnold". However, if the user continues to provide the utterance 556A of "call Arnold's Trattoria", one or more selectable graphical elements can be updated to include a second selectable graphical element, and when the second selectable graphical element is selected, it causes the automatic assistant to call the restaurant associated with "Arnold's Trattoria".In this example, assuming that user 101 does not provide any user selection of the first selectable graphical element or the second selectable graphical element within the threshold duration (with respect to the first selectable graphical element presented or the second selectable graphical element presented), the automatic assistant can initiate a call with the restaurant "Arnold's Trattoria" based on the NLU measurement associated with initiating a call with the restaurant "Arnold's Trattoria" better indicating the true intent of user 101 as compared to the NLU measurement associated with initiating a call with the contact entry "Arnold".
[0086] Turning now to FIG. 6, a block diagram of an exemplary computing device 610 that can optionally be utilized to execute one or more aspects of the techniques described herein is shown. In some implementations, one or more of a client device, a cloud-based automatic assistant component, and / or other components can comprise one or more components of the exemplary computing device 610.
[0087] The computing device 610 typically includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices can include, for example, a storage subsystem 624 that includes a memory subsystem 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable a user to interact with the computing device 610. The network interface subsystem 616 provides an interface to an external network and is coupled to a corresponding interface device within other computing devices.
[0088] The user interface input device 622 can include a keyboard, a mouse, a trackball, a touchpad, or a pointing device such as a graphics tablet, a scanner, a touch screen incorporated in a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into a computing device or a communication network.
[0089] The user interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 610 to a user or another machine or computing device.
[0090] The memory subsystem 624 stores programming structures and data structures that provide some or all of the functionality of some of the modules described herein. For example, the memory subsystem 624 can include logic for performing selected aspects of the methods disclosed herein and for implementing the various components shown in FIGS. 1 and 2.
[0091] These software modules are generally executed by the processor 614 alone or in combination with other processors. The memory 625 used within the memory subsystem 624 can include several memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution, and a read-only memory (ROM) 632 in which fixed instructions are stored. The file storage subsystem 626 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive with an associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functions of a particular implementation can be stored by the file storage subsystem 626 within the memory subsystem 624 or in other machines accessible by the processor 614.
[0092] The bus subsystem 612 provides a mechanism for enabling the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem 612 can use multiple buses.
[0093] The computing device 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 610 shown in FIG. 6 is intended only as a specific example for the purpose of illustrating some implementations. Many other configurations of the computing device 610 can have more or fewer components than the computing device shown in FIG. 6.
[0094] In situations where the systems described in this specification may collect or otherwise monitor personal information about a user, or may use personal information and / or monitoring information, the user may be provided with an opportunity to control whether the program or function collects user information (e.g., information regarding the user's social network, social behavior or activities, occupation, user preferences, or the user's current geographic location), or an opportunity to control whether and / or how the user receives content from a content server that may be relevant to the user. Also, certain data may be processed in one or more ways before it is stored or used so that information that could identify an individual is removed. For example, the user's identification information may be processed so that information that could identify an individual cannot be determined about the user, or the user's geographic location may be generalized at the location where the geographic location information is obtained (such as city, zip code, or state level) so that the user's specific geographic location cannot be determined. Thus, the user may control how information is collected and / or used about the user.
[0095] In some implementations, a method implemented by one or more processors is provided for processing a stream of audio data to generate a stream of ASR output using an automatic speech recognition (ASR) model, the stream of audio data being generated by one or more microphones of a user's client device, the stream of audio data capturing a portion of an utterance provided by the user directed to an automatic assistant at least partially implemented on the client device; processing the stream of ASR output to generate a stream of NLU output using a natural language understanding (NLU) model; determining audio-based features associated with a portion of the utterance based on processing the stream of audio data; determining whether the user has paused or completed providing the utterance based on the audio-based features associated with the portion of the utterance; determining natural conversation output to be provided for an audible prompt to the user in response to determining that the user has paused providing the utterance and in response to determining that the automatic assistant can start fulfilling the utterance based at least on the NLU output, the natural conversation output to be provided for an audible prompt to the user indicating that the automatic assistant is waiting for the user to complete providing the utterance; and causing the natural conversation output to be provided for an audible prompt to the user via one or more speakers of the client device.
[0096] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.
[0097] In some implementations, the step of causing the natural conversation output to be provided for an audible prompt to the user via one or more speakers of the client device can further be in response to determining that the user has paused providing the utterance for a threshold duration.
[0098] In some implementations, the step of determining whether the user has paused or completed the utterance based on audio-based features associated with a portion of the utterance may include processing the audio-based features associated with a portion of the utterance to generate an output using an audio-based classification machine learning (ML) model, and determining whether the user has paused or completed the utterance based on the output generated using the audio-based classification ML model.
[0099] In some implementations, the method may further include generating a stream of fulfillment data based on a stream of NLU outputs. The step of determining that the automated assistant can begin fulfillment of the utterance may further be based on the stream of fulfillment data. In some variations of those implementations, the method may further include causing the automated assistant to begin fulfillment of the utterance based on the stream of fulfillment data in response to determining that the user has completed providing the utterance. In additional or alternative variations of those implementations, the method may further include keeping one or more automated assistant components that utilize an ASR model active while providing natural conversation output for audible presentation to the user via one or more speakers of the client device. In additional or alternative variations of those implementations, the method may further include determining whether the utterance includes a particular word or phrase based on a stream of ASR outputs, refraining from determining whether the user has paused or completed providing the utterance based on audio-based features associated with a portion of the utterance in response to determining that the utterance includes the particular word or phrase, and causing the automated assistant to begin fulfillment of the utterance based on the stream of fulfillment data.In additional or alternative variations of those implementations, the method, after the step of providing a natural conversation output for audible presentation to the user via one or more speakers of a client device, includes determining whether the user continues to provide utterances within a threshold duration, and in response to determining that the user does not continue to provide one or more utterances within the threshold duration, determining whether the automatic assistant can start fulfillment of the utterance based on a stream of NLU data and / or a stream of fulfillment data, and in response to determining that the automatic assistant can start fulfillment of the utterance based on the stream of fulfillment data, causing the automatic assistant to start fulfillment of the utterance based on the stream of fulfillment data.
[0100] In some implementations, the method, after the step of providing a natural conversation output for audible presentation to the user via one or more speakers of a client device, includes determining whether the user continues to provide utterances within a threshold duration, and in response to determining that the user does not continue to provide utterances, determining additional natural conversation output to be provided for audible presentation to the user, where the additional natural conversation output to be provided for audible presentation to the user requests that the user complete providing an utterance, and causing the additional natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device.
[0101] In some implementations, the method includes providing, via one or more speakers of a client device, a natural conversation output for audible presentation to a user, while providing, via a display of the client device, one or more graphical elements for visual presentation to the user, wherein the one or more graphical elements to be provided for visual presentation to the user may further include a step indicating that the automatic assistant is waiting for the user to complete providing an utterance. In some variations of those implementations, the ASR output may include a streaming transcription corresponding to a portion of an utterance captured within a stream of audio data, and the method includes providing, via one or more speakers of the client device, a natural conversation output for audible presentation to the user, while providing, via a display of the client device, the streaming transcription for visual presentation to the user, wherein the one or more graphical elements are added to the beginning or end of the streaming transcription provided for visual presentation to the user via the display of the client device.
[0102] In some implementations, the method includes lighting one or more light-emitting diodes (LEDs) of a client device while providing, via one or more speakers of the client device, a natural conversation output for audible presentation to the user, wherein the one or more LEDs are lit to indicate that the automatic assistant is waiting for the user to complete providing an utterance.
[0103] In some implementations, audio-based features associated with a portion of an utterance may include one or more of intonation, tone, stress, rhythm, tempo, pitch, pauses, one or more grammars associated with the pauses, and elongated syllables.
[0104] In some implementations, the step of determining the natural conversation output to be provided for audible presentation to the user may include maintaining a set of natural conversation outputs in the on-device memory of the client device and selecting a natural conversation output from the set of natural conversation outputs based on audio-based features associated with a portion of the utterance.
[0105] In some implementations, the step of causing a natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device may include causing a natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device at a lower volume than other outputs provided for audible presentation to the user.
[0106] In some implementations, the step of causing a natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device may include processing the natural conversation output to generate synthetic voice audio data including the natural conversation output using a text-to-speech (TTS) model and causing the synthetic voice audio data to be provided for audible presentation to the user via one or more speakers of the client device.
[0107] In some implementations, the step of causing a natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device may include obtaining synthetic voice audio data including the natural conversation output from the on-device memory of the client device and causing the synthetic voice audio data to be provided for audible presentation to the user via one or more speakers of the client device.
[0108] In some implementations, one or more processors may be implemented locally by the user's client device.
[0109] In some implementations, a method implemented by one or more processors is provided for processing a stream of audio data to generate a stream of ASR output using an automatic speech recognition (ASR) model, the stream of audio data being generated by one or more microphones of a client device and capturing a portion of a user utterance directed to an automatic assistant at least partially implemented on the client device; processing the stream of ASR output to generate a stream of NLU output using a natural language understanding (NLU) model; determining, based at least on the stream of NLU output, whether the user has paused or completed providing an utterance; determining a natural conversation output to be provided for an audible prompt to the user in response to determining that the user has paused providing an utterance and has not completed providing an utterance, the natural conversation output to be provided for an audible prompt to the user indicating that the automatic assistant is waiting for the user to complete providing an utterance; and causing the natural conversation output to be provided for an audible prompt to the user via one or more speakers of the client device.
[0110] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.
[0111] In some implementations, the step of determining whether the user has paused or completed providing an utterance based on a stream of NLU outputs may include the step of determining whether the auto-assistant can start fulfilling the utterance based on the stream of NLU outputs. The step of determining that the user has paused providing an utterance may include the step of determining that the auto-assistant cannot start fulfilling the utterance based on the stream of NLU outputs. In some variations of those implementations, the method includes, after the step of providing natural conversation output for audible presentation to the user via one or more speakers of the client device, determining whether the user continues to provide an utterance within a threshold duration, and in response to determining that the user does not continue to provide an utterance, determining additional natural conversation output to be provided for audible presentation to the user, wherein the additional natural conversation output to be provided for audible presentation to the user requires that the user complete providing the utterance, and further including the step of providing the additional natural conversation output for audible presentation to the user via one or more speakers of the client device. In some further variations of those implementations, the additional natural conversation output to be provided for audible presentation to the user may require that an additional portion of the utterance includes specific data based on the stream of NLU data.
[0112] In some implementations, a method implemented by one or more processors is provided for processing a stream of audio data to generate a stream of ASR output using an automatic speech recognition (ASR) model, wherein the stream of audio data is generated by one or more microphones of a client device, and the stream of audio data captures a portion of a user's utterance directed to an automatic assistant at least partially implemented on the client device; processing the stream of ASR output to generate a stream of NLU output using a natural language understanding (NLU) model; determining whether the user has paused or completed providing an utterance; in response to determining that the user has paused providing an utterance and has not completed providing the utterance, determining a natural conversation output to be provided for an audible prompt to the user, wherein the natural conversation output to be provided for an audible prompt to the user indicates that the automatic assistant is waiting for the user to complete providing the utterance; causing the natural conversation output to be provided for an audible prompt to the user via one or more speakers of the client device; after causing the natural conversation output to be provided for an audible prompt to the user via one or more speakers of the client device, determining, in response to determining that the user has not completed providing the utterance within a threshold duration, whether the automatic assistant can initiate fulfillment of the utterance based at least on the stream of NLU data; and in response to determining that the automatic assistant can initiate fulfillment of the utterance based on the stream of NLU data, causing the automatic assistant to initiate fulfillment of the utterance.
[0113] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.
[0114] In some implementations, the method may further include determining audio-based features associated with a portion of the utterance based on processing a stream of audio data. The step of determining whether the user has paused or completed providing the utterance may be based on audio-based features associated with a portion of the utterance.
[0115] In some implementations, the step of determining whether the user has paused or completed providing the utterance may be based on a stream of NLU data.
[0116] In some implementations, the method is a step of determining a natural conversation output to be provided for an audible prompt to the user in response to determining that the automatic assistant cannot start fulfillment of the utterance based on a stream of NLU data, where the natural conversation output to be provided for an audible prompt to the user requires the user to complete providing the utterance, and may further include causing additional natural conversation output to be provided for an audible prompt to the user via one or more speakers of the client device. In some variations of those implementations, the natural conversation output to be provided for an audible prompt to the user may require that an additional portion of the utterance includes specific data based on a stream of NLU data.
[0117] In some implementations, the step of determining whether the automatic assistant can start fulfillment of the utterance may be further based on one or more computational costs associated with fulfillment of the utterance. In some variations of those implementations, the one or more computational costs associated with fulfillment of the utterance may include one or more of the computational costs associated with performing fulfillment of the utterance and the computational costs associated with canceling the performed fulfillment of the utterance.
[0118] In some implementations, the method may further include generating a stream of fulfillment data based on a stream of NLU outputs. The step of determining that the automated assistant can start fulfillment of the utterance may further be based on the stream of fulfillment data.
[0119] In some implementations, a method implemented by one or more processors is provided, the method including receiving a stream of audio data, the stream of audio data being generated by one or more microphones of a user's client device and capturing at least a portion of an utterance provided by the user directed to an automated assistant at least partially implemented on the client device; determining audio-based features associated with a portion of the utterance based on processing the stream of audio data; determining whether the user has paused or completed providing the utterance based on the audio-based features associated with the portion of the utterance; determining natural conversation output to be provided for an audible prompt to the user in response to determining that the user has paused providing the utterance and has not completed providing the utterance, the natural conversation output to be provided for an audible prompt to the user indicating that the automated assistant is waiting for the user to complete providing the utterance; and providing the natural conversation output for an audible prompt to the user via one or more speakers of the client device.
[0120] In addition, some implementations include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) of one or more computing devices, and the one or more processors are operable to execute instructions stored in an associated memory, and the instructions are configured to cause execution of any of the methods described above. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to execute any of the methods described above. Some implementations also include a computer program product including instructions executable by one or more processors to execute any of the methods described above.
Explanation of Signs
[0121] 101 User 110 Client Device 111 User Input Engine 112 Rendering Engine 113 Presence Sensor 114 Assistant Client 115 Assistant 115A Machine Learning (ML) Model Database, ML Model Database 120A1 Automatic Speech Recognition (ASR) Engine, ASR Engine 120A2 ASR Engine 130A1 Natural Language Understanding (NLU) Engine, NLU Engine 130A2 NLU Engine 140A1 Fulfillment Engine 140A2 Fulfillment Engine 150A1 Text-to-Speech (TTS) Engine, TTS Engine 150A2 TTS Engine 160 Natural Conversation Engine 161 Acoustic Engine 162 Pause Engine 163 Time Engine, Natural Conversation Output Engine, Natural Conversation Engine 164 Natural Conversation Output Engine, Fulfillment Output Engine 165 Fulfillment Output Engine, Time Engine 171 First Party (1P) Agent, 1P Agent 172 Third Party (3P) Agent, 3P Agent 180 Natural Conversation System 190 Display 191 Graphical Element 192 Dashed Line 199 Network 201A Stream of Audio Data 201B Stream of Non-Audio Data 220 Stream of ASR Output 230 Stream of NLU Output 240 Stream of Fulfillment Data 240A 1P Fulfillment Data 240B 3P Fulfillment Data 261 Audio-Based Feature 262 Block 263 Natural Conversation Output 264 Fulfillment Output 265 Duration of Pause 552A1 Utterance 552B1 Utterance 552C1 Utterance 552D1 Utterance 552E Streaming Transcription 554A Natural Conversation Output 554B1 Natural Conversation Output 554C1 Natural Conversation Output 554D1 Natural Conversation Output 556A Utterance 556B Additional Natural Conversation Output, Additional Natural Conversation 556C Additional Natural Conversation Output, Additional Natural Conversation, Natural Conversation Output 556D1 Synthetic Voice 556D2 Additional synthetic voice 558A Synthetic voice 558B Speech 558C Speech 560B Synthetic voice 610 Computing device 612 Bus subsystem 614 Processor 616 Network interface subsystem 620 User interface output device 622 User interface input device 624 Memory subsystem 625 Memory subsystem 626 File memory subsystem 630 Main random access memory (RAM) 632 Read-only memory (ROM)
Claims
1. A method implemented by one or more processors, the method comprising: processing a stream of audio data to generate a stream of ASR output using an automatic speech recognition (ASR) model, the stream of audio data being generated by one or more microphones of a user's client device, the stream of audio data capturing a portion of an utterance provided by the user directed to an automatic assistant at least partially implemented on the client device; processing the stream of ASR output to generate a stream of NLU output using a natural language understanding (NLU) model; determining, based on the stream of NLU output, that the automatic assistant can start fulfillment of the utterance; determining audio-based features associated with the portion of the utterance based on processing the stream of audio data; determining, based on the audio-based features associated with the portion of the utterance, whether the user has paused or completed providing the utterance; in response to determining that the user has paused providing the utterance and in response to determining that the automatic assistant can start fulfillment of the utterance based at least on the stream of NLU output, holding off on starting fulfillment of the utterance; determining a natural conversation output to be provided for an audible prompt to the user, the natural conversation output to be provided for the audible prompt to the user indicating that the automatic assistant is waiting for the user to complete providing the utterance; causing the natural conversation output to be provided for the audible prompt to the user via one or more speakers of the client device comprising a method.
2. The step of providing the natural conversation output for audible presentation to the user via the one or more speakers of the client device, further in response to determining that the user has paused the provision of the utterance for a threshold duration, the method according to claim 1.
3. The step of determining whether the user has paused or completed the provision of the utterance based on the audio-based features associated with the portion of the utterance, processing the audio-based features associated with the portion of the utterance to generate an output using an audio-based classification machine learning (ML) model; determining whether the user has paused or completed the provision of the utterance based on the output generated using the audio-based classification ML model; comprising the method according to claim 1 or 2.
4. further comprising the step of generating a stream of fulfillment data based on the stream of NLU output, the step of determining that the automated assistant can initiate fulfillment of the utterance is further based on the stream of fulfillment data, the method according to any one of claims 1 to 3.
5. further comprising the step of causing the automated assistant to initiate fulfillment of the utterance based on the stream of fulfillment data in response to determining that the user has completed the provision of the utterance, the method according to claim 4.
6. further comprising the step of keeping one or more automated assistant components utilizing the ASR model active while providing the natural conversation output for audible presentation to the user via one or more speakers of the client device, the method according to claim 4.
7. determining whether the utterance contains a specific word or phrase based on the stream of ASR output; in response to determining that the utterance contains the specific word or phrase, refraining from determining whether the user has paused or completed the provision of the utterance based on the audio-based features associated with the portion of the utterance; causing the automatic assistant to initiate fulfillment of the utterance based on the stream of fulfillment data; The method according to claim 4, further comprising.
8. After providing the natural conversation output for audible presentation to the user via one or more speakers of the client device, determining whether the user continued to provide the utterance within a threshold duration; In response to determining that the user did not continue to provide the one or more utterances within the threshold duration, determining whether the automatic assistant can initiate fulfillment of the utterance based on a stream of NLU data and / or a stream of the fulfillment data; In response to determining that the automatic assistant can initiate fulfillment of the utterance based on the stream of fulfillment data, causing the automatic assistant to initiate fulfillment of the utterance based on the stream of fulfillment data; The method according to claim 4, further comprising.
9. After providing the natural conversation output for audible presentation to the user via the one or more speakers of the client device, determining whether the user continued to provide the utterance within a threshold duration; In response to determining that the user did not continue to provide the utterance, determining additional natural conversation output to be provided for audible presentation to the user, wherein the additional natural conversation output to be provided for audible presentation to the user requests that the user complete providing the utterance; causing the additional natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device; The method according to any one of claims 1 to 8, further comprising.
10. While providing the natural conversation output for audible presentation to the user via one or more speakers of the client device, providing one or more graphical elements for visual presentation to the user via a display of the client device, the method according to any one of claims 1 to 9, further comprising a step in which the one or more graphical elements to be provided for visual presentation to the user indicate that the automatic assistant is waiting for the user to complete providing the utterance.
11. The ASR output includes a streaming transcription corresponding to the portion of the utterance captured within the stream of audio data, While providing the natural conversation output for audible presentation to the user via one or more speakers of the client device, providing the streaming transcription for visual presentation to the user via a display of the client device, the method further comprising a step in which the one or more graphical elements are added to the beginning or end of the streaming transcription provided for visual presentation to the user via the display of the client device. The method according to claim 10.
12. While providing the natural conversation output for audible presentation to the user via one or more speakers of the client device, lighting one or more light emitting diodes (LEDs) of the client device, the method according to any one of claims 1 to 11, further comprising a step in which the one or more LEDs are lit to indicate that the automatic assistant is waiting for the user to complete providing the utterance.
13. The audio-based features associated with the portion of the utterance include one or more of intonation, tone, stress, rhythm, tempo, pitch, pause, one or more grammars related to the pause, and elongated syllables. The method according to any one of claims 1 to 12.
14. The step of determining the natural conversation output to be provided for audible presentation to the user is maintaining a set of natural conversation outputs in on-device memory of the client device; selecting the natural conversation output from among the set of natural conversation outputs based on the audio-based features associated with the portion of the utterance comprising The method according to any one of claims 1 to 13.
15. causing the natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device, causing the natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device at a lower volume than other output provided for audible presentation to the user, The method according to any one of claims 1 to 14.
16. causing the natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device, processing the natural conversation output to generate synthetic audio data including the natural conversation output using a text-to-speech (TTS) model; causing the synthetic audio data to be provided for audible presentation to the user via one or more speakers of the client device comprising The method according to any one of claims 1 to 15.
17. causing the natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device, retrieving synthetic audio data including the natural conversation output from on-device memory of the client device; causing the synthetic audio data to be provided for audible presentation to the user via one or more speakers of the client device comprising The method according to any one of claims 1 to 16.
18. The method according to any one of claims 1 to 17, wherein the one or more processors are implemented locally by the client device of the user.
19. A method implemented by one or more processors on a client device of a user, the method comprising Processing a stream of audio data to generate a stream of ASR output using an automatic speech recognition (ASR) model, wherein the stream of audio data is generated by one or more microphones of the client device, and the stream of audio data captures a portion of the user's utterance directed to an automatic assistant at least partially implemented on the client device; Processing the stream of ASR output to generate a stream of NLU output using a natural language understanding (NLU) model; Based on the stream of the NLU output, determining that the automatic assistant can start fulfillment of the utterance; Based on at least the stream of the NLU output, determining whether the user has paused or completed providing the utterance; In response to determining that the user has paused providing the utterance and has not completed providing the utterance, Holding off on starting fulfillment of the utterance; Determining a natural conversation output to be provided for an audible prompt to the user, wherein the natural conversation output to be provided for the audible prompt to the user indicates that the automatic assistant is waiting for the user to complete providing the utterance; Causing the natural conversation output to be provided for the audible prompt to the user via one or more speakers of the client device comprising a method.
20. The step of determining whether the user has paused or completed providing the utterance based on the stream of the NLU output includes the step of determining whether the automatic assistant can start fulfillment of the utterance based on the stream of the NLU output, The step of determining that the user has paused providing the utterance includes the step of determining that the automatic assistant cannot start fulfillment of the utterance based on the stream of the NLU output. The method according to claim 19.
21. After the step of providing the natural conversation output for audible presentation to the user via the one or more speakers of the client device, determining whether the user continues to provide the utterance within a threshold duration; In response to determining that the user does not continue to provide the utterance, Determining additional natural conversation output to be provided for audible presentation to the user, wherein the additional natural conversation output to be provided for audible presentation to the user requires that the user complete providing the utterance; Providing the additional natural conversation output for audible presentation to the user via one or more speakers of the client device The method according to claim 20, further comprising. **Claim 22** The method according to claim 21, wherein the additional natural conversation output to be provided for audible presentation to the user requires that an additional part of the utterance contains specific data based on a stream of NLU data. **Claim 23** A method implemented by one or more processors on a user's client device, the method comprising: Processing a stream of audio data to generate a stream of ASR output using an automatic speech recognition (ASR) model, wherein the stream of audio data is generated by one or more microphones of a client device, and the stream of audio data captures a part of the user's utterance directed to an automatic assistant at least partially implemented in the client device; Processing the stream of ASR output to generate a stream of NLU output using a natural language understanding (NLU) model; Based on the stream of the NLU output, determining that the automatic assistant can start fulfilling the utterance; Determining whether the user pauses providing the utterance or completes providing the utterance; In response to determining that the user pauses providing the utterance and has not completed providing the utterance, and in response to determining that the automatic assistant can start fulfilling the utterance based on the stream of the NLU output, Steps for holding off the start of fulfillment of the utterance; Steps for determining a natural conversation output to be provided for audible presentation to the user, wherein the natural conversation output to be provided for audible presentation to the user indicates that the automatic assistant is waiting for the user to complete providing the utterance; Steps for causing the natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device; After the steps for causing the natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device, in response to determining that the user did not complete providing the utterance within a threshold duration, Steps for determining whether the automatic assistant can start fulfillment of the utterance based on at least a stream of NLU data; In response to determining that the automatic assistant can start fulfillment of the utterance based on the stream of NLU data, Steps for causing the automatic assistant to start fulfillment of the utterance Including A method.
24. Further including steps for determining audio-based features associated with the portion of the utterance based on processing the stream of audio data, wherein the steps for determining whether the user paused or completed providing the utterance are based on the audio-based features associated with the portion of the utterance. The method according to claim 23.
25. The method according to claim 23 or 24, wherein the steps for determining whether the user paused or completed providing the utterance are based on the stream of NLU data.
26. In response to determining that the automatic assistant cannot start fulfillment of the utterance based on the stream of NLU data, Steps for determining a natural conversation output to be provided for audible presentation to the user, wherein the natural conversation output to be provided for audible presentation to the user requests that the user complete providing the utterance. causing additional natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device The method according to any one of claims 23 to 25, further comprising. **Claim 27** The method according to claim 26, wherein the natural conversation output to be provided for audible presentation to the user requires that an additional portion of the utterance includes specific data based on a stream of the NLU data. **Claim 28** The method according to any one of claims 23 to 27, wherein determining whether the automatic assistant can start fulfillment of the utterance is further based on one or more computational costs associated with fulfillment of the utterance. **Claim 29** The method according to claim 28, wherein the one or more computational costs associated with fulfillment of the utterance include one or more of a computational cost associated with performing fulfillment of the utterance and a computational cost associated with canceling the performed fulfillment of the utterance. **Claim 30** further comprising generating a stream of fulfillment data based on the stream of NLU output, wherein determining that the automatic assistant can start fulfillment of the utterance is further based on the stream of fulfillment data, The method according to any one of claims 23 to 29. **Claim 31** A method implemented by one or more processors, the method comprising: receiving a stream of audio data, the stream of audio data being generated by one or more microphones of a user's client device, the stream of audio data capturing at least a portion of an utterance provided by the user directed to an automatic assistant at least partially implemented at the client device; determining audio-based features associated with the portion of the utterance based on processing the stream of audio data; determining whether the user has paused or completed providing the utterance based on the audio-based features associated with the portion of the utterance; In response to determining that the user has paused the provision of the utterance and has not completed the provision of the utterance, and in response to determining that the fulfillment of the utterance can be started, a step of holding off on starting the fulfillment of the utterance; a step of determining a natural conversation output to be provided for audible presentation to the user, wherein the natural conversation output to be provided for audible presentation to the user indicates that the automated assistant is waiting for the user to complete the provision of the utterance; a step of causing the natural conversation output to be provided for audible presentation to the user via one or more speakers of the client device comprising a method. **Claim 32** at least one processor; a memory storing instructions that, when executed, cause the at least one processor to perform operations corresponding to any one of claims 1 to 31 a system comprising. **Claim 33** a non-transitory computer-readable storage medium storing a program that, when executed, causes at least one processor to perform operations corresponding to any one of claims 1 to 31
Citation Information
Patent Citations
Conversation processing system, conversation processing method, and computer program
JP2004086001A
User interface / entertainment devices that simulate personal interactions and respond to the user's emotional state and / or personality
JP2004513445A
Speech interactive device and method
JP2008241890A
Enhanced utterance endpoint specification
JP2018504623A
Information processor, information processing method and information processing program
JP2021117371A