Semi-delegated calls made by automated assistants on behalf of human participants

Through assisted call technology, the automated assistant uses parameter parsing and entity identification to dynamically render notifications and optimize resource utilization, solving the problems of call extensions and resource waste caused by the automated assistant's lack of information, and achieving more efficient task completion.

CN114631300BActive Publication Date: 2025-09-23GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080073684.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2020-04-22
Publication Date
2025-09-23
Estimated Expiration
2040-04-22

AI Technical Summary

Technical Problem

Automated assistants may lack necessary information when performing tasks on behalf of users, leading to extended calls, wasted resources, and even failure.

Method used

Through assisted call technology, automated assistants utilize parameter parsing and entity identification, dynamically render notifications, and proactively provide output, reducing user interaction and optimizing resource utilization.

Benefits of technology

It improves call efficiency, reduces resource consumption, ensures task completion, and avoids user waiting and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114631300B_ABST
    Figure CN114631300B_ABST
Patent Text Reader

Abstract

Embodiments relate to using an automated assistant to initiate an assisted call on behalf of a given user. The assistant can receive a request for information unknown to the assistant from an additional user on the assisted call during the assisted call. In response, a prompt for information can be rendered and the assisted call can be continued using the resolved value(s) for the assisted call while awaiting responsive input from the given user. If responsive input is received within a threshold duration, synthesized speech corresponding to the responsive input is rendered as part of the assisted call. Embodiments additionally or alternatively relate to using an automated assistant to provide output during an ongoing call between the given user and the additional user based on the value requested by the additional user during the ongoing call.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The automated assistant can be interacted with by a user via various computing devices, such as smartphones, tablet computers, wearable devices, automotive systems, standalone personal assistant devices, etc. The automated assistant receives input (e.g., verbal, touch, and / or typed) from the user and responds with responsive output (e.g., visual and / or auditory).

[0002] A user can interact with an automated assistant to cause the automated assistant to perform action(s) on the user's behalf. As an example, the automated assistant can place a phone call on the user's behalf to perform a given action and can engage in a conversation with additional users to perform action(s). For example, a user can provide user input requesting the automated assistant to make a restaurant reservation on the user's behalf over the phone. The automated assistant can initiate a phone call with a specific restaurant and provide reservation information to the additional users associated with the specific restaurant to make the reservation. The automated assistant can then notify the user whether the restaurant reservation was successfully made on the user's behalf.

[0003] However, for some actions performed by an automated assistant on behalf of a user, the automated assistant may not know enough information to fully perform the action(s). As an example, assume that the automated assistant is making a restaurant reservation over the phone on behalf of the aforementioned user, and further assume that additional user requests associated with a particular restaurant are unknown to the automated assistant. Some automated assistants can determine that the requested information is unknown and provide a notification to the user requesting that the user proactively join the call to complete the restaurant reservation. However, waiting for the user to proactively join the call can prolong the call and the associated use of computing and / or network resources used during the call. Additionally or alternatively, the user may not be available to join the call. This can result in the call failing and requiring the automated assistant and / or user to perform the action(s) at a later time, thereby consuming more computing and / or network resources than if the action(s) had been successfully performed by the automated assistant during the initial call. Summary of the Invention

[0004] Some embodiments relate to using an automated assistant to conduct an assistance call with an entity to perform task(s) on behalf of a given user. The assistance call is between the automated assistant and an additional user associated with the entity. The automated assistant can conduct the assistance call using resolved values ​​for parameter(s) associated with the task(s) and / or the entity. Conducting the assistance call can include audibly rendering, by the automated assistant, instance(s) of synthesized speech audibly perceivable by the additional user in the assistance call. Rendering the instance(s) of synthesized speech in the assistance call can include injecting the synthesized speech into the assistance call so that the additional user (but not necessarily the given user) can audibly perceive the synthesized speech. The instance(s) of synthesized speech can each be generated based on one or more resolved values ​​and / or can be generated in response to utterance(s) of the additional user during the assistance call. Conducting the assistance call can also include performing, by the automated assistant, automatic speech recognition of audio data of the assistance call that captures the utterance(s) of the additional user to generate recognized text of the utterance(s), and using the recognized text in determining the instance(s) of synthesized speech audibly perceivable by the additional user for rendering in the assistance call.

[0005] Some of these embodiments further involve determining, during an assisted call, that an utterance of an additional user includes a request for information associated with an additional parameter, and that no value of the additional parameter is automatically determinable. In response, the automated assistant can cause an audio and / or visual notification (e.g., a prompt) to be rendered to the given user, wherein the notification requests additional user input related to resolving the value(s) of the additional parameter(s). In some embodiments, the notification can be rendered outside of the ongoing call (i.e., not injected as part of the ongoing call), but perceptible to the given user, so that the given user can ascertain the value(s) and communicate the value(s) during the ongoing call. The automated assistant can continue the assisted call before receiving any additional user input in response to the notification. For example, the automated assistant can proactively provide, during the call, instances of synthesized speech based on resolved value(s) that were not communicated in previous instances of synthesized speech, without awaiting additional user input in response to the notification. If additional user input is provided in response to the notification and the value(s) of the additional parameter(s) are resolvable based on the additional user input, the automated assistant can provide the additional synthesized speech conveying the resolved value(s) of the additional parameter(s) after continuing the assisted call. By continuing the assisted call without awaiting further user input in response to the notification, the value(s) necessary to complete the task can be communicated during the assisted call while awaiting additional value(s) for additional parameters, and the additional value(s) can be provided later, if received. In these and other ways, the assisted call can be concluded more quickly, thereby reducing the overall duration of computer and / or network resources utilized in executing the assisted call.

[0006] In some embodiments of rendering a notification requesting additional user input related to resolving additional parameter(s), the notification and / or one or more attributes used to render the notification can be dynamically determined based on the state of the auxiliary call and / or the state of the client device used in the auxiliary call. For example, if the state of the client device indicates that a given user is actively monitoring the auxiliary call, the notification can be visual-only and / or can include an audible component at a lower volume. On the other hand, if the state of the client device indicates that the given user is not actively monitoring the auxiliary call, the notification can include at least an audible component and / or the audible component can be rendered at a higher volume. The state of the client device can be based on, for example, sensor data from the client device's sensor(s) (e.g., gyroscope(s), accelerometer(s), presence sensor(s), and / or other sensor(s)) and / or sensor(s) of other client device(s) associated with the user. As another example, if the state of the auxiliary call indicates that multiple resolved values ​​must also be conveyed when providing the notification, the notification can be visual-only and / or can include an audible component at a lower volume. On the other hand, if the status indication of the assisted call must also convey only one (or no) resolution value when providing the notification, the notification can include at least an audible component and / or the audible component can be rendered at a higher volume. More generally, embodiments can seek to provide more intrusive notifications when the status of the client device indicates that the user is not actively monitoring the call and / or when the status of the conversation indicates that the duration for meaningfully continuing the conversation is relatively short. On the other hand, embodiments can seek to provide less intrusive notifications when the status of the client device indicates that the user is actively monitoring the call and / or when the status of the conversation indicates that the duration for meaningfully continuing the conversation is relatively long. Although more intrusive notifications can be more resource intensive to render, embodiments can still selectively render more intrusive notifications to seek to balance the increased resources needed to render the more intrusive notifications with the increased resources needed to excessively prolong the assisted call and / or end the assisted call without completing the task.

[0007] Some embodiments additionally or alternatively involve using an automated assistant to provide, during an ongoing call between a given user and an additional user, an output based on a value requested by the additional user during the ongoing call. The output can be provided proactively and can prevent the given user from launching and / or navigating within application(s) to independently seek the value. For example, assume that the given user is engaged in an ongoing call with a utility company representative, and further assume that the utility company representative requests the given user's address information and an account number associated with the utility company. In this example, the given user can provide the address information, but may not know the utility company's account number without searching through an email or message received from the utility company, searching through a website associated with the utility company, and / or performing other computing device interactions to locate the account number associated with the utility company. However, by using the techniques described herein, the automated assistant is able to readily identify an account number associated with a utility company independent of any user input from the given user requesting the automated assistant to identify the account number, and is able to visually and / or audibly provide the account number to the given user and / or additional users during an ongoing call in response to identifying that a utility company representative has requested the account number.

[0008] In these and other ways, client device resources can be conserved by preventing the launch of such application(s) and / or interaction with such application(s). Furthermore, the value indicated by the proactively provided output can be communicated more quickly within a given call than if the user had to independently seek the value, thereby shortening the overall duration of the ongoing call. In various embodiments, the automated assistant can process an audio data stream capturing at least one spoken utterance during an ongoing call to generate recognized text, wherein the at least one spoken utterance is of the given user or an additional user. Furthermore, the automated assistant can identify information requesting a parameter for the at least one spoken utterance based on processing the recognized text, and determine, for the parameter and using restricted-access data personal to the given user, whether a value for the parameter is resolvable. In response to determining that the value is resolvable, an output can be rendered. In some embodiments, the output can be rendered outside of the ongoing call (i.e., not injected as part of the ongoing call), but can be perceived by the given user so that the given user can ascertain and communicate the value within the ongoing call. In some additional or alternative embodiments, the output can be rendered as synthesized speech as part of the ongoing call. For example, synthesized speech can be rendered within an ongoing call automatically or upon receiving affirmative user interface input from a given user. In some embodiments, during an ongoing call between a given user and an additional user, an output based on a value requested by the additional user during the ongoing call is provided, the output is provided only in response to determining that no spoken input has been provided by the given user during the ongoing call within a threshold amount of time that includes the value. In these and other ways, instances of unnecessary rendering of output can be mitigated.

[0009] The above description is provided only as an overview of some embodiments disclosed herein. These and other embodiments are described in more detail herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 Depicted is a block diagram of an example environment illustrating various aspects of the present disclosure and in which implementations disclosed herein can be implemented.

[0011] Figure 2 Depicted is a flow chart illustrating an example method of performing an assisted call in accordance with various implementations.

[0012] Figure 3 Depicted is a flow chart illustrating an example method of providing auxiliary output during an ongoing non-assisted call in accordance with various implementations.

[0013] Figure 4A 、 4B4C and 4D depict various non-limiting examples of user interfaces for performing an assisted call according to various embodiments.

[0014] Figure 5A 、 5B 5C depict various non-limiting examples of user interfaces for providing auxiliary output during an ongoing non-assisted call in accordance with various embodiments.

[0015] Figure 6 Depicted is an example architecture for a computing device in accordance with various implementations. DETAILED DESCRIPTION

[0016] Figure 1 A block diagram illustrating an example environment for various aspects of the present disclosure is shown. Figure 1 , and in various embodiments includes a user input engine 111, a device state engine 112, a rendering engine 113, a scheduling engine 114, a speech recognition engine 120A1, a natural language understanding (“NLU”) engine 130A1, and a speech synthesis engine 140A1.

[0017] The user input engine 111 can detect various types of user input at the client device 110. The user input detected at the client device 110 can include spoken input detected via the microphone(s) of the client device 110 and / or additional spoken input transmitted to the client device 110 from additional client devices of additional users (e.g., during an assisted call and / or during other ongoing calls when an assisted call has not been invoked), touch input detected via a user interface input device (e.g., a touch screen) of the client device 110, and / or typed input detected via a user interface input device (e.g., via a virtual keyboard on the touch screen) of the client device 110. Additional users associated with an entity during an ongoing (assisted or non-assisted) call with the entity can be, for example, humans, additional human participants associated with additional client devices, additional automated assistants associated with additional client devices of additional users, and / or other additional users.

[0018] The assisted calls and / or ongoing calls described herein can be performed using various voice communication protocols (e.g., Voice over Internet Protocol (VoIP), Public Switched Telephone Network (PSTN), and / or other telephone communication protocols). As described herein, synthesized speech can be rendered as part of the assisted call and / or ongoing call, which can include injecting the synthesized speech into the call so that it is perceptible by at least one of the participants in the ongoing call and forms part of the audio data for the ongoing call. The synthesized speech can be generated and / or injected by a client device that is one of the endpoints of the call, and / or can be generated and / or injected by a server that is in communication with the client device and is also connected to the call. Also as described herein, although the output can be detected by a microphone of a client device connected to the call and, as a result, perceptible on the call, audible output can also be rendered outside of the assisted call, which does not include injecting the audible output into the call. In some embodiments, the call can optionally be muted and / or filtering can be used to mitigate the perception of audible output rendered outside of the call within the call.

[0019] In various embodiments, (in Figure 1 Automated assistant 115 (generally indicated by dashed lines in FIG) can use assisted call system 180 to perform assisted calls at client device 110 over network(s) 190 (e.g., Wi-Fi, Bluetooth, near field communication, local area network(s), wide area network(s), and / or other networks). In various embodiments, assisted call system 180 includes speech recognition engine 120A2, NLU engine 130A2, speech synthesis engine 140A2, and assisted call engine 150. Automated assistant 115 can utilize assisted call system 180 to perform task(s) on behalf of a given user of client device 110 during phone calls with additional users.

[0020] Furthermore, in some embodiments, automated assistant 115 can obtain consent from additional users to participate in a conversation with automated assistant 115 before performing any task(s) on behalf of a given user of client device 110. For example, automated assistant 115 can obtain consent when initiating an assisted call and before performing task(s). As another example, automated assistant 115 can obtain consent when a given user of client device 110 initiates an ongoing call, even if the ongoing call was not initiated by automated assistant 115. If automated assistant 115 obtains consent from the associated additional users, automated assistant 115 can perform the task(s) using assisted call system 180. However, if automated assistant 115 does not obtain consent from the additional users, automated assistant 115 can cause client device 110 (e.g., using rendering engine 113) to render a notification to the given user of client device 110 indicating that the given user is required to perform the task and / or end the call, and (e.g., using rendering engine 113) to render a notification to the given user of client device 110 indicating that the task(s) were not performed.

[0021] As described in more detail below, the automated assistant 115 can perform an assisted call using the assisted call system 180 in response to detecting user input from a given user of the client device 110 to initiate a call using assisted call and / or during an ongoing call (i.e., when the assisted call has not yet been invoked). In some embodiments, the automated assistant 115 can determine, on behalf of the given user of the client device 110, during the assisted call and / or during the ongoing call, values ​​for candidate parameters to be used when performing task(s). In some versions of those embodiments, the automated assistant 115 can engage in a conversation with the given user of the client device 110 to request values ​​for the candidate parameters before and / or during the assisted call. In some additional and / or alternative versions of those embodiments, the automated assistant 115 can determine the values ​​for the candidate parameters based on the user profile(s) associated with the given user of the client device 110 and not request values ​​for the candidate parameters before and / or during the assisted call. In some versions of those embodiments, automated assistant 115 is able to automatically use assisted call system 180 to effectuate an assisted call based on the conversation of the ongoing call, and without detecting any user input from a given user of client device 110 via user input engine 111 .

[0022] like Figure 1 As shown, the assisted call system 180 can be implemented remotely (e.g., via server(s) and / or other remote client device(s). Figure 1190, it should be understood that this is for illustrative purposes and is not intended to be limiting. For example, in various embodiments, the assisted call system 180 can be implemented locally on the client device 110. Furthermore, although the automated assistant 115 Figure 1 10 and remotely at assisted call system 180, but it should be understood that this is for illustrative purposes only and is not intended to be limiting. For example, in various embodiments, automated assistant 115 can be implemented locally on client device 110, or implemented locally on client device 110 and interact with a separate, cloud-based automated assistant.

[0023] In embodiments, when the user input engine 111 detects spoken input from a given user via the microphone(s) of the client device 110 and / or receives audio data capturing the additional spoken input transmitted from an additional client device to the client device 110 (e.g., during an auxiliary call and / or during an ongoing call), the speech recognition engine 120A1 of the client device 110 can process the captured spoken input and / or the audio data capturing the additional spoken input using the speech recognition model(s) 120A to generate recognized text corresponding to the spoken input and / or the additional spoken input. Furthermore, the NLU engine 130A1 of the client device 110 can process the recognized text generated by the speech recognition engine 120A1 using the NLU model(s) 130A to determine the intent(s) included in the spoken input and / or the additional spoken input. For example, if client device 110 detects spoken input from a given user of “call Example Café to make a reservation for tonight,” client device 110 can use speech recognition model(s) 120A to process audio data that captured the spoken input to generate recognized text corresponding to the spoken input of “call Example Café to make a reservation for tonight,” and can use NLU model(s) 130A to process the recognized text to determine at least a first intent of initiating a call and a second intent of making a restaurant reservation. As another example, if client device 110 detects “will any children bejoining the reservation?”, client device 110 can use speech recognition model 120A to process audio data that captured additional spoken input to generate recognized text corresponding to the additional spoken input of “will any children bejoining the reservation?” and can use NLU model(s) 130A to process the recognized text to determine the intent of a request for information associated with additional parameter(s), as described herein. In some versions of those embodiments, the client device 110 is capable of transmitting audio data, recognized text, and / or intent(s) to the assisted call system 180 .

[0024] In other embodiments, when the user input engine 111 detects spoken input from a given user via the microphone(s) of the client device 110 and / or detects audio data capturing additional spoken input from additional users transmitted from additional client devices to the client device 110 (e.g., during an assisted call and / or during an ongoing call), the automated assistant 115 can cause the client device 110 to transmit the audio data capturing the spoken input and / or the audio data capturing the additional user input to the assisted call system 180. The speech recognition engine 120A2 and / or the NLU engine 130A2 of the assisted call system 180 can process the audio data capturing the spoken input and / or the audio data capturing the additional spoken utterances in a manner similar to that described above with respect to the speech recognition engine 120A1 and / or the NLU engine 130A1 of the client device 110. In some additional and / or alternative embodiments, the speech recognition engine 120A1 and / or the NLU engine 130A1 of the client device 110 can be used in a distributed manner in conjunction with the speech recognition engine 120A2 and / or the NLU engine 130A2 of the assisted call system 180. Furthermore, the speech recognition model(s) 120A and / or the NLU model(s) 130A can be stored locally on the client device 110 and / or remotely on a server that communicates with the client device 110 and / or the assisted call system 180 via the network(s) 190.

[0025] In various embodiments, (multiple) speech recognition model 120A is (multiple) end-to-end speech recognition model, so that (multiple) speech recognition engine 120A1 and / or 120A2 can directly use the model to generate the recognition text corresponding to the oral input. For example, (multiple) speech recognition model 120A can be (multiple) end-to-end model for generating recognition text on the basis of character by character (or on the basis of other symbol by symbol). A non-limiting example of (multiple) such end-to-end model for generating recognition text on the basis of character by character is a cyclic neural network converter (RNN-T) model. The RNN-T model is a sequence-to-sequence model form that does not adopt an attention mechanism. In addition, for example, when (multiple) speech recognition model is not (multiple) end-to-end speech recognition model, (multiple) speech recognition engine 120A1 and / or 120A2 can generate (multiple) predicted phonemes (and / or other representations) on the contrary. For example, using such a model, the predicted phoneme(s) (and / or other representations) are then used by the speech recognition engine(s) 120A1 and / or 120A2 to determine recognized text that conforms to the predicted phoneme(s). In doing so, the speech recognition engine(s) 120A1 and / or 120A2 can optionally employ decoding graphs, dictionaries, and / or other resources.

[0026] In embodiments, when user input engine 111 detects touch and / or typed input via a user interface input device of client device 110, automated assistant 115 can cause an indication of the touch input and / or an indication of the typed input to be transmitted from client device 110 to assisted call system 180. In some versions of those embodiments, the indication of the touch input and / or the indication of the typed input can include underlying text of the touch input and / or text of the typed input, and the underlying text and / or the text can be processed using NLU model(s) 130A to determine intent(s) for the underlying text and / or text.

[0027] As described herein, the assisted call engine 150 of the assisted call system 180 can further process the recognized text generated by the speech recognition engine(s) 120A1 and / or 120A2, the underlying text of the touch input detected at the client device 110, the underlying text of the typed input detected at the client device 110, and / or the intent(s) determined by the NLU engine(s) 130A1 and / or 130A2. In various embodiments, the assisted call engine 150 includes an entity identification engine 151, a task determination engine 152, parameter engines(s) 153, a task execution engine 154, a feedback engine 155, and a recommendation engine 156.

[0028] The entity identification engine 151 can identify participating entities on behalf of a given user of the client device 110. The entity can be, for example, a personal entity, a business entity, a location entity, and / or other entities. In some embodiments, the entity identification engine 151 can also determine the specific type of the identified entity. For example, the type of personal entity can be a friend entity, a family member entity, a coworker entity, and / or other specific types of personal entities. Additionally, the type of business entity can be a restaurant entity, an airline entity, a hotel entity, a salon entity, a doctor's office entity, and / or other specific types of business entities. Additionally, the type of location entity can be a school entity, a museum entity, a library entity, a park entity, and / or other specific types of location entities. In some embodiments, the entity identification engine 151 can also determine the specific entity for the identified entity. For example, a specific entity for a person entity can be the name of a person (e.g., Jane Doe, etc.), a specific entity for a business entity can be the name of a business (e.g., Hypothetical Café, Example Café, Example Airlines, etc.), and a specific entity for a location entity can be the name of a location (e.g., Hypothetical University, Example National Park, etc.). Although the entities described herein can be defined at various levels of granularity, for simplicity, they are collectively referred to herein as "entities."

[0029] In some embodiments, the entity identification engine 151 can identify participating entities on behalf of a given user of the client device 110 based on user interactions with the client device 110 prior to initiating an assisted call using the automated assistant 115. In some versions of those embodiments, the entity can be identified in response to receiving user input to initiate an assisted call. For example, if the given user of the client device 110 directs (e.g., verbal or touch) input to a call interface element of a software application (e.g., to a contact in a contacts application, to a search result in a browser application, and / or other callable entities included in other software applications), the entity identification engine 151 can identify the entity associated with the call interface element. For example, if the user input is directed to a call interface element associated with "Example Cafe" in a browser application, the entity identification engine 151 can identify "Example Cafe" (or more generally, a business entity or restaurant entity) as a participating entity on behalf of the given user of the client device 110 during the assisted call.

[0030] In some embodiments, the entity identification engine 151 can identify participating entities on behalf of a given user of the client device 110 based on metadata associated with the ongoing call. The metadata can include, for example, a phone number associated with the additional user, a location associated with the additional user, an identifier identifying the additional user and / or an entity associated with the additional user, the time the ongoing call began, the duration of the ongoing call, and / or other metadata associated with the phone call. For example, if the given user of the client device 110 is participating in an ongoing call with an additional user, the entity identification engine 151 can analyze the metadata of the ongoing call between the given user of the client device 110 and the additional user to identify the phone number associated with the additional user participating during the ongoing call. In addition, the entity identification engine 151 can cross-reference the phone number with database(s) to identify the entity associated with the additional user, submit a search query for the phone number to identify corresponding search results associated with the phone number to identify the entity associated with the additional user, and / or perform other actions to identify the entity associated with the additional user. For example, if a given user of client device 110 is engaged in an ongoing call, entity identification engine 151 can analyze metadata associated with the ongoing call to identify an identification for "Example Airline."

[0031] Furthermore, the entity identification engine 151 can cause any identified entity to be stored in entity database(s) 151A. In some embodiments, the identified entities stored in entity database(s) 151A can be indexed by entity and / or specific entity type. For example, if the entity identification engine 151 identifies the "Example Cafe" entity, "Example Cafe" can be indexed in entity database(s) 151A as a business entity and can optionally be further indexed as a restaurant entity. Furthermore, if the entity identification engine 151 identifies the "Example Airline" entity, "Example Airline" can also be indexed in entity database(s) 151A as a business entity and can optionally be further indexed as an airline entity. By storing and indexing the identified entities in entity database(s) 151A, the entity identification engine 151 can easily identify and retrieve the entities, thereby reducing the subsequent processing of identifying the entities when encountering them in future assisted calls and / or ongoing calls. Furthermore, in various embodiments, each entity can be associated with task(s) in entity database(s) 151A.

[0032] Task determination engine 152 can determine task(s) to be performed on behalf of a given user of client device 110. In some embodiments, task determination engine 152 can determine the task(s) to be performed before initiating an assisted call using automated assistant 115. In some versions of those embodiments, task determination engine 152 can determine the task(s) to be performed on behalf of the given user of client device 110 based on user input initiating the assisted call. For example, if the given user of client device 110 provides verbal input of "call Example Café to make a reservation for tonight," task determination engine 152 can utilize the intent(s) (e.g., determined using NLU model(s) 130A) of initiating a call and making a restaurant reservation to determine, based on the verbal input, a task of making a restaurant reservation. As another example, if the given user of client device 110 provides touch input selecting a call interface element associated with "Example Café," and the call interface indicates that the given user wants to modify a restaurant reservation at Example Café, task determination engine 152 can determine, based on the touch input, a task of modifying an existing restaurant reservation.

[0033] In some additional and / or alternative versions of those embodiments, the task determination engine 152 can determine task(s) based on the identified entities participating during the assisted call. For example, a restaurant entity can be associated with a task to make a restaurant reservation, a task to modify a restaurant reservation, a task to cancel a restaurant reservation, and / or other tasks. As another example, a school entity can be associated with a task to inquire about closures, a task to report that a student / employee will not be attending school that day, and / or other tasks.

[0034] In other embodiments, task determination engine 152 can determine task(s) to be performed on behalf of a given user of client device 110 during an ongoing call. In some versions of those embodiments, an audio data stream corresponding to a conversation between the given user of client device 110 and additional users of additional client devices can be processed as described herein (e.g., with respect to speech recognition model(s) 120A and NLU model(s) 130A). For example, if task determination engine 152 identifies the recognized text "what is your frequent flier number" during an ongoing call between the given user of client device 110 and additional users associated with the additional client devices, task determination engine 152 can determine a task to provide the frequent flier number to the additional user. In some further versions of those embodiments, task determination engine 152 can also determine task(s) based on entities stored in entity database(s) 151A associated with the task(s). For example, if the entity identification engine 151 identifies an additional user associated with the example airline entity (e.g., based on metadata associated with the ongoing call as described above), the task determination engine 152 can determine the task of providing the additional user with the frequent flyer numbers associated with the example airline based on the task of providing the frequent flyer numbers stored in association with the airline entity.

[0035] Parameter engine(s) 153 can identify parameter(s) associated with the task(s) determined by task determination engine 152. Automated assistant 115 can use the value(s) for the parameter(s) to perform the task(s). In some embodiments, candidate parameter(s) can be stored in parameter database(s) 153A in association with the task(s). In some versions of those embodiments, candidate parameter(s) for a given task can be detected from the parameter database(s) in response to identifying a participating entity on behalf of a given user of client device 110 during an assisted call. For example, for a task of making a restaurant reservation, parameter engine(s) 153 can identify and retrieve one or more candidate parameters, including a name parameter, a date / time parameter, a parameter for the party size being reserved, a phone number parameter, various types of seating parameters (e.g., booth seating or table seating, indoor seating or outdoor seating, etc.), a children parameter (i.e., whether children will be included in the reservation), a special occasion parameter (e.g., birthdays, anniversaries, etc.), and / or other candidate parameters. In contrast, for the task of modifying a restaurant reservation, the parameter engine(s) 153 are capable of identifying and retrieving candidate parameters, including a name parameter, an original reservation date / time parameter, a modified reservation date / time parameter, a modified reservation party size parameter, and / or other candidate parameter(s).

[0036] In some additional and / or alternative embodiments, parameter engine(s) 153 can identify parameter(s) for a given task during an ongoing call between a given user of client device 110 and additional users of additional client devices. In some versions of those embodiments, the parameter(s) identified for a given task during the ongoing call may or may not be candidate parameter(s) stored in parameter database(s) 153A in association with the given task. As described above, the audio data stream corresponding to the conversation between the given user and the additional user can be processed to determine the task(s) to be performed during the ongoing call. Parameter engine(s) 153 can determine whether the additional user is requesting information associated with a given parameter that is unknown to automated assistant 115. For example, if, during an ongoing call, a representative of an example airline requests a frequent flyer number from the given user of the client device, parameter engine(s) 153 can identify the "example airline frequent flyer number" parameter based on the intent(s) included in the recognized text of the conversation and / or based on the identified "example airline" entity.

[0037] Furthermore, in some embodiments, the candidate parameters stored in the parameter database(s) 153A can be mapped to various entities stored in the entity database(s) 151A. By mapping the candidate parameters stored in the parameter database(s) 153A to the various entities stored in the entity database(s) 151A, the auxiliary call engine 150 can easily identify the parameter(s) for the task(s) in response to identifying a given entity. For example, in response to identifying the "Example Cafe" entity (or more generally, the restaurant entity), the auxiliary call engine 150 can identify (e.g., stored in the entity database 151A in association with the "Example Cafe" entity and / or the restaurant entity) (a plurality of) predefined tasks to determine a set of candidate parameters, and in response to identifying a particular task (e.g., based on the identified entity and / or user input), the auxiliary call engine 150 can determine the candidate parameters associated with the particular task for the identified entity. In other embodiments, the entity database(s) 151A and the parameter database(s) 153A can be combined into one database having various indexes (e.g., indexed by entity, indexed by task(s) and / or other indexes) such that the entity(ies), task(s) and candidate parameter(s) can each be stored in association with one another.

[0038] As described above, the automated assistant 115 can perform task(s) on behalf of a given user of the client device 110 using value(s) for parameter(s). The parameter(s) engine 153 can also determine value(s) for the parameter(s). In embodiments, when a given user of the client device 110 provides user input to initiate an assistance call, the parameter(s) engine 153 can cause the automated assistant 115 to engage in a conversation with the given user (e.g., visually and / or audibly via the client device 110) prior to initiating the assistance call to request additional user input for candidate parameter(s). As described herein, the automated assistant 115 can generate prompt(s) requesting information and can audibly (e.g., via the speaker(s) of the client device 110) and / or visually (e.g., via the display(s) of the client device 110) render the prompt(s) requesting corresponding value(s) (or a subset thereof). For example, in response to receiving user input (e.g., via touch input or verbal input) to initiate an assisted call to make a reservation at an example cafe, the automation device can generate prompt(s) requesting additional user input for candidate parameter(s) (or a subset thereof), including values ​​for a date / time parameter for the reservation, values ​​for a number of people parameter for the reservation, and the like.

[0039] In some versions of those embodiments, parameter engine(s) 153 can determine the value(s) for the candidate parameter(s) based on the user profile(s) stored in user profile database(s) 153B for a given user of client device 110. In some other versions of these embodiments, parameter engine(s) 153 can determine the value(s) for the candidate parameter(s) without requesting any additional user input from the given user. Parameter engine(s) 153 can access user profile database(s) 153B and retrieve the value(s) for the name parameter, phone number parameter, and date / time parameter from a software application (e.g., a calendar application, an email application, a contacts application, a reminder application, a notes application, an SMS or text messaging application, and / or other software applications). User profile(s) can include, for example and with the given user's permission, the given user's linked accounts, the given user's email accounts, the given user's photo albums, the given user's social media profile(s), the given user's contacts, user preferences, and / or other information. For example, if a given user of client device 110 is engaged in a text messaging conversation with a friend and discusses date / time information for a restaurant reservation at a given entity before providing user input to initiate an assisted call, parameter engine(s) 153 can utilize the date / time information from the text messaging conversation as the value(s) for the date / time parameter, and the automated assistant need not prompt the given user for the date / time information for the restaurant reservation. As discussed in more detail herein (e.g., with reference to Figure 4B ), a given user of the client device 110 is able to modify the value(s) of the candidate parameter(s) before the assisted call is initiated.

[0040] In addition, in various embodiments, some candidate parameters for a given task may be required parameters, while other candidate parameters for a given task may be optional parameters. In some versions of those embodiments, whether a given parameter is a required parameter or an optional parameter can be based on the task. In other words, the required parameters for a given task may be the minimum amount of information required to perform the given task. For example, for a restaurant reservation task, a name parameter and a time / date parameter may be the only required parameters. However, if values ​​for the optional parameters (e.g., to better accommodate the size of the banquet being reserved, a phone number to call if any additional communication is required, etc.) are known, the restaurant reservation task may be more beneficial to both the given user and the restaurant associated with the restaurant reservation task. Therefore, the automated assistant 115 can generate prompts for at least the required parameters and, optionally, the optional parameters.

[0041] In embodiments where the assisted call system 180 determines that an additional user participating in an ongoing call with a given user of the client device 110 is requesting information about parameter(s), the parameter(s) engine 153 can determine the value(s) based on the user profile(s) of the given user of the client device 110 stored on the user profile database(s) 153B. In some versions of those embodiments, the parameter(s) engine 153 can determine the value(s) of the parameter(s) in response to information identifying that the additional user is requesting the parameter(s). In other versions of those embodiments, the parameter(s) engine 153 can determine the value(s) of the parameter(s) in response to user input from the given user of the client device 110 requesting information about the parameter(s).

[0042] In various embodiments, the task execution engine 154 can enable the automated assistant 115 to participate in a conversation with an additional user associated with the identified entity during the assisted call using synthesized speech to perform task(s). The task execution engine 154 can provide text and / or phonemes including at least the value to the speech synthesis engine 140A1 of the client device 110 and / or the speech synthesis engine 140A2 of the assisted call system 180 to generate synthesized speech audio data. The synthesized speech audio data can be transmitted to an additional client device of the additional user for audible rendering at the additional client device. The speech synthesis engine(s) 140A1 and / or 140A2 can use the speech synthesis model(s) 140A to generate synthesized speech audio data that includes synthesized speech corresponding to the value(s) of at least the parameter(s). For example, speech synthesis engine(s) 140A1 and / or 140A2 can determine a sequence of phonemes that are determined to correspond to information of the additional user-requested parameter(s), and can process the sequence of phonemes using speech synthesis model 140A to generate synthesized speech audio data. The synthesized speech audio data can be, for example, in the form of an audio waveform. When determining the sequence of phonemes corresponding to the value(s) of at least the parameter(s), speech synthesis engine(s) 140A1 and / or 140A2 can access token-to-phoneme mappings stored locally at client device 110 or stored at server(s) (e.g., via network 190).

[0043] In some embodiments, the task execution engine 154 can cause the client device 110 to initiate an assisted call with a participating entity on behalf of a given user of the client device 110 and perform task(s) on behalf of the given user of the client device 110. Furthermore, the task execution engine 154 can utilize synthesized speech audio data comprising value(s) of at least parameter(s) to perform the task(s) on behalf of the given user of the client device 110. For example, for a task of making a restaurant reservation, the automated assistant 155 can cause synthesized speech to be rendered at an additional client device associated with an additional user, the additional user identifying the automated assistant 115 on behalf of the given user of the client device 110 and stating the task(s) to be performed on behalf of the given user during the assisted call (e.g., “This is Jane Doe's Automated Assistant calling to make a reservation on behalf of Jane Doe”).

[0044] In some versions of those embodiments, automated assistant 115 can cause corresponding values ​​for various candidate parameter(s) to be rendered (e.g., determined using parameter(s) engine 153 as described above) without being explicitly requested by additional users associated with the entity. Continuing with the above example, automated assistant 115 can provide, at the beginning of the conversation, a value for a date / time parameter for the reservation (e.g., "tonight at 7:00 PM," etc.), a value for a number of people for the reservation (e.g., "two," "three," "four," etc.), a value for a type of seating for the reservation (e.g., "booth," "room," etc.), and / or other values ​​for other candidate parameters. In other versions of these embodiments, automated assistant 115 can engage in a conversation with additional users of additional computing devices and provide specific value(s) for the parameter(s) explicitly requested by the additional users associated with the entity. Continuing with the above example, automated assistant 115 can process audio data capturing the additional user's speech (e.g., "For what time and for how many people?", etc.), and can, in response to receiving a request from the additional user, determine information regarding the parameter(s) requested by the additional user (e.g., "tonight at 7:00 PM for five people," etc.).

[0045] Furthermore, in some versions of those embodiments, task determination engine 154 can determine that the request from the additional user associated with the entity is a request for information associated with additional parameter(s) for which the additional value(s) are unknown to automated assistant 115. For example, for a restaurant reservation task, assume that automated assistant 115 knows the values ​​for the date / time information parameter, the number of people parameter, and the seat type parameter, but the additional user requests information for an unknown children parameter (i.e., whether any children are included in the reservation). In response to determining that the additional user is requesting information for the unknown children parameter, automated assistant 115 can cause client device 110 to render a notification indicating that the additional user (e.g., using rendering engine 113) has requested the additional value for the children parameter and prompting the given user of client device 110 to provide the additional value for the children parameter.

[0046] In some further versions of those embodiments, the type of notification rendered at client device 110 and / or one or more attributes of the rendered notification (e.g., volume, brightness, size) can be based on the state of client device 110 and / or the state of the ongoing call (e.g., as determined using device state engine 112). The state of the ongoing call can indicate, for example, which value(s) have been communicated and / or have not been communicated in the ongoing call, and / or can indicate which component(s) of the task of the ongoing call have been completed and / or have not been completed in the ongoing call. The state of client device 110 can be based on, for example, the software application(s) operating in the foreground of client device 110, the software application(s) operating in the background of client device 110, whether client device 110 is in a locked state, whether client device 110 is in a sleep state, whether client device 110 is in a powered-off state, sensor data from sensor(s) of client device 110, and / or other data. For example, if the state of client device 110 indicates that a software application (e.g., an automated assistant application, a call application, an assisted call application, and / or other software application) that displays a transcription of an assisted call is operating in the foreground of client device 110, the type of notification may be a banner notification, a pop-up notification, and / or other type of visual notification. As another example, if the state of client device 110 indicates that client device 110 is in a sleep or locked state, the type of notification may be an audible indication via speaker(s) and / or a vibration via speaker(s) or other hardware component of client device 110. As yet another example, if sensor data from presence sensor(s), accelerometer(s), and / or other sensor(s) of the client device indicates that a given user is not currently near the client device and / or is not currently holding the client device, a more intrusive notification (e.g., visual and audible at a first volume level) can be provided. On the other hand, if such sensor data indicates that a given user is currently near the client device and / or is currently holding the client device, a less intrusive notification (e.g., visual only, or visual and audible at a second volume level less than the first volume level) can be provided. As yet another example, when the status of the conversation indicates that the conversation is nearing completion, more intrusive notifications can be provided, while when the status of the conversation indicates that the conversation is not nearing completion, less intrusive notifications can be provided.

[0047] In some further versions of those embodiments, even if automated assistant 115 does not know the corresponding additional value(s) for the additional parameter(s) requested by the additional user, task determination engine 153 can cause automated assistant 115 to continue the conversation with the additional user. Automated assistant 115 can subsequently provide the additional value(s) in the conversation after receiving additional user input in response to the request from the additional user. For example, for a restaurant reservation task, assume that automated assistant 115 knows the values ​​of a date / time information parameter, a number of people parameter, and a seating type parameter, and further assume that the additional user requests the value(s) for the date / time information parameter and the number of people parameter. Further assume that the additional user then requests a value for a children parameter, for which automated assistant 115 does not know the value. In this example, task determination engine 153 can cause automated assistant 115 to render synthesized speech at the additional client device of the additional user, the synthesized speech including an indication that the requested information is not currently known whether children will be joining, and automated assistant 115 can continue assisting the call by providing additional value(s) (e.g., for seating type) until additional user input is detected at client device 110 in response to the notification, including a value for the children parameter. In response to receiving additional user input, automated assistant 115 can provide the additional value as a standalone value (e.g., “no children will be joining the reservation”) or as a continuation value (e.g., “Jane Doe prefers booth seating, by the way, no children will bejoining the reservation”).

[0048] Furthermore, in embodiments where an additional user requests information about unknown additional parameter(s), automated assistant 115 can end the assistance call if further user input in response to the notification requesting the information is not received within a threshold duration (e.g., 15 seconds, 30 seconds, 60 seconds, and / or other durations). In some versions of those embodiments, the threshold duration can begin when the notification requesting the information is rendered at a given user's client device 110. In other versions of those embodiments, the threshold duration can begin when the last value known to automated assistant 115 is requested by the additional user or proactively provided by automated assistant 115 (independent of the additional user request).

[0049] Furthermore, in embodiments where the additional user requests information about unknown additional parameter(s), the feedback engine 155 can store the additional parameter(s) as candidate parameter(s) for the task in the parameter database(s) 153A. In some versions of those embodiments, the feedback engine 153A can map the additional parameter(s) to entities associated with the additional user stored in the entity database(s) 151A. In some other versions of those embodiments, if the additional user(s) associated with the entity request the additional parameter(s) from multiple users a threshold number of times while the assistive voice is active, the feedback engine 155 can map the additional parameter(s) to the entity. For example, if a restaurant entity inquires whether a restaurant reservation will include children at least a threshold number of times (e.g., 100 times, 1,000 times, and / or other numerical thresholds) during a conversation between multiple assistive calls initiated by multiple users via respective client devices, the feedback engine 155 can map the children parameter to each restaurant entity in the entity database(s) 151A. In this example, the children parameter can be considered a new candidate parameter to request a value for prior to initiating future assistance calls for restaurant reservation(s) with various restaurant entities and / or a particular entity that frequently requests information on the children parameter. In these and other ways, the value can be requested prior to initiating future assistance calls, thereby shortening the duration of future assistance calls and / or preventing the need to utilize computing resources when rendering prompts for the value in future assistance calls.

[0050] In some embodiments, task execution engine 154 can cause automated assistant 115 to provide output related to an ongoing call to an additional user associated with the entity, even if the ongoing call was not initiated using an assisted voice (i.e., a non-assisted call), to perform task(s) on behalf of a given user of client device 110. For example, automated assistant 115 can interrupt an ongoing call to cause synthesized speech including value(s) to be rendered at an additional client device 110. In some versions of those embodiments, automated assistant 115 can provide output related to the ongoing call (e.g., such as regarding a user input from a given user of client device 110). Figure 5A(as described above). In some further versions of those embodiments, the automated assistant 115 may not process audio data corresponding to the ongoing call, thereby eliminating the need to obtain consent from the additional user. Instead, the automated assistant 115 can analyze metadata associated with the ongoing call and determine (multiple) corresponding values ​​for the parameter requested by the additional user based on the metadata and / or user input. For example, the client device 110 can detect user input that activates the auxiliary call and queries the automated assistant 115 to retrieve (multiple) corresponding values ​​requested by the additional user (e.g., a value for the frequent flyer number parameter), and can determine based on the metadata associated with the ongoing call that the value is associated with a particular entity (e.g., example airline), even if the particular entity is not explicitly identified in the request. The automated assistant can then search the user's access-restricted data using the search parameter (e.g., term) based on both the request of the additional user (e.g., "frequent flyer") and the metadata (e.g., "example airline"). Furthermore, in some further versions of those embodiments, the user input for the automated assistant 115 to provide output related to the ongoing call is responsive to a notification generated by the automated assistant 115 and rendered at the client device 110 (e.g., using the rendering engine 113) indicating that the assisted call system 180 is capable of providing value(s) to the additional user (e.g., such as regarding the Figure 5B For example, the automated assistant 115 can proactively notify a given user of the client device 110 that the assisted call system 180 can provide corresponding value(s) for additional user-requested parameter(s).

[0051] In other versions of those embodiments, automated assistant 115 provides output related to the ongoing call (e.g., synthesized speech within the ongoing call) in response to determining that information is being appended to a user request, and does not receive any user input from a given user of client device 110 (e.g., as described in relation to the user request). Figure 5CAs described above). In some versions of those embodiments, automated assistant 115 can automatically provide the value(s) based on a confidence metric for the value that satisfies a confidence threshold. The confidence metric can be based on, for example, whether the given user previously provided the determined value in response to receiving a previous request for the same value, whether the parameter(s) identified during the ongoing call are candidate parameter(s) stored in association with the task identified during the ongoing call, the source from which the value was determined (e.g., an email / calendar application versus a text messaging application), and / or the manner in which the confidence metric was determined. For example, for an ongoing call between a given user of client device 110 and an additional user associated with an airline entity, assisted call system 180 can determine a frequent flyer number requested by the additional user, and assisted call system 180 can cause automated assistant 115 to automatically provide synthesized speech including the frequent flyer number as part of the ongoing call—and can automatically provide the synthesized speech without receiving any user input requesting automated assistant 115 to provide the frequent flyer number. In some versions of these embodiments, a given user of client device 110 may have to authorize an assisted call to automatically interrupt an ongoing call in settings associated with the assisted call.

[0052] In various embodiments, the recommendation engine 156 can determine candidate value(s) to be communicated in the conversation and can cause the automated assistant 115 to provide the candidate value(s) as recommendation(s) to a given user of the client device 110. In some embodiments, the candidate value(s) can be transmitted to the client device 110 via the network(s) 190. Furthermore, the candidate value(s) can be visually rendered as recommendation(s) on a display of the client device (e.g., using the rendering engine 113) in response to requests from additional users. In some versions of those embodiments, the recommendation(s) can be selectable, such that when user input is directed to a given recommendation (e.g., as determined by the user input engine 111), the given recommendation can be incorporated into the synthesized speech that is audibly rendered at an additional client device of the additional user.

[0053] In some embodiments, the recommendation engine 156 can determine candidate value(s) based on a request for information from additional users. For example, if the request is a yes / no question (e.g., “will any children be included in the reservation”), the recommendation engine 156 can determine a first recommendation including a “yes” value and a second recommendation including a “no” value. In other embodiments, the recommendation engine 156 can determine candidate value(s) based on user profile(s) of a given user associated with the client device 110 stored in the user profile database(s) 153B. For example, if the request solicits specific information (e.g., “how many children there are”), the recommendation engine 156 can determine, based on the given user profile(s) and with the given user's permission, that the given user of the client device 110 has three children—and can determine a first recommendation including a “three” value, a second recommendation including a “two” value, a third recommendation including a “one” value, and / or other recommendations including other value(s). The various recommendations described herein can be rendered visually and / or audibly at the client device 110 for presentation to a given user associated with the client device.

[0054] As described herein, rendering engine 113 can render various notifications or other outputs at client device 110. Rendering engine 113 can audibly and / or visually render the various notifications described herein. Additionally, rendering engine 113 can cause a transcript of the conversation to be rendered on a user interface of client device 110. In some implementations, the transcript can correspond to a conversation between a given user of client device 110 and automated assistant 115 (e.g., as described with respect to a particular user). Figure 4B In other embodiments, the transcription can correspond to a conversation between additional users of additional client devices and automated assistant 115 (e.g., as described with respect to Figure 4C and 4D In still other embodiments, the transcription can correspond to a conversation between a given user of client device 110, additional users of additional client devices, and automated assistant 115 (e.g., as described with respect to Figures 5A-5C described above).

[0055] In some embodiments, scheduling engine 114 can cause automated assistant 115 to include, along with and / or included in a notification indicating the results of performing the task(s), recommendations for the automated assistant to perform additional (or subsequent) tasks based on the results of performing the task(s). In some versions of those embodiments, the recommendations can be selectable, such that when user input is directed to a given recommendation (e.g., as determined by user input engine 111), the given recommendation can cause the additional task(s) to be performed. For example, for a successful restaurant reservation task, automated assistant 115 can render a selectable element via a user interface of a display of client device 110 that, when selected by a given user of client device 110, causes scheduling engine 114 to create a calendar entry for the successful restaurant reservation. As another example, for a successful restaurant reservation task, automated assistant 115 can send an SMS or text message to the other user(s) who are joining the restaurant reservation indicating that the restaurant reservation task has been successfully performed. In contrast, for an unsuccessful restaurant reservation task, automated assistant 115 can render a selectable element that, when selected by a given user of client device 110, causes scheduling engine 114 to create a reminder and / or calendar entry to perform the restaurant reservation task again at a later time and before the time / date value of the attempted restaurant reservation task (e.g., automatically by automated assistant 115 at a later time or by automated assistant 115 in response to user selection of the reminder and / or calendar entry).

[0056] In other embodiments, the scheduling engine 114 can, in response to determining the results of performing the task, cause the automated assistant 115 to automatically perform(s) additional (or subsequent) tasks based on the results of performing the task(s). For example, for a successful restaurant reservation task, the automated assistant 115 can automatically create a calendar entry for the successful restaurant reservation, automatically send an SMS or text message to the other user(s) who are joining the restaurant reservation indicating that the restaurant reservation task has been successfully performed, and / or other additional tasks that can be performed by the automated assistant 115 in response to the successful performance of the restaurant reservation task. In contrast, for an unsuccessful restaurant reservation task, the automated assistant 115 can automatically create a reminder and / or calendar entry to perform the restaurant reservation task again at a later time and before the time / date value of the restaurant reservation task (e.g., automatically performed by the automated assistant 115 at a later time or performed by the automated assistant 115 in response to the user selecting the reminder and / or calendar entry).

[0057] By using the techniques described herein, various technical advantages can be achieved. As a non-limiting example, the automated assistant 115 can end assisted calls more quickly because, when additional users request information currently unknown to the automated assistant 115, the assisted call conversation is not paused to await the value(s) of the additional parameter(s). Because the length of assisted calls can be reduced by using the techniques disclosed herein, network and computing resources can be conserved. As another non-limiting example, the automated assistant 115 can provide corresponding value(s) for parameter(s) during the performance of task(s) by a given user during an ongoing phone call. By providing the corresponding value(s) automatically or in response to explicit user input as described above, the client device 110 receives less input from the given user of the client device 110 because the user does not need to navigate to various applications with different user interfaces to determine the corresponding value(s), thereby conserving computing resources at the given client device. Furthermore, because the user does not need to navigate to these various applications, the system conserves computing and network resources by ending ongoing calls more quickly.

[0058] Figure 2 A flow chart illustrating an example method 200 for performing an assisted call according to various embodiments is depicted. For convenience, the operations of the method 200 are described with reference to a system performing the operations. The system of the method 200 includes (a plurality of) computing devices (e.g., Figure 1 Client device 110, Figures 4A-4D Client device 410, Figures 5A-5C Client device 510, Figure 6 The one or more processors and / or other components of the computing device 610, one or more servers, and / or other computing devices (e.g., a processor 610 ...

[0059] At block 252, the system receives user input from a given user via a client device associated with the given user to initiate an assisted call. In some embodiments, the user input is a spoken input detected via microphone(s) of the client device. For example, the spoken input can include "call Example Café," or specifically, "use assistedcall to call Example Café." In other embodiments, the user input is a touch input detected at the client device. For example, the touch input can be detected while various software applications (e.g., a browser application, a messaging application, an email application, a note application, a reminder application, and / or other software applications) are operating on the client device.

[0060] At block 254, the system identifies an entity that is participating on behalf of a given user during the assisted call in response to the user input initiating the assisted call. The entity may be identified based on the user input. In embodiments where the user input is verbal, the entity may be identified based on (e.g., using Figure 1 The speech recognition model(s) 120A and / or the NLU model(s) 130A of the client device process the audio data of the captured spoken input to identify entities (e.g., a business entity, a specific business entity, a location entity, and / or other entities) included in the spoken input. In embodiments where the user input is touch input, the entity can be identified based on user interaction(s) with the client device (e.g., touch input selecting a contact entry associated with the entity, a search result associated with the entity, an advertisement associated with the entity, and / or other user interaction(s)).

[0061] At block 256, the system determines at least one task to be performed on behalf of the given user during the assisted call based on the user input and / or the entity. In various embodiments, the predefined task(s) can be stored in one or more databases (e.g., Figure 1 1A). For example, tasks for booking a flight, changing a flight, canceling a flight, lost luggage inquiry, and / or other tasks can be stored in association with multiple different airline entities. In some embodiments, at least one task to be performed can be determined based on user input. For example, if verbal input of "call Example Café to make a reservation for tonight at 7:00 PM" is received at a client device, the system can determine that the verbal input includes a restaurant reservation task. In other embodiments, at least one task to be performed can be determined based on the entity identified at box 254. For example, if verbal input of "call Example Café" is received at a client device (i.e., without specifying a restaurant reservation task), the system can infer the task of making a reservation based on the fact that it is a predefined task associated with the restaurant entity.

[0062] At block 258, the system identifies candidate parameter(s) associated with the at least one task. In various embodiments, the candidate parameter(s) can be stored in one or more databases (e.g., Figure 1Parameter database(s) 153A). For example, tasks for booking a flight, changing a flight, canceling a flight, lost luggage inquiry, and / or other tasks can be stored in association with corresponding candidate parameters. Furthermore, for example, a task for changing a flight associated with airline entity 1 can be stored in association with a first corresponding parameter, and a task for changing a flight associated with airline entity 2 can be stored in association with a second corresponding parameter. As another example, restaurant entity 1 can be stored in association with a first corresponding parameter, and restaurant entity 2 can be stored in association with a second corresponding parameter.

[0063] At block 260, the system determines corresponding value(s) for the candidate parameter(s) to be used when performing at least one task. In some implementations, the system can determine corresponding value(s) for the candidate parameter(s) to be used when performing at least one task. Figure 1 The system may determine the value(s) of the candidate parameter(s) based on the user profile(s) of the given user in the user profile database(s) 153B). The user profile(s) may include, for example, the given user's linked accounts, the given user's email accounts, the given user's photo albums, the given user's social media profile(s), the given user's contacts, user preferences, and / or other information. For example, for a salon booking task, the system may determine the given user's name parameter and phone number parameter based on a contacts application, and may determine the preferred hairstylist at the salon based on previous communications with a particular hairstylist (e.g., email messages, text or SMS messages, phone calls, and / or other communications). In some additional and / or alternative embodiments, the value(s) of the candidate parameter(s) may be determined additionally or alternatively based on additional user input in response to prompt(s) visually and / or audibly rendered at the client device and requesting information about the parameters. In some versions of those embodiments, the system may generate prompt(s) only for corresponding values ​​of the candidate parameter(s) that the system cannot determine based on the user profile(s). For example, for the above-mentioned hair salon reservation task, the system already knows the name parameter, the phone number parameter, and the value(s) of the phone number parameter, so the system can only generate a prompt to request the value of the date / time parameter. Figure 4B ), the system is able to provide a given user of a client device with an opportunity to modify the value(s) of the candidate parameter(s) prior to initiating an assisted call.

[0064] At block 262, the system initiates an auxiliary call with an entity on behalf of the given user using a client device associated with the given user to perform at least one task using values ​​for candidate parameter(s). The system can process audio data received at the client device from the additional computing device to determine the value(s) for the parameter(s) being requested by the additional user. Furthermore, the system can generate synthesized speech audio that is transmitted to the additional client device of the additional user associated with the entity identified at block 254 and includes at least the value(s) for the parameter(s) (e.g., proactively or in response to a request by the additional user). In some embodiments, the system can engage in a conversation with the additional user and generate synthesized speech audio that includes specific values ​​included in the information requested by the additional user. For example, for a restaurant reservation task, the additional user can request information for a date / time parameter, and the system can generate synthesized speech audio data that includes the date / time value in response to determining that the request from the additional user is for information for the date / time parameter. Furthermore, the system can cause the synthesized speech audio data to be transmitted to an additional client device of an additional user, and cause the synthesized speech included in the synthesized speech audio data to be audibly rendered at the additional client device. The additional user can request additional information of various parameter(s), and the system can provide the additional user with the value(s) to perform the restaurant reservation task.

[0065] In some embodiments, method 200 can include an optional sub-block 262A. If included, at optional sub-block 262A, the system obtains consent from additional users associated with the entity to monitor the auxiliary call. For example, the system can obtain consent when initiating the auxiliary call and before performing the task(s). If the system obtains consent from the associated additional users, the system can perform the task(s). However, if the system does not obtain consent from the additional users, the system can cause the client device to render a notification to the given user indicating that the given user is required to perform the task and / or end the call, and to render a notification to the given user indicating that the task(s) are not to be performed.

[0066] At block 264, the system determines whether an additional user associated with the entity during the auxiliary call requested any information associated with (multiple) additional parameters. As described above, the system can process audio data received at the client device from the additional computing device to determine the value(s) requested by the additional user. In addition, the system can determine whether the requested value is for (multiple) additional parameters that the system has not previously resolved, such that the value(s) requested by the additional user are currently unknown to the system. For example, if the system has not previously determined a value for the type of seating parameter for a restaurant reservation task, the type of seating parameter can be considered to be an additional parameter with a value currently unknown to the system. If, at an iteration of block 264, the system determines that an additional user associated with the entity during the auxiliary call did not request information associated with (multiple) additional parameters, the system can proceed to block 272, which is discussed in more detail below.

[0067] If, at an iteration of block 264 that includes optional block 266, the system determines that value(s) for additional parameter(s) are to be included in requests from additional users associated with the entity during the assisted call, the system may proceed to optional block 266. In embodiments that include optional block 266, the system may proceed directly from block 264 to block 266, which will be discussed in greater detail below.

[0068] If so, then at optional block 266, the system determines the state of the client device associated with the given user. The state of the client device can be based on, for example, the software application(s) operating in the foreground of the client device, the software application(s) operating in the background of the client device, whether the client device 110 is in a locked state, whether the client device is in a sleep state, whether the client device 110 is in a powered-off state, sensor data from the sensor(s) of the client device, and / or other data. In some embodiments, the system additionally or alternatively determines the state of the ongoing call at block 266.

[0069] In box 268, the system causes the client device associated with a given user to render a notification identifying the additional parameter(s). The notification can further request the value(s) of the additional parameter(s) included in the information. In an embodiment including optional box 268, the type of notification rendered by the client device and / or one or more attributes for rendering can be based on the state of the client device determined in optional box 268 and / or the state of the ongoing call. For example, if the state of the client device indicates that a software application (e.g., automated assistant application, call application, auxiliary call application and / or other software application) that displays the transcription of the auxiliary call is operating in the foreground of the client device, the type of notification can be a banner notification, a pop-up notification and / or other types of visual notifications. As another example, if the state of the client device indicates that the client device is in a sleep or locked state, the type of notification can be an audible indication via (multiple) speakers and / or a vibration via (multiple) speakers or other hardware components of the client device.

[0070] At block 270, the system determines whether any additional user input has been received at the client device associated with the given user within a threshold duration. The additional user input can be, for example, additional verbal input, additional typed input, and / or additional touch input in response to a notification requesting information. In some embodiments, the threshold duration can begin when a notification requesting information is rendered at the given user's client device. In other embodiments, the threshold duration can begin when a final value is requested by an additional user. If, at an iteration of block 270, the system determines that additional user input has been received within the threshold duration, the system can proceed to block 272. Additional user input can be received in response to a notification indicating that information is being requested by an additional user, and the additional user input can include an indication of the value(s) in response to the request.

[0071] At block 272, the system completes at least one task based on the values ​​of the candidate parameter(s) and / or the additional parameter(s). In embodiments where the system determines that no additional value(s) were included in the request for information from the additional user at block 264 during the assisted call, the system can use the corresponding value(s) of the candidate parameter(s) determined at block 260 to complete the at least one task. In these embodiments, the system can complete the assisted call without necessarily involving the given user of the client device. In embodiments where the system determines that information associated with the additional parameter(s) was requested by the additional user at block 264 during the assisted call, the system can use the corresponding value(s) of the candidate parameter(s) determined at block 260 and the value(s) of the additional parameter(s) received at block 270 to complete the at least one task. In these embodiments, the system can complete the assisted call with minimal involvement by the given user of the client device. From block 272, the system can proceed to block 276, which will be discussed in more detail below.

[0072] Notably, even though the system may determine that the additional user is requesting information currently unknown to the system at block 264, the system can continue to perform at least one task without the value(s) for the additional parameter(s). For example, the system can cause synthesized speech to be provided as part of the assistance call, thereby causing it to be rendered at the additional client device of the additional user. Furthermore, the synthesized speech can indicate that the system currently does not know the value(s) for the additional parameter(s), but the system can request the value(s) associated with the information from the user and provide other value(s) for the candidate parameter(s) determined at block 260 while the system is requesting the value(s) for the additional parameter(s) from the given user of the client device. In this manner, the system can continue the conversation with the additional user to perform the task. Furthermore, if additional user input including the value(s) is provided in response to the request for information from the additional user, the system can provide the value(s) as a follow-up to providing one of the known value(s) or as a separate value when there is a break in the conversation. In this way, the system can successfully perform tasks on behalf of a given user of a client device in a faster and more efficient manner because the conversation is not paused to wait for the value(s) of(s) additional parameters. By performing tasks in a faster and more efficient manner, both network and computing resources can be conserved because the length of the session can be reduced by using the techniques disclosed herein.

[0073] If, at an iteration of block 270, the system determines that no additional user input has been received within the threshold duration, the system may proceed to block 274. At block 274, the system terminates execution of the at least one task. Additionally, the system may terminate the ongoing call with the entity. By terminating the ongoing call, as opposed to waiting for additional user input for a duration exceeding the threshold, the conversation can be concluded more quickly, achieving the technical advantages described above, even if the task cannot be fully executed. From block 274, the system may proceed to block 276.

[0074] At block 276, the system renders a notification via the client device indicating the results of performing at least one task. In embodiments where the system completes execution of the task from block 272, the notification can include an indication that the task was completed on behalf of a given user of the client device, and can include confirmation information associated with the completion of the task (e.g., date / time information, a monetary cost associated with the task, a confirmation number, information associated with the entity, and / or other confirmation information). In embodiments where the system ends execution of the task from block 274, the notification can include an indication that the task was not completed, and can include task information associated with the termination of the task (e.g., value(s) of required(s) particular parameter(s), the entity cannot accommodate the corresponding value(s) determined at block 260 and / or the value(s) received at block 270, the entity is closed, and / or other task information). In various embodiments, the notification can include (multiple) selectable graphical elements that, when selected, can cause the system to create a calendar entry based on the results of the task, create a reminder based on the results of the task, send a message (e.g., text, SMS, email, and / or other message) including the results of the task, and / or send other additional tasks in response to user selection of the other additional tasks.

[0075] Figure 3 A flow chart illustrating an example method 300 of assistive output during an ongoing non-assisted call according to various embodiments is depicted. For convenience, the operations of the method 300 are described with reference to a system performing the operations. The system of the method 300 includes computing device(s) (e.g., Figure 1 Client device 110, Figures 4A-4D Client device 410, Figures 5A-5C Client device 510, Figure 6 The method 300 may be implemented as a single process or as a combination of one or more processors and / or other components of a computing device 610, one or more servers, and / or other computing devices. Furthermore, while the operations of method 300 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.

[0076] At block 352, the system detects, at a client device, an ongoing call between a given user associated with the client device and an additional user associated with an additional client device. The system can also optionally identify an entity associated with the additional user. The system can identify the entity based on metadata associated with the ongoing call. In some embodiments, method 300 can include optional sub-block 352A. If included, at optional sub-block 352A, the system obtains consent from the additional user associated with the entity to monitor the ongoing call. The system can Figure 2 Obtain consent from the additional user in the same manner as described for optional sub-box 260A.

[0077] At block 354, the system processes the audio data stream corresponding to the ongoing call to generate recognized text. The audio data stream corresponding to the ongoing call can include at least additional spoken input from additional users that is transmitted to the given user's client device. The audio data stream corresponding to the ongoing call can also include spoken input from the given user. In addition, the system can use speech recognition model(s) (e.g., Figure 1 The audio data stream is processed by the speech recognition model(s) 120A) to generate the recognized text. It should be understood that, assuming the additional user agrees to monitor the call, the system can continuously process the audio data stream corresponding to the ongoing call.

[0078] At block 356, the system identifies parameters of at least one task to be performed by the given user during the ongoing call based on the recognized text. The system can use NLU models (e.g., Figure 1 The recognized text from block 354 is processed by the NLU model(s) 130A to determine the intent(s) included in the audio data stream. In some implementations, the system can determine that the additional user input from the additional user includes a request for information about parameter(s) of at least one task. For example, if the additional spoken input from the additional user is "do you have a quality assurance case number for this matter," the system can identify the quality assurance case number parameter and can determine that the user input includes a request for a value for the quality assurance case number parameter. In this example, the task can be any task associated with an airline entity and / or a specific task of providing a quality assurance case number, regardless of other parameter(s) associated with the airline entity.

[0079] At block 358, the system determines corresponding value(s) for the parameter(s) to be used when performing the at least one task. The corresponding value(s) can be determined after identifying the parameter(s) for the at least one task in block 356. In some embodiments, the corresponding value(s) can be automatically determined in response to identifying the parameter(s) for the at least one task based on the user profile(s) associated with a given user of the client device. For example, in response to identifying the quality assurance case number parameter, the system can, with permission from the given user (e.g., prior permission), access an email account associated with the given user and search for emails that include the corresponding value for the quality assurance case number parameter. Furthermore, in embodiments that identify an entity engaged with the given user during an ongoing call, the system can limit the search to only emails associated with the identified entity. In other embodiments, as opposed to being automatically identified, the corresponding value(s) can be determined in response to receiving user input including information requesting the corresponding value(s) for the parameter(s). The corresponding value(s) can be determined in response to receiving user input in the same or similar manner as described above.

[0080] In some embodiments, method 300 may include optional blocks 360, 362, and / or 364. If included, at optional block 360, the system can determine whether any user input for activating a secondary call was received at the client device associated with the given user. The system can determine whether the user input activated the secondary call based on spoken input, typed input, and / or touch input invoking the secondary call during an ongoing call between the given user of the client device and additional users of additional client devices in any manner described herein. If, at an iteration of optional block 360, the system determines that user input for activating the secondary call was received, the system can proceed to block 366, which will be discussed in more detail below. If, at an iteration of optional block 360, the system determines that no user input for activating the secondary call was received, the system can proceed to optional block 362.

[0081] If included, at optional block 362, the system can render a notification via the client device associated with the given user indicating that the auxiliary call can perform at least one task. The notification can include, for example, an indication that the additional user is requesting information about the corresponding value(s) determined in block 358 for the parameter(s) identified in block 356, and can also include an indication that the system can provide the corresponding value(s) to the additional user on behalf of the given user. The notification can be rendered visually and / or audibly. In embodiments where the notification is rendered audibly, the notification can be rendered audibly only on the client device so that the additional user of the client device is not aware of the notification (i.e., outside of the call). Optionally, to reduce the chance that the additional user will be aware of the notification, the ongoing call can be temporarily muted during the audible rendering of the notification, or the notification can be filtered using acoustic echo cancellation or other filtering and prevented from being provided as part of the ongoing call. In other embodiments where the notification is rendered audibly, the notification can be rendered audibly on both the given user's client device and the additional user's additional client device so that the notification interrupts the ongoing call between the given user and the additional user.

[0082] If so, at optional block 364, the system can determine whether any user input for activating an auxiliary call was received on the client device associated with the given user. The user input received at block 364 can be in response to rendering a notification indicating that the auxiliary call can perform at least one task. The system can determine whether the user input activates the auxiliary call based on spoken input, typed input, and / or touch input invoking the auxiliary call during an ongoing call between the given user of the client device and an additional user of an additional client device in any manner described herein. If, at an iteration of optional block 360, the system determines that no user input for activating the auxiliary call was received, the system can return to block 354 to process additional audio data corresponding to the ongoing call. For example, the system can determine that the given user of the client device provided spoken input for the corresponding value(s) included in the notification rendered at block 362, and the system can return to block 354 to continue processing the audio data stream to monitor any additional parameter(s) being requested by the additional user. If, at an iteration of optional block 364, the system determines that user input for activating the auxiliary call was received, the system can proceed to block 366.

[0083] At block 366, the system causes the value(s) to be rendered at the additional client device for presentation to the additional user. The system can cause synthesized speech including the corresponding value(s) to be rendered at the additional client device of the additional user and / or at the client device of the given user in response to receiving user input for activating an auxiliary call to provide the corresponding values ​​to the additional user.

[0084] In embodiments that include optional blocks 360, 362, and / or 364, the system can cause corresponding values ​​to be rendered on additional client devices of additional users and / or on the client device of a given user in response to receiving explicit user input to invoke an assisted call. In some versions of those embodiments, the user input to activate the assisted call and interrupt the ongoing call can be proactive. In other words, if the system receives user input at block 360, the system can render the corresponding value(s) on the additional client device of the additional user and / or on the client device of the given user in response to receiving explicit user input to invoke an assisted call. Figure 5A In other versions of those embodiments, the user input that activates the auxiliary call can be reactive. In other words, if the system receives user input at block 364, then at block 362, the system renders an indication that the system can perform at least one task (e.g., as described in reference to FIG. 364 ). Figure 5B After notification of the (described) task, an auxiliary call can be activated to provide the corresponding value(s) for the task. In embodiments that do not include optional blocks 360, 362, and / or 364, the system can proceed directly from block 258 to block 366. In some versions of those embodiments, the system can determine the corresponding value(s) at block 358 (e.g., as described with respect to the task). Figure 5C The active assisted call is automatically interrupted without any explicit user input being received (i.e., without receiving any explicit user input for the active assisted call).

[0085] At block 368, the system determines whether any user input for continuing the auxiliary call is received at the client device associated with the given user. If at an iteration of block 368, the system determines that no user input is received to continue the auxiliary call, the system can return to block 354 to process additional audio data corresponding to the ongoing call. If at an iteration of block 368, the system determines that user input is received to continue the auxiliary call, the system can proceed to Figure 2, and determines whether any value(s) for the additional parameter(s) are required during the secondary call. In this manner, the system is able to provide corresponding value(s) for the parameter(s) during the performance of the task(s) by a given user during the telephone call. By providing the corresponding value(s) automatically or in response to explicit user input as described above, the system receives less input from the given user of the client device because the user does not need to navigate to various applications with different user interfaces to determine the corresponding value(s), thereby conserving computing resources at the given client device. Furthermore, because the user does not need to navigate to these various applications, the system conserves computing and network resources by ending ongoing calls more quickly. As a non-limiting example, by using the techniques described herein, a given user does not need to pause a conversation or place it on hold while the user searches for the corresponding value(s) in an email application, an airline application, and / or other applications.

[0086] Now refer to Figures 4A-4D , describes various non-limiting examples of user interfaces for performing assisted calls. Figures 4A-4D , and 480 each depict a client device 410 having a graphical user interface 480 that displays examples of interactions for a given user of the client device 410. The interactions can include, for example, interactions with one or more software applications (e.g., a web browser application, an automated assistant application, a contacts application, an email application, a calendar application, and / or other software-based applications accessible to the client device 410), as well as interactions with additional users (e.g., additional human participants associated with additional client devices, additional automated assistants associated with additional client devices of additional users, and / or other additional users). Figure 1 One or more aspects of the automated assistant 115 may be implemented locally on the client device 410 and / or in a distributed manner (e.g., via Figure 1 (multiple) networks 190) and other client devices that communicate with the client device 410. For simplicity, Figures 4A-4D The operations are described herein as being performed by an automated assistant. Figures 4A-4D The client device 410 is depicted as a mobile phone, but it should be understood that this is not meant to be limiting. The client device 410 can be, for example, a standalone assistant device (e.g., with speaker(s) and / or display), a laptop computer, a desktop computer, and / or any other client device capable of making phone calls.

[0087] Figures 4A-4D4. The graphical user interface 480 of FIG. 4 further includes a text reply interface element 484 that the user can select to generate user input via a virtual keyboard or other touch and / or typing input, and a voice reply interface element 485 that the user can select to generate user input via the microphone(s) of the client device 410. In some embodiments, the user can generate user input via the microphone(s) without selecting the voice reply interface element 485. For example, active monitoring of audible user input via the microphone(s) can occur to avoid the need for the user to select the voice reply interface element 485. In some of those embodiments and / or in other embodiments, the voice reply interface element 485 can be omitted. Furthermore, in some embodiments, the text reply interface element 484 can additionally and / or alternatively be omitted (e.g., the user can provide only audible user input). Figures 4A-4D The graphical user interface 480 also includes system interface elements 481 , 482 , 483 , which can be interacted with by a user to cause the computing device 410 to perform one or more actions.

[0088] In various embodiments described herein, user input can be received to initiate a phone call (e.g., an assisted call) with an entity using an automated assistant. The user input can be a spoken input, a touch input, and / or a typed input, including an indication to initiate an assisted call. Furthermore, the automated assistant can perform task(s) with respect to an entity on behalf of a given user of the client device 410. Figure 4A As shown, user interface 480 includes search results for restaurant entities (e.g., as indicated by URL 411 of “www.exampleuril0.com / ”) from a browser application accessible on client device 410. In addition, the search results include a first search result 420 for “imaginary cafe” and a second search result 430 for “example cafe.”

[0089] In some embodiments, the search results 420 and / or 430 can be associated with various selectable graphical elements that, when selected, cause the client device 410 to perform corresponding actions. For example, when the call graphical element 421 and / or 431 associated with a given one of the search results 420 and / or 430 is selected, the user input can indicate that a phone call action should be performed to the restaurant entity associated with the search results 420 and / or 430. As another example, when the direction graphical element 422 and / or 432 associated with a given one of the search results 420 and / or 430 is selected, the user input can indicate that a navigation action should be performed to the restaurant entity associated with the search results 420 and / or 430. As yet another example, when the menu graphical element 423 and / or 433 associated with a given one of the search results 420 and / or 430 is selected, the user input can indicate that a browser-based action for displaying a menu for the restaurant entity associated with the search results 420 and / or 430 should be performed. Although in Figure 4A Assisted calls can be initiated from a browser application in the example embodiment, but it should be understood that this is for illustrative purposes and is not intended to be limiting. For example, assisted calls can be initiated from various software applications accessible on the client device 410 (e.g., a contacts application, an email application, a text or SMS messaging application, and / or other software applications), and if verbal input is used, assisted calls can be initiated from the home screen of the client device 410, from the lock screen of the client device 410, and / or from other states of the client device 410.

[0090] For purposes of example, assume that user input is detected on client device 410 to initiate a phone call using second search result 430 for "Example Cafe." The user input can be, for example, a spoken input of "Call Example Cafe" or a touch input directed to call graphical element 431. In some embodiments, a call details interface 470 can be rendered on client device 410 in response to receiving the user input to initiate a phone call using "Example Cafe." In some versions of those embodiments, call details interface 470 can be rendered on client device 410 as part of user interface 480. In some other versions of those embodiments, call details interface 470 can be a separate interface from user interface 480, which overlays user interface 480, and can include call details interface element 486 that allows the user to expand call details interface 470 to display additional call details (e.g., by swiping up on call details interface element 486) and / or dismiss call details interface 470 (e.g., by swiping down on call details interface element 486). While call details interface 470 is depicted at the bottom of user interface 480, it should be understood that this is for illustrative purposes and is not intended to be limiting. For example, call details interface 470 can be rendered on top of user interface 480 , to the side of user interface 480 , or as an interface completely separate from user interface 480 .

[0091] In various embodiments, call details interface 470 can include multiple graphical elements. In some embodiments, the graphical elements can be selectable such that when a given one of the graphical elements is selected, client device 410 can perform a corresponding action. Figure 4AAs shown in FIG, call details interface 470 includes a first graphical element 471 for "Assisted Call," a second graphical element 472 for "Regular Call," and a third graphical element 473 for "Save Contact Example 'Example Cafe'." Furthermore, first graphical element 471, when selected, can provide an indication to the automated assistant that an assisted call is desired to be initiated using the automated assistant; second graphical element 472, when selected, can cause the automated assistant to initiate a call without using the assisted call; and third graphical element 473, when selected, can cause the automated assistant to create a contact associated with the Example Cafe. Notably, in some versions of those embodiments, graphical elements can include sub-elements to provide instructions for tasks to be performed. For example, the first graphical element 471 of “Assisted Call” may include a first sub-element 471A of “Make Reservation” associated with the task of making a restaurant reservation at the example cafe, a second sub-element 471B of “Modify Reservation” associated with the task of modifying the reservation at the example cafe restaurant, and a third sub-element 471C of “Cancel Reservation” associated with the task of canceling the reservation at the example cafe restaurant.

[0092] For purposes of example, assume that user input is detected at client device 410 to initiate an assisted call with Example Café to make a restaurant reservation at Example Café. The user input can be, for example, a spoken input of "call Example Café to make a restaurant reservation" or a touch input directed to first sub-element 471A. In response to detecting the user input, the automated assistant can determine the task of "make a restaurant reservation at Example Café" and can identify candidate parameter(s) associated with the identified task, as described herein (e.g., with respect to Figure 1 (a plurality of) parameter engines 153). In some embodiments, and as Figure 4B As shown in , the automated assistant is able to determine the value(s) of the candidate parameter(s). Figures 4A to 4BDuring the call, call details interface 470 can be updated to include the identified candidate parameter(s). In some versions of those embodiments, the automated assistant can determine the value(s) for the candidate parameter(s) based on the user profile(s) associated with a given user of client device 410. For example, the automated assistant can determine a value 474A for name parameter 474 (e.g., Jane Doe) and a value 475A for phone number parameter 475 (e.g., (502) 123-4567). Although described herein with respect to specific candidate parameter(s) identified by the automated assistant, Figure 4B , however it should be understood that this is for illustrative purposes.

[0093] In some versions of those embodiments, the automated assistant can engage in a (e.g., audible and / or visual) dialog with a given user of client device 410 to request corresponding value(s) for candidate parameter(s) that are not identified based on the user profile(s) of the given user of client device 410. In some other versions of those embodiments, the automated assistant may only request parameters that are considered as described herein (e.g., with respect to Figure 1 1 ). The automated assistant can generate a prompt 452B1 for a task of making a restaurant reservation, and receive a user input 454B1 (e.g., typed or spoken) of “March 1st, 7:00 PM with booth seating”. As a result, the call detail interface 470 can be updated to include a value 476A for the date / time parameter 476 (e.g., March 1, 2020, at 7:00 PM). Notably, the user input 454B1 also includes a value 478A of “booth seating” for a type of seating parameter 478 that was not requested in the prompt 452B1. Even though the automated assistant did not request value 478A in prompt 452B1, the automated assistant can determine that value 478A corresponds to a type of seating parameter 478, which can be considered an optional parameter as described herein (e.g., with respect to Figure 1Parameter(s) engine 153). In addition, the automated assistant can generate an additional prompt (e.g., prompt 452B2) for the additional parameter (e.g., headcount parameter 477) to determine a corresponding value for the additional parameter (e.g., 477A with a value of five). In this way, the automated assistant can determine the value(s) of the candidate parameter(s) to be used when performing a task on behalf of a given user of client device 410 before initiating an assistance call.

[0094] Furthermore, in various embodiments, when determining the value(s) of the candidate parameter(s), and as Figure 4B As shown, the call details interface 470 can include various graphical elements. For example, the call details interface 470 can include an edit graphical element 441B, which, when selected, allows the user to modify the values ​​474A-478A before initiating the assisted call, a cancel graphical element 442B, which, when selected, ends the assisted call and the user interface 480 can optionally return the user interface 480 to a state before detecting the user input for initiating the assisted call (e.g., Figure 4A ), and a call interface element 443B, which, when selected, allows the user to initiate a regular call without the assistance of an automated assistant (e.g., similar to the Figure 4A selection of graphic element 472).

[0095] After determining the corresponding value(s) 474A-478A for the candidate parameter(s) 474-478, the automated assistant can initiate an assistance call on behalf of a given user of the client device 410 to perform a task with respect to the entity. The automated assistant can initiate the assistance call using a calling application accessible at the client device 410. In various embodiments, and as Figure 4C As shown in , in response to a call being initiated by the automated assistant, the call details interface 470 can be updated to include various graphical elements. For example, the call details interface 470 can include an end call graphical element 441C, which, when selected, causes the automated assistant to end the assisted call after initiating the assisted call, a join call graphical element 442C, which, when selected, allows a given user of the client device 410 to take over the call and / or perform tasks from the automated assistant, and a speaker interface element 443C, which, when selected, causes the client device 410 to audibly render the conversation between the additional user and the automated assistant. These graphical elements 442C, 443C, and 444C can be selected throughout the duration of the assisted call.

[0096] Furthermore, in various embodiments, the automated assistant can obtain consent from additional users of the client device after initiating the assistance call and before performing the task. Figure 4CAs shown, the automated assistant can cause synthesized speech 452C1 to be rendered on an additional client device associated with the example cafe representative, the additional client device requesting the example cafe representative to provide consent for interacting with the automated assistant. The automated assistant can process audio data 456C1 corresponding to the additional spoken input from the example cafe representative to determine that the example cafe representative provided consent (e.g., "yes" in audio data 456C1). Furthermore, the automated assistant can cause additional synthesized speech 456C2 including value 476A for date / time parameter 476 to be rendered on the additional client device of the example cafe representative in response to determining that audio data 456C1 requests information for value 476A for date / time parameter 476. In this manner, the automated assistant can perform the task of making a restaurant reservation at the example cafe by providing synthesized speech that includes value(s) responsive to the example cafe representative's request for information.

[0097] In some implementations, the automated assistant can process audio data corresponding to the additional spoken input from the example cafe representative and can determine that the additional spoken input is requesting information about additional value(s) associated with parameters that are not currently known to the automated assistant. Figure 4C As shown, audio data 456C2 captures additional verbal input requesting information about a child parameter that is currently unknown to the automated assistant. In response to determining that the additional verbal input from the example cafe representative is requesting information that is currently unknown to the automated assistant, the automated assistant can cause rendering at the additional client device of yet another synthesized speech indicating that the requested information is currently unknown. Although the automated assistant may not currently know the information being requested, the automated assistant can utilize the other value(s) that the automated assistant currently knows to continue performing the task. For example, the automated assistant can cause rendering of yet another synthesized speech 454C3 stating, “I’m not sure, I’ll have to ask Jane Doe and get back with you can we continue making the reservation until I hear back?” in response to determining that audio data 456C2 requests information about additional value(s) associated with parameter(s) that are currently unknown to the automated assistant.

[0098] In some versions of those embodiments, when the automated assistant determines that the additional user's audio data includes a request for information, the automated assistant can cause a notification to be rendered at client device 410 indicating that the additional user is requesting information associated with a parameter that is not currently known to the automated assistant. In some other versions of those embodiments, the notification can further include suggested values ​​as recommendation(s) in response to the request for information. The recommendation(s) can be selectable such that upon selecting a given one of the recommendations, the automated assistant can utilize the value(s) included in the given one of the recommendations as the value in response to the additional user's request. For example, Figure 4D As shown, the automated assistant can cause a notification 479 to be visually rendered in the call detail interface 470. Notification 479 includes an indication that "Example Café wants to know if any children will be joining the reservation" and also includes a first suggestion 479A of "Yes" and a second suggestion 479B of "No" provided as recommendation values ​​in response to the request for the additional user. As described in more detail below, various other types of notifications can be rendered on the client device 410, and the type of notification can be based on the state of the client device 410 when the system determines that the audio data includes a request for information that is not currently known to the automated assistant.

[0099] As described above, even if the automated assistant determines that an additional user is requesting information that is currently unknown to the automated assistant, the automated assistant can continue to perform the task using the other value(s) currently known to the automated assistant. The additional value(s) can be provided later in the conversation after receiving additional user input from a given user of client device 410 in response to notification 479. For example, Figure 4DAs depicted in FIG, in response to rendering yet another synthesized speech 452C3 indicating that the value of the children parameter is currently unknown to the automated assistant, but the automated assistant knows other value(s) and wants to continue performing the task using those other value(s). Furthermore, assume that audio data 456D1 is received, "Sure, what type of seating?" (e.g., a request for information regarding value 478A for the type of "Booth" for seating parameter 478). Further assume that a given user of client device 410 provides verbal input and / or touch input directed to second suggestion 479B indicating that no children will be added to the restaurant reservation at the example coffee shop while audio data 456D1 is being processed by the automated assistant. In this example, the automated assistant can cause synthesized speech 452D1 including value 478A for the type of "Booth" for seating parameter 478 and including that value (e.g., based on a user selection of "No" for second suggestion 479B) to be audibly rendered on an additional client device of the additional user. The automated assistant can process the additional audio data 456D2 of “Perfect, the reservation is complete” to determine that the restaurant reservation task is completed, and can result in additional synthesized audio data 456D2 of “I’ll let Jane Doe know, have a nice day!” and can end the call.

[0100] Notably, in some embodiments, the automated assistant is capable of including additional value(s) in the synthesized speech in response to a request for additional information from an additional user (i.e., previously unknown at the time of the request by the additional user but now known based on additional user interface input). In other words, the automated assistant is capable of including additional value in the synthesized speech even if the previous request from the additional user was not a request for information. For example, Figure 4DAs depicted, synthesized speech 452D1 includes a "Booth" value 478A in response to an immediately previous request for information associated with the type of seating parameter 478 included in audio data 456D1, and the synthesized speech also includes a "No" value for the child parameter in response to a previous request for information associated with the child parameter included in audio data 456C2. In this manner, the automated assistant is able to continue performing a task while waiting for additional user input including additional value(s), and is able to provide the additional value(s) to the additional user as the additional value(s) become known to the automated assistant in a logical and conversational manner. By continuing to perform a task while waiting for additional user input, the task can be performed in a faster and more efficient manner because execution of the task is not paused until the additional user input is received. By performing tasks in a faster and more efficient manner, the techniques described herein can conserve computing and network resources when performing tasks using an assisted call.

[0101] In various embodiments, and although not depicted, the automated assistant may determine during the assistance call that no further audio data has been received from an additional user associated with the entity within a threshold duration. In some versions of those embodiments, the automated assistant can render a further synthesized speech based on one or more corresponding values ​​that have not yet been included in the request for information from the additional user. For example, if the automated assistant determines that the additional user has not spoken anything for ten seconds and the automated assistant knows the value of the party size for a restaurant reservation task that the additional user has not yet requested, the assistant can render a synthesized speech that reads, "In case you wanted to know, five people will be joining thereservation." In some additional and / or alternative versions of those embodiments, the automated assistant can render a further synthesized speech based on one or more parameters that have not yet been included in the request for information from the additional user. For example, if the automated assistant determines that the additional user has not spoken anything for ten seconds and the automated assistant knows the value of the party size for a restaurant reservation task that the additional user has not yet requested, the assistant can render a synthesized speech that reads, "Do you want to know how many people will be joining thereservation?"

[0102] Furthermore, in various embodiments, the automated assistant can cause transcripts of various conversations to be visually rendered in the user interface 480 of the client device 410. For example, the transcripts can be displayed in various software applications (e.g., the assistant application, the calling application, and / or other applications) within the home of the client device 410. In some embodiments, the transcripts can include a conversation between the automated assistant and a given user of the client device 410 (e.g., Figure 4B In some additional and / or alternative embodiments, the transcription can include a conversation between the automated assistant and an additional user (e.g., as Figure 4C and 4D as depicted in ).

[0103] although Figures 4B-4D Each of the are depicted as including a transcription of the conversation, but it should be noted that this is for purposes of example and is not meant to be limiting. It should be understood that the above-described assisted calls can be performed while the client device 410 is in a sleep state, a locked state, while (multiple) other software applications are operating in the foreground, and / or in other states. Furthermore, in embodiments where the automated assistant causes notification(s) to be rendered at the client device 410, the type of notification(s) rendered at the client device is based on the state of the client device 410, as described herein. Furthermore, although Figures 4A-4D While described herein with respect to the task of making restaurant reservations, it should be understood that this is not meant to be limiting and that the techniques described herein can be used for a number of different tasks that can be performed with respect to a number of different entities.

[0104] Furthermore, in various embodiments, the automated assistant can be placed on hold by the additional user at the start of an assisted call and / or during an assisted call. In some versions of those embodiments, the automated assistant can be considered on hold when it is not participating in a conversation with an additional human participant. For example, the automated assistant can be considered on hold if it is occupied by a hold system associated with the entity, an interactive voice response (IVR) system associated with the entity, and / or other systems associated with the entity. Furthermore, when the assisted call is placed on hold and / or when the assisted call is resumed after being placed on hold, the automated assistant can cause a notification to be rendered on the client device 410. The notification can indicate, for example, that the assisted call was placed on hold, that the assisted call has been resumed after being placed on hold, that the user was requested to join the assisted call when it was resumed after being placed on hold, and / or other information related to the assisted call. Furthermore, the automated assistant can determine that the assisted call has been resumed after being placed on hold based on processing of an audio data stream transmitted from the additional client devices of the additional users to the client device 410 of a given user.

[0105] In some versions of those embodiments, the automated assistant can determine that the additional user has requested all information associated with the parameter(s) known to the automated assistant. In some other versions of those embodiments, if the remainder of the task requires the given user of client device 410, the automated assistant can proactively request the given user of client device 410 to join the assisted call when the assisted call is resumed after being placed on hold. For example, certain tasks may require the given user of client device 410 but not the automated assistant. For example, assume that the entity is a bank, the additional user is a bank representative, and the task is to dispute a debit card charge. In this example, the automated assistant can initially call the bank, providing synthesized speech and / or simulated button presses, such as synthesized speech including the given user's name, simulated button presses including the user's bank account number, synthesized speech providing a reason for the call, and / or simulated button presses for navigating through an automated system (e.g., an IVR system) used to route calls associated with the bank. However, the automated assistant can be aware that when the assisted call is transferred to the bank representative, the given user will be asked to join the assisted call to verify the given user's identity and explain the disputed debit card charge, and can cause a notification to be rendered on the user's client device 410 when the assisted call is transferred to the bank representative. Thus, the automated assistant can handle the initial portion of the assisted call and, when the bank representative is available to discuss the disputed charge, request the given user of the client device 410 to take over the assisted call.

[0106] In some other additional versions of those embodiments, the automated assistant can proactively request that a given user of client device 410 join the assisted call when the assisted call is resumed after being placed on hold, because the automated assistant does not know any additional information about the restaurant reservation. Figures 4B-4D The additional user on the assisted call depicted in audio data 456D2 places the automated assistant on hold rather than indicating that the restaurant reservation is complete. In this example, the automated assistant has already provided values ​​for all the information known to the automated assistant regarding the restaurant reservation, and any additional information requested by the additional user when the assisted call is resumed will be information currently unknown to the automated assistant. Thus, the automated assistant can proactively request that the given user of client device 410 take over the assisted call when it is resumed.

[0107] In still further versions of those embodiments, the automated assistant can retain a notification indicating that the auxiliary call was placed on hold and / or that the auxiliary call has been resumed after being placed on hold, resume the auxiliary call once the auxiliary call has been resumed after being placed on hold, and have that notification along with or in lieu of a notification indicating that the additional user is requesting information associated with parameter(s) that the automated assistant is not currently aware of (e.g., Figure 4DContinuing with the above example, rather than proactively requesting the given user of client device 410 to take over the assisted call when the assisted call is resumed, the automated assistant can process additional audio data corresponding to the additional user's additional verbal input to determine whether the additional user is requesting additional information associated with parameter(s) that are currently unknown to the automated assistant. In this example, the automated assistant can process additional audio data corresponding to the additional user's additional verbal input in addition to or in lieu of a parameter similar to Figure 4D The notification of notification 479 provides a notification requesting a given user of client device 410 to take over the assistance call, and the value of which requests additional information. In this way, the automated assistant is able to Figures 4B-4D Continuing the assisted call and / or passing control of the assisted call to a given user of the client device 410 as described in .

[0108] Now refer to Figures 5A-5C , describes various non-limiting examples of user interfaces for providing auxiliary output during an ongoing non-assisted call. Figures 5A-5C Each depicts a client device 510 having a graphical user interface 580 that displays an example of interactions for a given user of the client device 510. The interactions can include, for example, interactions with one or more software applications (e.g., a web browser application, an automated assistant application, a contacts application, an email application, a calendar application, and / or other software-based applications accessible to the client device 510), as well as interactions with additional users (e.g., an automated assistant associated with the client device 510, additional human participants associated with additional client devices, additional automated assistants associated with additional client devices of additional users, and / or other additional users). Figure 1 One or more aspects of the automated assistant 115 may be implemented locally on the client device 510 and / or implemented in a distributed manner (e.g., via Figure 1 For simplicity, Figures 5A-5C The operations are described herein as being performed by an automated assistant. Figures 5A-5C The client device 510 is depicted as a mobile phone, but it should be understood that this is not meant to be limiting. The client device 510 can be, for example, a standalone speaker, a speaker connected to a graphical user interface, a laptop computer, a desktop computer, and / or any other client device capable of making a phone call.

[0109] Figures 5A-5CThe graphical user interface 580 further includes a text reply interface element 584 that the user can select to generate user input via a virtual keyboard or other touch and / or typing input, and a voice reply interface element 585 that the user can select to generate user input via the microphone(s) of the client device 510. In some embodiments, the user can generate user input via the microphone(s) without selecting the voice reply interface element 585. For example, active monitoring of audible user input via the microphone(s) can occur to eliminate the need for the user to select the voice reply interface element 585. In some of those embodiments and / or in other embodiments, the voice reply interface element 585 can be omitted. Furthermore, in some embodiments, the text reply interface element 584 can additionally and / or alternatively be omitted (e.g., the user can provide only audible user interface input). Figures 5A-5C Graphical user interface 580 also includes system interface elements 581, 582, 583 that can be interacted with by a user to cause computing device 510 to perform one or more actions. In some implementations, call details interface 570 can be rendered on client device 510, and call details interface 510 can include graphical element 542 that, when selected, can cause an ongoing call to end, and can also include graphical element 543 that, when selected, can cause the automated assistant to take over a call from a given user of client device 510 using an assisted call. In some versions of those implementations, assisted call interface 570 can also include call details interface element 586 that allows a user to expand call details interface 570 to display additional call details (e.g., by swiping up on call details interface element 586) and / or dismiss call details interface 570 (e.g., by swiping down on call details interface element 586).

[0110] In various embodiments, the automated assistant can interrupt an ongoing call (i.e., not an auxiliary call) between a given user of client device 410 and an additional user of an additional client device. In some embodiments, the automated assistant can process audio data corresponding to the ongoing call to identify an entity associated with the additional user, task(s) to be performed during the ongoing call, and / or parameter(s) for the task(s) to be performed during the ongoing call. Furthermore, the automated assistant can determine value(s) for the identified parameter(s). For example, Figures 5A-5CAs depicted, assume that audio data 552A1, 552B1, and / or 552C1 capturing spoken input from an additional user, "Example Airlines representative, how may I help you?" is received at client device 510. In this example, the automated assistant can process audio data 552A1, 552B1, and / or 552C1 and, based on the processing, can determine that the additional user is an Example Airlines representative and / or identify Example Airlines as an entity associated with the Example Airlines representative. In some additional and / or alternative embodiments, the automated assistant can additionally and / or alternatively identify the entity based on metadata associated with the ongoing call, as described herein (e.g., with respect to Figure 1 In some versions of those embodiments, the automated assistant may not process the audio data stream corresponding to the ongoing call and identify entities based solely on metadata associated with the ongoing call, thereby eliminating the need to obtain consent from additional users of the ongoing call. As described below, the automated assistant can still perform task(s) on behalf of a given user of client device 510 without processing the audio data stream corresponding to the ongoing call.

[0111] Assume further that audio data 554A1, 554B1, and / or 554C1 capturing spoken input of “Hello, I need to change my flight” from a given user of client device 510 (e.g., Jane Doe) is detected at client device 510 and transmitted to additional client devices of additional users. In this example, the automated assistant can process audio data 554A1, 554B1, and / or 554C1 and, based on the processing, can determine a task to change the flight. In other examples, the automated assistant can additionally and / or alternatively determine task(s) associated with an entity as described herein (e.g., regarding Figure 1assisted call engine 150). Further assume that audio data 552A2, 552B2, and / or 552C2 capturing spoken input from an additional user, "Alright, do you have a frequent flyer number?" is received at client device 510. In this example, the automated assistant can process audio data 554A2, 554B2, and / or 552C2 and, based on the processing, can identify a frequent flyer number parameter for a task of changing a flight (or a task that considers providing a frequent flyer number). In other examples, the automated assistant can additionally and / or alternatively identify parameter(s) based on parameter(s) stored in association with a task and / or entity, as described herein (e.g., with respect to Figure 1 Additionally, the automated assistant can determine a value for a common flyer number parameter and can provide the value for the common flyer number parameter to the additional user.

[0112] In some implementations, the automated assistant can determine the value(s) of the parameter(s) identified during an ongoing call between the given user of client device 510 and the additional user in response to receiving user input from the given user of client device 510 including a request for information associated with the parameter(s). The automated assistant can determine the value(s) of the parameter(s) identified during an ongoing call between the given user of client device 510 and the additional user(s). Figure 1 In some versions of those embodiments, the automated assistant can cause the value(s) of the parameter(s) to be rendered on additional client devices of additional users and / or the given user's client device 510 in response to determining the value(s) of the parameter(s). For example, Figure 5AAs shown, assume that audio data 554A2 capturing spoken input “Assistant, what’s my frequent flier number?” from a given user of client device 510 (e.g., Jane Doe) is detected at client device 510. In response to receiving the spoken input captured in audio data 554A2, the automated assistant can determine a value for a frequent flyer parameter (e.g., based on example airline accounts associated with the given user being included in the user profile(s)) and can cause synthesized speech 556A1 of “Jane Doe’s Example Airlines frequent flier number is: 0112358” to be rendered on additional client devices of additional users and / or the given user’s client device 510.

[0113] In some additional and / or alternative versions of those embodiments, rather than receiving verbal input included in audio data 554A2, the automated assistant can receive a selection of graphical element 543 to take over the call using assisted calling for a given user of client device 510. For example, in response to receiving selection of graphical element 543, the automated assistant can determine values ​​for common flyer parameters (e.g., based on example airline accounts associated with the given user included in the user profile(s)) and can cause synthesized speech 556A1 to be rendered on additional client devices of additional users and / or on the given user's client device 510.

[0114] Notably, audio data 554A2 includes "what's my frequent flier number" without identifying an entity associated with the frequent flier number. In embodiments where the automated assistant does not obtain consent from the additional user and / or provide an indication of an entity associated with the frequent flier number parameter, the automated assistant can determine, based on metadata associated with the ongoing call, that "my frequent flier number" refers to a value for the frequent flier number parameter associated with the example airline. In this manner, the automated assistant can still provide value(s) in response to a request from a given user of client device 510 without processing the audio data stream corresponding to the ongoing call.

[0115] In other embodiments, the automated assistant can proactively determine the value(s) of the parameter(s) identified during an ongoing call between the given user of client device 510 and the additional user without receiving any user input from the given user of client device 510 including a request for information associated with the parameter(s). Figure 1 The (multiple) values ​​of the (multiple) parameters are determined based on the (multiple) user profiles in the (multiple) user profile database 153B.

[0116] In some versions of those embodiments, the automated assistant can cause a notification to be rendered at client device 510 that includes an indication that the assistance call can perform the task and / or can provide value(s) of the parameter(s) to the additional user. The automated assistant can cause synthesized speech including the value(s) to be rendered at an additional client device of the additional user in response to receiving user input from a given user of client device 510 invoking the assistance speech. For example, Figure 5B As shown, assume that the automated assistant determines the value of the frequent flyer number parameter in response to a request to identify the value of the frequent flyer number parameter in audio data 552B2. Further assume that the automated assistant causes notification 579 to be rendered in call detail interface 570 of client device 510: "Your Example Airlines frequent flier number is 0112358, would you like me to provide it to Example Airlines Representative?" The automated assistant can then cause synthesized speech 556B1 to be rendered on additional client devices of additional users and / or the given user's client device 510: "Jane Doe's Example Airlines frequent flier number is: 0112358" in response to receiving selection of graphical element 579A and / or graphical element 543 indicating that the automated assistant should provide the value of the frequent flyer number parameter to the additional user. In some additional and / or alternative embodiments, the automated assistant can detect spoken input from a given user of client device 510 that includes value(s) for parameter(s) identified during an ongoing conversation. In some versions of those embodiments, the automated assistant can automatically dismiss notification 579 included in call details interface 579.

[0117] In some other versions of those embodiments, the automated assistant can also proactively provide the value(s) of the parameter(s) identified during the ongoing call in response to determining the value(s) of the parameter(s), and without receiving any user input including a request for information associated with the parameter(s). The automated assistant can cause synthesized speech including the value(s) to be rendered at an additional client device of the additional user in response to determining the value(s). For example, Figure 5C As shown, assume that the automated assistant determines the value of the frequent flyer number parameter in audio data 552C2 in response to a request from an additional user to identify a value for the frequent flyer number parameter. The automated assistant can then cause synthesized speech 556C1 of "Jane Doe's Example Airlines frequent flier number is: 0112358" to be rendered on the additional client device of the additional user and / or the given user's client device 510 in response to determining the value of the frequent flyer number parameter and without receiving any user input including a request for information associated with the frequent flyer number parameter. In some further versions of those embodiments, the automated assistant may only proactively provide the identified value(s) of the parameter(s) during an ongoing call if a confidence metric associated with the determined value(s) of the parameter(s) satisfies a confidence threshold. The confidence metric can be based on, for example, whether a given user previously provided the determined value in response to receiving a previous request for the same information, whether the parameter(s) identified during the ongoing call are candidate parameter(s) stored in association with the task identified during the ongoing call, the source from which the value was determined (e.g., an email / calendar application versus a text messaging application), and / or the manner in which the confidence metric was determined.

[0118] Furthermore, in various embodiments, after the automated assistant provides the value(s) for the identified parameter(s) during the ongoing call, the automated assistant can take over the remainder of the ongoing call. In some versions of those embodiments, the automated assistant can take over the remainder of the ongoing call in response to a user selection of graphical element 543. Thus, the automated assistant can continue to identify the parameter(s) for the task based on the audio data transmitted to the client device 510 and determine the value(s) for the identified parameter(s). Furthermore, if the automated assistant determines that a given request is for information that cannot be determined by the automated assistant, the automated assistant can provide a notification requesting additional user input in response to the request, and the automated assistant can then provide values ​​based on the additional user input in response to the request, as described above (e.g., with respect to Figure 4C and 4D ).

[0119] although Figures 5A-5C While described herein with respect to the task of providing a common flyer number, it should be understood that this is not meant to be limiting and that the techniques described herein can be used for a variety of different tasks that can be performed with respect to a variety of different entities. Figures 5A-5C , it should be understood that the automated assistant can obtain consent from the additional user when the ongoing call is initiated, even if the ongoing call is not initiated by the automated assistant using the assisted call. The automated assistant can obtain consent from the additional user using any of the methods described herein.

[0120] Figure 6 6 is a block diagram of an example computing device 610 that can optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client device, cloud-based automated assistant(s), and / or other components can include one or more components of the example computing device 610.

[0121] The computing device 610 typically includes at least one processor 614 that communicates with a number of peripheral devices via a bus subsystem 612. These peripheral devices may include a storage subsystem 624, including, for example, a memory subsystem 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow a user to interact with the computing device 610. The network interface subsystem 616 provides an interface to an external network and couples to corresponding interface devices in other computing devices.

[0122] The user interface input devices 622 may include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphic tablet, a scanner, a touch screen incorporated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 610 or onto a communication network.

[0123] User interface output device 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. Generally, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from computing device 610 to a user or to another machine or computing device.

[0124] The storage subsystem 624 stores programs and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 may include programs for performing selected aspects of the methods disclosed herein and implementing Figure 1 The logic of the various components is depicted in .

[0125] These software modules are typically executed by processor 614 alone or in combination with other processors. The memory 625 used in storage subsystem 624 can include many memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 for storing fixed instructions. File storage subsystem 626 can provide persistent storage for program and data files and can include a hard drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of certain embodiments can be stored in storage subsystem 624 by file storage subsystem 626, or in other machines accessible by processor(s) 614.

[0126] The bus subsystem 612 provides a mechanism for the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.

[0127] The computing device 610 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Figure 6 The description of the computing device 610 depicted in FIG is intended only as a specific example for purposes of illustrating some embodiments. Many other configurations of the computing device 66 are possible with more Figure 6 The computing devices depicted in the drawings may have more or fewer components.

[0128] In cases where the systems described herein collect or otherwise monitor personal information about a user or can make use of the personal information and / or monitored information, the user may be provided with an opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social behavior or activities, occupation, user preferences, or the user's current geographic location), or to control whether and / or how content that may be more relevant to the user is received from a content server. In addition, before certain data is stored or used, it may be processed in one or more ways to remove personally identifiable information. For example, the user's identity may be processed so that the user's personally identifiable information cannot be determined, or the user's geographic location may be summarized (such as to a city, ZIP code, or state level) in the case where geographic location information is obtained so that the user's specific geographic location cannot be determined. Thus, the user may have control over how information about the user is collected and / or used.

[0129] In some embodiments, a method implemented by one or more processes is provided and includes receiving user input from a given user via a client device associated with the given user to initiate an assistance call and determining, based on the user input, entities to be engaged on behalf of the given user during the assistance call and tasks to be performed on behalf of the given user during the assistance call. The method also includes determining, for one or more candidate parameters stored in association with the tasks and / or entities, one or more corresponding values ​​to be used when automatically generating synthesized speech during the assistance call when performing the task, initiating the assistance call using the client device associated with the given user, and, during the assistance call and based on processing audio data from the assistance call that captures utterances of the additional user associated with the entity, determining that information associated with the additional parameters is requested by the additional user. The method further includes, in response to determining that information associated with the additional parameters is requested, causing the client device to render a notification outside of the assistance call identifying the additional parameters and requesting additional user input for the information. The method also includes continuing the assistance call before receiving any additional input responsive to the notification. Continuing the assistance call includes rendering one or more instances of synthesized speech based on one or more of the corresponding values ​​of the candidate parameters. The method further includes determining, during the continuing assistance call, whether additional user input responsive to the notification and identifying a specific value for the additional parameter is received within a threshold duration, and rendering additional synthesized speech based on the specific value as part of the assistance call in response to determining that the additional user input is received within the threshold duration.

[0130] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.

[0131] In some embodiments, determining a given value among the corresponding values ​​includes generating a prompt that identifies a given candidate parameter among the candidate parameters and requests additional information associated with the given candidate parameter before initiating the auxiliary call, causing the client device to render the prompt, and identifying a given value for the given candidate parameter based on additional user input in response to the prompt.

[0132] In some embodiments, determining the additional ones of the corresponding values ​​includes identifying the additional values ​​based on a user profile associated with the given user prior to initiating the secondary call.

[0133] In some embodiments, continuing the auxiliary call includes processing additional audio data of the auxiliary call to determine that an additional utterance of the additional user includes a request for a given one of the candidate parameters, and, in response to determining that the additional utterance includes the request for the given candidate parameter, causing the client device to render a given one of the one or more instances of synthesized speech during the call. In some versions of those embodiments, the given instance includes the given one of the corresponding values ​​based on determining the given value for the given candidate parameter, and the given instance is rendered without requesting any additional user input from the given user.

[0134] In some embodiments, continuing the assistance call includes processing additional audio data to determine whether additional utterances from the additional user are received within an additional threshold duration, and in response to determining that additional utterances are not received from the additional user within the additional threshold duration, rendering an additional instance of one or more instances of synthesized speech during the assistance call, the additional instance being based on one or more corresponding values ​​of the corresponding values ​​that have not been requested by the additional user.

[0135] In some implementations, the method further includes updating one or more candidate parameters stored in association with the entity to include the additional parameter in response to determining that information associated with the additional parameter is requested by the additional user.

[0136] In some implementations, the method further includes determining a state of the client device when the additional user requests information associated with the additional parameters, and determining the notification and / or one or more attributes for rendering the notification based on the state of the client device.

[0137] In some versions of those embodiments, the state of the client device indicates that the given user is actively monitoring the auxiliary call, and determining the notification based on the state of the client device includes determining a notification including a visual component visually rendered via a display of the client device along with one or more selectable graphical elements based on the state of the client device indicating that the given user is actively monitoring the auxiliary call. In some versions of those embodiments, in response to the notification, the further user input includes a selection of a given selectable graphical element of the one or more selectable graphical elements.

[0138] In some versions of those embodiments, the state of the client device indicates that the given user is not actively monitoring the auxiliary call, and determining the notification based on the state of the client device includes determining the notification to include an audible component audibly rendered via one or more speakers of the client device based on the state of the client device indicating that the given user is actively monitoring the auxiliary call.

[0139] In some embodiments, the method further includes terminating the assisted call and, after terminating the assisted call, causing the client device to render an additional notification including an indication of a result of the assisted call. In some versions of those embodiments, the additional notification including an indication of the assisted call includes an indication of an additional task to be performed on behalf of the user in response to terminating the assisted call, or includes one or more selectable graphical elements that, when selected, cause the client device to perform the additional task on behalf of the user.

[0140] In some implementations, the method further includes terminating the assisted call in response to determining that no additional user input is received within the threshold duration, and after terminating the assisted call, causing the client device to render an additional notification including an indication of an outcome of the assisted call.

[0141] In some implementations, the threshold duration is: a fixed duration from the time a notification identifying additional parameters and requesting additional user input for information is rendered, or a dynamic duration based on when a last one or more of the corresponding values ​​was rendered for presentation to an additional user via an additional client device.

[0142] In some embodiments, the method further includes obtaining, after initiating the assistance call, consent from an additional user associated with the entity to monitor the assistance call.

[0143] In some embodiments, a method implemented by one or more processors is provided and includes: detecting, at a client device, an ongoing call between a given user of the client device and additional users of additional client devices, processing an audio data stream capturing at least one spoken utterance during the ongoing call to generate recognition text. The at least one spoken utterance is of the given user or the additional user. The method further includes identifying information of at least one spoken utterance request parameter based on processing the recognition text, and determining, for the parameter and using access-restricted data personal to the given user, that a value for the parameter is resolvable. The method further includes rendering output based on the value during the ongoing call in response to determining that the value is resolvable.

[0144] These and other implementations of the technology disclosed herein can optionally include one or more of the following features.

[0145] In some embodiments, the method further includes resolving a value of the parameter. In some versions of those embodiments, rendering the output during the ongoing call is further responsive to the resolved parameter value. In some versions of those embodiments, resolving the value of the parameter includes analyzing metadata of the ongoing call between the given user and the additional user, identifying an entity associated with the additional user based on the analysis, and resolving the value based on the value being stored in association with the entity and the parameter.

[0146] In some embodiments, the output includes synthesized speech, and rendering the value-based output during the ongoing call includes rendering the synthesized speech as part of the ongoing call. In some versions of those embodiments, the method further includes receiving user input from the given user to activate assistance during the ongoing call before rendering the synthesized speech as part of the ongoing call. In some versions of those embodiments, rendering the synthesized speech as part of the ongoing call is further responsive to receiving user input to activate assistance.

[0147] In some embodiments, the output includes a notification rendered at the client device and outside of the ongoing call. In some versions of those embodiments, the output further includes synthesized speech rendered as part of the ongoing call, and the method further includes rendering the synthesized speech after rendering the notification and in response to receiving affirmative user input in response to the notification.

[0148] In some embodiments, the method further includes determining, based on processing the audio data stream within a threshold duration following the at least one spoken utterance requesting information about the parameter, whether any additional spoken utterance by the given user and received within the threshold duration includes the value. In some versions of those embodiments, providing the output is contingent upon determining that the additional spoken utterance by the given user and received within the threshold duration does not include the value.

[0149] Furthermore, some embodiments include one or more processors (e.g., central processing unit(s) (CPUs), graphics processing unit(s) (GPUs), and / or tensor processing unit(s) (TPUs)) of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to result in performance of any of the above-described methods. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions executable by the one or more processors to perform any of the above-described methods. Some embodiments also include a computer program product comprising instructions executable by the one or more processors to perform any of the above-described methods.

[0150] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein.

Claims

1. A method implemented by one or more processors, the method comprising: receiving user input from a given user via a client device associated with the given user to initiate an auxiliary call; Based on the user input, determine: an entity participating on behalf of the given user during the assisted call, and tasks performed on behalf of said given user during said assisted call; determining, for one or more candidate parameters stored in association with the task and / or the entity and based at least in part on the conversation with the given user, one or more corresponding values ​​to be used when automatically generating synthesized speech during the assisted call when performing the task; initiating execution of the assisted call using a client device associated with a given user; during execution of the assistance call and based on processing audio data of the assistance call capturing an utterance of an additional user associated with the entity, determining that a particular value associated with an additional parameter and not previously determined during the conversation with the given user is requested by the additional user; as well as In response to determining to request the particular value associated with the additional parameter: causing the client device to render a notification outside of the secondary call identifying the additional parameter and requesting further user input for the particular value; continuing the auxiliary call before receiving any further input responsive to the notification, wherein continuing the auxiliary call comprises rendering, as part of the auxiliary call and for presentation to the additional user, one or more instances of synthesized speech based on one or more of the corresponding values ​​of the candidate parameters that were previously determined during the conversation with the given user but not yet provided to the additional user during the auxiliary call; determining, while continuing the secondary call, whether additional user input responsive to the notification and identifying the particular value of the additional parameter is received within a threshold duration; as well as In response to determining that the additional user input is received within the threshold duration: Additional synthesized speech based on the particular value is rendered as part of the assistance call.

2. The method according to claim 1, wherein Determining a given value in the corresponding values ​​based on the conversation with the given user includes: Before initiating the auxiliary call: generating a prompt identifying a given one of the candidate parameters and requesting additional information associated with the given candidate parameter; causing the client device to render the prompt; and The given value for the given candidate parameter is identified based on additional user input in response to the prompt.

3. The method according to claim 2, wherein: Determining additional ones of the corresponding values ​​includes: Before initiating the auxiliary call: The further value is identified based on a user profile associated with the given user.

4. The method according to claim 1, wherein Continuing the auxiliary call includes: processing additional audio data of the auxiliary call to determine that additional utterances of the additional user include a request for a given one of the candidate parameters; and In response to determining that the additional utterance includes the request for the given candidate parameter: causing the client device to render a given instance of the one or more instances of synthesized speech during the assisted call, wherein the given instance includes a given value among the corresponding values ​​based on determining a given value for the given candidate parameter, and Wherein, the given instance is rendered without requesting any additional user input from the given user.

5. The method according to claim 1, wherein Continuing the auxiliary call includes: processing the additional audio data to determine whether additional utterances of the additional user are received within additional threshold durations; and In response to determining that no further utterance is received from the additional user within the additional threshold duration: Another instance of the one or more instances of the synthesized speech is rendered during the auxiliary call, the another instance being based on one or more of the corresponding values ​​that has not been requested by the additional user.

6. The method according to claim 1, further comprising: In response to determining that the particular value associated with the additional parameter is requested by the additional user: The one or more candidate parameters stored in association with the entity are updated to include the additional parameter.

7. The method according to claim 1, further comprising: determining a state of the client device when the additional user requests the particular value associated with the additional parameter, and Based on the state of the client device, the notification and / or one or more properties for rendering the notification are determined.

8. The method according to claim 7, wherein: The status of the client device indicates that the given user is actively monitoring the auxiliary call, and wherein determining the notification based on the status of the client device comprises: determining, based on the status of the client device indicating that the given user is actively monitoring the assisted call, the notification including a visual component visually rendered via a display of the client device along with one or more selectable graphical elements, and Wherein, in response to the notification, the further user input comprises a selection of a given selectable graphical element of the one or more selectable graphical elements.

9. The method according to claim 7, wherein: The status of the client device indicates that the given user is not actively monitoring the secondary call, and wherein determining the notification based on the status of the client device comprises: Based on the status of the client device indicating that the given user is not actively monitoring the auxiliary call, determining the notification includes an audible component that is audibly rendered via one or more speakers of the client device.

10. The method according to claim 1, further comprising: terminating the auxiliary call; as well as After terminating the auxiliary call: The client device is caused to render an additional notification including an indication of an outcome of the assisted call.

11. The method according to claim 10, wherein: the additional notification comprising the indication of the secondary call: including an indication of additional tasks to be performed on behalf of the user in response to terminating the auxiliary call, or One or more selectable graphical elements are included that, when selected, cause the client device to perform additional tasks on behalf of the user.

12. The method according to claim 1, further comprising: In response to determining that the additional user input has not been received within the threshold duration: terminating the auxiliary call; as well as After terminating the auxiliary call: The client device is caused to render an additional notification including an indication of an outcome of the assisted call.

13. The method according to claim 1, wherein The threshold duration is: a fixed duration from the time a notification is rendered identifying the additional parameter and requesting further user input for the specific value, or A dynamic duration based on when a last corresponding value of one or more of the corresponding values ​​was rendered for presentation to the additional user via an additional client device.

14. The method according to claim 1, further comprising: After initiating the secondary call: Consent is obtained from the additional user associated with the entity to monitor the secondary call.

15. At least one computing device comprising: at least one processor; as well as At least one memory storing instructions which, when executed, cause the at least one processor to perform the method according to any one of claims 1 to 14.

16. A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor to perform the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Operation method of dialog agent and apparatus thereof

    EP3618062A1