Sub-delegated calls with automated assistants acting on behalf of human participants

The automated assistant addresses the challenge of incomplete information by using synthetic speech and notifications to efficiently complete tasks during calls, minimizing resource usage and call duration.

JP7819172B2Active Publication Date: 2026-02-24GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023196592
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2023-11-20
Publication Date
2026-02-24
Estimated Expiration
2040-04-22

AI Technical Summary

Technical Problem

Automated assistants may struggle to complete tasks when they lack sufficient information, leading to prolonged calls and increased resource consumption due to user intervention or call failures.

Method used

An automated assistant performs assisted calls using resolved parameters and synthetic speech, rendering instances based on user utterances, and provides notifications for unresolved parameters, allowing tasks to be completed efficiently without waiting for further user input.

Benefits of technology

Reduces call duration and resource utilization by enabling quick task completion through proactive parameter resolution and intelligent notification strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000046_0000
    Figure 00000046_0000
  • Figure 00000046_0001
    Figure 00000046_0001
  • Figure 00000047_0000
    Figure 00000047_0000
Patent Text Reader

Abstract

To use an automated assistant to initiate an assisted call on behalf of a given user.SOLUTION: An assistant can, during an assisted call, receive a request for information that is not known to the assistant from an additional user of the assisted call. In response, the assistant can render a prompt for the information and, while waiting for a responsive input from a given user, continue the assisted call by using already resolved values for the assisted call. If the responsive input is received within a threshold duration of time, a synthesized speech corresponding to the responsive input is rendered as part of the assisted call. Implementation is additionally or alternatively directed to using an automated assistant to provide, during an ongoing call between the given user and the additional user, output that is based on a value requested by the additional user during the ongoing call.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Automated assistants can be interacted with by users through a variety of computing devices, such as smartphones, tablet computers, wearable devices, automotive systems, standalone personal assistant devices, etc. The automated assistant receives input (e.g., spoken, touched, and / or typed) from the user and responds with responsive output (e.g., visual and / or audible).

[0002] A user can interact with an automated assistant to have the automated assistant perform actions on behalf of the user. As an example, the automated assistant can make a phone call on behalf of the user to perform a given action and can interact with an additional user to perform the action. For example, a user can provide user input requesting that the automated assistant make a restaurant reservation over the phone on behalf of the user. The automated assistant can initiate a phone call with a particular restaurant and provide reservation information to an additional user associated with the particular restaurant to make the reservation. The automated assistant can then notify the user whether the restaurant reservation on behalf of the user was successfully made.

[0003] However, for some actions performed by an automated assistant on behalf of a user, the automated assistant may not know enough information to fully perform the action. As an example, assume that an automated assistant makes a restaurant reservation over the phone on behalf of the user described above, and further assume that an additional user associated with the particular restaurant requests information not known to the automated assistant. Some automated assistants may determine that the requested information is not known and provide a notification to the user requesting the user to actively join the call to complete the restaurant reservation. However, waiting for the user to actively join the call may prolong the call and the associated use of computational and / or network resources used during the call. Additionally or alternatively, the user may not be able to join the call. This may lead to a call failure, requiring the automated assistant and / or the user to perform the action at a later time, thereby consuming more computational and / or network resources than if the action had been successfully performed by the automated assistant during the initial call. Summary of the Invention [Means for solving the problem]

[0004] Some implementations are directed to using an automated assistant to perform an assisted call with an entity to perform a task on behalf of a given user. The assisted call is between the automated assistant and an additional user associated with the entity. The automated assistant can perform the assisted call using resolved values ​​of parameters associated with the task and / or the entity. Performing the assisted call can include audibly rendering, by the automated assistant, an instance of synthetic speech in the assisted call that is audibly perceptible to the additional user. Rendering the instance of synthetic speech in the assisted call can include inserting synthetic speech into the assisted call that is audibly perceptible to the additional user (but not necessarily the given user). Each instance of synthetic speech can be generated based on one or more of the resolved values ​​and / or can be generated to respond to utterances of the additional user during the assisted call. Conducting the assisted call may also include performing, by the automated assistant, automatic speech recognition of audio data of the assisted call capturing further user utterances to generate recognized text of the utterances, and using the recognized text in determining an instance of synthetic speech responsive to the utterances for rendering in the assisted call.

[0005] Some of these implementations are further directed to determining that, during the assisted call, a further user utterance includes a request for information related to a further parameter, and that a value is not automatically determinable for the further parameter. In response, the automated assistant can cause an audio and / or visual notification (e.g., a prompt) to be rendered to the given user, the notification requesting further user input related to resolving the value of the further parameter. In some implementations, the notification is rendered outside the ongoing call (i.e., not inserted as part of the ongoing call) but can be perceptible to the given user to allow the given user to confirm the value and communicate the value within the ongoing call. The automated assistant can continue the assisted call until it receives some further user input in response to the notification. For example, the automated assistant may proactively provide an instance of synthetic speech during the call based on an already resolved value that has not yet been conveyed in a previous instance of synthetic speech, without waiting for further user input in response to the notification. If further user input is provided in response to the notification and the value of the further parameter is resolvable based on the further user input, the automated assistant can provide further synthetic speech conveying the resolved value of the further parameter after continuing the assisted call. By continuing the assisted call without waiting for further user input in response to the notification, the value needed to complete the task can be conveyed during the assisted call while waiting for further values ​​of further parameters, and the further values ​​(if received) can be provided later. In these and other ways, the assisted call can be completed more quickly, thereby reducing the overall duration that computer and / or network resources are utilized in executing the assisted call.

[0006] In some implementations that render a notification requesting further user input related to resolving values ​​for further parameters, the notification and / or one or more properties for rendering the notification may be dynamically determined based on the state of the assisted call and / or the state of the client device being used in the assisted call. For example, if the state of the client device indicates that a given user is actively monitoring the assisted call, the notification may be a visual-only notification and / or may include a lower-volume auditory component. On the other hand, if the state of the client device indicates that a given user is not actively monitoring the assisted call, the notification may include at least an auditory component, and / or the auditory component may be rendered at a higher volume. The state of the client device may be based, for example, on sensor data from sensors (e.g., gyroscope, accelerometer, presence sensor, and / or other sensors) of the client device and / or sensors of other client devices associated with the user. As yet another example, if the state of the assisted call indicates that multiple resolved values ​​have not yet been communicated at the time of providing the notification, the notification may be a visual-only notification and / or may include a lower-volume auditory component. On the other hand, if the state of the assisted call indicates that only one resolved value has not yet been communicated (or all resolved values ​​have already been communicated) at the time of providing the notification, the notification may include at least an auditory component, and / or the auditory component may be rendered at a higher volume. More broadly, when the state of the client device indicates that the user is not actively monitoring the call, and / or when the state of the conversation indicates that the duration for meaningfully continuing the conversation is relatively short, implementations may attempt to provide more intrusive notifications.On the other hand, an implementation may attempt to provide a less intrusive notification when the state of the client device indicates that the user is actively monitoring the call and / or when the state of the conversation indicates that the conversation will continue meaningfully for a relatively long duration. While more intrusive notifications may require more resources to render, an implementation may still selectively render more intrusive notifications, attempting to balance the increased resources for rendering a more intrusive notification with the increased resources that would be required to unnecessarily prolong the assisted call and / or terminate the assisted call without completing the task.

[0007] Some implementations are additionally or alternatively directed to using an automated assistant during an ongoing call between a given user and an additional user to provide output based on values ​​requested by the additional user during the ongoing call. The output can be provided proactively, preventing the given user from launching and / or navigating within an application to find the value on their own. For example, assume that a given user is engaged in an ongoing call with a utility company representative, and further assume that the utility company representative requests the given user's address information and an account number associated with the utility company. In this example, the given user may provide the address information but may not know the account number associated with the utility company without searching emails or messages received from the utility company, searching websites associated with the utility company, and / or performing other computing device interactions to find the account number associated with the utility company. However, by using the techniques described herein, the automated assistant can readily identify an account number associated with a utility company independent of any user input from a given user requesting that the automated assistant identify the account number, and can visually and / or audibly provide the account number to the given user and / or additional users during an ongoing call in response to determining that a utility company representative has requested the account number.

[0008] In these and other ways, client device resources may be conserved by preventing the launch of and / or interaction with such applications. Furthermore, values ​​indicated by proactively provided output may be communicated more quickly in a given call than if the user had to find the value independently, thereby reducing the overall duration of the ongoing call. In various implementations, the automated assistant may process a stream of audio data capturing at least one spoken utterance during the ongoing call to generate recognized text, the at least one spoken utterance being of the given user or an additional user. Furthermore, the automated assistant may identify, based on processing the recognized text, that the at least one spoken utterance requests parameter information and, using the given user's personal, access-limited data for the parameter, determine that a value for the parameter is resolvable. In response to determining that the value is resolvable, an output may be rendered. In some implementations, the output may be rendered outside the ongoing call (i.e., not inserted as part of the ongoing call) but may be perceptible to the given user to allow the given user to see the value and communicate the value within the ongoing call. In some additional or alternative implementations, the output may be rendered as synthetic speech as part of an ongoing call. For example, the synthetic speech may be rendered automatically within the ongoing call or upon receiving affirmative user interface input from a given user. In some implementations that provide output during an ongoing call between a given user and an additional user based on a value requested by the additional user during the ongoing call, the output is provided only in response to a determination that speech input including the value is not provided by the given user within a threshold amount of time during the ongoing call. In these and other ways, instances of rendering output unnecessarily are reduced.

[0009] The above description is provided as a summary of only some of the implementations disclosed herein. These and other implementations are described in more detail herein. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 illustrates various aspects of the present disclosure and is a block diagram of an example environment in which implementations disclosed herein may be implemented. [Figure 2] 1 is a flow diagram illustrating an example method for conducting an assisted call, according to various implementations. [Figure 3] 1 is a flow diagram illustrating an example method for providing an assistive output during an ongoing non-assisted call, according to various implementations. [Figure 4A] 1A-1C illustrate various non-limiting examples of user interfaces associated with making assisted phone calls, according to various implementations. [Figure 4B] 1A-1C illustrate various non-limiting examples of user interfaces associated with making assisted phone calls, according to various implementations. [Figure 4C] 1A-1C illustrate various non-limiting examples of user interfaces associated with making assisted phone calls, according to various implementations. [Figure 4D] 1A-1C illustrate various non-limiting examples of user interfaces associated with making assisted phone calls, according to various implementations. [Figure 5A] 10A-10C illustrate various non-limiting examples of user interfaces associated with providing auxiliary output during an ongoing unassisted call, according to various implementations. [Figure 5B] 10A-10C illustrate various non-limiting examples of user interfaces associated with providing auxiliary output during an ongoing unassisted call, according to various implementations. [Figure 5C] 10A-10C illustrate various non-limiting examples of user interfaces associated with providing auxiliary output during an ongoing unassisted call, according to various implementations. [Figure 6] FIG. 1 illustrates an exemplary architecture of a computing device according to various implementations. DETAILED DESCRIPTION OF THE INVENTION

[0011] 1 shows a block diagram of an exemplary environment illustrating various aspects of the present disclosure. A client device 110 is shown in FIG. 1 and, in various implementations, includes a user input engine 111, a device state engine 112, a rendering engine 113, a scheduling engine 114, a speech recognition engine 120A1, a natural language understanding (“NLU”) engine 130A1, and a speech synthesis engine 140A1.

[0012] The user input engine 111 can detect various types of user input at the client device 110. The user input detected at the client device 110 may include spoken input detected by a microphone of the client device 110 and / or further spoken input sent to the client device 110 from further client devices of further users (e.g., during the assisted call and / or during other ongoing calls when the assisted call has not yet called), touch input detected by a user interface input device (e.g., a touchscreen) of the client device 110, and / or typed input detected by a user interface input device (e.g., by a virtual keyboard on the touchscreen) of the client device 110. The further user associated with the entity in an ongoing call (assisted or unassisted) with the entity can be, for example, a human, a further human participant associated with the further client device, a further automated assistant associated with a further client device of the further user, and / or other further users.

[0013] The assisted call and / or ongoing call described herein may be performed using various voice communication protocols (e.g., Voice over Internet Protocol (VoIP), Public Switched Telephone Network (PSTN), and / or other telephony communication protocols). As described herein, synthetic audio may be rendered as part of the assisted call and / or ongoing call, which may include inserting synthetic audio into the call such that the synthetic audio is perceptible by at least one of the participants in the ongoing call and forms part of the audio data of the ongoing call. The synthetic audio may be generated and / or inserted by a client device that is one of the endpoints of the call and / or by a server that communicates with the client device and is also connected to the call. Also as described herein, auditory output may be rendered outside the assisted call, which does not include inserting auditory output into the call, although the auditory output may be detected by a microphone of a client device connected to the call and thereby be perceptible in the call. In some implementations, calls can be optionally muted and / or filtering can be used to reduce the perception within a call of auditory output rendered outside the call.

[0014] In various implementations, automated assistant 115 (generally indicated by a dashed line in FIG. 1 ) can use assisted call system 180 to perform assisted calls at client device 110 over network 190 (e.g., Wi-Fi, Bluetooth, near-field communication, a local area network, a wide area network, and / or other networks). Assisted call system 180, in various implementations, includes speech recognition engine 120A2, NLU engine 130A2, speech synthesis engine 140A2, and assisted call engine 150. Automated assistant 115 can utilize assisted call system 180 to perform tasks on behalf of a given user of client device 110 during a call with an additional user.

[0015] Additionally, in some implementations, before performing any task on behalf of a given user of client device 110, automated assistant 115 can obtain consent from an additional user to interact with automated assistant 115. For example, automated assistant 115 can obtain consent before performing a task when initiating an assisted call. As another example, automated assistant 115 can obtain consent for a given user of client device 110 when the given user initiates an ongoing call, even if the ongoing call is not initiated by automated assistant 115. If automated assistant 115 obtains consent from the associated additional user, automated assistant 115 can perform the task using assisted call system 180. However, if the automated assistant 115 does not obtain consent from the further user, the automated assistant 115 may cause the client device 110 to render (e.g., using the rendering engine 113) a notification to the given user of the client device 110 indicating that the given user is required to perform a task and / or end the call, and may cause the client device 110 to render (e.g., using the rendering engine 113) a notification to the given user of the client device 110 indicating that the task was not performed.

[0016] As described in more detail below, the automated assistant 115 can perform an assisted call using the assisted call system 180 in response to detecting user input from a given user of the client device 110 to initiate a call using the assisted call and / or during an ongoing call (i.e., when the assisted call has not yet been invoked). In some implementations, the automated assistant 115 can determine values ​​of candidate parameters to be used in performing a task on behalf of the given user of the client device 110 during the assisted call and / or during the ongoing call. In some versions of those implementations, the automated assistant 115 can interact with the given user of the client device 110 to solicit values ​​of the candidate parameters before and / or during the assisted call. In some additional and / or alternative versions of those implementations, the automated assistant 115 can determine values ​​of the candidate parameters based on a user profile associated with the given user of the client device 110 without soliciting values ​​of the candidate parameters before and / or during the assisted call. In some versions of their implementation, the automated assistant 115 can conduct an assisted call using the assisted call system 180 automatically based on the ongoing call dialogue, without detecting any user input from a given user of the client device 110 by the user input engine 111.

[0017] As shown in FIG. 1 , assisted call system 180 may be implemented remotely (e.g., by a server and / or other remote client device). While assisted call system 180 is shown in FIG. 1 as being implemented remotely via network 190, it should be understood that this is for purposes of illustration and not intended to be limiting. For example, in various implementations, assisted call system 180 may be implemented locally on client device 110. Furthermore, while automated assistant 115 is shown in FIG. 1 as being implemented both locally on client device 110 and on remote assisted call system 180, it should be understood that this is also for purposes of illustration and not intended to be limiting. For example, in various implementations, automated assistant 115 can be implemented locally on client device 110, or can be implemented locally on client device 110 and interact with a separate cloud-based automated assistant.

[0018] In implementation, when user input engine 111 detects a spoken input of a given user via the microphone of client device 110 and / or receives audio data capturing further spoken input from the further user transmitted to client device 110 from a further client device (e.g., during an assisted call and / or an ongoing call), speech recognition engine 120A1 of client device 110 can use speech recognition model 120A to process the audio data capturing the spoken input and / or capturing the further spoken input to generate recognized text corresponding to the spoken input and / or the further spoken input. Additionally, NLU engine 130A1 of client device 110 can use NLU model 130A to process the recognized text generated by speech recognition engine 120A1 to determine the intent contained in the spoken input and / or the further spoken input. For example, if client device 110 detects a spoken input of "Call Example Cafe and make a reservation for tonight" from a given user, client device 110 can use speech recognition model 120A to process audio data capturing the spoken input to generate recognized text corresponding to the spoken input of "Call Example Cafe and make a reservation for tonight," and can use NLU model 130A to process the recognized text to determine at least a first intent to initiate a phone call and a second intent to make a restaurant reservation.As another example, if client device 110 detects the further spoken input of "Do you have children?", client device 110 can use speech recognition model 120A to process audio data capturing the further spoken input to generate recognized text corresponding to the further spoken input of "Do you have children?" and can use NLU model 130A to process the recognized text to determine the intent of the request for information related to the additional parameters, as described herein. In some versions of these implementations, client device 110 can send the audio data, the recognized text, and / or the intent to assistant call system 180.

[0019] In other implementations, when user input engine 111 detects audio data capturing spoken input of a given user by the microphone of client device 110 and / or further spoken input from a further user transmitted from a further client device to client device 110 (e.g., during the assisted call and / or during an ongoing call), automated assistant 115 can cause client device 110 to transmit the audio data capturing the spoken input and / or the audio data capturing the further spoken input to assisted call system 180. Speech recognition engine 120A2 and / or NLU engine 130A2 of assisted call system 180 can process the audio data capturing the spoken input and / or the audio data capturing the further spoken utterances in a manner similar to that described above in connection with speech recognition engine 120A1 and / or NLU engine 130A1 of client device 110. In some additional and / or alternative embodiments, speech recognition engine 120A1 and / or NLU engine 130A1 of client device 110 may be used in a distributed manner in conjunction with speech recognition engine 120A2 and / or NLU engine 130A2 of assisted call system 180. Furthermore, speech recognition model 120A and / or NLU model 130A may be stored locally on client device 110 and / or on a remote server that communicates with client device 110 and / or assisted call system 180 via network 190.

[0020] In various implementations, speech recognition model 120A is an end-to-end speech recognition model such that speech recognition engines 120A1 and / or 120A2 can use the model directly to generate recognized text corresponding to spoken input. For example, speech recognition model 120A can be an end-to-end model used to generate recognized text character-by-character (or other token-by-token). One non-limiting example of such an end-to-end model used to generate recognized text character-by-character is a recurrent neural network transducer (RNN-T) model. RNN-T models are a form of sequence-to-sequence model that does not use an attention mechanism. Also, for example, when the speech recognition model is not an end-to-end speech recognition model, speech recognition engines 120A1 and / or 120A2 may instead generate predicted phonemes (and / or other representations). For example, using such a model, the predicted phonemes (and / or other representations) are then utilized by speech recognition engines 120A1 and / or 120A2 to determine recognized text that matches the predicted phonemes. In doing so, speech recognition engines 120A1 and / or 120A2 may optionally use a decoding graph, a dictionary, and / or other resources.

[0021] In implementations, when user input engine 111 detects touch input and / or typed input via a user interface input device of client device 110, automated assistant 115 can cause an indication of the touch input and / or an indication of the typed input to be sent from client device 110 to assisted call system 180. In some versions of these implementations, the indication of the touch input and / or the indication of the typed input can include the underlying text of the touch input and / or the text of the typed input, and the underlying text and / or the text can be processed using NLU model 130A to determine the intent of the underlying text and / or the text.

[0022] As described herein, assisted call engine 150 of assisted call system 180 can further process the recognized text generated by speech recognition engines 120A1 and / or 120A2, the text underlying touch input detected at client device 110, the text underlying typed input detected at client device 110, and / or the intent determined by NLU engines 130A1 and / or 130A2. Assisted call engine 150, in various implementations, includes entity identification engine 151, task determination engine 152, parameter engine 153, task execution engine 154, feedback engine 155, and recommendation engine 156.

[0023] The entity identification engine 151 can identify entities that interact on behalf of a given user of the client device 110. The entities may be, for example, person entities, business entities, location entities, and / or other entities. In some implementations, the entity identification engine 151 can also determine a particular entity type for the identified entity. For example, a person entity type may be a friend entity, a family entity, a coworker entity, and / or other particular type of person entity. Further, a business entity type may be a restaurant entity, an airline entity, a hotel entity, a salon entity, a clinic entity, and / or other particular type of business entity. Further, a location entity type may be a school entity, a museum entity, a library entity, a park entity, and / or other particular type of location entity. In some implementations, the entity identification engine 151 can also determine a particular entity for the identified entity. For example, a particular entity for a person entity may be a person's name (e.g., Jane Doe, etc.), a particular entity for a business entity may be a business name (e.g., Hypothetical Cafe, Example Cafe, Example Airlines, etc.), and a particular entity for a place entity may be a place name (e.g., Hypothetical University, Example National Park, etc.) Although the entities described herein may be defined at various levels of granularity, for simplicity, they are collectively referred to herein as "entities."

[0024] In some implementations, the entity identification engine 151 can identify an entity to engage on behalf of a given user of the client device 110 based on the user's interactions with the client device 110 prior to initiating an assisted call using the automated assistant 115. In some versions of those implementations, the entity can be identified in response to receiving user input to initiate the assisted call. For example, if a given user of the client device 110 sends input (e.g., spoken or touch) to a call interface element of a software application (e.g., regarding a contact in a contacts application, regarding a search result in a browser application, and / or regarding other callable entities included in other software applications), the entity identification engine 151 can identify an entity associated with the call interface element. For example, if the user input is targeted to a call interface element associated with “Example Cafe” in the browser application, the entity identification engine 151 can identify “Example Cafe” (or, more broadly, a business entity or a restaurant entity) as the entity to engage on behalf of the given user of the client device 110 during the assisted call.

[0025] In some implementations, the entity identification engine 151 can identify entities to engage on behalf of a given user of the client device 110 based on metadata associated with the ongoing call. The metadata may include, for example, a phone number associated with the additional user, a location associated with the additional user, an identifier identifying the additional user and / or an entity associated with the additional user, a time the ongoing call began, a duration of the ongoing call, and / or other metadata associated with the call. For example, if a given user of the client device 110 is engaged in an ongoing call with an additional user, the entity identification engine 151 can analyze the metadata of the ongoing call between the given user of the client device 110 and the additional user to identify phone numbers associated with the additional user engaged during the ongoing call. Additionally, the entity identification engine 151 can cross-reference the phone number with a database to identify entities associated with the additional user, submit search queries for the phone number to identify corresponding search results related to the phone number to identify entities associated with the additional user, and / or perform other actions to identify entities associated with the additional user. For example, if a given user of client device 110 is engaged in an ongoing call, entity identification engine 151 can analyze metadata associated with the ongoing call to identify the identity of "Example Airlines."

[0026] Further, the entity identification engine 151 may store any identified entities in the entity database 151A. In some implementations, the identified entities stored in the entity database 151A may be indexed by entity and / or particular type of entity. For example, if the entity identification engine 151 identifies an “Example Cafe” entity, “Example Cafe” may be indexed in the entity database 151A as a business entity and, optionally, may be further indexed as a restaurant entity. Further, if the entity identification engine 151 identifies an “Example Airlines” entity, “Example Airlines” may also be indexed in the entity database 151A as a business entity and, optionally, may be further indexed as an airline entity. By storing and indexing the identified entities in the entity database 151A, the entity identification engine 151 can easily identify and retrieve the entities, thereby reducing subsequent processing to identify the entities when encountering those entities in future assisted calls and / or ongoing calls. Further, in various implementations, each entity may be associated with a task in the entity database 151A.

[0027] The task determination engine 152 can determine a task to be performed on behalf of a given user of the client device 110. In some implementations, the task determination engine 152 can determine a task before initiating an assisted call using the automated assistant 115. In some versions of those implementations, the task determination engine 152 can determine a task to be performed on behalf of a given user of the client device 110 based on the user input for initiating the assisted call. For example, if a given user of the client device 110 provides a spoken input of "Call Example Cafe and make a reservation for tonight," the task determination engine 152 can determine a task to make a restaurant reservation based on the spoken input, utilizing the intent to initiate a call and make a restaurant reservation (e.g., determined using the NLU model 130A). As another example, if a given user of the client device 110 provides a touch input that selects a call interface element related to "Example Cafe" and the call interface indicates that the given user wants to modify a restaurant reservation at Example Cafe, the task determination engine 152 can determine a task to modify an existing restaurant reservation based on the touch input.

[0028] In some additional and / or alternative versions of these implementations, the task determination engine 152 can determine a task based on the identified entities engaged during the assisted call. For example, a restaurant entity can be associated with tasks such as making a restaurant reservation, modifying a restaurant reservation, canceling a restaurant reservation, and / or other tasks. As another example, a school entity can be associated with tasks such as inquiring about school closures, reporting that students / staff will not be attending school that day, and / or other tasks.

[0029] In other implementations, task determination engine 152 may determine a task to be performed on behalf of a given user of client device 110 during an ongoing phone call. In some versions of those implementations, a stream of audio data corresponding to an interaction between a given user of client device 110 and an additional user of an additional client device may be processed as described herein (e.g., in conjunction with speech recognition model 120A and NLU model 130A). For example, if task determination engine 152 identifies the recognized text "What's your frequent flier number?" during an ongoing phone call between a given user of client device 110 and an additional user associated with the additional client device, task determination engine 152 may determine a task to provide the additional user with the frequent flier number. In some further versions of those implementations, task determination engine 152 may also determine a task based on entities stored in association with the task in entity database 151A. For example, if the entity identification engine 151 identifies an additional user associated with the Example Airlines entity (e.g., based on metadata associated with the ongoing call, as described above), the task determination engine 152 may determine a task to provide the additional user with a frequent flyer number associated with Example Airlines based on a task for providing a frequent flyer number stored in association with the airline entity.

[0030] The parameter engine 153 can identify parameters associated with the task determined by the task determination engine 152. The automated assistant 115 can use the values ​​of the parameters to perform the task. In some implementations, candidate parameters can be stored in association with the task in a parameter database 153A. In some versions of these implementations, candidate parameters for a given task can be retrieved from the parameter database in response to an identification of an entity to be engaged on behalf of a given user of the client device 110 during the assisted call. For example, for the task of making a restaurant reservation, the parameter engine 153 can identify and retrieve one or more candidate parameters including a name parameter, a date / time parameter, a number of people in the reservation parameter, a phone number parameter, various seating type parameters (e.g., booth or table seating, indoor or outdoor seating, etc.), a children parameter (i.e., whether children are attending the reservation), a special occasion parameter (e.g., birthday, anniversary, etc.), and / or other candidate parameters. In contrast, for the task of modifying a restaurant reservation, the parameter engine 153 may identify and retrieve candidate parameters including a name parameter, a date / time parameter of the original reservation, a date / time parameter of the modified reservation, a number of people parameter of the modified reservation, and / or other candidate parameters.

[0031] In some additional and / or alternative implementations, parameter engine 153 may identify parameters for a given task during an ongoing call between a given user of client device 110 and an additional user of an additional client device. In some versions of these implementations, the parameters identified for a given task during the ongoing call may or may not be candidate parameters stored in parameter database 153A in association with the given task. As described above, a stream of audio data corresponding to an interaction between the given user and the additional user may be processed to determine a task to be performed during the ongoing call. Parameter engine 153 may determine whether the additional user requests information related to a given parameter that is unknown to automated assistant 115. For example, if, during the ongoing call, an Example Airlines representative requests a frequent flyer number from a given user of a client device, parameter engine 153 may identify the “Example Airlines frequent flyer number” parameter based on the intent included in the recognized text of the conversation and / or based on the identified “Example Airlines” entity.

[0032] Additionally, in some implementations, the candidate parameters stored in parameter database 153A may be mapped to various entities stored in entity database 151A. By mapping the candidate parameters stored in parameter database 153A to various entities stored in entity database 151A, assisted calling engine 150 can readily identify parameters for a task in response to an identification of a given entity. For example, in response to an identification of an “Example Cafe” entity (or, more broadly, a restaurant entity), assisted calling engine 150 can identify predefined tasks (e.g., stored in association with the “Example Cafe” entity and / or the restaurant entity in entity database 151A) to determine a set of candidate parameters, and in response to an identification of a particular task (e.g., based on the identified entity and / or user input), assisted calling engine 150 can determine candidate parameters related to the particular task for the identified entity. In other implementations, entity database 151A and parameter database 153A may be combined into a single database using various indexes (e.g., indexed by entity, indexed by task, and / or indexed in other ways) so that entities, tasks, and candidate parameters may each be stored in association with one another.

[0033] As described above, the automated assistant 115 can use the values ​​of the parameters to perform a task on behalf of a given user of the client device 110. The parameter engine 153 can also determine the values ​​of the parameters. In an implementation, when a given user of the client device 110 provides user input to initiate an assisted call, the parameter engine 153 can cause the automated assistant 115 to interact with the given user (e.g., visually and / or audibly via the client device 110) before initiating the assisted call to solicit further user input requesting information about the candidate parameters. As described herein, the automated assistant 115 can generate prompts requesting the information and can render the prompts to solicit the corresponding values ​​(or a subset thereof) audibly (e.g., via a speaker of the client device 110) and / or visually (e.g., via a display of the client device 110). For example, in response to receiving user input (e.g., by touch input or spoken input) to initiate an assisted call to make a reservation at Example Cafe, the automated assistant 115 may generate prompts requesting further user input regarding the candidate parameters (or a subset thereof), including values ​​for date / time parameters for the reservation, values ​​for number of people parameters for the reservation, etc.

[0034] In some versions of these implementations, parameter engine 153 can determine values ​​for candidate parameters based on a user profile of a given user of client device 110 stored in user profile database 153B. In some further versions of these implementations, parameter engine 153 can determine values ​​for candidate parameters without requiring any further user input from the given user. Parameter engine 153 can access user profile database 153B and retrieve values ​​for name parameters, phone number parameters, date / time parameters from software applications (e.g., calendar applications, email applications, contact applications, reminder applications, notes applications, SMS or text messaging applications, and / or other software applications). The user profile can include, for example, with permission from the given user, the given user's linked accounts, the given user's email accounts, the given user's photo albums, the given user's social media profiles, the given user's contacts, user preferences, and / or other information. For example, if a given user of client device 110 is engaged in a text messaging conversation with a friend and reviewing date / time information for a restaurant reservation at a given entity before providing user input to initiate an assisted call, parameter engine 153 can utilize the date / time information from the text messaging conversation as the value of the date / time parameter, and the automated assistant need not prompt the given user regarding the date / time information for the restaurant reservation. As discussed in more detail herein (e.g., in connection with FIG. 4B ), a given user of client device 110 can modify the values ​​of the candidate parameters before the assisted call is initiated.

[0035] Furthermore, in various implementations, some candidate parameters for a given task may be required parameters, while other candidate parameters for a given task may be optional parameters. In some versions of these implementations, whether a given parameter is required or optional may be task-based. In other words, required parameters for a given task may be the minimum amount of information that needs to be known to perform the given task. For example, for a restaurant reservation task, only the name parameter and the time / date parameter may be required parameters. However, if values ​​for optional parameters (e.g., the number of people to more fully accommodate the reservation, a phone number to call if any additional communication is needed, etc.) are known, the restaurant reservation task may provide greater benefit to both the given user and the restaurant associated with the restaurant reservation task. Thus, the automated assistant 115 may generate prompts for at least the required parameters and, optionally, for the optional parameters.

[0036] In implementations, when assisted call system 180 determines that an additional user engaged in an ongoing call with a given user of client device 110 requests information for a parameter, parameter engine 153 can determine the value of the parameter based on the user profile of the given user of client device 110 stored in user profile database 153B. In some versions of these implementations, parameter engine 153 can determine the value of the parameter in response to determining that an additional user requests information for the parameter. In other versions of these implementations, parameter engine 153 can determine the value of the parameter in response to user input from the given user of client device 110 requesting information for the parameter.

[0037] In various implementations, the task execution engine 154 can cause the automated assistant 115 to interact with the additional user associated with the identified entity using synthesized speech during an assisted call to perform the task. The task execution engine 154 can provide text and / or phonemes including at least values ​​to the speech synthesis engine 140A1 of the client device 110 and / or the speech synthesis engine 140A2 of the assisted call system 180 to generate synthesized speech audio data. The synthesized speech audio data can be transmitted to an additional client device of an additional user for audible rendering at the additional client device. The speech synthesis engine 140A1 and / or 140A2 can use the speech synthesis model 140A to generate synthesized speech audio data including synthesized speech corresponding to at least the parameter values. For example, the speech synthesis engine 140A1 and / or 140A2 can determine a sequence of phonemes determined to correspond to the parameter information requested by the additional user and process the sequence of phonemes using the speech synthesis model 140A to generate the synthesized speech audio data. The synthesized speech audio data can be in the form of an audio waveform, for example. In determining the sequence of phonemes corresponding to at least the values ​​of the parameters, speech synthesis engine 140A1 and / or 140A2 may access token-to-phoneme mappings stored locally on client device 110 or stored on a server (e.g., via network 190).

[0038] In some implementations, the task execution engine 154 can cause the client device 110 to initiate an assisted call with an entity interacting on behalf of a given user of the client device 110 and perform a task on behalf of the given user of the client device 110. Additionally, the task execution engine 154 can utilize synthesized voice audio data including at least the values ​​of the parameters to perform a task on behalf of the given user of the client device 110. For example, for a task of making a restaurant reservation, the automated assistant 115 can cause a further client device associated with the further user to identify itself as the automated assistant 115 on behalf of the given user of the client device 110 and render a synthesized voice stating the task to be performed on behalf of the given user during the assisted call (e.g., "This is Jane Doe's automated assistant calling to make a reservation on behalf of Jane Doe").

[0039] In some versions of these implementations, the automated assistant 115 can render corresponding values ​​for various candidate parameters (e.g., determined using the parameter engine 153 as described above) without being explicitly requested by a further user associated with the entity. Continuing with the example above, the automated assistant 115 can provide, at the beginning of a dialogue, a value for a date / time parameter for a reservation (e.g., "tonight at 7 o'clock"), a value for a number of people parameter for a reservation (e.g., "2," "3," "4," etc.), a value for a seating type parameter for a reservation (e.g., "booth seating," "indoor seating," etc.), and / or other values ​​for other candidate parameters. In other versions of these implementations, the automated assistant 115 can interact with a further user of a further computing device and provide specific values ​​for parameters explicitly requested by a further user associated with the entity. Continuing with the example above, the automated assistant 115 can process audio data capturing the further user's speech (e.g., "What time, how many people?"), and, in response to receiving a request from the further user, can determine the parameter information requested by the further user (e.g., "5 people at 7 o'clock tonight").

[0040] Additionally, in some versions of those implementations, the task determination engine 154 may determine that a request from a further user associated with the entity is a request for information related to a further parameter for which the automated assistant 115 does not know a corresponding further value. For example, for a restaurant reservation task, assume that the automated assistant 115 knows the values ​​of a date / time information parameter, a number of people parameter, and a seating type parameter, but the further user requests information for an unknown child parameter (i.e., whether children will be attending the reservation). In response to determining that the further user is requesting information for the unknown child parameter, the automated assistant 115 may cause the client device 110 to render (e.g., using the rendering engine 113) a notification indicating that a further value for the child parameter has been requested by the further user and prompting the given user of the client device 110 to provide a further value for the child parameter.

[0041] In some further versions of those implementations, the type of notification rendered at client device 110 and / or one or more properties (e.g., volume, brightness, size) for rendering the notification may be based on the state of client device 110 and / or the state of an ongoing call (e.g., determined using device state engine 112). The state of an ongoing call may indicate, for example, which values ​​have been communicated and / or have not yet been communicated in the ongoing call, and / or which components of tasks of the ongoing call have been completed and / or have not yet been completed in the ongoing call. The state of client device 110 may be based, for example, on software applications running in the foreground of client device 110, software applications running in the background of client device 110, whether client device 110 is in a locked state, whether client device 110 is in a sleep state, whether client device 110 is in an off state, sensor data from sensors of client device 110, and / or other data. For example, if the state of client device 110 indicates that a software application (e.g., an automated assistant application, a calling application, an assisted calling application, and / or other software application) that displays a transcription of an assisted call is running in the foreground of client device 110, the type of notification may be a banner notification, a pop-up notification, and / or other type of visual notification. As another example, if the state of client device 110 indicates that client device 110 is asleep or locked, the type of notification may be an audible indication by a speaker and / or a vibration by a speaker or other hardware component of client device 110.As yet another example, if sensor data from a presence sensor, accelerometer, and / or other sensor of a client device indicates that a given user is not currently near and / or holding the client device, a more intrusive notification (e.g., visual and audible at a first volume level) may be provided. On the other hand, if such sensor data indicates that a given user is currently near and / or holding the client device, a less intrusive notification (e.g., visual only, or visual and audible at a second volume level less than the first volume level) may be provided. As yet another example, if the state of the interaction indicates that the interaction is nearing an end, a more intrusive notification may be provided, while a less intrusive notification may be provided if the state of the interaction indicates that the interaction is not nearing an end.

[0042] In some further versions of these implementations, the task decision engine 153 can cause the automated assistant 115 to continue the dialogue with the further user even if the automated assistant 115 does not know the corresponding further values ​​of the further parameters requested by the further user. The automated assistant 115 can then provide the further values ​​later in the dialogue after receiving further user input responsive to the request from the further user. For example, for a restaurant reservation task, assume that the automated assistant 115 knows the values ​​of the date / time information parameter, the number of people parameter, and the seating type parameter, and further assume that the further user requests values ​​for the date / time information parameter and the number of people parameter. Further assume that the further user then requests the value of a children parameter for which the automated assistant 115 does not know the value. In this example, task decision engine 153 can cause automated assistant 115 to render a synthesized speech at a further client device of a further user that includes an indication that the requested information regarding whether children will be attending is not currently known and that automated assistant 115 can continue the assisted call by providing other values ​​(e.g., type of seat) until further user input responsive to the notification, including the value of the children parameter, is detected at client device 110. In response to receiving further user input, automated assistant 115 can provide the further values ​​as standalone values ​​(e.g., "No children") or as subsequent values ​​(e.g., "Jane Doe wants a booth seat. No children").

[0043] Additionally, in implementations, if an additional user requests information for an unknown additional parameter, the automated assistant 115 may terminate the assisted call if no additional user input in response to the notification requesting the information is received within a threshold duration (e.g., 15 seconds, 30 seconds, 60 seconds, and / or other duration). In some versions of these implementations, the threshold duration may begin when the notification requesting the information is rendered at the given user's client device 110. In other versions of these implementations, the threshold duration may begin when the last value known to the automated assistant 115 is requested by the additional user or proactively provided by the automated assistant 115 (independent of the additional user requesting it).

[0044] Further, in implementations, if an additional user requests information on an unknown additional parameter, feedback engine 155 can store the additional parameter as a candidate parameter for the task in parameter database 153A. In some versions of these implementations, feedback engine 155 can map the additional parameter to an entity associated with the additional user stored in entity database 151A. In some further versions of these implementations, feedback engine 155 may map the additional parameter to an entity if an additional user associated with the entity requests the additional parameter from multiple users a threshold number of times while assisted speech is active. For example, if a restaurant entity asks whether a restaurant reservation includes children at least a threshold number of times (e.g., 100 times, 1000 times, and / or other number threshold) during interactions across multiple assisted calls initiated by multiple users via their respective client devices, feedback engine 155 may map a children parameter to various restaurant entities in entity database 151A. In this example, the child parameter may be considered a new candidate parameter for which a value is sought before initiating future assisted calls for the task of making a restaurant reservation with various restaurant entities and / or a particular entity that frequently requests information for the child parameter. In these and other ways, values ​​can be sought before initiating future assisted calls, thereby shortening the duration of future assisted calls and / or preventing the need to utilize computational resources to render prompts regarding values ​​in future assisted calls.

[0045] In some implementations, the task execution engine 154 can cause the automated assistant 115 to provide output appropriate for an ongoing call with an additional user associated with the entity to perform a task on behalf of a given user of the client device 110, even if the ongoing call was not initiated using assisted speech (i.e., an unassisted call). For example, the automated assistant 115 can interrupt the ongoing call to generate synthesized speech including values ​​that are rendered at the additional client device 110. In some versions of those implementations, the automated assistant 115 can provide output appropriate for the ongoing call in response to user input from a given user of the client device 110 (e.g., as described in connection with FIG. 5A ). In some further versions of those implementations, the automated assistant 115 may not process audio data corresponding to the ongoing call, thereby eliminating the need to obtain consent from the additional user. Rather, the automated assistant 115 can analyze metadata associated with the ongoing call and determine corresponding values ​​for parameters requested by the additional user based on the metadata and / or user input. For example, client device 110 can detect user input that activates an assisted call and prompts automated assistant 115 to retrieve a corresponding value requested by the additional user (e.g., the value of a frequent flyer number parameter), and can determine that the value is associated with a particular entity (e.g., Example Airlines) based on metadata associated with the ongoing call, even though the particular entity was not explicitly identified in the request. The user's restricted-access data can then be searched by the automated assistant using search parameters (e.g., terms) based on both the additional user's request (e.g., "frequent flyer") and the metadata (e.g., "Example Airlines").Moreover, in some further versions of those implementations, the user input for the automated assistant 115 to provide an output appropriate for an ongoing call is in response to a notification that the assisted call system 180 can provide values ​​to the additional user, generated by the automated assistant 115 (e.g., as described in connection with FIG. 5B ) and rendered at the client device 110 (e.g., using the rendering engine 113). For example, the automated assistant 115 may proactively notify a given user of the client device 110 that the assisted call system 180 can provide corresponding values ​​for parameters requested by the additional user.

[0046] In other versions of those implementations, the automated assistant 115 provides an output appropriate for the ongoing call (e.g., synthesized speech within the ongoing call) without receiving any user input from a given user of the client device 110 in response to determining that information is requested by an additional user (e.g., as described in connection with FIG. 5C ). In some versions of those implementations, the automated assistant 115 may automatically provide a value based on a reliability metric for the value satisfying a reliability threshold. The reliability metric may be based, for example, on whether a given user previously provided the determined value in response to receiving a previous request for the same value, whether the parameter identified during the ongoing call is a candidate parameter stored in association with the task identified during the ongoing call, the source from which the value is determined (e.g., email / calendar application vs. text messaging application), and / or the method for determining the reliability metric. For example, with respect to an ongoing call between a given user of client device 110 and an additional user associated with an airline entity, assisted call system 180 can determine a frequent flyer number requested by the additional user, and assisted call system 180 can cause automated assistant 115 to automatically provide synthesized speech including the frequent flyer number as part of the ongoing call—without the automated assistant 115 receiving any user input requesting the frequent flyer number. In some versions of these implementations, a given user of client device 110 may have to grant the assisted call permission to automatically interrupt the ongoing call in settings associated with the assisted call.

[0047] In various implementations, recommendation engine 156 can determine candidate values ​​to be communicated in the dialogue and can cause automated assistant 115 to provide the candidate values ​​as recommendations for a given user of client device 110. In some implementations, the candidate values ​​can be transmitted to client device 110 over network 190. Additionally, the candidate values ​​can be visually rendered on a display of the client device (e.g., using rendering engine 113) as a recommendation upon request from a further user. In some versions of those implementations, the recommendation can be selectable such that when user input (e.g., as determined by user input engine 111) targets a given recommendation, the given recommendation can be incorporated into synthesized speech that is audibly rendered at a further client device of a further user.

[0048] In some implementations, the recommendation engine 156 can determine candidate values ​​based on a request for additional information from the user. For example, if the request is a yes / no question (e.g., "Are children included in the reservation?"), the recommendation engine 156 can determine a first recommendation including a value of "yes" and a second recommendation including a value of "no." In other implementations, the recommendation engine 156 can determine candidate values ​​based on a user profile of a given user associated with the client device 110 stored in the user profile database 153B. For example, if the request asks for specific information (e.g., "How many children do you have?"), the recommendation engine 156, with permission from the given user, can determine, based on the user profile of the given user, that the given user of the client device 110 has three children—and can determine a first recommendation including a value of "3," a second recommendation including a value of "2," a third recommendation including a value of "1," and / or other recommendations including other values. The various recommendations described herein may be visually and / or audibly rendered at client device 110 for presentation to a given user associated with the client device.

[0049] As described herein, rendering engine 113 can render various notifications or other output at client device 110. Rendering engine 113 can audibly and / or visually render the various notifications described herein. Additionally, rendering engine 113 can cause a transcription of an interaction to be rendered on the user interface of client device 110. In some implementations, the transcription may correspond to an interaction between a given user of client device 110 and automated assistant 115 (e.g., as described in connection with FIG. 4B ). In other implementations, the transcription may correspond to an interaction between an additional user of an additional client device and automated assistant 115 (e.g., as described in connection with FIGS. 4C and 4D ). In still other implementations, the transcription may correspond to an interaction between a given user of client device 110, an additional user of an additional client device, and automated assistant 115 (e.g., as described in connection with FIGS. 5A-5C ).

[0050] In some implementations, the scheduling engine 114 may cause the automated assistant 115 to include, along with and / or in a notification indicating the result of the task execution, a recommendation for the automated assistant to perform a further (or subsequent) task based on the result of the task execution. In some versions of those implementations, the recommendations may be selectable such that a given recommendation may cause the further task to be performed when user input (e.g., as determined by the user input engine 111) targets the given recommendation. For example, for a successful restaurant reservation task, the automated assistant 115 may render, via a user interface on the display of the client device 110, a selectable element that, when selected by a given user of the client device 110, causes the scheduling engine 114 to create a calendar entry for the successful restaurant reservation. As another example, for a successful restaurant reservation task, the automated assistant 115 may send an SMS or text message to the other users participating in the restaurant reservation indicating that the restaurant reservation task was successfully performed. In contrast, with respect to a failed restaurant reservation task, the automated assistant 115 may render a selectable element that, when selected by a given user of the client device 110, causes the scheduling engine 114 to create a reminder and / or calendar item to perform the restaurant reservation task again later and before the time / date value of the attempted restaurant reservation task (e.g., automatically performed by the automated assistant 115 at a later time, or performed by the automated assistant 115 in response to a user selection of the reminder and / or calendar item).

[0051] In other implementations, the scheduling engine 114 may cause the automated assistant 115, in response to a determination of the result of the task execution, to automatically perform further (or subsequent) tasks based on the result of the task execution. For example, for a successful restaurant reservation task, the automated assistant 115 may automatically create a calendar entry for the successful restaurant reservation, automatically send an SMS or text message to other users participating in the restaurant reservation indicating that the restaurant reservation task was successfully performed, and / or automatically perform other further tasks that may be performed by the automated assistant 115 in response to the successful performance of the restaurant reservation task. In contrast, for a failed restaurant reservation task, the automated assistant 115 may automatically create a reminder and / or a calendar entry to perform the restaurant reservation task again later and before the time / date value of the restaurant reservation task (e.g., automatically performed by the automated assistant 115 later or performed by the automated assistant 115 in response to a user selection of the reminder and / or calendar entry).

[0052] By using the techniques described herein, various technical advantages may be realized. As one non-limiting example, the automated assistant 115 may end an assisted call more quickly because the assisted call interaction is not stalled waiting for additional parameter values ​​when additional information not currently known to the automated assistant 115 is requested by the user. By using the techniques disclosed herein, the length of the assisted call may be reduced, thereby saving both network and computational resources. As another non-limiting example, the automated assistant 115 may provide corresponding values ​​of parameters during the performance of a task by a given user during an ongoing call. By providing corresponding values ​​either automatically or in response to explicit user input as described above, the client device 110 receives less input from a given user of the client device 110 because the user does not need to navigate to various applications with entirely different user interfaces to determine the corresponding values, thereby saving computational resources on the given client device. Furthermore, because the user does not need to navigate to these various applications, the system saves both computational and network resources by ending an ongoing call more quickly.

[0053] FIG. 2 shows a flow diagram illustrating an example method 200 for performing an assisted call, according to various implementations. For convenience, the operations of method 200 are described with reference to a system that performs the operations. This system of method 200 includes one or more processors and / or other components of a computing device (e.g., client device 110 of FIG. 1 , client device 410 of FIGS. 4A-4D , client device 510 of FIGS. 5A-5C , computing device 610 of FIG. 6 , one or more servers, and / or other computing devices). Furthermore, while the operations of method 200 are shown in a particular order, this is not intended to be limiting. One or more operations could be reordered, omitted, or added.

[0054] At block 252, the system receives user input to initiate an assisted call from a given user via a client device associated with the given user. In some implementations, the user input is spoken input detected by a microphone on the client device. For example, the spoken input may include "Call Example Cafe," or specifically "Call Example Cafe using assisted calling." In other implementations, the user input is touch input detected on the client device. For example, the touch input may be detected while various software applications (e.g., a browser application, a messaging application, an email application, a notes application, a reminder application, and / or other software applications) are running on the client device.

[0055] In block 254, the system identifies entities to engage on behalf of the given user during the assisted call in response to the user input to initiate the assisted call. The entities may be identified based on the user input. In an implementation, when the user input is a spoken input, the entities may be identified based on processing audio data capturing the spoken input (e.g., using speech recognition model 120A and / or NLU model 130A of FIG. 1 ) to identify entities included in the spoken input (e.g., a business entity, a specific business entity, a location entity, and / or other entities). In an implementation, when the user input is a touch input, the entities may be identified based on the user's interaction with the client device (e.g., a contact entry related to the entity, a search result related to the entity, a touch input selecting an advertisement related to the entity, and / or other user interaction).

[0056] In block 256, the system determines at least one task to be performed on behalf of the given user during the assisted call based on the user input and / or the entities. In various implementations, predefined tasks may be stored in association with multiple corresponding entities in one or more databases (e.g., entity database 151A of FIG. 1). For example, tasks for booking a flight, changing a flight, canceling a flight, inquiring about lost baggage, and / or other tasks may be stored in association with multiple different airline entities. In some implementations, the at least one task to be performed may be determined based on the user input. For example, if spoken input of "Call Example Cafe and make a reservation for tonight at 7 o'clock" is received at the client device, the system may determine that the spoken input includes a restaurant reservation task. In other implementations, the at least one task to be performed may be determined based on the entities identified in block 254. For example, if the spoken input "Call Example Cafe" is received at a client device (i.e., without explicitly stating the task of making a restaurant reservation), the system can infer the task of making a reservation based on the fact that it is a predefined task associated with the restaurant entity.

[0057] In block 258, the system identifies candidate parameters associated with at least one task. In various implementations, the candidate parameters may be stored in one or more databases (e.g., parameter database 153A of FIG. 1) in association with at least one task (determined in block 256) and / or in association with at least one entity (determined in block 254). For example, tasks of booking a flight, changing a flight, canceling a flight, inquiring about lost baggage, and / or other tasks may be stored in association with corresponding candidate parameters. Also, for example, a task of changing a flight associated with Airline Entity 1 may be stored in association with a first corresponding parameter, and a task of changing a flight associated with Airline Entity 2 may be stored in association with a second corresponding parameter. As another example, Restaurant Entity 1 may be stored in association with a first corresponding parameter, and Restaurant Entity 2 may be stored in association with a second corresponding parameter.

[0058] At block 260, the system determines corresponding values ​​for the candidate parameters to be used in performing at least one task. In some implementations, the values ​​of the candidate parameters may be determined based on a user profile for the given user stored in one or more databases (e.g., user profile database 153B of FIG. 1). The user profile may include, for example, the given user's linked accounts, the given user's email account, the given user's photo albums, the given user's social media profiles, the given user's contacts, user preferences, and / or other information. For example, for a salon appointment task, the system may determine the given user's name and phone number parameters based on a contacts application and determine the salon's preferred stylists based on previous communications with particular stylists (e.g., email messages, text or SMS messages, phone calls, and / or other communications). In some additional and / or alternative implementations, the values ​​of the candidate parameters may additionally or alternatively be determined based on further user input in response to prompts visually and / or audibly rendered at the client device and requesting information for the parameters. In some versions of those implementations, the system may generate prompts only for corresponding values ​​of candidate parameters that the system was unable to determine based on the user profile. For example, for the salon appointment task described above, the system already knows the values ​​of the name parameter, the phone number parameter, and the phone number parameter, so the system need only generate a prompt requesting the value of the date / time parameter. As described herein (e.g., in connection with FIG. 4B), the system may provide an opportunity for a given user of a client device to modify the values ​​of the candidate parameters before initiating the assisted call.

[0059] In block 262, the system initiates an assisted call with the entity on behalf of the given user to perform at least one task using the candidate parameter values ​​using a client device associated with the given user. The system may process audio data received at the client device from the additional computing device to determine the parameter values ​​requested by the additional user. Furthermore, the system may generate (e.g., proactively or in response to a request by the additional user) synthesized speech audio to be sent to an additional client device of the additional user associated with the entity identified in block 254, the synthesized speech audio including at least the parameter values. In some implementations, the system may interact with the additional user and generate synthesized speech audio including specific values ​​included in the information requested by the additional user. For example, for a restaurant reservation task, the additional user may request date / time parameter information, and the system may generate synthesized speech audio data including the date / time value in response to determining that the request from the additional user is for date / time parameter information. Furthermore, the system may cause the synthesized speech audio data to be sent to an additional client device of the additional user, and cause the synthesized speech included in the synthesized speech audio data to be audibly rendered at the additional client device. Additional users may request additional information for various parameters, and the system may provide values ​​to the additional users to perform the task of making a restaurant reservation.

[0060] In some implementations, method 200 may include optional sub-block 262A. If included, in optional sub-block 262A, the system may obtain consent from a further user associated with the entity to monitor the assisted call. For example, the system may obtain consent when initiating the assisted call and before performing the task. If the system obtains consent from the associated further user, the system may perform the task. However, if the system does not obtain consent from the further user, the system may cause the client device to render a notification to the given user indicating that the given user is required to perform the task and / or end the call, and a notification to the given user indicating that the task was not performed.

[0061] In block 264, the system determines whether any information related to additional parameters is requested by an additional user associated with the entity during the assisted call. As described above, the system may process audio data received at the client device from the additional computing device to determine the value requested by the additional user. Additionally, the system may determine whether the requested value is for an additional parameter for which the system has not previously resolved, such that the value requested by the additional user is not currently known to the system. For example, if the system has not previously determined a value for a seating type parameter for a restaurant reservation task, the seating type parameter may be considered an additional parameter having a value currently unknown to the system. If, in a repetition of block 264, the system determines that information related to additional parameters is not requested by an additional user associated with the entity during the assisted call, the system may proceed to block 272, which will be discussed in more detail below.

[0062] If, in an iteration of block 264 that includes optional block 266, the system determines that a request from an additional user associated with the entity during the assisted call includes values ​​for additional parameters, the system may proceed to optional block 266. In implementations that include optional block 266, the system may proceed directly from block 264 to block 266, which is discussed in more detail below.

[0063] If included, in optional block 266, the system determines the state of a client device associated with the given user. The state of the client device may be based on, for example, software applications running in the foreground of the client device, software applications running in the background of the client device, whether the client device 110 is locked, whether the client device is asleep, whether the client device 110 is off, sensor data from sensors in the client device, and / or other data. In some implementations, the system additionally or alternatively determines the state of an ongoing call in block 266.

[0064] In block 268, the system causes the client device associated with the given user to render a notification identifying the further parameter. The notification may further request values ​​for the further parameter included in the information. In implementations that include optional block 268, the type of notification rendered by the client device and / or one or more properties for the rendering may be based on the state of the client device and / or the state of the ongoing call determined in optional block 268. For example, if the state of the client device indicates that a software application (e.g., an automated assistant application, a calling application, an assisted calling application, and / or other software application) that displays a transcription of the assisted call is running in the foreground of the client device, the type of notification may be a banner notification, a pop-up notification, and / or other type of visual notification. As another example, if the state of the client device indicates that the client device is asleep or locked, the type of notification may be an audible indication via a speaker and / or a vibration via the speaker or other hardware component of the client device.

[0065] At block 270, the system determines whether any additional user input is received at a client device associated with the given user within a threshold duration. The additional user input may be, for example, additional spoken input, additional typed input, and / or additional touch input in response to a notification requesting information. In some implementations, the threshold duration may begin when a notification requesting information is rendered at the given user's client device. In other implementations, the threshold duration may begin when the last value is requested by the additional user. If, at a repetition of block 270, the system determines that additional user input is received within the threshold duration, the system may proceed to block 272. The additional user input may be received in response to a notification indicating information being requested by the additional user and may include an indication of the value requested.

[0066] In block 272, the system completes at least one task based on the values ​​of the candidate parameters and / or the further parameters. In implementations, during the assisted call, if the system determines in block 264 that a request for information from the further user does not include further values, the system may complete at least one task using the corresponding values ​​of the candidate parameters determined in block 260. In these implementations, the system may complete the assisted call without having to involve the given user of the client device. In implementations, during the assisted call, if the system determines in block 264 that information related to further parameters is requested by the further user, the system may complete at least one task using the corresponding values ​​of the candidate parameters determined in block 260 and the values ​​of the further parameters received in block 270. In these implementations, the system may complete the assisted call with minimal involvement from the given user of the client device. From block 272, the system may proceed to block 276, which will be discussed in more detail below.

[0067] In particular, in block 264, the system may determine that an additional user requests information not currently known to the system, but the system may continue to perform at least one task without values ​​for the additional parameters. For example, the system may cause synthesized speech to be provided as part of an assisted call, thereby causing the synthesized speech to be rendered at an additional client device of the additional user. Furthermore, the synthesized speech may indicate that the system does not currently know the values ​​of the additional parameters, but that the system can prompt the user for values ​​related to the information, and can provide other values ​​of the candidate parameters determined in block 260 while the system prompts the given user of the client device for values ​​of the additional parameters. In this manner, the system can continue to interact with the additional user to perform the task. Furthermore, if additional user input is received from the additional user, including a value responsive to the request for information, the system can provide the value as a continuation of the provision of one of the known values ​​or as a standalone value when there is a break in the interaction. In this manner, because the interaction is not stopped to wait for values ​​for the additional parameters, the system can more quickly and efficiently successfully perform tasks on behalf of the given user of the client device. By performing tasks more quickly and efficiently, the length of conversations may be reduced using the techniques disclosed herein, thereby conserving both network and computational resources.

[0068] If, in an iteration of block 270, the system determines that no further user input is received within the threshold duration, the system may proceed to block 274. In block 274, the system terminates execution of at least one task. Additionally, the system may terminate an ongoing call with the entity. By terminating an ongoing call, as opposed to waiting for further user input beyond the threshold duration, even if the task cannot be fully performed, the interaction may still be more quickly concluded to achieve the technical advantages described above. From block 274, the system may proceed to block 276.

[0069] In block 276, the system renders a notification indicating the results of the execution of the at least one task by the client device. In an implementation, if the system completes execution of the task from block 272, the notification can include an indication that the task was completed on behalf of the given user of the client device and can include confirmation information related to the completion of the task (e.g., date / time information, monetary costs associated with the task, a confirmation number, information related to the entity, and / or other confirmation information). In an implementation, if the system finishes execution of the task from block 274, the notification can include an indication that the task was not completed and can include task information related to the completion of the task (e.g., values ​​for certain parameters required, the inability of the entity to accommodate the corresponding values ​​determined in block 260 and / or received in block 270, that the entity is closed, and / or other task information). In various implementations, the notification may include a selectable graphical element that, when selected, causes the system to create a calendar item based on the task results, create a reminder based on the task results, send a message (e.g., a text, SMS, email, and / or other message) including the task results, and / or perform other further tasks depending on the user's selection.

[0070] FIG. 3 shows a flow diagram illustrating an example method 300 of auxiliary output during an ongoing unassisted call, according to various implementations. For convenience, the operations of method 300 are described with reference to a system that performs the operations. This system of method 300 includes one or more processors and / or other components of a computing device (e.g., client device 110 of FIG. 1 , client device 410 of FIGS. 4A-4D , client device 510 of FIGS. 5A-5C , computing device 610 of FIG. 6 , one or more servers, and / or other computing devices). Furthermore, while the operations of method 300 are shown in a particular order, this is not intended to be limiting. One or more operations could be reordered, omitted, or added.

[0071] In block 352, the system detects, at the client device, an ongoing call between a given user associated with the client device and an additional user associated with an additional client device. Optionally, the system may also identify an entity associated with the additional user. The system may identify the entity based on metadata associated with the ongoing call. In some implementations, method 300 may include optional sub-block 352A. If included, in optional sub-block 352A, the system obtains consent from the additional user associated with the entity to monitor the ongoing call. The system may obtain consent from the additional user in the same manner as described in connection with optional sub-block 260A of FIG. 2 .

[0072] At block 354, the system processes the stream of audio data corresponding to the ongoing call to generate recognized text. The stream of audio data corresponding to the ongoing call may include at least additional spoken input of an additional user that is sent to the given user's client device. The stream of audio data corresponding to the ongoing call may also include additional spoken input of the given user. Further, the system may process the stream of audio data using a speech recognition model (e.g., speech recognition model 120A of FIG. 1) to generate recognized text. It should be understood that, assuming the additional user consents to call monitoring, the system may continue to process the stream of audio data corresponding to the ongoing call.

[0073] In block 356, the system identifies parameters for at least one task to be performed by the given user during the ongoing call based on the recognized text. The system may process the recognized text from block 354 using an NLU model (e.g., NLU model 130A of FIG. 1 ) to determine the intent contained in the stream of audio data. In some implementations, the system may determine that further user input from a further user includes a request for information of at least one parameter of the task. For example, if further spoken input from a further user is received saying, "Do you have a quality assurance case number for this?", the system may identify a quality assurance case number parameter and determine that the user input includes a request for a value for the quality assurance case number parameter. In this example, the task may be any task related to an airline entity and / or a specific task that provides a quality assurance case number regardless of other parameters associated with the airline entity.

[0074] At block 358, the system determines corresponding values ​​for the parameters to be used in performing the at least one task. The corresponding values ​​may be determined after identifying parameters for the at least one task at block 356. In some implementations, the corresponding values ​​may be automatically determined in response to identifying parameters for the at least one task based on a user profile associated with a given user of the client device. For example, in response to identifying a quality assurance case number parameter, the system may, with permission (e.g., prior permission) from the given user, access an email account associated with the given user and search for emails containing the corresponding value of the quality assurance case number parameter. Further, in implementations, if an entity engaging with the given user during an ongoing call is identified, the system may limit the search to only emails related to the identified entity. In other implementations, the corresponding values ​​may be determined in response to receiving user input including information requesting the corresponding values ​​of the parameters, as opposed to being automatically identified. The corresponding values ​​may be determined in response to receiving user input in the same or similar manner as described above.

[0075] In some implementations, method 300 may include optional blocks 360, 362, and / or 364. If included, in optional block 360, the system may determine whether any user input for activating an assisted call is received at a client device associated with a given user. The system may determine whether user input activates an assisted call based on spoken, typed, and / or touch input that invokes an assisted call during an ongoing call between a given user of a client device and an additional user of an additional client device in any manner described herein. If, in an iteration of optional block 360, the system determines that user input for activating an assisted call is received, the system may proceed to block 366, which will be discussed in more detail below. If, in an iteration of optional block 360, the system determines that user input for activating an assisted call is not received, the system may proceed to optional block 362.

[0076] If included, in optional block 362, the system may render a notification by a client device associated with the given user indicating that the assisted call can perform at least one task. The notification may include, for example, an indication that additional users have requested information on the corresponding values ​​determined in block 358 of the parameters identified in block 356, and may also include an indication that the system can provide the corresponding values ​​to additional users on behalf of the given user. The notification may be rendered visually and / or audibly. In implementations, if the notification is rendered audibly, the notification may be rendered audibly only at the client device so that additional users of the client device do not perceive the notification (i.e., outside the call). Optionally, to reduce the likelihood that additional users will perceive the notification, the ongoing call may be temporarily muted during the audible rendering of the notification, or acoustic echo cancellation or other filtering may be utilized to filter the notification and prevent it from being provided as part of the ongoing call. In other implementations, if the notification is rendered audibly, the notification may be rendered audibly on both the client device of the given user and the further client device of the further user, such that the notification interrupts an ongoing call between the given user and the further user.

[0077] If included, in optional block 364, the system may determine whether any user input for activating an assisted call is received at a client device associated with the given user. The user input received in block 364 may be in response to rendering a notification indicating that the assisted call can perform at least one task. The system may determine whether the user input activates the assisted call based on spoken, typed, and / or touch input that invokes the assisted call during an ongoing call between the given user of the client device and an additional user of an additional client device in any manner described herein. If, in an iteration of optional block 360, the system determines that no user input for activating the assisted call is received, the system may return to block 354 to process further audio data corresponding to the ongoing call. For example, the system may determine that the given user of the client device has provided spoken input including the corresponding value rendered in the notification in block 362, and the system may return to block 354 and continue processing the stream of audio data to monitor any further parameters requested by the additional user. If, in any iteration of block 364 , the system determines that user input to activate assisted calling is received, the system can proceed to block 366 .

[0078] In block 366, the system causes the value to be rendered at the additional client device for presentation to the additional user. In response to receiving user input to activate assisted speech to provide the corresponding value to the additional user, the system may cause synthesized speech including the corresponding value to be rendered at the additional client device of the additional user and / or at the client device of the given user.

[0079] In implementations including any of blocks 360, 362, and / or 364, the system may, in response to receiving explicit user input invoking an assisted call, cause corresponding values ​​to be rendered on additional client devices of additional users and / or on the client device of the given user. In some versions of these implementations, the user input to activate the assisted call and interrupt the ongoing call may be proactive. In other words, if the system receives user input at block 360, the assisted call may be activated to provide corresponding values ​​for the task even if the system did not render any notification indicating that the system can perform at least one task (e.g., as described in connection with FIG. 5A). In other versions of these implementations, the user input to activate the assisted call may be reactive. In other words, if the system receives user input at block 364, the assisted call may be activated to provide corresponding values ​​for the task after rendering at block 362 of a notification indicating that the system can perform at least one task (e.g., as described in connection with FIG. 5B). In implementations that do not include optional blocks 360, 362, and / or 364, the system may proceed directly from block 258 to block 366. In some versions of those implementations, the system may automatically interrupt the ongoing call (i.e., without receiving any explicit user input to activate the assisted call) in response to a determination of the corresponding value in block 358 (e.g., as described in connection with FIG. 5C ).

[0080] In block 368, the system determines whether any user input to continue the assisted call is received at the client device associated with the given user. If, in the iteration of block 368, the system determines that user input to continue the assisted call is not received, the system may return to block 354 to process further audio data corresponding to the ongoing call. If, in the iteration of block 368, the system determines that user input to continue the assisted call is received, the system may proceed to block 264 of FIG. 2 to determine whether values ​​for any additional parameters are required during the assisted call. In this manner, the system may provide corresponding values ​​for parameters during the performance of tasks by a given user during the call. By providing corresponding values ​​either automatically or in response to explicit user input as described above, the system receives less input from a given user of a client device, thereby conserving computational resources on the given client device, because the user does not need to navigate to various applications with entirely different user interfaces to determine corresponding values. Furthermore, because the user does not need to navigate to these various applications, the system conserves both computational and network resources by more quickly terminating ongoing calls. As one non-limiting example, by using the techniques described herein, a given user does not have to pause or put an interaction on hold while searching an email application, an airline application, and / or other applications for corresponding values.

[0081] 4A-4D, various non-limiting examples of user interfaces associated with making assisted phone calls are described. Each of FIGS. 4A-4D illustrates a client device 410 having a graphical user interface 480 displaying example interactions of a given user of the client device 410. The interactions may include, for example, interactions with one or more software applications (e.g., a web browser application, an automated assistant application, a contacts application, an email application, a calendar application, and / or other software-based applications accessible by the client device 410) as well as interactions with additional users (e.g., additional human participants associated with additional client devices, additional automated assistants associated with additional client devices of additional users, and / or other additional users). One or more aspects of an automated assistant associated with the client device 410 (e.g., automated assistant 115 of FIG. 1) may be implemented locally on the client device 410 and / or distributed across other client devices in network communication with the client device 410 (e.g., via network 190 of FIG. 1). For simplicity, the operations of Figures 4A-4D are described herein as being performed by an automated assistant. While client device 410 in Figures 4A-4D is depicted as a mobile phone, it should be understood that this is not intended to be limiting. Client device 410 could be, for example, a standalone assistant device (e.g., with a speaker and / or display), a laptop, a desktop computer, and / or any other client device capable of making phone calls.

[0082] 4A-4D further includes a text response interface element 484 that a user may select to generate user input via a virtual keyboard or other touch and / or typed input, and a voice response interface element 485 that a user may select to generate user input via a microphone of the client device 410. In some implementations, a user may generate user input via a microphone without selecting the voice response interface element 485. For example, active monitoring of audible user input via a microphone may be performed to eliminate the need for a user to select the voice response interface element 485. In some of these and / or other implementations, the voice response interface element 485 may be omitted. Furthermore, in some implementations, the text response interface element 484 may additionally and / or alternatively be omitted (e.g., the user may provide only audible user input). The graphical user interface 480 of FIGS. 4A-4D also includes system interface elements 481, 482, 483 that may be interacted with by a user to cause the computing device 410 to perform one or more actions.

[0083] In various implementations described herein, user input may be received to initiate a call (e.g., an assisted call) with an entity using an automated assistant. The user input can be spoken, touch, and / or typed input that includes an indication to initiate an assisted call. Furthermore, the automated assistant can perform tasks on behalf of a given user of client device 410 in association with the entity. As shown in FIG. 4A , user interface 480 includes search results for a restaurant entity from a browser application accessible at client device 410 (e.g., as indicated by URL 411, “www.exampleurl0.com / ”). Furthermore, the search results include a first search result 420 for “Hypothetical Cafe” and a second search result 430 for “Example Cafe.”

[0084] In some implementations, search results 420 and / or 430 may be associated with various selectable graphical elements that, when selected, cause client device 410 to perform a corresponding action. For example, when call graphical element 421 and / or 431 associated with a given one of search results 420 and / or 430 is selected, user input may indicate that a phone call action to the restaurant entity associated with search result 420 and / or 430 should be performed. As another example, when directions graphical element 422 and / or 432 associated with a given one of search results 420 and / or 430 is selected, user input may indicate that a navigation action to the restaurant entity associated with search result 420 and / or 430 should be performed. As yet another example, when menu graphical element 423 and / or 433 associated with a given one of search results 420 and / or 430 is selected, user input may indicate that a browser-based action of displaying a menu for the restaurant entity associated with search result 420 and / or 430 should be performed. 4A , the assisted call is initiated from a browser application, it should be understood that this is for purposes of illustration and not intended to be limiting. For example, the assisted call can be initiated from various software applications accessible on the client device 410 (e.g., a contacts application, an email application, a text or SMS messaging application, and / or other software applications), and if the assisted call is initiated using spoken input, can be initiated from the home screen of the client device 410, from the lock screen of the client device 410, and / or from other states of the client device 410.

[0085] For illustrative purposes, assume that user input to initiate a call using the second search result 430 for “Example Cafe” is detected at client device 410. The user input can be, for example, a spoken input of “Call Example Cafe” or a touch input directed at call graphical element 431. In some implementations, in response to receiving the user input to initiate a call with “Example Cafe,” a call details interface 470 can be rendered at client device 410. In some versions of those implementations, call details interface 470 can be rendered at client device 410 as part of user interface 480. In some other versions of those implementations, call details interface 470 can be a separate interface from user interface 480 that overlays the user interface and can include a call details interface element 486 that allows the user to expand call details interface 470 to view additional call details (e.g., by swiping up on call details interface element 486) and / or dismiss call details interface 470 (e.g., by swiping down on call details interface element 486). While call details interface 470 is depicted as being at the bottom of user interface 480, it should be understood that this is for purposes of illustration and not intended to be limiting. For example, call details interface 470 may be rendered at the top of user interface 480, to the side of user interface 480, or in an interface that is entirely separate from user interface 480.

[0086] The call details interface 470 may include multiple graphical elements in various implementations. In some implementations, the graphical elements may be selectable such that when a given one of the graphical elements is selected, the client device 410 can perform a corresponding action. As shown in FIG. 4A , the call details interface 470 includes a first graphical element 471 for “Assisted Call,” a second graphical element 472 for “Normal Call,” and a third graphical element 473 for “Save Contact ‘Example Cafe’.” Furthermore, the first graphical element 471, when selected, can provide an indication to the automated assistant of a desire to initiate an assisted call using the automated assistant, the second graphical element 472, when selected, can cause the automated assistant to initiate a call without using an assisted call, and the third graphical element 473, when selected, can cause the automated assistant to create a contact associated with Example Cafe. Notably, in some versions of these implementations, the graphical elements may include sub-elements to provide an indication of the task to be performed. For example, a first graphical element 471 of "Assisted Call" may include a first sub-element 471A of "Make a Reservation" related to the task of making a restaurant reservation at Example Cafe, a second sub-element 471B of "Modify Reservation" related to the task of modifying the restaurant reservation at Example Cafe, and a third sub-element 471C of "Cancel Reservation" related to the task of canceling the restaurant reservation at Example Cafe.

[0087] For illustrative purposes, assume that user input is detected at client device 410 to initiate an assisted call with Example Cafe to make a restaurant reservation at Example Cafe. The user input can be, for example, a spoken input such as "Call Example Cafe to make a restaurant reservation" or a touch input directed to first subelement 471A. In response to detecting the user input, the automated assistant can determine the task of "making a restaurant reservation at Example Cafe" and identify candidate parameters associated with the identified task as described herein (e.g., in connection with parameter engine 153 of FIG. 1). In some implementations, as shown in FIG. 4B, the automated assistant can determine values ​​for the candidate parameters. In particular, when proceeding from FIG. 4A to FIG. 4B, call details interface 470 can be updated to include the identified candidate parameters. In some versions of those implementations, the automated assistant can determine values ​​for the candidate parameters based on a user profile associated with a given user of client device 410. For example, the automated assistant can determine value 474A (e.g., Jane Doe) for name parameter 474 and value 475A (e.g., (502)123-4567) for phone number parameter 475 without having to ask a given user for values ​​474A and 475A based on the automated assistant having access to values ​​474A and 475A via the given user's user profile on client device 410. While Figure 4B is described herein with reference to particular candidate parameters identified by the automated assistant, it should be understood that this is for illustrative purposes.

[0088] In some versions of those implementations, the automated assistant can interact (e.g., audibly and / or visually) with a given user of client device 410 to solicit corresponding values ​​for candidate parameters not identified based on the user profile of the given user of client device 410. In some further versions of those implementations, the automated assistant may solicit values ​​only for candidate parameters that are deemed essential parameters as described herein (e.g., in connection with parameter engine 153 of FIG. 1 ). For example, for a task of making a restaurant reservation, the automated assistant can generate prompt 452B1, "What day and time would you like to make a reservation at Example Cafe?" and receive user input 454B1 (e.g., typed or spoken) of "Booth on March 1 at 7 PM." Accordingly, call details interface 470 can be updated to include value 476A of date / time parameter 476 (e.g., March 1, 2020, 7 PM). Notably, user input 454B1 also includes value 478A of seat type parameter 478, “booth,” which was not requested in prompt 452B1. Even though the automated assistant did not request value 478A in prompt 452B1, the automated assistant can determine that value 478A corresponds to seat type parameter 478, which may be considered an optional parameter as described herein (e.g., in connection with parameter engine 153 of FIG. 1 ). Furthermore, the automated assistant can generate a further prompt (e.g., prompt 452B2) for an additional parameter (e.g., number of people parameter 477) to determine a corresponding value (e.g., 477A having a value of 5) for the additional parameter. In this manner, the automated assistant can determine values ​​of candidate parameters to be used in performing a task on behalf of a given user of client device 410 before initiating an assisted call.

[0089] Further, in various implementations, when determining values ​​for the candidate parameters, call details interface 470 may include various graphical elements, as shown in FIG. 4B. For example, call details interface 470 may include edit graphical element 441B that, when selected, allows the user to modify values ​​474A-478A before initiating the assisted call, cancel graphical element 442B that, when selected, terminates the assisted call and optionally returns user interface 480 to the state prior to detecting the user input to initiate the assisted call (e.g., user interface 480 of FIG. 4A), and call interface element 443B that, when selected, allows the user to initiate a regular call without the assistance of an automated assistant (e.g., similar to selecting graphical element 472 of FIG. 4A).

[0090] After determining corresponding values ​​474A-478A of candidate parameters 474-478, the automated assistant can initiate an assisted call in association with the entity to perform a task on behalf of a given user of client device 410. The automated assistant can initiate the assisted call using a calling application accessible at client device 410. In various implementations, as shown in FIG. 4C , call details interface 470 can be updated to include various graphical elements in response to the initiation of a call by the automated assistant. For example, call details interface 470 can include a call end graphical element 441C that, when selected, causes the automated assistant to end the assisted call after initiating the assisted call, a call join graphical element 442C that, when selected, allows a given user of client device 410 to take over the call and / or task execution from the automated assistant, and a speaker interface element 443C that, when selected, causes client device 410 to audibly render an interaction between the user and the automated assistant. These graphical elements 442C, 443C, and 444C may be selected throughout the duration of the assisted call.

[0091] Further, in various implementations, the automated assistant can obtain consent from an additional user of the client device after initiating the assisted call and before performing the task. As shown in FIG. 4C , the automated assistant can cause synthesized speech 452C1 to be rendered at an additional client device associated with the Example Cafe representative, requesting that the Example Cafe representative provide consent to interact with the automated assistant. The automated assistant can process audio data 456C1 corresponding to additional spoken input from the Example Cafe representative to determine that consent has been provided by the Example Cafe representative (e.g., "Yes" in audio data 456C1). Furthermore, in response to determining that audio data 456C1 requests information about value 476A of date / time parameter 476, the automated assistant can cause further synthesized speech 456C2 including value 476A of date / time parameter 476 to be rendered at an additional client device of the Example Cafe representative. In this manner, the automated assistant can perform the task of making a restaurant reservation at Example Cafe by providing synthesized speech including a value responsive to the request for information from the Example Cafe representative.

[0092] In some implementations, the automated assistant can process audio data corresponding to further spoken input from the Example Cafe attendant and can determine that the further spoken input requests further value information related to a parameter not currently known to the automated assistant. As shown in FIG. 4C , audio data 456C2 captures further spoken input requesting information for a child parameter not currently known to the automated assistant. In response to determining that the further spoken input from the Example Cafe attendant requests information not currently known to the automated assistant, the automated assistant can cause yet another synthesized speech to be rendered at a further client device indicating that the requested information is not currently known. Although the automated assistant may not currently know the requested information, it can continue to perform the task using other values ​​that the automated assistant currently knows. For example, in response to determining that audio data 456C2 requests further value information related to a parameter not currently known to the automated assistant, the automated assistant can generate yet another synthesized speech 454C3 stating, "I don't know. I'll have to ask Jane Doe. Would you like to continue with the reservation until I get an answer?"

[0093] In some versions of those implementations, when the automated assistant determines that the additional user's audio data includes a request for information, the automated assistant may cause a notification to be rendered at the client device 410 indicating that the additional user is requesting information related to parameters that the automated assistant does not currently know. In some further versions of those implementations, the notification may further include suggested values ​​as recommendations in response to the request for information. The recommendations may be selectable such that selecting a given one of the recommendations allows the automated assistant to utilize the value included in the given one of the recommendations as a value in response to the additional user's request. For example, as shown in FIG. 4D , the automated assistant may cause a notification 479 to be visually rendered in the call details interface 470. The notification 479 includes an indication that "Example Cafe wants to know if children are included in the reservation" and also includes a first suggestion 479A of "yes" and a second suggestion 479B of "no" that are offered as recommended values ​​in response to the additional user's request. As described in more detail below, various other types of notifications can be rendered at the client device 410, and the type of notification can be based on the state of the client device 410 when the system determines that the audio data includes a request for information that the automated assistant does not currently know.

[0094] As described above, the automated assistant can continue to perform the task using other values ​​that the automated assistant currently knows, even if the automated assistant determines that additional users are requesting information that the automated assistant does not currently know. The additional values ​​may be provided later in the interaction after receiving additional user input from a given user of client device 410 in response to notification 479. For example, as shown in FIG. 4D , in response to rendering yet another synthesized speech 452C3 indicating that the value of the child parameter is not currently known to the automated assistant, but that the automated assistant knows other values ​​and wishes to continue performing the task using those other values. Further assume that audio data 456D1 is received saying, "That's fine. What type of seating would you like?" (e.g., a request for information for value 478A of "booth" for seating type parameter 478). Further assume that while audio data 456D1 is being processed by the automated assistant, the given user of client device 410 provides spoken and / or touch input directed to second suggestion 479B, indicating that no children will be participating in the restaurant reservation at Example Cafe. In this example, the automated assistant can include value 478A of "booth" for seat type parameter 478 and can cause synthesized speech 452D1 including a value (e.g., "No" based on the user's selection of second suggestion 479B) to be audibly rendered at a further client device of a further user. The automated assistant can process further audio data 456D2 of "Perfect. Reservation complete," determine that the task of making a restaurant reservation is complete, can produce further synthesized speech 452D2 of "Notifying Jane Doe. Have a nice day," and can end the call.

[0095] In particular, in some implementations, the automated assistant can include an additional value (i.e., not yet known at the time of the request by the additional user, but now known based on additional user interface input) in the synthesized speech in response to a request for more information from the additional user. In other words, the automated assistant can include an additional value in the synthesized speech even if the immediately preceding request from the additional user was not a request for information. For example, as shown in FIG. 4D , synthesized speech 452D1 includes a value 478A of “Booth Seat” in response to a immediately preceding request for information related to seat type parameter 478 included in audio data 456D1, and the synthesized speech also includes a value of “No” for the children parameter in response to a previous request for information related to the children parameter included in audio data 456C2. In this way, the automated assistant can continue executing the task while waiting for further user input, including the additional value, and can logically and conversationally provide the additional value to the additional user as the additional value becomes known to the automated assistant. By continuing to execute the task while waiting for further user input, the task can be executed more quickly and efficiently because execution of the task is not paused until further user input is received. By performing tasks more quickly and efficiently, the techniques described herein can conserve computational and network resources when performing tasks using assisted calls.

[0096] In various implementations, not depicted, the automated assistant may determine that no further audio data associated with the entity is received from the further user during the assisted call within a threshold duration. In some versions of those implementations, the automated assistant may render a further synthesized speech based on one or more corresponding values ​​that were not included in the request for information from the further user. For example, if the automated assistant determines that the further user has not said anything for 10 seconds and the automated assistant knows the value of a number of people parameter for a restaurant reservation task that was not requested by the further user, the assistant may render a synthesized speech that says, "Just to be clear, five people will be attending the reservation." In some additional and / or alternative versions of those implementations, the automated assistant can render a further synthesized speech based on one or more parameters that were not included in the request for information from the further user. For example, if the automated assistant determines that the further user has not said anything for 10 seconds and the automated assistant knows the value of a number of people parameter for a restaurant reservation task that was not requested by the further user, the assistant may render a synthesized speech that says, "Would you like to know how many people will be attending the reservation?"

[0097] Further, in various implementations, the automated assistant may cause transcriptions of various interactions to be visually rendered in the user interface 480 of the client device 410. The transcriptions may be displayed, for example, in various software applications (e.g., an assistant application, a calling application, and / or other applications) in the home of the client device 410. In some implementations, the transcriptions may include interactions between the automated assistant and a given user of the client device 410 (e.g., as shown in FIG. 4B). In some additional and / or alternative implementations, the transcriptions may include interactions between the automated assistant and additional users (e.g., as shown in FIGS. 4C and 4D).

[0098] 4B-4D are each shown as including a transcript of the interaction, it should be noted that this is for illustrative purposes only and is not intended to be limiting. It should be understood that the assisted calls described above may be performed while the client device 410 is asleep, locked, has other software applications running in the foreground, and / or is in other states. Furthermore, in implementations, if the automated assistant causes a notification to be rendered at the client device 410, the type of notification rendered at the client device is based on the state of the client device 410, as described herein. Furthermore, it should be understood that while FIGS. 4A-4D are described herein in connection with the task of making a restaurant reservation, this is again not intended to be limiting, and that the techniques described herein can be utilized for a number of different tasks that may be performed in connection with a number of different entities.

[0099] Further, in various implementations, the automated assistant may be placed on hold by an additional user at the beginning of the assisted call and / or during the assisted call. In some versions of those implementations, the automated assistant may be considered on hold when the automated assistant is not engaged in a conversation with an additional human participant. For example, the automated assistant may be considered on hold if it is serving a hold system associated with the entity, an interactive voice response (IVR) system associated with the entity, and / or other systems associated with the entity. Further, the automated assistant may cause a notification to be rendered at the client device 410 when the assisted call is placed on hold and / or when the assisted call is resumed after being placed on hold. The notification may indicate, for example, that the assisted call was placed on hold, that the assisted call resumed after being placed on hold, that the user is requested to join the assisted call when it is resumed after being placed on hold, and / or other information related to the assisted call. Additionally, the automated assistant can determine that the assisted call has resumed after being placed on hold based on processing a stream of audio data transmitted to the given user's client device 410 from an additional client device of an additional user.

[0100] In some versions of those implementations, the automated assistant can determine that the additional user has already requested all of the information related to the parameters that the automated assistant knows. In some further versions of those implementations, the automated assistant can proactively request that a given user of the client device 410 join the assisted call when the assisted call is resumed after being put on hold if the remainder of the task requires the given user of the client device 410. For example, a particular task may require a given user of the client device 410 rather than the automated assistant. For example, assume the entity is a bank, the additional user is a bank representative, and the task is disputing a debit card charge. In this example, the automated assistant can first call the bank and provide synthesized voice and / or emulated button presses, such as a synthesized voice including the given user's name, simulated button presses including the user's bank account number, a synthesized voice providing the reason for the call, and / or simulated button presses navigating an automated system (e.g., an IVR system) to transfer calls related to the bank. However, the automated assistant may know that when the assisted call is transferred to a bank representative, a given user will be asked to join the assisted call to verify the given user's identity and explain the disputed debit card charge, and may cause a notification to be rendered at the user's client device 410 when the assisted call is transferred to the bank representative. Thus, the automated assistant can handle the initial part of the assisted call and request that the given user of client device 410 take over the assisted call when a bank representative is available to discuss the disputed charge.

[0101] In some other further versions of these implementations, the automated assistant may proactively request that a given user of client device 410 join the assisted call when the assisted call is resumed after being put on hold because the automated assistant does not know any further information related to the restaurant reservation. For example, assume that an additional user of the assisted call shown in FIGS. 4B-4D puts the automated assistant on hold in audio data 456D2 rather than indicating that the restaurant reservation is completed. In this example, the automated assistant has already provided all values ​​of information known to the automated assistant related to the restaurant reservation, and any further information requested by the additional user when the assisted call is resumed will be information not currently known to the automated assistant. Thus, the automated assistant may proactively request that a given user of client device 410 take over the assisted call when the assisted call is resumed.

[0102] In yet other further versions of these implementations, the automated assistant can withhold notifications indicating that the assisted call has been put on hold and / or that the assisted call has been resumed after being put on hold, continue the assisted call when the assisted call is resumed after being put on hold, and have a notification rendered at the given user's client device 410 along with or in place of a notification (e.g., notification 479 of FIG. 4D ) indicating that an additional user has requested information related to a parameter that the automated assistant does not currently know. Continuing with the above example, rather than proactively requesting that a given user of client device 410 take over the assisted call when the assisted call is resumed, the automated assistant can process additional audio data corresponding to additional spoken input of the additional user to determine whether the additional user has requested more information related to a parameter that the automated assistant does not currently know. In this example, the automated assistant can provide a notification requesting that a given user of client device 410 take over the assisted call along with or in place of a notification requesting more information values, similar to notification 479 of FIG. 4D . In this manner, the automated assistant can continue the assisted call as described in Figures 4B-4D and / or hand over control of the assisted call to a given user of the client device 410.

[0103] 5A-5C, various non-limiting examples of user interfaces associated with providing auxiliary output during an ongoing, unassisted call are described. Each of FIGS. 5A-5C illustrates a client device 510 having a graphical user interface 580 displaying example interactions of a given user of the client device 510. The interactions may include, for example, interactions with one or more software applications (e.g., a web browser application, an automated assistant application, a contacts application, an email application, a calendar application, and / or other software-based applications accessible by the client device 510) as well as interactions with additional users (e.g., an automated assistant associated with the client device 510, an additional human participant associated with an additional client device, an additional automated assistant associated with an additional client device of an additional user, and / or other additional users). One or more aspects of an automated assistant associated with client device 510 (e.g., automated assistant 115 of FIG. 1 ) may be implemented locally on client device 510 and / or distributed across other client devices in network communication with client device 510 (e.g., via network 190 of FIG. 1 ). For simplicity, the operations of FIGS. 5A-5C are described herein as being performed by an automated assistant. While client device 510 in FIGS. 5A-5C is depicted as a mobile phone, it should be understood that this is not intended to be limiting. Client device 510 could be, for example, a standalone speaker, a speaker connected to a graphical user interface, a laptop, a desktop computer, and / or any other client device capable of making phone calls.

[0104] 5A-5C further includes a text response interface element 584 that a user may select to generate user input via a virtual keyboard or other touch and / or typed input, and a voice response interface element 585 that a user may select to generate user input via a microphone of the client device 510. In some implementations, a user may generate user input via a microphone without selecting the voice response interface element 585. For example, active monitoring of audible user input via a microphone may be performed to eliminate the need for a user to select the voice response interface element 585. In some of these and / or other implementations, the voice response interface element 585 may be omitted. Furthermore, in some implementations, the text response interface element 584 may additionally and / or alternatively be omitted (e.g., the user may provide only audible user interface input). 5A-5C also includes system interface elements 581, 582, 583 that may be interacted with by a user to cause computing device 510 to perform one or more actions. In some implementations, call details interface 570 may be rendered at client device 510, and call details interface 570 may include graphical element 542 that, when selected, can terminate an ongoing call and may also include graphical element 543 that, when selected, can cause an automated assistant to take over a call from a given user of client device 510 using assisted calling.In some versions of those implementations, the assisted call interface 570 may also include a call details interface element 586 that allows the user to expand the call details interface 570 to view additional call details (e.g., by swiping up on the call details interface element 586) and / or dismiss the call details interface 570 (e.g., by swiping down on the call details interface element 586).

[0105] In various implementations, the automated assistant can interrupt an ongoing call (i.e., a non-assisted call) between a given user of client device 410 and an additional user of an additional client device. In some implementations, the automated assistant can process audio data corresponding to the ongoing call to identify an entity associated with the additional user, a task performed during the ongoing call, and / or parameters for the task performed during the ongoing call. Additionally, the automated assistant can determine values ​​for the identified parameters. For example, as shown in FIGS. 5A-5C , assume that audio data 552A1, 552B1, and / or 552C1 capturing spoken input from an additional user, "This is Example Airlines. How can I help you?" is received at client device 510. In this example, the automated assistant can process audio data 552A1, 552B1, and / or 552C1 and, based on the processing, determine that the additional user is a representative of Example Airlines and / or identify Example Airlines as an entity associated with a representative of Example Airlines. In some additional and / or alternative implementations, the automated assistant can additionally and / or alternatively identify the entity based on metadata associated with the ongoing call as described herein (e.g., in connection with assisted call engine 150 of FIG. 1 ). In some versions of those implementations, the automated assistant may identify the entity based solely on metadata associated with the ongoing call without processing the stream of audio data corresponding to the ongoing call, thereby eliminating the need to obtain consent from an additional user of the ongoing call. As described below, the automated assistant can still perform tasks on behalf of a given user of client device 510 without processing the stream of audio data corresponding to the ongoing call.

[0106] Further assume that audio data 554A1, 554B1, and / or 554C1 capturing a spoken input of "Hello, I need to change my flight" from a given user (e.g., Jane Doe) of client device 510 is detected at client device 510 and transmitted to an additional client device of an additional user. In this example, the automated assistant can process audio data 554A1, 554B1, and / or 554C1 and determine a task to change the flight based on the processing. In other examples, the automated assistant can additionally and / or alternatively determine a task associated with an entity as described herein (e.g., in connection with assisted call engine 150 of FIG. 1 ). Further assume that audio data 552A2, 552B2, and / or 552C2 capturing a spoken input of "OK, do you have a frequent flyer number" from an additional user is received at client device 510. In this example, the automated assistant can process audio data 554A2, 554B2, and / or 552C2 and, based on the processing, identify a frequent flyer number parameter for the task of changing a flight (or consider the task of providing a frequent flyer number). In other examples, the automated assistant can additionally and / or alternatively identify the parameter based on the parameter being stored in association with the task and / or entity, as described herein (e.g., in connection with assisted call engine 150 of FIG. 1). Furthermore, the automated assistant can determine a value for the frequent flyer number parameter and provide the value of the frequent flyer number parameter to a further user.

[0107] In some implementations, the automated assistant can determine a value for a parameter identified during an ongoing call between a given user of client device 510 and an additional user in response to receiving user input from the given user of client device 510, including a request for information related to the parameter. The automated assistant can determine the value of the parameter based on a user profile associated with the given user of client device 510 (e.g., stored in user profile database 153B of FIG. 1 ). In some versions of these implementations, the automated assistant can cause the value of the parameter to be rendered on an additional client device of the additional user and / or on the given user's client device 510 in response to determining the value of the parameter. For example, as shown in FIG. 5A , assume that audio data 554A2 capturing spoken input from a given user (e.g., Jane Doe) of client device 510, "Assistant, what's my frequent flyer number," is detected at client device 510. In response to receiving the spoken input captured in audio data 554A2, the automated assistant can determine a value for a frequent flyer parameter (e.g., based on the inclusion of an Example Airlines account associated with the given user in the user profile) and cause a synthesized speech 556A1 to be rendered at the additional user's client device and / or the given user's client device 510: "Jane Doe's Example Airlines frequent flyer number is 0112358."

[0108] In some additional and / or alternative versions of these implementations, rather than receiving spoken input contained in audio data 554A2, the automated assistant may receive a selection of graphical element 543 to take over a call from a given user of client device 510 using an assisted call. For example, in response to receiving a selection of graphical element 543, the automated assistant may determine a value for a frequent flyer parameter (e.g., based on the inclusion of an Example Airlines account associated with the given user in the user profile) and cause synthesized speech 556A1 to be rendered at an additional client device of an additional user and / or at client device 510 of the given user.

[0109] In particular, audio data 554A2 includes "What's my frequent flyer number," which does not identify an entity associated with the frequent flyer number. In an implementation, if the automated assistant does not obtain further user consent and / or provide an indication of an entity associated with the frequent flyer number parameter, the automated assistant can determine, based on metadata associated with the ongoing call, that "My frequent flyer number" refers to the value of the frequent flyer number parameter associated with Example Airlines. In this manner, the automated assistant can still provide a value in response to a request from a given user of client device 510 without processing a stream of audio data corresponding to the ongoing call.

[0110] In other implementations, the automated assistant can proactively determine values ​​for parameters identified during an ongoing call between a given user of client device 510 and an additional user without receiving any user input, including a request for information related to the parameters, from the given user of client device 510. The automated assistant can determine values ​​for the parameters based on a user profile associated with the given user of client device 510 (e.g., stored in user profile database 153B of FIG. 1).

[0111] In some versions of those implementations, the automated assistant may cause a notification to be rendered at client device 510 that includes an indication that the assisted call can perform the task and / or provide the value of the parameter to an additional user. In response to receiving user input invoking assisted speech from a given user of client device 510, the automated assistant may cause synthesized speech including the value to be rendered at an additional client device of the additional user. For example, as shown in FIG. 5B , assume that the automated assistant determines the value of a frequent flyer number parameter in response to identifying a request for the value of the frequent flyer number parameter in audio data 552B2. Further assume that the automated assistant causes a notification 579 to be rendered in call details interface 570 of client device 510 that reads, "Your Example Airlines frequent flyer number is 0112358. Would you like to provide your frequent flyer number to an Example Airlines representative?" The automated assistant can then, in response to receiving a selection of graphical element 579A and / or graphical element 543 indicating that the automated assistant should provide the value of the frequent flyer number parameter to the further user, cause synthesized speech 556B1 to be rendered at the further client device of the further user and / or at the given user's client device 510: "Jane Doe's Example Airlines frequent flyer number is 0112358." In some additional and / or alternative implementations, the automated assistant can detect a spoken input of the given user of client device 510 that includes a value for the identified parameter during the ongoing dialogue. In some versions of those implementations, the automated assistant can automatically terminate notification 579 included in call details interface 570.

[0112] In some other versions of these implementations, the automated assistant can proactively provide the value of an identified parameter during an ongoing call in response to determining the value of the parameter without receiving any user input that includes a request for information related to the parameter. The automated assistance can cause a synthesized speech including the value to be rendered at a further client device of a further user in response to determining the value. For example, as shown in FIG. 5C , assume that the automated assistant determines the value of a frequent flyer number parameter in response to identifying a request for the value of the frequent flyer number parameter from a further user in audio data 552C2. The automated assistant can then cause a synthesized speech 556C1 to be rendered at a further client device of a further user and / or at a given user's client device 510 that says, "Jane Doe's Example Airlines frequent flyer number is 0112358" in response to determining the value of the frequent flyer number parameter without receiving any user input that includes a request for information related to the frequent flyer number parameter. In some further versions of those implementations, the automated assistant may only proactively provide a value for a parameter identified during an ongoing call if a reliability metric associated with the determined value of the parameter meets a reliability threshold. The reliability metric may be based, for example, on whether a given user has previously provided the determined value in response to receiving a previous request for the same information, whether the parameter identified during the ongoing call is a candidate parameter stored in association with the task identified during the ongoing call, the source from which the value is determined (e.g., email / calendar application vs. text messaging application), and / or the method for determining the reliability metric.

[0113] Further, in various implementations, after the automated assistant provides values ​​for identified parameters during an ongoing call, the automated assistant can take over the remainder of the ongoing call. In some versions of those implementations, the automated assistant can take over the remainder of the ongoing call in response to a user selection of graphical element 543. Furthermore, the automated assistant can continue to identify parameters for the task based on audio data sent to client device 510 and determine values ​​for the identified parameters. Furthermore, if the automated assistant determines that a given request is for information that cannot be determined by the automated assistant, the automated assistant can provide a notification requesting further user input to comply with the request, as described above (e.g., in connection with FIGS. 4C and 4D ), and the automated assistant can then provide values ​​based on the further user input to comply with the request.

[0114] 5A-5C are described herein in connection with the task of providing a frequent flyer number, it should be understood that this is not intended to be limiting, and the techniques described herein can be utilized for a number of different tasks that may be performed in connection with a number of different entities. Additionally, although not shown in FIGURES 5A-5C, it should be understood that the automated assistant can obtain consent from an additional user when the ongoing call is initiated, even if the ongoing call is not initiated by the automated assistant using an assisted call. The automated assistant can obtain consent from the additional user using any of the methods described herein.

[0115] 6 is a block diagram of an example computing device 610 that may optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client device, cloud-based automated assistant component, and / or other component may include one or more components of the example computing device 610.

[0116] Generally, computing device 610 includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices may include, for example, a storage subsystem 624 including a memory subsystem 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with computing device 610. The network interface subsystem 616 provides an interface to external networks and is coupled to corresponding interface devices of other computing devices.

[0117] The user interface input devices 622 may include a keyboard, a pointing device such as a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touchscreen integrated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 610 or a communications network.

[0118] The user interface output devices 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as through an audio output device. In general, use of the term "output device" is intended to encompass all possible types of devices and ways of outputting information from the computing device 610 to a user or to another machine or computing device.

[0119] Storage subsystem 624 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 624 may include logic for performing selected aspects of the methods disclosed herein and for implementing the various components shown in FIG.

[0120] These software modules are generally executed by the processor 614 alone or in combination with other processors. The memory 625 used in the storage subsystem 624 may include several memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution, and a read-only memory (ROM) 632 in which certain instructions are stored. The file storage subsystem 626 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular implementation may be stored by the file storage subsystem 626 in the storage subsystem 624 or on other machines that can be accessed by the processor 614.

[0121] The bus subsystem 612 provides a mechanism for allowing the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem 612 may use multiple buses.

[0122] Computing device 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 610 shown in Figure 6 is intended only as a specific example intended to illustrate some implementations. Many other configurations of computing device 610 are possible, having more or fewer components than the computing device shown in Figure 6.

[0123] In situations where the systems described herein may collect or otherwise monitor personal information about users or utilize personal information and / or monitored information, users may be given the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how content that may be more relevant to the user should be received from a content server. Also, certain data may be processed in one or more ways before being stored or used such that personally identifiable information is removed. For example, a user's identity may be processed such that personally identifiable information cannot be determined about the user, or such that if geographic location information is obtained, the user's geographic location may be generalized (such as to the city, zip code, or state level) so that the user's specific geographic location cannot be determined. Thus, users may be able to control how information is collected and / or used about them.

[0124] In some implementations, a method implemented by one or more processors is provided, including receiving user input to initiate an assisted call from a given user via a client device associated with the given user, and determining, based on the user input, entities to engage on behalf of the given user during the assisted call and tasks to be performed on behalf of the given user during the assisted call. The method further includes determining, for one or more candidate parameters stored in association with the tasks and / or entities, one or more corresponding values ​​to be used to automatically generate synthetic speech when performing the tasks during the assisted call, initiating execution of the assisted call using a client device associated with the given user, and determining, based on processing assisted call audio data capturing utterances of the further user associated with the entities during execution of the assisted call, that information related to the further parameters is requested by the further user. In response to determining that information related to the further parameters is requested, the method further includes causing the client device to render a notification outside the assisted call identifying the further parameters and requesting further user input regarding the information. The method further includes continuing the assisted call prior to receiving any further input responsive to the notification. Continuing the assisted call includes rendering one or more instances of synthetic speech based on one or more of the corresponding values ​​of the candidate parameters. The method further includes, while continuing the assisted call, determining, in response to the notification, whether further user input specifying a particular value of the further parameter is received within a threshold duration, and, in response to determining that the further user input is received within the threshold duration, rendering the further synthetic speech based on the particular value as part of the assisted call.

[0125] These and other implementations of the technology disclosed herein may optionally include one or more of the following features.

[0126] In some implementations, determining the given value of the corresponding values ​​includes, before initiating the assisted call, identifying a given candidate parameter of the candidate parameters and generating a prompt requesting further information related to the given candidate parameter, causing the client device to render the prompt, and identifying the given value of the given candidate parameter based on further user input in response to the prompt.

[0127] In some implementations, determining the further ones of the corresponding values ​​includes identifying the further values ​​based on a user profile associated with the given user prior to initiating the assisted call.

[0128] In some implementations, continuing the assisted call includes processing further audio data of the assisted call to determine that a further utterance of the further user includes a request for a given one of the candidate parameters, and, in response to determining that the further utterance includes a request for the given candidate parameter, causing the client device to render a given instance of the one or more instances of synthesized speech in the call. In some versions of these implementations, the given instance includes a given value of the corresponding values ​​based on the given value determined for the given candidate parameter, and the given instance is rendered without requiring any further user input from the given user.

[0129] In some implementations, continuing the assisted call includes processing the further audio data to determine whether further utterances of the further user are received within a further threshold duration, and, in response to determining that further utterances are not received from the further user within the further threshold duration, rendering another instance during the assisted call of the one or more instances of synthetic speech that is based on one or more of the corresponding values ​​not requested by the further user.

[0130] In some implementations, the method further includes, in response to determining that information related to the additional parameters is requested by an additional user, updating one or more candidate parameters stored in association with the entity to include the additional parameters.

[0131] In some implementations, the method further includes determining a state of the client device when information related to the further parameter is requested by the further user, and determining, based on the state of the client device, the notification and / or one or more properties for rendering the notification.

[0132] In some versions of those implementations, a state of the client device indicates that a given user is actively monitoring the assisted call, and determining the notification based on the state of the client device includes determining the notification to include a visual component visually rendered by a display of the client device together with one or more selectable graphical elements based on the state of the client device indicating that the given user is actively monitoring the assisted call. In some further versions of those implementations, a further user input responsive to the notification includes selecting a given one of the one or more selectable graphical elements.

[0133] In some versions of those implementations, the state of the client device indicates that the given user is not actively monitoring the assisted call, and determining the notification based on the state of the client device includes determining the notification to include an auditory component that is audibly rendered by one or more speakers of the client device based on the state of the client device indicating that the given user is not actively monitoring the assisted call.

[0134] In some implementations, the method further includes terminating the assisted call and, after terminating the assisted call, causing the client device to render a further notification that includes an indication of the outcome of the assisted call. In some versions of these implementations, the further notification that includes the indication of the assisted call includes an indication of a further task to be performed on behalf of the user in response to terminating the assisted call, or includes one or more selectable graphical elements that, when selected, cause the client device to perform a further task on behalf of the user.

[0135] In some implementations, the method further includes, in response to determining that no further user input is received within the threshold duration, terminating the assisted call, and, after terminating the assisted call, causing the client device to render a further notification including an indication of the outcome of the assisted call.

[0136] In some implementations, the threshold duration is a fixed duration from when a notification identifying the further parameters and requesting further user input regarding the information is rendered, or a dynamic duration based on when the last one or more of the corresponding values ​​is rendered for presentation to a further user via a further client device.

[0137] In some implementations, the method further includes, after initiating the assisted call, obtaining consent from a further user associated with the entity to monitor the assisted call.

[0138] In some implementations, a method implemented by one or more processors is provided, including detecting, at a client device, an ongoing call between a given user of the client device and an additional user of an additional client device; and processing a stream of audio data capturing at least one spoken utterance during the ongoing call to generate recognized text. The at least one spoken utterance is of the given user or the additional user. The method further includes identifying, based on processing the recognized text, that the at least one spoken utterance requests parameter information, and determining, using personal access-limited data of the given user with respect to the parameter, that a value for the parameter is resolvable. The method further includes, in response to determining that the value is resolvable, rendering an output based on the value during the ongoing call.

[0139] These and other implementations of the technology disclosed herein may optionally include one or more of the following features.

[0140] In some implementations, the method further includes resolving a value for the parameter. In some versions of those implementations, rendering the output during the ongoing call is further responsive to resolving the value of the parameter. In some further versions of those implementations, resolving the value of the parameter includes analyzing metadata of the ongoing call between the given user and an additional user, identifying an entity associated with the additional user based on the analysis, and resolving the value based on the value stored in association with the entity and the parameter.

[0141] In some implementations, the output includes synthesized speech, and rendering the output based on the value during the ongoing call includes rendering the synthesized speech as part of the ongoing call. In some versions of those implementations, the method further includes receiving user input from a given user to activate assistance during the ongoing call before rendering the synthesized speech as part of the ongoing call. In some versions of those implementations, rendering the synthesized speech as part of the ongoing call is further responsive to receiving user input to activate assistance.

[0142] In some implementations, the output includes a notification rendered at the client device outside the ongoing call. In some versions of those implementations, the output further includes synthesized speech rendered as part of the ongoing call, and the method further includes, after rendering the notification, rendering the synthesized speech in response to receiving affirmative user input in response to the notification.

[0143] In some implementations, the method further includes determining, based on processing the stream of audio data for a threshold duration after at least one spoken utterance requesting information for the parameter, whether any further spoken utterances of the given user received within the threshold duration include the value. In some versions of these implementations, providing the output is conditioned on determining that any further spoken utterances of the given user received within the threshold duration do not include the value.

[0144] Additionally, some implementations include one or more processors (e.g., central processing units (CPUs), graphics processing units (GPUs), and / or tensor processing units (TPUs)) of one or more computing devices, the one or more processors operable to execute instructions stored in associated memory, the instructions configured to cause performance of any of the above-described methods. Some implementations also include one or more non-transitory computer-readable storage media that store computer instructions that can be executed by the one or more processors to perform any of the above-described methods. Some implementations also include computer program products that include instructions that can be executed by the one or more processors to perform any of the above-described methods.

[0145] It is understood that all combinations of the above concepts, and additional concepts described in more detail herein, are considered to be part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. [Explanation of symbols]

[0146] 110 client devices 111 User Input Engine 112 Device State Engine 113 Rendering Engine 114 Scheduling Engine 115 Automated Assistants 120A1 voice recognition engine 120A2 voice recognition engine 130A1 NLU engine 130A2 NLU engine 140A1 speech synthesis engine 140A2 speech synthesis engine 150 assisted call engine 151 Entity Identification Engine 151A Entity Database 152 Task Decision Engine 153 Parameter Engine 153A Parameter Database 153B User Profile Database 154 Task Execution Engine 155 Feedback Engine 156 Recommendation Engine 180 Assisted Call System 190 Network 200 ways 410 Client Device 411 URL 420 first search result 421 Call Graphical Elements 422 Directions Graphical Elements 423 Menu Graphical Elements 430 Second Search Result 431 Call Graphical Elements 432 Directions Graphical Elements 433 Menu Graphical Elements 441B Editing Graphical Elements 441C End Call Graphical Element 442B Cancel Graphical Elements 442C Call Join Graphical Elements 443B Call Interface Elements 443C Speaker Interface Elements 452B1 Prompt 452B2 prompt 452C1 Synthetic Speech 452C3 Synthetic Speech 452D1 Synthetic Speech 452D2 Synthetic Voice 454B1 User Input 454C3 Synthetic Speech 456C1 Audio Data 456C2 Audio Data 456D1 Audio Data 456D2 Audio Data 470 Call Detail Interface 471 First Graphical Element 471A First Subelement 471B Second Subelement 471C Third Subelement 472 Secondary Graphical Element 473 Third Graphical Element 474 Name Parameter 474A Value 475 Phone Number Parameter 475A value 476 Date / Time Parameters 476A value 477 Number of people parameter Value 477A 478 Seat Type Parameters 478A value 479 Notification 479A First Proposal 479B Second Proposal 480 Graphical User Interface 481 System Interface Elements 482 System Interface Elements 483 System Interface Elements 484 Text-Responsive Interface Elements 485 Voice Response Interface Elements 486 Call Details Interface Elements 510 client device 542 Graphical Elements 543 Graphical Elements 552A1 Audio Data 552A2 Audio Data 552B1 Audio Data 552B2 Audio Data 552C1 Audio Data 552C2 Audio Data 554A1 Audio Data 554A2 Audio Data 554B1 Audio Data 554C1 Audio Data 556A1 Synthetic Voice 556B1 Synthetic Voice 556C1 Synthetic Speech 570 Call Details Interface 579 Notifications 579A Graphical Elements 580 Graphical User Interface 581 System Interface Elements 582 System Interface Elements 583 System Interface Elements 584 Text-Responsive Interface Elements 585 Voice Response Interface Elements 586 Call Details Interface Elements 610 Computing Devices 612 Bus Subsystem 614 processor 616 Network Interface Subsystem 620 User Interface Output Device 622 User Interface Input Devices 624 Storage Subsystem 625 Memory Subsystem 626 File Storage Subsystem 630 Main Random Access Memory (RAM) 632 Read-Only Memory (ROM)

Claims

1. 1. A method implemented by one or more processors, comprising: detecting, at a client device, an ongoing call between a given user of said client device and a further user of a further client device; processing a stream of audio data capturing at least one spoken utterance during the ongoing call to generate recognized text, wherein the at least one spoken utterance is of the given user or the further user; identifying the at least one spoken utterance as including a request for parametric information based on processing the recognized text; determining that a value for said parameter can be determined using personal, limited-access data of said given user; in response to determining that the value can be determined, without receiving user input from the given user or the further user; automatically determining the values ​​of the parameters, analyzing metadata of the ongoing call between the given user and the further user; Identifying entities associated with the further user based on the analysis; and determining that the stored value is the value of the parameter based on the value being stored in a database in association with the entity and the parameter; and automatically rendering an output corresponding to said value during said ongoing call; A method comprising:

2. the output comprises synthesized speech, and automatically rendering the output corresponding to the value during the ongoing call; The method of claim 1 , comprising rendering the synthesized speech as part of the ongoing call.

3. before rendering the synthesized speech as part of the ongoing call; receiving a user input from the given user to activate assistance during the ongoing call; 3. The method of claim 2, wherein rendering the synthetic speech as part of the ongoing call further comprises rendering the synthetic speech as part of the ongoing call in response to receiving the user input to activate the assistance.

4. The method of claim 1 , wherein the output comprises a notification that is rendered at the client device outside the ongoing call.

5. wherein the output further comprises synthesized speech rendered as part of the ongoing call, and the method further comprises: The method of claim 4 , further comprising the step of, after rendering the notification, rendering the synthesized speech in response to receiving affirmative user input in response to the notification.

6. further comprising determining, based on processing the stream of audio data for a threshold duration after the at least one spoken utterance requesting information for the parameter, whether any further spoken utterances of the given user received within the threshold duration include the value; 6. The method of claim 1, wherein providing the output is conditioned on determining that further spoken utterances of the given user received within the threshold duration do not include the value.

7. at least one processor; at least one memory storing instructions that, when executed, cause said at least one processor to perform the method of any one of claims 1 to 6; At least one computing device, including:

8. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause at least one processor to perform operations corresponding to the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Simple automated system for inspecting abnormality of plane

    JP1988018255A

  • Electronic device, control method of the same, control program, and recording medium

    JP2018101847A

  • Escalation to a human operator

    JP2019522914A

  • Operation method of dialog agent and apparatus thereof

    JP2020034914A

  • Handling calls on a shared speech-enabled device

    US20180337962A1