Digital assistant interaction in communication session

The digital assistant system improves interaction efficiency and reduces user inputs by initiating tasks based on the most recently invoked assistant's language, addressing language diversity in communication sessions and enhancing usability and battery life.

JP2025114525APending Publication Date: 2025-08-05APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025032971
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-18
Filing Date
2025-03-03
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing digital assistant technologies lack efficient mechanisms for interacting with multiple participants in a communication session, particularly when different digital assistants operate in different languages, leading to confusion and increased user inputs, which affects usability and battery life.

Method used

A method and system that enables a digital assistant on an electronic device to receive user inputs, generate prompts for further input, transmit these prompts to external devices, and initiate tasks based on responses, providing feedback and reducing the need for additional user inputs by utilizing the most recently invoked digital assistant's language.

Benefits of technology

Enhances user-device interaction efficiency by reducing user inputs, minimizing errors, and conserving battery life by allowing participants to confirm task initiation using the correct participant's information, even when multiple digital assistants with different languages are involved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114525000001_ABST
    Figure 2025114525000001_ABST
Patent Text Reader

Abstract

To realize an intelligent automated assistant.SOLUTION: A process includes: receiving input for invoking a first digital assistant from a first user; receiving natural language input corresponding to a task from the first user; generating a prompt for further user input relating to the task by the first digital assistant according to invoking the first digital assistant; transmitting the prompt for further user input relating to the task to an external device; after transmitting the prompt for further user input, receiving a response for the prompt for further user input from the external device; starting the task on the basis of the response and information corresponding to the first user stored in an electronic device by the first digital assistant; and transmitting output indicating the started task to the external device.SELECTED DRAWING: Figure 8B
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 63 / 233,001, entitled "DIGITAL ASSISTANT INTERACTION IN A COMMUNICATION SESSION," filed August 13, 2021, and U.S. Non-Provisional Application No. 17 / 866,984, entitled "DIGITAL ASSISTANT INTERACTION IN A COMMUNICATION SESSION," filed July 18, 2022, the contents of which are incorporated herein by reference for all purposes.

[0002] The present invention relates generally to intelligent automated assistants, and more particularly to interacting with an intelligent automated assistant in a communication session. [Background technology]

[0003] Intelligent automated assistants (or digital assistants) can provide a useful interface between human users and electronic devices. Such assistants can enable users to interact with devices or systems using natural language, either verbally and / or in text form. For example, a user can provide speech input containing a user request to a digital assistant running on an electronic device. The digital assistant can interpret the user's intent from the speech input and activate the user's intent into a task. The task can then be performed by executing one or more services on the electronic device and can return an associated output response to the user request to the user. Summary of the Invention

[0004] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having one or more processors and memory, receiving an input from a first user of the electronic device to invoke a first digital assistant running on the electronic device while the electronic device is engaged in a communication session with one or more external devices, receiving a natural language input corresponding to a task from the first user, and in response to invoking the first digital assistant, generating a prompt for further user input related to the task by the first digital assistant, transmitting the prompt for further user input related to the task to one or more external devices, receiving a response to the prompt for further user input from an external device of the one or more external devices after transmitting the prompt for further user input, and starting the task by the first digital assistant based on the response and information corresponding to the first user stored on the electronic device, and transmitting an output indicating the started task to the one or more external devices.

[0005] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. When executed by one or more processors of the electronic device, the one or more programs cause the electronic device to receive, from a first user of the electronic device, an input to invoke a first digital assistant running on the electronic device, receive from the first user a natural language input corresponding to a task, and, in accordance with invoking the first digital assistant, generate a prompt for further user input related to the task by the first digital assistant, transmit the prompt for further user input related to the task to one or more external devices, and after transmitting the prompt for further user input, receive a response to the prompt for further user input from an external device of the one or more external devices, and, based on the response and information corresponding to the first user stored on the electronic device, start a task by the first digital assistant, and transmit an output indicating the started task to the one or more external devices.

[0006] An exemplary electronic device is disclosed herein. The exemplary electronic device includes one or more processors, a memory, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: receiving, while the electronic device is engaged in a communication session with one or more external devices, an input from a first user of the electronic device to invoke a first digital assistant running on the electronic device; receiving from the first user a natural language input corresponding to a task; generating, by the first digital assistant, a prompt for further user input related to the task in accordance with invoking the first digital assistant; transmitting the prompt for further user input related to the task to one or more external devices; receiving, after transmitting the prompt for further user input, a response to the prompt for further user input from an external device of the one or more external devices; starting, by the first digital assistant, a task based on the response and information corresponding to the first user stored on the electronic device; and transmitting an output indicating the started task to the one or more external devices.

[0007] An exemplary electronic device includes means for receiving, from a first user of the electronic device while the electronic device is engaged in a communication session with one or more external devices, an input to invoke a first digital assistant running on the electronic device; receiving from the first user natural language input corresponding to a task; generating, by the first digital assistant, a prompt for further user input related to the task in accordance with invoking the first digital assistant; transmitting the prompt for further user input related to the task to one or more external devices; receiving, after transmitting the prompt for further user input, a response to the prompt for further user input from an external device of the one or more external devices; initiating, by the first digital assistant, the task based on the response and information corresponding to the first user stored on the electronic device; and transmitting an output indicating the initiated task to the one or more external devices.

[0008] Starting a task in the above manner and transmitting an output indicating the started task provides participants in the communication session with feedback that any participant can respond to a prompt for further user input and that the task has started. Thus, by allowing any participant, such as a participant with a correct and / or optimal response to a prompt, to cause the digital assistant to correctly start the task, user-device interaction becomes more efficient and flexible. Providing participants with improved feedback improves device usability and makes the user-device interface more efficient (e.g., by reducing the number of user inputs required by the device to perform a task, by helping the user provide appropriate inputs, and by reducing user errors when interacting with the device), and further improves device battery life by reducing power usage and allowing users to use the device more quickly and efficiently.

[0009] Furthermore, by initiating a task based on information corresponding to the first user and transmitting an output, participants in the communication session are provided with feedback that the digital assistant initiated the task based on information corresponding to the participant who most recently invoked the digital assistant, even though another participant responded to the digital assistant's prompt. Thus, by avoiding confusion about which information the digital assistant uses to initiate a task and by allowing participants to confirm that the task will be initiated using the correct participant's information (or by notifying participants that the task will be initiated using incorrect participant's information so that the participant can provide corrective input), more consistent and efficient user-device interaction can be provided. Providing improved feedback to participants and automatically initiating a task based on the first user's information (e.g., without requiring further user input after the device receives a response from an external device) improves device usability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when interacting with the device), and further improves device battery life by reducing power usage and allowing users to use the device more quickly and efficiently.

[0010] An exemplary method is disclosed herein in an electronic device having one or more processors and memory, the exemplary method including receiving, while the electronic device is engaged in a communication session with one or more external devices, an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; generating, by the first digital assistant in response to the first natural language input in the first language in accordance with the invocation of the first digital assistant; and generating the first response as a single response. and transmitting the first response to the external device; and after transmitting the first response, receiving a second natural language input in a second language from an external device of the one or more external devices, the second natural language input being received by the external device after receiving the transmitted first response without the external device receiving a second input for invoking a second digital assistant running on the external device, the second digital assistant being configured to operate in the second language; and receiving a second response in the second language to the second natural language input from the external device, the second response being generated by the second digital assistant.

[0011] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive, from a first user of the electronic device, an input for invoking a first digital assistant operating on the electronic device while the electronic device is engaged in a communication session with one or more external devices, the first digital assistant being configured to operate in a first language; receive, from the first user, a first natural language input in the first language; and, pursuant to invoking the first digital assistant, perform a response in the first language to the first natural language input by the first digital assistant. and generating a first response in the second language from an external device of the one or more external devices after transmitting the first response, the second natural language input being received by the external device after receiving the transmitted first response without the external device receiving a second input for invoking a second digital assistant running on the external device, the second digital assistant being configured to operate in the second language; and receiving from the external device a second response in the second language to the second natural language input, the second response being generated by the second digital assistant.

[0012] An exemplary electronic device is disclosed herein, the exemplary electronic device comprising one or more processors, a memory, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for receiving, from a first user of the electronic device, an input to invoke a first digital assistant running on the electronic device while the electronic device is engaged in a communication session with one or more external devices, the first digital assistant being configured to operate in a first language; receiving, from the first user, a first natural language input in the first language; and executing, in accordance with the invocation of the first digital assistant, the first digital assistant. generating a first response in the first language to the first natural language input by the digital assistant, transmitting the first response to one or more external devices, and after transmitting the first response, receiving a second natural language input in a second language from an external device of the one or more external devices, the second natural language input being received by the external device after receiving the transmitted first response without the external device receiving a second input for invoking a second digital assistant running on the external device, the second digital assistant being configured to operate in the second language; and receiving a second response in the second language to the second natural language input from the external device, the second response being generated by the second digital assistant.

[0013] An exemplary electronic device includes: receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language, while the electronic device is engaged in a communication session with one or more external devices; receiving a first natural language input in the first language from the first user; generating, by the first digital assistant, a first response in the first language to the first natural language input pursuant to the invocation of the first digital assistant; and communicating the first response to the one or more external devices. and after transmitting the first response, receiving a second natural language input in a second language from an external device of the one or more external devices, the second natural language input being received by the external device after receiving the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device, the second digital assistant being configured to operate in the second language; and means for receiving a second response in the second language to the second natural language input from the external device, the second response being generated by the second digital assistant.

[0014] Receiving the second response generated by the second digital assistant in the second language provides participants in the communication session with feedback that any participant can successfully issue a follow-up request (e.g., a second natural language input) to the digital assistant in the language in which their digital assistant is configured to operate. Thus, participants are informed of the correct language for interacting with the digital assistant within the communication session, even if, for example, multiple digital assistants within the communication session are each configured to operate in different languages. Additionally, the second digital assistant may generate a second response to the follow-up request without requiring input to invoke the second digital assistant (e.g., after receiving the transmitted first response), thereby reducing the number of user inputs required to successfully perform the task. Providing participants with improved feedback and reducing the number of inputs required for the digital assistant to perform the task improves device usability and makes the user-device interface more efficient (e.g., by helping users provide appropriate inputs and reducing user errors when interacting with the device), as well as improving device battery life by reducing power usage and allowing users to use the device more quickly and efficiently.

[0015] An exemplary method is disclosed herein in an electronic device having one or more processors and memory, the exemplary method including receiving, while the electronic device is engaged in a communication session with one or more external devices, an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; generating, by the first digital assistant in response to the first natural language input in the first language in accordance with invoking the first digital assistant; transmitting a response to the one or more external devices; after transmitting the first response, receiving a second natural language input in the first language from an external device of the one or more external devices, the second natural language input being received by the external device after receiving the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device, the second digital assistant being configured to operate in the second language; generating, by the first digital assistant, a second response in the first language to the second natural language input; and transmitting the second response to the one or more external devices.

[0016] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive, from a first user of the electronic device, an input for invoking a first digital assistant running on the electronic device while the electronic device is engaged in a communication session with one or more external devices, the first digital assistant being configured to operate in a first language; receive, from the first user, a first natural language input in the first language; and, in response to the invocation of the first digital assistant, perform a first natural language response in the first language by the first digital assistant. generating a first response, transmitting the first response to one or more external devices; receiving, after transmitting the first response, a second natural language input in the first language from an external device of the one or more external devices, the second natural language input being received by the external device after receiving the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device, the second digital assistant being configured to operate in the second language; generating, by the first digital assistant, a second response in the first language to the second natural language input; and transmitting the second response to the one or more external devices.

[0017] An exemplary electronic device is disclosed herein, the exemplary electronic device comprising one or more processors, a memory, and one or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions including instructions for receiving an input from a first user of the electronic device to invoke a first digital assistant running on the electronic device, the first digital assistant being configured to operate in a first language, receiving a first natural language input in the first language from the first user, and, pursuant to invoking the first digital assistant, providing a first natural language input in the first language by the first digital assistant. generating a first response in the first language to the word input, transmitting the first response to one or more external devices, and after transmitting the first response, receiving a second natural language input in the first language from an external device of the one or more external devices, the second natural language input being received by the external device after receiving the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device, the second digital assistant being configured to operate in the second language; generating, by the first digital assistant, a second response in the first language to the second natural language input; and transmitting the second response to the one or more external devices.

[0018] An exemplary electronic device includes: receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant configured to operate in a first language; receiving a first natural language input in the first language from the first user; generating, by the first digital assistant, a first response in the first language to the first natural language input in accordance with invoking the first digital assistant; transmitting the first response to one or more external devices; and transmitting the first response to one or more external devices after transmitting the first response. The system includes means for receiving a second natural language input in a first language from one of the above external devices, the second natural language input being received by the external device after receiving the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device, the second digital assistant being configured to operate in the second language, and means for receiving the second natural language input and generating, by the first digital assistant, a second response in the first language to the second natural language input, and transmitting the second response to one or more external devices.

[0019] Transmitting the second response generated by the first digital assistant in the first language provides participants in the communication session with feedback that any participant can successfully issue a follow-up request (e.g., a second natural language input) to the digital assistant in the language in which the most recently invoked digital assistant is configured to operate. Thus, participants are informed of the correct language for interacting with the digital assistant within the communication session, even if, for example, multiple digital assistants within the communication session are each configured to operate in different languages. Additionally, the first digital assistant may generate a second response to the follow-up request without requiring input to invoke the digital assistant (e.g., after the external device receives the transmitted first response), thereby reducing the number of user inputs required to successfully perform the task. Providing participants with improved feedback and reducing the number of inputs required for the digital assistant to perform the task improves device usability and makes the user-device interface more efficient (e.g., by helping users provide appropriate inputs and reducing user errors when interacting with the device), as well as improving device battery life by reducing power usage and allowing users to use the device more quickly and efficiently. [Brief explanation of the drawings]

[0020] [Figure 1] FIG. 1 is a block diagram illustrating a system and environment for implementing a digital assistant, according to various embodiments. [Figure 2A] FIG. 1 is a block diagram illustrating a portable multifunction device running a client-side portion of a digital assistant, according to various embodiments. [Figure 2B] FIG. 2 is a block diagram illustrating example components for event processing, in accordance with various embodiments. [Figure 3] FIG. 1 illustrates a portable multifunction device running a client-side portion of a digital assistant, according to various embodiments. [Figure 4] FIG. 1 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface in accordance with various embodiments. [Figure 5A] 1A-1C illustrate exemplary user interfaces for a menu of applications on a portable multifunction device in accordance with various embodiments. [Figure 5B] 1A-1C illustrate exemplary user interfaces for a multifunction device having a touch-sensitive surface separate from the display, in accordance with various embodiments. [Figure 6A] FIG. 1 illustrates a personal electronic device according to various embodiments. [Figure 6B] FIG. 1 is a block diagram illustrating a personal electronic device according to various embodiments. [Figure 7A] FIG. 1 is a block diagram illustrating a digital assistant system or a server portion thereof, according to various embodiments. [Figure 7B] 7B illustrates the functionality of the digital assistant shown in FIG. 7A, according to various embodiments. [Figure 7C] FIG. 2 illustrates a portion of an ontology, according to various embodiments. [Figure 8A] 1 illustrates systems and techniques for digital assistant interaction in a communication session, according to various embodiments. [Figure 8B] 1 illustrates systems and techniques for digital assistant interaction in a communication session, according to various embodiments. [Figure 8C] 1 illustrates systems and techniques for digital assistant interaction in a communication session, according to various embodiments. [Figure 8D] 1 illustrates systems and techniques for digital assistant interaction in a communication session, according to various embodiments. [Figure 8E] 1 illustrates systems and techniques for digital assistant interaction in a communication session, according to various embodiments. [Figure 8F]1 illustrates systems and techniques for digital assistant interaction in a communication session, according to various embodiments. [Figure 8G] 1 illustrates systems and techniques for digital assistant interaction in a communication session, according to various embodiments. [Figure 8H] 1 illustrates systems and techniques for digital assistant interaction in a communication session, according to various embodiments. [Figure 9A] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 9B] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 9C] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 9D] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 9E] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 9F] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 10A] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 10B]1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 10C] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 10D] 1 illustrates techniques for digital assistant interaction in a communication session when the DAs of participants in the communication session are configured to operate in different languages, according to various embodiments. [Figure 11A] 1 illustrates a process for digital assistant interaction in a communication session, according to various embodiments. [Figure 11B] 1 illustrates a process for digital assistant interaction in a communication session, according to various embodiments. [Figure 11C] 1 illustrates a process for digital assistant interaction in a communication session, according to various embodiments. [Figure 12A] 1 illustrates a process for digital assistant interaction in a communication session, according to various embodiments. [Figure 12B] 1 illustrates a process for digital assistant interaction in a communication session, according to various embodiments. [Figure 13A] 1 illustrates a process for digital assistant interaction in a communication session, according to various embodiments. [Figure 13B] 1 illustrates a process for digital assistant interaction in a communication session, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0021] In the following description of the embodiments, reference is made to the accompanying drawings, which show, by way of illustration, specific embodiments that may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the various embodiments.

[0022] In the following description, terms such as "first" and "second" are used to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first input can be referred to as a second input, and similarly, a second input can be referred to as a first input, without departing from the scope of various embodiments described. The first input and the second input are both inputs, and in some cases, are separate and distinct inputs.

[0023] The terminology used in the description of the various embodiments set forth herein is for the purpose of describing particular embodiments only and is not intended to be limiting. When used in the description of the various embodiments set forth and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" should be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes," "including," "comprises," and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0024] The term "if" can be interpreted to mean "when" or "upon," or "in response to determining," or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "upon determining," or "in response to determining," or "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]," depending on the context. 1. System and Environment

[0025] FIG. 1 illustrates a block diagram of a system 100 according to various embodiments. In some embodiments, the system 100 implements a digital assistant. The terms “digital assistant,” “virtual assistant,” “intelligent automated assistant,” or “automated digital assistant” refer to any information processing system that infers user intent by interpreting spoken and / or textual natural language input and performs actions based on the inferred user intent. For example, to act based on the inferred user intent, the system performs one or more of the following: identifies a task flow having steps and parameters designed to fulfill the inferred user intent, inputs specific requirements from the inferred user intent into the task flow, executes the task flow by invoking programs, methods, services, APIs, or the like, and generates an output response to the user in an audible (e.g., spoken) and / or visual form.

[0026] Specifically, a digital assistant can accept user requests, at least in part, in the form of natural language commands, requests, opinions, discourse, and / or inquiries. Typically, a user request seeks either an informational answer or task performance by the digital assistant. A satisfactory response to a user request includes providing the requested informational answer, performing the requested task, or a combination of the two. For example, a user asks a digital assistant a question such as, "Where am I right now?" Based on the user's current location, the digital assistant responds, "You're in Central Park near the West Gate." The user also requests a task to be performed, such as, "Please invite my friends to my girlfriend's birthday party next week." In response, the digital assistant can acknowledge the request by stating, "Yes, right now," and then send appropriate calendar invitations on behalf of the user to each of the user's friends listed in the user's electronic address book. During the performance of a requested task, the digital assistant may interact with the user in a continuous conversation involving multiple information exchanges over time. There are many other ways to interact with a digital assistant to request information or to perform various tasks. In addition to providing verbal responses and taking programmed actions, digital assistants also provide responses in other visual or audio formats, such as text, alerts, music, videos, animations, etc.

[0027] 1, in some embodiments, the digital assistant is implemented according to a client-server model. The digital assistant includes a client-side portion 102 (hereinafter, "DA client 102") that runs on a user device 104 and a server-side portion 106 (hereinafter, "DA server 106") that runs on a server system 108. The DA client 102 communicates with the DA server 106 over one or more networks 110. The DA client 102 provides client-side functionality, such as user-responsive input and output processing and communication with the DA server 106. The DA server 106 provides server-side functionality to any number of DA clients 102, each residing on a separate user device 104.

[0028] In some embodiments, the DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, and an I / O interface to external services 118. The client-facing I / O interface 112 facilitates client-facing input and output processing of the DA server 106. The one or more processing modules 114 utilize the data and models 116 to process speech input and determine user intent based on natural language input. Furthermore, the one or more processing modules 114 perform task execution based on the inferred user intent. In some embodiments, the DA server 106 communicates with external services 120 over network(s) 110 to complete tasks or obtain information. The I / O interface to external services 118 facilitates such communication.

[0029] User device 104 can be any suitable electronic device. In some examples, user device 104 is a portable multifunction device (e.g., device 200 described below in connection with FIG. 2A), a multifunction device (e.g., device 400 described below in connection with FIG. 4), or a personal electronic device (e.g., device 600 described below in connection with FIGS. 6A-6B). A portable multifunction device is, for example, a mobile phone that also includes other functions, such as PDA and / or music player functionality. Specific examples of portable multifunction devices include the Apple Watch®, iPhone®, iPod Touch®, and iPad® devices by Apple Inc. (Cupertino, California). Other examples of portable multifunction devices include, but are not limited to, earphones / headphones, speakers, and laptop or tablet computers. Furthermore, in some examples, user device 104 is a non-portable multifunction device. Specifically, user device 104 is a desktop computer, a game console, a speaker, a television, or a television set-top box. In some examples, user device 104 includes a touch-sensitive surface (e.g., a touchscreen display and / or a touchpad). Additionally, user device 104 optionally includes one or more other physical user interface devices, such as a physical keyboard, a mouse, and / or a joystick. Various examples of electronic devices, such as multifunction devices, are described in further detail below.

[0030] Examples of communication network(s) 110 include a local area network (LAN) and a wide area network (WAN), such as the Internet. Communication network(s) 110 may be implemented using any known network protocol, including various wired or wireless protocols, such as, for example, Ethernet, Universal Serial Bus (USB), FIREWIRE, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth®, Wi-Fi®, voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.

[0031] The server system 108 may be implemented on one or more standalone data processing devices or on a distributed computer network. In some embodiments, the server system 108 may also employ various virtual devices and / or the services of third-party service providers (e.g., third-party cloud service providers) to provide the underlying computing and / or infrastructure resources of the server system 108.

[0032] In some embodiments, the user device 104 communicates with the DA server 106 via a second user device 122. The second user device 122 is similar to or identical to the user device 104. For example, the second user device 122 is similar to devices 200, 400, or 600 described below in connection with FIGS. 2A, 4, and 6A-6B. The user device 104 is configured to be communicatively coupled to the second user device 122 via a direct communication connection, such as Bluetooth, NFC, or BTLE, or via a wired or wireless network, such as a local Wi-Fi network. In some embodiments, the second user device 122 is configured to act as a proxy between the user device 104 and the DA server 106. For example, the DA client 102 of the user device 104 is configured to send information (e.g., a user request received at the user device 104) to the DA server 106 via the second user device 122. The DA server 106 processes the information and returns relevant data (eg, data content responsive to the user request) to the user device 104 via the second user device 122.

[0033] In some embodiments, the user device 104 is configured to reduce the amount of information transmitted from the user device 104 by communicating with the second user device 122 via an abbreviated request for data. The second user device 122 is configured to determine supplemental information to add to the abbreviated request and generate a complete request to send to the DA server 106. This system architecture can advantageously allow a user device 104 (e.g., a watch or similar small electronic device) with limited communication capabilities and / or limited battery power to access services provided by the DA server 106 by using a second user device 122 (e.g., a mobile phone, laptop computer, tablet computer, etc.) with greater communication capabilities and / or battery power as a proxy to the DA server 106. While only two user devices 104 and 122 are shown in FIG. 1 , it should be understood that the system 100, in some embodiments, includes any number and type of user devices configured to communicate with the DA server system 106 in this proxy configuration.

[0034] 1 includes both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106), but in some embodiments, the digital assistant's functionality is implemented as an independent application installed on a user device. Furthermore, the allocation of functionality between the client and server portions of the digital assistant may vary depending on the implementation. For example, in some embodiments, the DA client is a thin client that provides only user-facing input and output processing functionality and delegates all other digital assistant functionality to a back-end server. 2. Electronic Devices

[0035] Attention now turns to embodiments of electronic devices for executing the client-side portion of a digital assistant. FIG. 2A is a block diagram illustrating portable multifunction device 200 with touch-sensitive display system 212, according to some embodiments. Touch-sensitive display 212 may conveniently be referred to as a "touch screen" and may also be known or referred to as a "touch-sensitive display system." Device 200 includes memory 202 (optionally including one or more computer-readable storage media), a memory controller 222, one or more processing units (CPUs) 220, a peripherals interface 218, RF circuitry 208, audio circuitry 210, a speaker 211, a microphone 213, an input / output (I / O) subsystem 206, other input control devices 216, and an external port 224. Device 200 optionally includes one or more optical sensors 264. Device 200 optionally includes one or more contact intensity sensors 265 that detect the intensity of a contact on device 200 (e.g., a touch-sensitive surface such as touch-sensitive display system 212 of device 200). Device 200 optionally includes one or more tactile output generators 267 that generate a tactile output on device 200 (e.g., generate a tactile output on a touch-sensitive surface such as touch-sensitive display system 212 of device 200 or touchpad 455 of device 400). These components optionally communicate via one or more communication buses or signal lines 203.

[0036] As used herein and in the claims, the term “intensity” of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or a proxy for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four distinct values and more typically includes hundreds (e.g., at least 256) distinct values. The intensity of a contact is optionally determined (or measured) using various techniques and various sensors or combinations of sensors. For example, one or more force sensors under or adjacent to the touch-sensitive surface are optionally used to measure force at various points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine an estimated force of the contact. Similarly, a pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change in the contact area detected on the touch-sensitive surface, the capacitance and / or change in the capacitance of the touch-sensitive surface proximate the contact, and / or the resistance and / or change in the capacitance of the touch-sensitive surface proximate the contact are optionally used as a surrogate for the force or pressure of the contact on the touch-sensitive surface. In some implementations, the surrogate measure of the force or pressure of the contact is used directly to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measure). In some implementations, the surrogate measure of the contact force or pressure is converted to an estimate of the force or pressure, and the estimate of the force or pressure is used to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using contact intensity as an attribute of user input allows users to access additional device functionality (e.g., on a touch-sensitive display) and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls such as knobs or buttons) that may not otherwise be accessible to users on devices of reduced size that have limited footprint for displaying affordances.

[0037] As used herein and in the claims, the term “tactile output” refers to a physical displacement of a device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device, that will be detected by a user with the user's sense of touch. For example, in a situation where a device or a component of a device is in contact with a touch-sensitive surface of a user (e.g., the fingers, palm, or other part of the user's hand), the tactile output produced by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in a physical property of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is optionally interpreted by the user as a “downclick” or “upclick” of a physical actuator button. In some cases, a user feels a tactile sensation such as a “downclick” or “upclick” even when there is no movement of a physical actuator button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's action. As another example, movement of a touch-sensitive surface is optionally interpreted or perceived by a user as "roughness" of the touch-sensitive surface, even if there is no change in the smoothness of the touch-sensitive surface. While such user interpretation of touch depends on the user's personal sensory perception, there are many sensory perceptions of touch that are common to the majority of users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., "upclick," "downclick," "roughness"), unless otherwise specified, the generated tactile output corresponds to a physical displacement of the device, or a component of the device, that produces the described sensory perception for a typical (or average) user.

[0038] It should be understood that device 200 is only one example of a portable multifunction device, and that device 200 optionally has more or fewer components than those shown, optionally combines two or more components, or optionally has a different configuration or arrangement of its components. The various components shown in Figure 2A are implemented as hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.

[0039] Memory 202 includes one or more computer-readable storage media that are, for example, tangible and non-transitory. Memory 202 includes high-speed random-access memory and also includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 222 controls access to memory 202 by other components of device 200.

[0040] In some embodiments, the non-transitory computer-readable storage medium of memory 202 is used to store instructions (e.g., to perform aspects of the processes described below) for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other system capable of fetching instructions from the instruction execution system, apparatus, or device and executing the instructions. In other examples, the instructions (e.g., to perform aspects of the processes described below) are stored in a non-transitory computer-readable storage medium (not shown) of server system 108 or are split between the non-transitory computer-readable storage medium of memory 202 and the non-transitory computer-readable storage medium of server system 108.

[0041] Peripheral interface 218 is used to couple input and output peripherals of the device to CPU 220 and memory 202. One or more processors 220 operate or execute various software programs and / or instruction sets stored in memory 202 to perform various functions and process data for device 200. In some embodiments, peripheral interface 218, CPU 220, and memory controller 222 are implemented on a single chip, such as chip 204. In some other embodiments, they are implemented on separate chips.

[0042] RF (radio frequency) circuitry 208 transmits and receives RF signals, also called electromagnetic signals. RF circuitry 208 converts electrical signals to electromagnetic signals and vice versa, and communicates with communication networks and other communication devices via electromagnetic signals. RF circuitry 208 optionally includes well-known circuitry for performing these functions, including, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, etc. RF circuitry 208 optionally communicates via wireless communication with networks, such as the Internet, also known as the World Wide Web (WWW), an intranet, and / or wireless networks, such as cellular telephone networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs), and with other devices. RF circuitry 208 optionally includes well-known circuitry for detecting near field communication (NFC) fields, such as by short-range radios. Wireless communication is optionally supported by, but is not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPA), Long Term Evolution (LTE), and other standards.Wireless technology includes, but is not limited to, wireless technology such as LTE evolution, near field communications (NFC), wideband code division multiple access (W-CDMA), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet message access protocol (IMAP) and / or post office protocol (POP)), instant messaging (e.g., extensible messaging and presence protocol), and the like. The present invention may use any of a number of communication standards, protocols, and technologies, including the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (XMPP), the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), the Instant Messaging and Presence Service (IMPS), and / or the Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the filing date of this application.

[0043] The audio circuit 210, speaker 211, and microphone 213 provide an audio interface between a user and device 200. Audio circuit 210 receives audio data from peripherals interface 218, converts the audio data into electrical signals, and transmits the electrical signals to speaker 211. Speaker 211 converts the electrical signals into sound waves audible to humans. Audio circuit 210 also receives electrical signals converted from sound waves by microphone 213. Audio circuit 210 converts the electrical signals into audio data and transmits the audio data to peripherals interface 218 for processing. The audio data is retrieved from and / or transmitted to memory 202 and / or RF circuit 208 by peripherals interface 218. In some embodiments, audio circuit 210 also includes a headset jack (e.g., 312 in FIG. 3 ). The headset jack provides an interface between audio circuitry 210 and a detachable audio input / output peripheral such as an output-only headphone or a headset with both an output (e.g., single or double ear headphones) and an input (e.g., a microphone).

[0044] I / O subsystem 206 couples input / output peripherals on device 200, such as touchscreen 212 and other input control devices 216, to peripheral interface 218. I / O subsystem 206 optionally includes one or more input controllers 260 for display controller 256, optical sensor controller 258, intensity sensor controller 259, haptic feedback controller 261, and other input or control devices. One or more input controllers 260 receive / send electrical signals from / to other input control devices 216. Other input control devices 216 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some alternative embodiments, input controller(s) 260 are optionally coupled to any (or none) of a keyboard, infrared port, USB port, and pointer device such as a mouse. The one or more buttons (e.g., 308 in FIG. 3) optionally include up / down buttons for volume control of speaker 211 and / or microphone 213. The one or more buttons optionally include a push button (e.g., 306 in FIG. 3).

[0045] A quick press of a push button unlocks the touch screen 212 or initiates the process of using gestures on the touch screen to unlock the device, as described in U.S. Patent Application No. 11 / 322,549, filed December 23, 2005, and U.S. Patent No. 7,657,849, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," both of which are incorporated herein by reference in their entireties. A long press of a push button (e.g., 306) powers the device 200 on or off. The user can customize the functionality of one or more buttons. The touch screen 212 can be used to implement virtual or soft buttons and one or more soft keyboards.

[0046] The touch-sensitive display 212 provides an input and output interface between the device and a user. The display controller 256 receives and / or sends electrical signals to and from the touchscreen 212. The touchscreen 212 displays visual output to the user. The visual output includes graphics, text, icons, animation, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output corresponds to user interface objects.

[0047] Touchscreen 212 has a touch-sensitive surface, sensor, or set of sensors that accepts input from a user based on haptic and / or tactile contact. Touchscreen 212 and display controller 256 (along with any associated modules and / or instruction sets in memory 202) detects contacts (and any movement or cessation of contact) on touchscreen 212 and translates the detected contacts into interactions with user interface objects (e.g., one or more softkeys, icons, web pages, or images) displayed on touchscreen 212. In an exemplary embodiment, the point of contact between touchscreen 212 and the user corresponds to the user's finger.

[0048] Touchscreen 212 uses LCD (liquid crystal display), LPD (light emitting polymer display), or LED (light emitting diode) technology, although other display technologies may be used in other embodiments. Touchscreen 212 and display controller 256 detect contact and its movement or breaking using any of several now known or later developed touch sensing technologies, including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements that determine one or more points of contact using touchscreen 212. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that found in the iPhone® and iPod Touch® from Apple Inc. of Cupertino, California.

[0049] The touch-sensitive display of some embodiments of touchscreen 212 is similar to the multi-touch-sensing touchpad described in U.S. Patents 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman) and / or U.S. Patent Publication 2002 / 0015024 (A1), each of which is incorporated by reference in its entirety. However, touchscreen 212 displays visual output from device 200, whereas touch-sensitive touchpads do not provide visual output.

[0050] The touch-sensitive display in some embodiments of touchscreen 212 is described in the following applications: (1) U.S. patent application Ser. No. 11 / 381,313, filed May 2, 2006, entitled "Multipoint Touch Surface Controller"; (2) U.S. patent application Ser. No. 10 / 840,862, filed May 6, 2004, entitled "Multipoint Touchscreen"; (3) U.S. patent application Ser. No. 10 / 903,964, filed July 30, 2004, entitled "Gestures For Touch Sensitive Input Devices"; (4) U.S. patent application Ser. No. 11 / 048,264, filed January 31, 2005, entitled "Gestures For Touch Sensitive Input Devices"; and (5) U.S. patent application Ser. No. 11 / 038,590, filed January 18, 2005, entitled "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices." No. 11 / 228,758, filed September 16, 2005, entitled "Virtual Input Device Placement On A Touch Screen User Interface," (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, entitled "Operation Of A Computer With A Touch Screen Interface," (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, entitled "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard," and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, entitled "Multi-Functional Hand-Held Device," all of which are incorporated herein by reference in their entireties.

[0051] The touchscreen 212 has a video resolution of, for example, greater than 100 dpi. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. A user contacts the touchscreen 212 using a suitable object or accessory, such as a stylus, finger, or the like. In some embodiments, the user interface is designed to operate primarily using finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of a finger on the touchscreen. In some embodiments, the device translates the coarse finger input into precise pointer / cursor positions or commands to perform the action desired by the user.

[0052] In some embodiments, in addition to the touchscreen, device 200 includes a touchpad (not shown) for activating or deactivating certain functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touchscreen, does not display visual output. The touchpad may be a touch-sensitive surface separate from touchscreen 212 or may be an extension of the touch-sensitive surface formed by the touchscreen.

[0053] Device 200 also includes a power system 262 that provides power to the various components. Power system 262 includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a charging system, power failure detection circuitry, power converters or inverters, power status indicators (e.g., light emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in a portable device.

[0054] Device 200 also includes one or more optical sensors 264. FIG. 2A shows the optical sensor coupled to optical sensor controller 258 in I / O subsystem 206. Optical sensor 264 includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. Optical sensor 264 receives light from the environment projected through one or more lenses and converts the light into data representing an image. In conjunction with imaging module 243 (also referred to as a camera module), optical sensor 264 captures still images or video. In some embodiments, the optical sensor is located on the back of device 200, opposite touchscreen display 212 on the front of the device, so that the touchscreen display is used as a viewfinder for still image and / or video capture. In some embodiments, the optical sensor is located on the front of the device so that an image of the user for a video conference is captured while the user views other video conference participants on the touchscreen display. In some embodiments, the position of the optical sensor 264 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), so that a single optical sensor 264 is used for both video conferencing and capturing still images and / or video, along with a touchscreen display.

[0055] Device 200 also optionally includes one or more contact intensity sensors 265. FIG. 2A shows a contact intensity sensor coupled to intensity sensor controller 259 in I / O subsystem 206. Contact intensity sensor 265 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of a contact on a touch-sensitive surface). Contact intensity sensor 265 receives contact intensity information (e.g., pressure information or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed with or proximate to the touch-sensitive surface (e.g., touch-sensitive display system 212). In some embodiments, at least one contact intensity sensor is located on the back of device 200, opposite touchscreen display 212, which is located on the front of device 200.

[0056] Device 200 also includes one or more proximity sensors 266. Figure 2A shows proximity sensor 266 coupled to peripheral interface 218. Alternatively, proximity sensor 266 is coupled to input controller 260 within I / O subsystem 206. Proximity sensor 266 functions as described in U.S. patent application Ser. Nos. 11 / 241,839, "Proximity Detector In Handheld Device," 11 / 240,788, "Proximity Detector In Handheld Device," 11 / 620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output," 11 / 586,862, "Automated Response To And Sensing Of User Activity In Portable Devices," and 11 / 638,251, "Methods And Systems For Automatic Configuration Of Peripherals." In some embodiments, when the multifunction device is placed near the user's ear (eg, when the user is making a phone call), the proximity sensor turns off and disables touchscreen 212.

[0057] Device 200 also optionally includes one or more tactile output generators 267. FIG. 2A shows tactile output generators 267 coupled to haptic feedback controller 261 in I / O subsystem 206. Tactile output generator 267 optionally includes one or more electroacoustic devices, such as speakers or other audio components, and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components that convert electrical signals into tactile output on the device). Contact intensity sensor 265 receives tactile feedback generation instructions from haptic feedback module 233 and generates a tactile output on device 200 that can be sensed by a user of device 200. In some embodiments, at least one tactile output generator is juxtaposed with or proximate to a touch-sensitive surface (e.g., touch-sensitive display system 212) and, optionally, generates a tactile output by moving the touch-sensitive surface vertically (e.g., in / out of the surface of device 200) or horizontally (e.g., back and forth in the same plane as the surface of device 200). In some embodiments, at least one tactile output generator sensor is located on the back of device 200, opposite touchscreen display 212, which is located on the front of device 200.

[0058] Device 200 also includes one or more accelerometers 268. FIG. 2A shows accelerometer 268 coupled to peripherals interface 218. Alternatively, accelerometer 268 is coupled to input controller 260 within I / O subsystem 206. Accelerometer 268 operates as described, for example, in U.S. Patent Publication No. 20050190059, entitled "Acceleration-Based Theft Detection for Portable Electronic Devices," and U.S. Patent Publication No. 20060017692, entitled "Acceleration-Based Theft Detection System for Portable Electronic Devices," both of which are incorporated herein by reference in their entireties. In some embodiments, information is displayed on the touchscreen display in portrait or landscape orientation based on an analysis of data received from the one or more accelerometers. Device 200 optionally includes, in addition to accelerometer(s) 268, a magnetometer (not shown), and a GPS (or GLONASS or other global navigation system) receiver (not shown) for obtaining information regarding the location and orientation (e.g., portrait or landscape) of device 200.

[0059] In some embodiments, software components stored in memory 202 include operating system 226, communication module (or instruction set) 228, touch / motion module (or instruction set) 230, graphics module (or instruction set) 232, text input module (or instruction set) 234, Global Positioning System (GPS) module (or instruction set) 235, digital assistant client module 229, and applications (or instruction sets) 236. Additionally, memory 202 stores data and models, such as user data and models 231. Additionally, in some embodiments, as shown in FIGS. 2A and 4, memory 202 (FIG. 2A) or memory 470 (FIG. 4) stores device / global internal state 257. The device / global internal state 257 includes one or more of an active application state indicating which applications, if any, are currently active; a display state indicating which applications, views, or other information occupy various areas of the touchscreen display 212; a sensor state including information obtained from the device's various sensors and input control devices 216; and location information regarding the device's location and / or orientation.

[0060] Operating system 226 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers that control and manage general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware and software components.

[0061] Communications module 228 facilitates communication with other devices via one or more external ports 224 and also includes various software components for processing data received by RF circuitry 208 and / or external port 224. External port 224 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) is adapted to couple to other devices directly or indirectly via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as, similar to, and / or compatible with the 30-pin connector used on iPod® (trademark of Apple Inc.) devices.

[0062] Contact / motion module 230, optionally in conjunction with display controller 256, detects contact with touchscreen 212 and other touch-sensing devices (e.g., a touchpad or physical click wheel). Contact / motion module 230 includes various software components for performing various operations related to contact detection, such as determining whether contact occurs (e.g., detecting a finger-down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there is contact movement and tracking the movement across the touch-sensitive surface (e.g., detecting one or more finger-drag events), and determining whether the contact has ceased (e.g., detecting a finger-up event or an interruption of contact). Contact / motion module 230 receives contact data from the touch-sensitive surface. Determining the movement of the contact, as represented by the series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact. These actions are optionally applied to a single contact (e.g., a single finger contact) or multiple simultaneous contacts (e.g., "multi-touch" / multiple finger contacts). In some embodiments, contact / motion module 230 and display controller 256 detect contacts on the touchpad.

[0063] In some embodiments, contact / motion module 230 uses a set of one or more intensity thresholds to determine whether an action has been performed by a user (e.g., to determine whether a user has “clicked” on an icon). In some embodiments, at least a subset of the intensity thresholds are determined according to software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator, but can be adjusted without modifying the physical hardware of device 200). For example, the mouse “click” threshold of a trackpad or touchscreen display can be set to any of a wide range of predefined thresholds without modifying the trackpad or touchscreen display hardware. Additionally, in some implementations, a user of the device is provided with a software setting to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once via a system-level click “intensity” parameter).

[0064] Contact / motion module 230 optionally detects gesture input by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different movements, timing, and / or intensities of detected contacts). Thus, gestures are optionally detected by detecting particular contact patterns. For example, detecting a finger tap gesture includes detecting a finger down event, followed by detecting a finger up (lift off) event at the same position (or substantially the same position) as the finger down event (e.g., the position of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger down event, followed by one or more finger drag events, followed by detecting a finger up (lift off) event.

[0065] Graphics module 232 includes various known software components that render and display graphics on touchscreen 212 or other display, including components that modify the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual properties) of the displayed graphics. As used herein, the term "graphics" includes any object that can be displayed to a user, including, but not limited to, text, web pages, icons (e.g., user interface objects, including softkeys), digital images, video, animation, etc.

[0066] In some embodiments, graphics module 232 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. Graphics module 232 receives one or more codes specifying the graphics to be displayed, including coordinate data and other graphic property data, as needed, from an application or the like, and then generates screen image data to output to display controller 256.

[0067] The tactile feedback module 233 includes various software components for generating instructions used by the tactile output generator(s) 267 to generate tactile outputs at one or more locations on the device 200 in response to a user's interaction with the device 200.

[0068] A text input module 234, a component of graphics module 232, in some embodiments provides a soft keyboard for entering text in various applications (e.g., contacts 237, email 240, IM 241, browser 247, and any other application requiring text input).

[0069] The GPS module 235 determines the location of the device and provides this information for use within various applications (e.g., to the phone 238 for use in location-based dialing, to the camera 243 as picture / video metadata, and to applications that provide location-based services such as weather widgets, local yellow pages widgets, and map / navigation widgets).

[0070] Digital assistant client module 229 includes various client-side digital assistant instructions for providing client-side functionality of the digital assistant. For example, digital assistant client module 229 can accept audio input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces of portable multifunction device 200 (e.g., microphone 213, accelerometer(s) 268, touch-sensitive display system 212, optical sensor(s) 264, other input control devices 216, etc.). Digital assistant client module 229 can also provide audio (e.g., speech output), visual, and / or tactile output, etc., through various output interfaces of portable multifunction device 200 (e.g., speaker 211, touch-sensitive display system 212, tactile output generator(s) 267, etc.). For example, output may be provided as voice, sound, an alert, a text message, a menu, a graphic, a video, an animation, a vibration, and / or a combination of two or more of the above. During operation, the digital assistant client module 229 communicates with the DA server 106 using the RF circuitry 208.

[0071] User data and models 231 includes various data associated with a user (e.g., user-specific vocabulary data, user preference data, user-specified name pronunciations, data from the user's electronic address book, to-do lists, shopping lists, etc.) for providing the client-side functionality of the digital assistant. Additionally, user data and models 231 includes various models (e.g., speech recognition models, statistical language models, natural language processing models, ontologies, task flow models, service models, etc.) for processing user input and determining user intent.

[0072] In some embodiments, digital assistant client module 229 establishes a context associated with the user, the current user interaction, and / or the current user input by utilizing various sensors, subsystems, and peripherals of portable multifunction device 200 to gather additional information from the environment surrounding portable multifunction device 200. In some embodiments, digital assistant client module 229 provides context information, or a subset thereof, along with the user input to DA server 106 to assist in inferring the user's intent. In some embodiments, the digital assistant also uses the context information to determine how to prepare and deliver output to the user. The context information is referred to as context data.

[0073] In some embodiments, the context information accompanying the user input includes sensor information, such as lighting, ambient noise, ambient temperature, images or videos of the surrounding environment, etc. In some embodiments, the context information may also include the physical state of the device, such as device orientation, device location, device temperature, power level, speed, acceleration, motion patterns, cellular signal strength, etc. In some embodiments, information about the software state of DA server 106, such as running processes, installed programs, past and present network activity, background services, error logs, resource usage, etc., as well as information about the software state of portable multifunction device 200, is provided to DA server 106 as context information associated with the user input.

[0074] In some embodiments, digital assistant client module 229 selectively provides information stored on portable multifunction device 200 (e.g., user data 231) in response to a request from DA server 106. In some embodiments, digital assistant client module 229 also elicits additional input from the user via a natural language dialog or other user interface in response to a request by DA server 106. Digital assistant client module 229 passes the additional input to DA server 106 to assist DA server 106 in intent inference and / or fulfillment of the user's intent expressed in the user request.

[0075] A more detailed description of the digital assistant is provided below with reference to Figures 7A-7C. It should be appreciated that the digital assistant client module 229 can include any number of sub-modules of the digital assistant module 726 described below.

[0076] The application 236 includes the following modules (or sets of instructions), or a subset or superset thereof: • a contacts module 237 (sometimes called an address book or contact list); ●Telephone module 238, ●Video conferencing module 239, ● an email client module 240; ● Instant messaging (IM) module 241; ●Training support module 242, camera module 243 for still images and / or video; ● Image management module 244, ●Video player module, ●Music player module, ● Browser module 247, ●Calendar module 248, widget module 249, which in some embodiments includes one or more of a weather widget 249-1, a stock price widget 249-2, a calculator widget 249-3, an alarm clock widget 249-4, a dictionary widget 249-5, and other widgets acquired by the user and user-created widgets 249-6; a widget creator module 250 for creating user-created widgets 249-6; ● Search module 251, A video and music player module 252 that integrates a video player module and a music player module; ● Memo module 253, Map module 254, and / or ●Online video module 255.

[0077] Examples of other applications 236 stored in memory 202 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice duplication.

[0078] In conjunction with touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, contacts module 237 is used to manage an address book or contact list (e.g., stored in memory 202 or in the application internal state 292 of contacts module 237 in memory 470), including adding name(s) to the address book, deleting name(s) from the address book, associating phone number(s), email address(es), physical address(es), or other information with names, associating pictures with names, categorizing and sorting names, providing phone numbers or email addresses to initiate and / or facilitate communication via telephone 238, videoconferencing module 239, email 240, or IM 241, and the like.

[0079] In conjunction with RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, telephone module 238 is used to enter character strings corresponding to telephone numbers, access one or more telephone numbers in contacts module 237, modify entered telephone numbers, dial individual telephone numbers, conduct conversations, and disconnect or hang up when the conversation is completed. Wireless communication thus uses any of a number of communication standards, protocols, and technologies.

[0080] Videoconferencing module 239 includes executable instructions for cooperating with RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touchscreen 212, display controller 256, optical sensor 264, optical sensor controller 258, contact / motion module 230, graphics module 232, text input module 234, contact module 237, and telephone module 238 to initiate, conduct, and end a videoconference between a user and one or more other participants in accordance with user commands.

[0081] Email client module 240, in conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, includes executable instructions for composing, sending, receiving, and managing emails in response to user commands. In conjunction with image management module 244, email client module 240 greatly facilitates the creation and sending of emails with still or video images captured by camera module 243.

[0082] Instant messaging module 241, in conjunction with RF circuitry 208, touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, includes executable instructions for entering a series of characters corresponding to an instant message, modifying previously entered characters, transmitting individual instant messages (e.g., using Short Message Service (SMS) or Multimedia Message Service (MMS) protocols for telephony-based instant messaging, or XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, sent and / or received instant messages include graphics, photos, audio files, video files, and / or other attachments supported by MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both telephony-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).

[0083] The training support module 242 includes executable instructions to work with the RF circuitry 208, touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, GPS module 235, map module 254, and music player module to create workouts (e.g., with time, distance, and / or calorie burn goals), communicate with training sensors (sports devices), receive training sensor data, calibrate sensors used to monitor workouts, select and play music for workouts, and display, store, and transmit workout data.

[0084] Camera module 243, in conjunction with touchscreen 212, display controller 256, optical sensor(s) 264, optical sensor controller 258, contact / motion module 230, graphics module 232, and image management module 244, includes executable instructions to capture and store still images or video (including video streams) in memory 202, modify characteristics of still images or video, or delete still images or video from memory 202.

[0085] Image management module 244 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slideshow or album), and storing still and / or video images in conjunction with touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and camera module 243.

[0086] Browser module 247, in conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, contains executable instructions for browsing the Internet according to user commands, including retrieving, linking to, receiving, and displaying web pages or portions thereof, as well as attachments and other files linked to web pages.

[0087] The calendar module 248 includes executable instructions to cooperate with the RF circuitry 208, the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, the text input module 234, the email client module 240, and the browser module 247 to create, display, modify, and store calendars and data associated with the calendars (e.g., calendar items, to-do lists, etc.) according to user instructions.

[0088] In conjunction with RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, and browser module 247, widget modules 249 are mini-applications that can be downloaded and used by a user (e.g., weather widget 249-1, stock quotes widget 249-2, calculator widget 249-3, alarm clock widget 249-4, and dictionary widget 249-5) or created by a user (e.g., user-created widget 249-6). In some embodiments, a widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, a widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! Widgets).

[0089] In conjunction with the RF circuitry 208, touch screen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, and browser module 247, widget creation module 250 is used by a user to create widgets (e.g., turn user-specified portions of a web page into widgets).

[0090] The search module 251 includes executable instructions for working in conjunction with the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, and the text input module 234 to search for text, music, sound, images, video, and / or other files in the memory 202 that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with a user's commands.

[0091] Video and music player module 252 includes executable instructions that, in conjunction with touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, and browser module 247, enable a user to download and play pre-recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing videos (e.g., on touchscreen 212 or on an external display connected via external port 224). In some embodiments, device 200 optionally includes the functionality of an MP3 player, such as an iPod (a trademark of Apple Inc.).

[0092] The notes module 253 includes executable instructions for cooperating with the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, and the text input module 234 to create and manage notes, to-do lists, and the like according to user commands.

[0093] In conjunction with RF circuitry 208, touch screen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, GPS module 235, and browser module 247, map module 254 is used to receive, display, modify, and store maps and data associated with maps (e.g., driving directions, data about stores, other locations at or near a particular place, and other location-based data) in accordance with user instructions.

[0094] Online video module 255, in conjunction with touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, text input module 234, email client module 240, and browser module 247, contains instructions that enable a user to access, browse for, receive (e.g., by streaming and / or downloading), and play (e.g., on the touchscreen or on an external display connected via external port 224) particular online videos, send emails with links to particular online videos, and otherwise manage online videos in one or more file formats, such as H.264. In some embodiments, instant messaging module 241 is used to send links to particular online videos, rather than email client module 240. For additional description of online video applications, see U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," the contents of which are incorporated herein by reference in their entireties.

[0095] The above-identified modules and applications each correspond to sets of executable instructions that perform one or more of the functions and methods described herein (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules; various embodiments may combine or otherwise rearrange various subsets of these modules. For example, a video player module may be combined with a music player module into a single module (e.g., video and music player module 252, FIG. 2A). In some embodiments, memory 202 stores a subset of the above-identified modules and data structures. Additionally, memory 202 stores additional modules and data structures not described above.

[0096] In some embodiments, device 200 is a device in which operation of a predetermined set of functions on the device is performed exclusively via a touchscreen and / or touchpad. By using the touchscreen and / or touchpad as the primary input control device for operation of device 200, the number of physical input control devices (push buttons, dials, etc.) on device 200 is reduced.

[0097] The set of predefined functions performed only through the touchscreen and / or touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 200 to a main menu, home menu, or root menu from any user interface displayed on device 200. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is a physical push button or other physical input control device rather than a touchpad.

[0098] 2B is a block diagram illustrating exemplary components for event processing, according to some embodiments. In some embodiments, memory 202 (FIG. 2A) or memory 470 (FIG. 4) includes event sorter 270 (e.g., in operating system 226) and individual application 236-1 (e.g., any of applications 237-251, 255, 480-490 described above).

[0099] Event sorter 270 receives the event information and determines which application 236-1 to deliver the event information to and application view 291 for application 236-1. Event sorter 270 includes an event monitor 271 and an event dispatcher module 274. In some embodiments, application 236-1 includes application internal state 292 that indicates the current application view(s) that are displayed on touch-sensitive display 212 when the application is active or running. In some embodiments, device / global internal state 257 is used by event sorter 270 to determine which application(s) are currently active, and application internal state 292 is used by event sorter 270 to determine which application(s) to deliver the event information to.

[0100] In some embodiments, application internal state 292 includes additional information such as one or more of resume information to be used when application 236-1 resumes execution, user interface state information indicating or ready to display information being displayed by application 236-1, state cues that allow the user to return to a previous state or view of application 236-1, and redo / undo cues of previous actions taken by the user.

[0101] Event monitor 271 receives event information from peripherals interface 218. The event information includes information about sub-events (e.g., a user touch as part of a multi-touch gesture on touch-sensitive display 212). Peripherals interface 218 transmits information it receives from I / O subsystem 206 or sensors such as proximity sensor 266, accelerometer(s) 268, and / or microphone 213 (via audio circuitry 210). The information that peripherals interface 218 receives from I / O subsystem 206 includes information from touch-sensitive display 212 or a touch-sensitive surface.

[0102] In some embodiments, event monitor 271 sends requests to peripherals interface 218 at predetermined intervals. In response, peripherals interface 218 transmits event information. In other embodiments, peripherals interface 218 transmits event information only when there is a significant event (e.g., receipt of an input above a predetermined noise threshold and / or for more than a predetermined duration).

[0103] In some embodiments, the event sorter 270 also includes a hit view determination module 272 and / or an active event recognizer determination module 273 .

[0104] Hit view determination module 272 provides a software procedure that determines where a sub-event occurred within one or more views when touch-sensitive display 212 is displaying more than one view. A view consists of the controls and other elements that a user can see on the display.

[0105] Another aspect of a user interface associated with an application is the set of views, sometimes referred to herein as application views or user interface windows, within which information is displayed and touch-based gestures occur. The application view (of a particular application) at which a touch is detected corresponds to a programmatic level within the application's program or view hierarchy. For example, the lowest-level view at which a touch is detected is called a hit view, and the set of events that are recognized as valid inputs is determined based at least in part on the hit view of the initial touch that initiates a touch-based gesture.

[0106] The hit view determination module 272 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized in a hierarchy, the hit view determination module 272 identifies the hit view as the lowest view in the hierarchy that should process the sub-events. In most situations, the hit view is the lowest-level view in which an initiating sub-event occurs (e.g., the first sub-event in a series of sub-events that form an event or potential event). Once a hit view is identified by the hit view determination module 272, the hit view typically receives all sub-events related to the same touch or input source identified as the hit view.

[0107] Active event recognizer determination module 273 determines which view(s) in the view hierarchy should receive a particular sequence of sub-events. In some embodiments, active event recognizer determination module 273 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, active event recognizer determination module 273 determines that all views that contain the physical location of the sub-events are actively participating views, and therefore determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to the area associated with one particular view, views higher in the hierarchy still remain actively participating views.

[0108] Event dispatcher module 274 dispatches event information to event recognizers (e.g., event recognizer 280). In embodiments that include active event recognizer determination module 273, event dispatcher module 274 delivers event information to the event recognizers determined by active event recognizer determination module 273. In some embodiments, event dispatcher module 274 stores event information in an event queue, which is retrieved by individual event receivers 282.

[0109] In some embodiments, operating system 226 includes event sorter 270. Alternatively, application 236-1 includes event sorter 270. In still other embodiments, event sorter 270 is a stand-alone module or is part of another module stored in memory 202, such as contact / motion module 230.

[0110] In some embodiments, application 236-1 includes multiple event handlers 290 and one or more application views 291, each containing instructions for processing touch events that occur within a respective view of the application's user interface. Each application view 291 of application 236-1 includes one or more event recognizers 280. Typically, an individual application view 291 includes multiple event recognizers 280. In other embodiments, one or more of the event recognizers 280 are part of a separate module, such as a user interface kit (not shown) or a higher-level object from which application 236-1 inherits methods and other properties. In some embodiments, an individual event handler 290 includes one or more of a data updater 276, an object updater 277, a GUI updater 278, and / or event data 279 received from event sorter 270. Event handler 290 utilizes or calls data updater 276, object updater 277, or GUI updater 278 to update application internal state 292. Alternatively, one or more of the application views 291 include one or more respective event handlers 290. Also, in some embodiments, one or more of the data updater 276, object updater 277, and GUI updater 278 are included in individual application views 291.

[0111] A separate event recognizer 280 receives event information (e.g., event data 279) from the event sorter 270 and identifies events from the event information. The event recognizer 280 includes an event receiver 282 and an event comparator 284. In some embodiments, the event recognizer 280 includes at least a subset of metadata 283 and event delivery instructions 288 (including sub-event delivery instructions).

[0112] The event receiver 282 receives event information from the event sorter 270. The event information includes information about a sub-event, e.g., a touch or a movement of a touch. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. If the sub-event is related to a movement of a touch, the event information also includes the speed and direction of the sub-event. In some embodiments, the event includes a rotation of the device from one orientation to another (e.g., from portrait to landscape or vice versa), and the event information includes corresponding information about the current orientation of the device (also called the device's attitude).

[0113] The event comparator 284 compares the event information to predefined event or sub-event definitions and determines the event or sub-event, or determines or updates the state of the event or sub-event, based on the comparison. In some embodiments, the event comparator 284 includes an event definition 286. The event definition 286 includes definitions of events (e.g., a predefined set of sub-events), such as Event 1 (287-1) and Event 2 (287-2). In some embodiments, sub-events within an event (287) include, for example, touch start, touch end, touch movement, touch cancellation, and multiple touches. In one example, the definition for Event 1 (287-1) is a double tap on a displayed object. A double tap includes, for example, a first touch on a displayed object relative to a predetermined phase (touch start), a first lift-off (touch end) relative to the predetermined phase, a second touch on a displayed object relative to the predetermined phase (touch start), and a second lift-off (touch end) relative to the predetermined phase. In another example, a definition of event 2 (287-2) is a drag on a displayed object. Drag includes, for example, a touch (or contact) on the displayed object to a predetermined stage, a movement of the touch across the touch-sensitive display 212, and a lift-off of the touch (touch end). In some embodiments, the event also includes information about one or more associated event handlers 290.

[0114] In some embodiments, event definition 287 includes definitions of events for individual user interface objects. In some embodiments, event comparator 284 performs a hit test to determine which user interface objects are associated with the sub-event. For example, if a touch is detected on touch-sensitive display 212 in an application view in which three user interface objects are displayed on touch-sensitive display 212, event comparator 284 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a separate event handler 290, event comparator 284 uses the results of the hit test to determine which event handler 290 to activate. For example, event comparator 284 selects the event handler associated with the sub-event and object that triggers the hit test.

[0115] In some embodiments, the definition of an individual event 287 also includes a delay action that delays delivery of the event information until it is determined whether a set of sub-events corresponds to the event type of the event recognizer.

[0116] If the individual event recognizer 280 determines that the sequence of sub-events does not match any of the events in the event definition 286, the individual event recognizer 280 enters an event disabled, event failed, or event finished state and thereafter ignores the next sub-event of the touch-based gesture. In this situation, any other event recognizers that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.

[0117] In some embodiments, individual event recognizers 280 include metadata 283 with configurable properties, flags, and / or lists that indicate to actively participating event recognizers how the event delivery system should perform sub-event delivery. In some embodiments, metadata 283 includes configurable properties, flags, and / or lists that indicate how event recognizers interact with each other or how event recognizers are allowed to interact with each other. In some embodiments, metadata 283 includes configurable properties, flags, and / or lists that indicate how sub-events are delivered to various levels in the view or programmatic hierarchy.

[0118] In some embodiments, the individual event recognizer 280 activates the event handler 290 associated with an event when one or more specific sub-events of the event are recognized. In some embodiments, the individual event recognizer 280 delivers event information associated with the event to the event handler 290. Activating the event handler 290 is separate from sending (and postponing sending) sub-events to the individual hit view. In some embodiments, the event recognizer 280 pops a flag associated with the recognized event, and the event handler 290 associated with the flag captures the flag and performs a predetermined process.

[0119] In some embodiments, the event delivery instructions 288 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver the event information to an event handler associated with a set of sub-events or to an actively participating view. The event handler associated with the set of sub-events or the actively participating view receives the event information and performs a predetermined process.

[0120] In some embodiments, data updater 276 creates and updates data used by application 236-1. For example, data updater 276 updates phone numbers used by contacts module 237 or stores video files used by video player module. In some embodiments, object updater 277 creates and updates objects used by application 236-1. For example, object updater 277 creates new user interface objects or updates the positions of user interface objects. GUI updater 278 updates the GUI. For example, GUI updater 278 prepares display information and sends the display information to graphics module 232 for display on the touch-sensitive display.

[0121] In some embodiments, event handler(s) 290 include or have access to data updater 276, object updater 277, and GUI updater 278. In some embodiments, data updater 276, object updater 277, and GUI updater 278 are included in a single module of an individual application 236-1 or application view 291. In other embodiments, they are included in two or more software modules.

[0122] It should be understood that the foregoing description of event processing of a user's touch on a touch-sensitive display also applies to other forms of user input for operating multifunction device 200 using input devices, although not all of them are initiated on a touchscreen. For example, mouse movements and mouse button presses, contact movements such as tapping, dragging, scrolling on a touchpad, optionally coordinated with single or multiple keyboard presses or holds, pen stylus input, device movement, verbal commands, detected eye movements, biometric input, and / or any combination thereof, are optionally utilized as inputs corresponding to sub-events that define the recognized event.

[0123] FIG. 3 illustrates portable multifunction device 200 having touchscreen 212, according to some embodiments. The touchscreen optionally displays one or more graphics within user interface (UI) 300. In this embodiment, as well as other embodiments described below, a user may select one or more of the graphics by performing a gesture on the graphics, for example, using one or more fingers 302 (not drawn to scale) or one or more styluses 303 (not drawn to scale). In some embodiments, selection of one or more graphics is performed when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (left to right, right to left, upward and / or downward), and / or rolling (right to left, left to right, upward and / or downward) of a finger in contact with device 200. In some implementations or situations, accidental contact with a graphic does not select the graphic, for example, if the gesture corresponding to selection is a tap, a swipe gesture sweeping over an application icon optionally does not select the corresponding application.

[0124] Device 200 also includes one or more physical buttons, such as a "home" or menu button 304. As described above, menu button 304 is used to navigate to any application 236 within a set of applications running on device 200. Alternatively, in some embodiments, the menu button is implemented as a soft key within a GUI displayed on touchscreen 212.

[0125] In one embodiment, device 200 includes touchscreen 212, menu button 304, pushbutton 306 for powering the device on / off and locking the device, volume control button(s) 308, subscriber identity module (SIM) card slot 310, headset jack 312, and external docking / charging port 224. Pushbutton 306 is optionally used to power the device on / off by pressing and holding the button down for a predetermined period of time, to lock the device by pressing and releasing the button before the predetermined time has elapsed, and / or to unlock the device or initiate the unlocking process. In an alternative embodiment, device 200 also accepts verbal input via microphone 213 for activating or deactivating certain functions. Device 200 also optionally includes one or more contact intensity sensors 265 for detecting the intensity of a contact on touchscreen 212 and / or one or more tactile output generators 267 for generating a tactile output for a user of device 200.

[0126] FIG. 4 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface, according to some embodiments. Device 400 need not be portable. In some embodiments, device 400 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a child's learning toy), a gaming system, or a control device (e.g., a home or commercial controller). Device 400 typically includes one or more processing units (CPUs) 410, one or more network or other communication interfaces 460, memory 470, and one or more communication buses 420 interconnecting these components. Communication bus 420 optionally includes circuitry (sometimes referred to as a chipset) that interconnects and controls communication between system components. Device 400 includes input / output (I / O) interface 430, including display 440, which is typically a touchscreen display. I / O interface 430 also optionally includes a keyboard and / or mouse (or other pointing device) 450, as well as a touchpad 455, a tactile output generator 457 (e.g., similar to tactile output generator(s) 267 described above with reference to FIG. 2A ) for generating tactile output on device 400, sensors 459 (e.g., optical sensors, acceleration sensors, proximity sensors, touch-sensitive sensors, and / or contact intensity sensors similar to contact intensity sensor(s) 265 described above with reference to FIG. 2A ). Memory 470 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 470 optionally includes one or more storage devices located remotely from CPU(s) 410.In some embodiments, memory 470 stores programs, modules, and data structures similar to, or a subset of, programs, modules, and data structures stored in memory 202 of portable multifunction device 200 (FIG. 2A). Additionally, memory 470 optionally stores additional programs, modules, and data structures not present in memory 202 of portable multifunction device 200. For example, memory 470 of device 400 optionally stores drawing module 480, presentation module 482, word processing module 484, website creation module 486, disc authoring module 488, and / or spreadsheet module 490, while memory 202 of portable multifunction device 200 (FIG. 2A) optionally does not store these modules.

[0127] Each of the above-identified elements in FIG. 4 may, in some embodiments, be stored in any one or more of the above-mentioned memory devices. Each of the above-identified modules corresponds to a set of instructions that perform the functions described above. The above-identified modules or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules; thus, various subsets of these modules may be combined or otherwise reconfigured in various embodiments. In some embodiments, memory 470 stores a subset of the above-identified modules and data structures. Additionally, memory 470 stores additional modules and data structures not described above.

[0128] Attention is now directed to user interface embodiments that may be implemented in portable multifunction device 200, for example.

[0129] 5A shows an exemplary user interface for a menu of applications on portable multifunction device 200, according to some embodiments. A similar user interface is implemented on device 400. In some embodiments, user interface 500 includes the following elements, or a subset or superset thereof:

[0130] signal strength indicator(s) 502 for wireless communication(s), such as cellular and Wi-Fi signals; ●Time 504, ●Bluetooth indicator 505, ● Battery status indicator 506, Tray 508 with icons of frequently used applications, such as: An icon 516 for the phone module 238, labeled "Phone," optionally including an indicator 514 of the number of missed calls or voicemail messages; An icon 518 for the email client module 240, labeled "Mail," optionally including an indicator 510 of the number of unread emails; ○ An icon 520 for the browser module 247, labeled "Browser"; and ○ An icon 522 for the video and music player module 252, also called the iPod (trademark of Apple Inc.) module 252, labeled "iPod"; and ● Icons of other applications, such as: ○ Icon 524 of IM module 241, labeled "Messages" ○ Icon 526 of the calendar module 248, labeled "Calendar" ○ Icon 528 of the image management module 244, labeled "Photos" ○ An icon 530 for the camera module 243, labeled "camera"; ○ Icon 532 of the online video module 255, labeled "Online Video"; Icon 534 of Stock Price Widget 249-2, labeled "Stock Price" ○ Icon 536 of map module 254, labeled "Map" ○ Icon 538 of weather widget 249-1, labeled "Weather" ○ Icon 540 of alarm clock widget 249-4, labeled "Clock" ○ Icon 542 of Training Support Module 242, labeled "Training Support"; ○ An icon 544 in the Notes module 253 labeled "Notes," and A settings application or module icon 546 labeled "Settings" that provides access to settings for the device 200 and its various applications 236.

[0131] 5A are merely exemplary. For example, icon 522 for video and music player module 252 is optionally labeled "Music" or "Music Player." Other labels are optionally used for various application icons. In some embodiments, the label for an individual application icon includes the name of the application that corresponds to the individual application icon. In some embodiments, the label for a particular application icon is different from the name of the application that corresponds to that particular application icon.

[0132] 5B shows an exemplary user interface on a device (e.g., device 400 of FIG. 4) that has touch-sensitive surface 551 (e.g., tablet or touchpad 455 of FIG. 4) that is separate from display 550 (e.g., touchscreen display 212). Device 400 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 459) that detect the intensity of a contact on touch-sensitive surface 551, and / or one or more tactile output generators 457 that generate a tactile output for a user of device 400.

[0133] Although some of the following examples are described with reference to input on touchscreen display 212 (when the touch-sensitive surface and display are combined), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, as shown in FIG. 5B . In some embodiments, this touch-sensitive surface (e.g., 551 in FIG. 5B ) has a major axis (e.g., 552 in FIG. 5B ) that corresponds to a major axis (e.g., 553 in FIG. 5B ) on the display (e.g., 550). According to these embodiments, the device detects contact with touch-sensitive surface 551 (e.g., 560 and 562 in FIG. 5B ) at locations that correspond to respective locations on the display (e.g., in FIG. 5B , 560 corresponds to 568 and 562 corresponds to 570). In this manner, when the touch-sensitive surface is separate from the display, user input (e.g., contacts 560 and 562 and their movement) detected by the device on the touch-sensitive surface (e.g., 551 in FIG. 5B ) is used by the device to operate a user interface on the display (e.g., 550 in FIG. 5B ) of the multifunction device. It should be understood that similar methods are optionally used for the other user interfaces described herein.

[0134] Additionally, while the following examples are given primarily with reference to finger input (e.g., finger contact, finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of the finger inputs are replaced with input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a contact) followed by movement of a cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is optionally replaced by a mouse click (e.g., instead of detecting a contact and then ceasing contact detection) while the cursor is positioned over the location of the tap gesture. Similarly, it should be understood that when multiple user inputs are detected simultaneously, multiple computer mice are optionally used simultaneously, or a mouse and finger contacts are optionally used simultaneously.

[0135] FIG. 6A shows an exemplary personal electronic device 600. Device 600 includes a main body 602. In some embodiments, device 600 includes some or all of the features described in connection with devices 200 and 400 (e.g., FIGS. 2A-4 ). In some embodiments, device 600 includes a touch-sensitive display screen 604, hereafter touchscreen 604. Alternatively, or in addition to touchscreen 604, device 600 includes a display and a touch-sensitive surface. As with devices 200 and 400, in some embodiments, touchscreen 604 (or the touch-sensitive surface) includes one or more intensity sensors that detect the intensity of an applied contact (e.g., a touch). The one or more intensity sensors in touchscreen 604 (or the touch-sensitive surface) provide output data that represents the intensity of the touch. The user interface of device 600 responds to touches based on the intensity of the touch, meaning that touches of different intensities can invoke different user interface actions on device 600.

[0136] Techniques for detecting and processing touch intensity can be found, for example, in related applications: International Patent Application No. PCT / US2013 / 040061, filed May 8, 2013, entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," and International Patent Application No. PCT / US2013 / 069483, filed November 11, 2013, entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," each of which is incorporated herein by reference in its entirety.

[0137] In some embodiments, device 600 has one or more input mechanisms 606 and 608. Input mechanisms 606 and 608, if included, are physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 600 has one or more attachment mechanisms. Such attachment mechanisms, if included, can allow device 600 to be attached to, for example, hats, eyewear, earrings, necklaces, shirts, jackets, bracelets, watch bands, chains, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow device 600 to be worn by a user.

[0138] FIG. 6B illustrates an exemplary personal electronic device 600. In some embodiments, device 600 includes some or all of the components described in connection with FIGS. 2A, 2B, and 4. Device 600 includes a bus 612 operably coupling an I / O unit 614 to one or more computer processors 616 and memory 618. The I / O unit 614 is connected to a display 604, which may have touch-sensing components 622 and, optionally, touch-intensity-sensing components 624. Additionally, the I / O unit 614 is connected to a communication unit 630 that receives application and operating system data using Wi-Fi, Bluetooth, near-field communication (NFC), cellular, and / or other wireless communication technologies. Device 600 includes input mechanisms 606 and / or 608. Input mechanism 606 is, for example, a rotatable input device or a depressible and rotatable input device. Input mechanism 608 is, in some embodiments, a button.

[0139] The input mechanism 608, in some embodiments, is a microphone. The personal electronic device 600 includes various sensors, such as a GPS sensor 632, an accelerometer 634, an orientation sensor 640 (e.g., a compass), a gyroscope 636, a motion sensor 638, and / or combinations thereof, all of which are operably connected to the I / O section 614.

[0140] The memory 618 of the personal electronic device 600 is a non-transitory computer-readable storage medium that stores computer-executable instructions that, when executed by one or more computer processors 616, for example, cause the computer processors to perform the following techniques and processes. Those computer-executable instructions may also be stored and / or transmitted in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as, for example, a computer-based system, a system including a processor, or other system capable of fetching instructions from and executing those instructions. The personal electronic device 600 is not limited to the components and configuration of FIG. 6B and may include other or additional components in multiple configurations.

[0141] As used herein, the term "affordance" refers to a user-interactive graphical user interface object displayed on a display screen of, for example, device 200, 400, 600, 800, and / or 900 (FIGS. 2A, 4, 6A-6B, 8A-8H, 9A-9F, and 10A-10D). For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each constitute an affordance.

[0142] As used herein, the term “focus selector” refers to an input element that indicates the current portion of a user interface with which a user is interacting. In some implementations including a cursor or other position marker, the cursor serves as a “focus selector” such that when input (e.g., a press input) is detected on a touch-sensitive surface (e.g., touchpad 455 in FIG. 4 or touch-sensitive surface 551 in FIG. 5B ) while the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted according to the detected input. In some implementations including a touchscreen display (e.g., touch-sensitive display system 212 in FIG. 2A or touchscreen 212 in FIG. 5A ) that allows direct interaction with user interface elements on the touchscreen display, a contact detected on the touchscreen serves as a “focus selector” such that when input (e.g., a press input by a contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, the particular user interface element is adjusted according to the detected input. In some implementations, focus is moved from one region of the user interface to another region of the user interface without a corresponding cursor movement or contact movement on the touchscreen display (e.g., by using the tab key or arrow keys to move focus from one button to another), and in these implementations, the focus selector moves to follow the movement of focus between various regions of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or contact on a touchscreen display) that is controlled by the user to communicate the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface through which the user intends to interact).For example, the location of a focus selector (e.g., a cursor, touch, or selection box) over an individual button while a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen) indicates that the user intends to activate that individual button (and not other user interface elements shown on the device's display).

[0143] As used herein and in the claims, the term "characteristic intensity" of a contact refers to a characteristic of that contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on a plurality of intensity samples. The characteristic intensity is optionally based on a predetermined number of intensity samples, i.e., a set of intensity samples collected during a predetermined time period (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined event (e.g., after detecting the contact, before detecting lift-off of the contact, before or after detecting the start of contact movement, before detecting the end of the contact, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The characteristic intensity of the contact is optionally based on one or more of the maximum intensity of the contact, the median intensity of the contact, the average intensity of the contact, the top 10 percent of the intensity of the contact, half the maximum intensity of the contact, 90 percent of the maximum intensity of the contact, etc. In some embodiments, the duration of the contact is used in determining the characteristic intensity (e.g., when the characteristic intensity is an average of the intensity of the contact over time). In some embodiments, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether an action is performed by the user. For example, the set of one or more intensity thresholds includes a first intensity threshold and a second intensity threshold. In this example, a contact having a characteristic intensity that does not exceed the first threshold results in a first action, a contact having a characteristic intensity above the first intensity threshold but not above the second intensity threshold results in a second action, and a contact having a characteristic intensity above the second threshold results in a third action. In some embodiments, the comparison of the characteristic intensity to one or more thresholds is not used to determine whether to perform the first action or the second action, but rather to determine whether to perform one or more actions (e.g., whether to perform an individual action or to forgo performing an individual action).

[0144] In some embodiments, a portion of the gesture is identified for purposes of determining the characteristic intensity. For example, the touch-sensitive surface receives a continuous swipe contact transitioning from a start location point to an end location point of increasing contact intensity. In this example, the characteristic intensity of the contact at the end location is based on only a portion of the continuous swipe contact, rather than the entire swipe contact (e.g., only the portion of the swipe contact at the end location). In some embodiments, a smoothing algorithm is applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of an unweighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some situations, these smoothing algorithms eliminate narrow spikes or dips in the swipe contact intensity for purposes of determining the characteristic intensity.

[0145] The intensity of a contact on the touch-sensitive surface is characterized with respect to one or more intensity thresholds, such as a contact-detection intensity threshold, a light pressure intensity threshold, a deep pressure intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light pressure intensity threshold corresponds to an intensity at which the device performs an action typically associated with clicking a physical mouse button or trackpad. In some embodiments, the deep pressure intensity threshold corresponds to an intensity at which the device performs an action different from an action typically associated with clicking a physical mouse button or trackpad. In some embodiments, when a contact is detected having a characteristic intensity below the light pressure intensity threshold (e.g., and above a nominal contact-detection intensity threshold below which the contact is not detected), the device moves the focus selector according to the movement of the contact on the touch-sensitive surface without performing an action associated with the light pressure intensity threshold or the deep pressure intensity threshold. In general, unless otherwise specified, these intensity thresholds are consistent across various sets of values for a user interface.

[0146] An increase in the characteristic intensity of a contact from an intensity below the light pressure intensity threshold to an intensity between the light pressure intensity threshold and the deep pressure intensity threshold may be referred to as inputting a "light press." An increase in the characteristic intensity of a contact from an intensity below the deep pressure intensity threshold to an intensity above the deep pressure intensity threshold may be referred to as inputting a "deep press." An increase in the characteristic intensity of a contact from an intensity below the contact-detection intensity threshold to an intensity between the contact-detection intensity threshold and the light pressure intensity threshold may be referred to as detecting a contact on the touch surface. A decrease in the characteristic intensity of a contact from an intensity above the contact-detection intensity threshold to an intensity below the contact-detection intensity threshold may be referred to as detecting a lift-off of the contact from the touch surface. In some embodiments, the contact-detection intensity threshold is zero. In some embodiments, the contact-detection intensity threshold is greater than zero.

[0147] In some embodiments described herein, one or more actions are performed in response to detecting a gesture including an individual pressure input or in response to detecting an individual pressure input performed by an individual contact (or multiple contacts), where the individual pressure input is detected based at least in part on detecting an increase in intensity of the contact (or multiple contacts) above a pressure input intensity threshold. In some embodiments, the individual action is performed in response to detecting an increase in intensity of the individual contact above the pressure input intensity threshold (e.g., a "downstroke" of the individual pressure input). In some embodiments, the pressure input includes an increase in intensity of the individual contact above the pressure input intensity threshold followed by a decrease in intensity of the contact below the pressure input intensity threshold, and the individual action is performed in response to detecting a subsequent decrease in intensity of the individual contact below the pressure input threshold (e.g., an "upstroke" of the individual pressure input).

[0148] In some embodiments, the device employs intensity hysteresis to avoid accidental input, sometimes referred to as “jitter,” and the device defines or selects a hysteresis intensity threshold that has a predetermined relationship to the pressure input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units below the pressure input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the pressure input intensity threshold). Thus, in some embodiments, the pressure input includes an increase in the intensity of a discrete contact above the pressure input intensity threshold followed by a decrease in the intensity of the contact below the hysteresis intensity threshold corresponding to the pressure input intensity threshold, and a discrete action is performed in response to detecting a subsequent decrease in the intensity of the discrete contact below the hysteresis intensity threshold (e.g., an “upstroke” of the discrete pressure input). Similarly, in some embodiments, a pressure input is detected only when the device detects an increase in the intensity of the contact from an intensity below the hysteresis intensity threshold to an intensity above the pressure input intensity threshold, and optionally a subsequent decrease in the intensity of the contact to an intensity below the hysteresis intensity, and a distinct action is performed in response to detecting the pressure input (e.g., an increase in the intensity of the contact or a decrease in the intensity of the contact, as the case may be).

[0149] For ease of explanation, descriptions of operations performed in response to a pressure input associated with a pressure input intensity threshold, or a gesture including a pressure input, are optionally triggered in response to detecting any of: an increase in the intensity of the contact above the pressure input intensity threshold; an increase in the intensity of the contact from an intensity below a hysteresis intensity threshold to an intensity above the pressure input intensity threshold; a decrease in the intensity of the contact below the pressure input intensity threshold; and / or a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to the pressure input intensity threshold. Further, in examples where an operation is described as being performed in response to detecting a decrease in the intensity of the contact below a pressure input intensity threshold, the operation is optionally performed in response to detecting a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to and lower than the pressure input intensity threshold. 3. Digital Assistant System

[0150] FIG. 7A shows a block diagram of a digital assistant system 700 according to various embodiments. In some embodiments, digital assistant system 700 is implemented on a standalone computer system. In some embodiments, digital assistant system 700 is distributed across multiple computers. In some embodiments, some of the modules and functionality of the digital assistant are allocated to a server portion and a client portion, where the client portion resides on one or more user devices (e.g., devices 104, 122, 200, 400, 600, 800, or 900), for example, as shown in FIG. 1, and communicates with the server portion (e.g., server system 108) through one or more networks. In some embodiments, digital assistant system 700 is an implementation of server system 108 (and / or DA server 106) shown in FIG. 1. It should be noted that digital assistant system 700 is only one example of a digital assistant system, and that digital assistant system 700 may have more or fewer components than those shown, may combine two or more components, or may have a different configuration or arrangement of the components. The various components shown in FIG. 7A may be implemented as hardware, including one or more signal processing circuits and / or application specific integrated circuits, software instructions executed by one or more processors, firmware, or a combination thereof.

[0151] Digital assistant system 700 includes memory 702, one or more processors 704, an input / output (I / O) interface 706, and a network communication interface 708. These components can communicate with each other via one or more communication buses or signal lines 710.

[0152] In some embodiments, memory 702 includes a non-transitory computer-readable medium, such as high-speed random access memory and / or a non-volatile computer-readable storage medium (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).

[0153] In some embodiments, I / O interface 706 couples input / output devices 716 of digital assistant system 700, such as a display, keyboard, touchscreen, and microphone, to user interface module 722. I / O interface 706 interfaces with user interface module 722 to receive user inputs (e.g., voice input, keyboard input, touch input, etc.) and process them accordingly. In some embodiments, for example, when the digital assistant is implemented on a standalone user device, digital assistant system 700 includes any of the components and I / O communication interfaces described with respect to devices 200, 400, 600, 800, or 900 of FIGS. 2A, 4, 6A-6B, 8A-8H, 9A-9F, and 10A-10D. In some embodiments, digital assistant system 700 represents the server portion of a digital assistant implementation and can interact with a user through a client-side portion that resides on a user device (e.g., device 104, 200, 400, 600, 800, or 900).

[0154] In some embodiments, network communication interface 708 includes wired communication port(s) 712 and / or wireless transceiver circuitry 714. The wired communication port(s) transmit and receive communication signals via one or more wired interfaces, such as Ethernet, Universal Serial Bus (USB), FIREWIRE, etc. The wireless circuitry 714 transmits and receives RF and / or optical signals to and from communication networks and other communication devices. Wireless communication uses any of a number of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. Network communication interface 708 enables communication between digital assistant system 700 and networks, such as the Internet, intranets, and / or wireless networks, such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs), and with other devices.

[0155] In some embodiments, memory 702, or the computer-readable storage medium of memory 702, stores programs, modules, instructions, and data structures, including all or a subset of an operating system 718, a communications module 720, a user interface module 722, one or more applications 724, and a digital assistant module 726. In particular, memory 702, or the computer-readable storage medium of memory 702, stores instructions for performing the processes described below. One or more processors 704 execute these programs, modules, and instructions and read / write from / to the data structures.

[0156] An operating system 718 (e.g., Darwin, RTXC, LINUX, UNIX, iOS, OS X, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitating communication between various hardware, firmware, and software components.

[0157] Communications module 720 facilitates communication between digital assistant system 700 and other devices via network communications interface 708. For example, communications module 720 communicates with RF circuitry 208 of electronic devices such as devices 200, 400, and 600 shown in FIGS. 2A, 4, and 6A-6B, respectively. Communications module 720 also includes various components for processing data received by wireless circuitry 714 and / or wired communications port 712.

[0158] The user interface module 722 receives commands and / or input from a user via the I / O interface 706 (e.g., from a keyboard, touchscreen, pointing device, controller, and / or microphone) and generates user interface objects on the display. The user interface module 722 also prepares and delivers output to the user (e.g., speech, sound, animation, text, icons, vibration, haptic feedback, light, etc.) via the I / O interface 706 (e.g., through a display, audio channel, speaker, touchpad, etc.).

[0159] Applications 724 include programs and / or modules configured to be executed by one or more processors 704. For example, if the digital assistant system is implemented on a standalone user device, applications 724 include user applications such as games, calendar applications, navigation applications, or email applications. If the digital assistant system 700 is implemented on a server, applications 724 include, for example, resource management applications, diagnostic applications, or scheduling applications.

[0160] Memory 702 also stores digital assistant module 726 (or the server portion of the digital assistant). In some embodiments, digital assistant module 726 includes the following submodules, or a subset or superset thereof: input / output processing module 728, speech-to-text (STT) processing module 730, natural language processing module 732, dialog flow processing module 734, task flow processing module 736, service processing module 738, and speech synthesis processing module 740. Each of these modules has access to one or more of the following systems or data and models of digital assistant module 726, or a subset or superset thereof: ontology 760, vocabulary index 744, user data 748, task flow model 754, service model 756, and ASR system 758.

[0161] In some examples, using the processing modules, data, and models embodied in digital assistant module 726, the digital assistant can perform at least some of the following: converting speech input to text and identifying a user's intent expressed in natural language input received from the user; actively eliciting and obtaining the information necessary to fully infer the user's intent (e.g., by disambiguating words, games, intent, etc.); determining a task flow to satisfy the inferred intent; and executing the task flow to satisfy the inferred intent.

[0162] In some embodiments, as shown in FIG. 7B , I / O processing module 728 interacts with a user through I / O device 716 of FIG. 7A or with a user device (e.g., device 104, 200, 400, or device 600) through network communication interface 708 of FIG. 7A to obtain user input (e.g., speech input) and to provide responses to the user input (e.g., as speech output). I / O processing module 728 optionally obtains contextual information associated with the user input from the user device along with or immediately after receiving the user input. The contextual information includes user-specific data, vocabulary, and / or preferences related to the user input. In some embodiments, the contextual information also includes information about the software and hardware state of the user device at the time the user request is received and / or the user's ambient environment at the time the user request is received. In some embodiments, I / O processing module 728 also sends follow-up questions to the user regarding the user request and receives answers from the user. When a user request is received by the I / O processing module 728 and the user request includes speech input, the I / O processing module 728 forwards the speech input to the STT processing module 730 (or speech recognizer) for speech-to-text conversion.

[0163] The STT processing module 730 includes one or more ASR systems 758. The one or more ASR systems 758 can process speech input received via the I / O processing module 728 to generate recognition results. Each ASR system 758 includes a front-end speech preprocessor. The front-end speech preprocessor extracts representative features from the speech input. For example, the front-end speech preprocessor performs a Fourier transform on the speech input to extract spectral features that characterize the speech input as a sequence of representative multi-dimensional vectors. Furthermore, each ASR system 758 includes one or more speech recognition models (e.g., acoustic models and / or language models) and implements one or more speech recognition engines. Examples of speech recognition models include hidden Markov models, Gaussian mixture models, deep neural network models, n-gram language models, and other statistical models. Examples of speech recognition engines include dynamic time warping-based engines and weighted finite-state transducer (WFST)-based engines. One or more speech recognition models and one or more speech recognition engines are used to process the extracted representative features of the front-end speech preprocessor to generate intermediate recognition results (e.g., phonemes, phoneme strings, subwords) and ultimately text recognition results (words, word strings, sequences of tokens). In some implementations, the speech input is at least partially processed by a third-party service or on the user's device (e.g., device 104, 200, 400, or device 600) to generate the recognition results. Once the STT processing module 730 generates a recognition result including a text string (e.g., a word, a string of words, or a string of tokens), the recognition result is passed to the natural language processing module 732 for intent inference. In some implementations, the STT processing module 730 generates multiple candidate text representations of the speech input. Each candidate text representation is a sequence of words or tokens corresponding to the speech input. In some implementations, each candidate text representation is associated with a speech recognition confidence score.Based on the speech recognition confidence scores, the STT processing module 730 ranks the candidate text representations and provides the n best (e.g., the n highest-ranked) candidate text representation(s) to the natural language processing module 732 for intent inference, where n is a predetermined integer greater than zero. For example, in one embodiment, only the highest-ranked (n=1) candidate text representation is passed to the natural language processing module 732 for intent inference. In another embodiment, the five highest-ranked (n=5) candidate text representations are passed to the natural language processing module 732 for intent inference.

[0164] Further details regarding speech-to-text processing are described in U.S. Utility Patent Application No. 13 / 236,942, filed September 20, 2011, for "Consolidating Speech Recognition Results," the entire disclosure of which is incorporated herein by reference.

[0165] In some implementations, the STT processing module 730 includes a vocabulary of recognizable words and / or has access to that vocabulary via the phonetic alphabet conversion module 731. Each vocabulary word is associated with one or more candidate pronunciations of that word, represented in a speech recognition phonetic alphabet. Specifically, the vocabulary of recognizable words includes words associated with multiple candidate pronunciations. For example, the vocabulary may include: [Table 1] The vocabulary words include the word "tomato," which is associated with a pronunciation candidate based on previous speech input from the user. Additionally, vocabulary words are associated with custom pronunciation candidates based on previous speech input from the user. Such custom pronunciation candidates are stored within the STT processing module 730 and associated with a particular user via the user's profile on the device. In some implementations, pronunciation candidates for a word are determined based on the spelling of the word and one or more linguistic and / or phonetic rules. In some implementations, pronunciation candidates are generated manually, for example, based on known canonical pronunciations.

[0166] In some implementations, the pronunciation candidates are ranked based on the commonality of the pronunciation candidate. [Table 2] teeth, [Table 3] , because the former is a more commonly used pronunciation (e.g., among all users, for users in a particular geographic region, or for any other suitable subset of users). In some implementations, pronunciation candidates are ranked based on whether they are custom pronunciation candidates associated with the user. For example, custom pronunciation candidates are ranked higher than regular pronunciation candidates. This can be useful for recognizing proper nouns that have unique pronunciations that deviate from the regular pronunciation. In some implementations, pronunciation candidates are associated with one or more speech characteristics, such as place of origin, nationality, or ethnicity. For example, pronunciation candidate [Table 4] is associated with the United States, while the pronunciation candidate [Table 5] is associated with the United Kingdom. Furthermore, the ranking of the pronunciation candidates is based on one or more characteristics of the user (e.g., place of origin, nationality, ethnicity, etc.) stored in the user's profile on the device. For example, it may be determined from the user's profile that the user is associated with the United States. Based on the user's association with the United States, the pronunciation candidates (associated with the United States) [Table 6] is a possible pronunciation (associated with the UK) [Table 7] In some implementations, one of the ranked pronunciation candidates is selected as the predicted pronunciation (e.g., the most likely pronunciation).

[0167] When a speech input is received, the STT processing module 730 is used to determine the phonemes that correspond to the speech input (e.g., using an acoustic model) and then attempts to determine a word that matches the phonemes (e.g., using a language model). For example, the STT processing module 730 first determines a sequence of phonemes that correspond to a portion of the speech input. [Table 8] , then based on lexical index 744, it may be determined that this string corresponds to the word "tomato."

[0168] In some implementations, the STT processing module 730 uses proximity matching techniques to determine words in an utterance. Thus, for example, the STT processing module 730 may use a sequence of phonemes: [Table 9] corresponds to the word "tomato" even if that particular sequence of phonemes is not one of the possible sequences of phonemes for that word.

[0169] The digital assistant's natural language processing module 732 ("natural language processor") takes the n best candidate text representation(s) ("word string(s)" or "token string(s)") generated by the STT processing module 730 and attempts to associate each of those candidate text representations with one or more "actionable intents" recognized by the digital assistant. An "actionable intent" (or "user intent") represents a task that can be performed by the digital assistant and may have an associated task flow realized in task flow model 754. This associated task flow is a sequence of programmed actions and steps that the digital assistant performs to accomplish the task. The scope of the digital assistant's capabilities is determined according to the number and variety of task flows realized and stored in task flow model 754, or in other words, the number and variety of "actionable intents" recognized by the digital assistant. However, the effectiveness of a digital assistant is also judged according to its ability to infer the correct "actionable intent(s)" from user requests expressed in natural language.

[0170] In some embodiments, in addition to the string of words or tokens obtained from STT processing module 730, natural language processing module 732 also receives contextual information associated with the user request, e.g., from I / O processing module 728. Natural language processing module 732 optionally uses the contextual information to clarify, complement, and / or further define the information contained in the candidate text representation received from STT processing module 730. Contextual information includes, for example, user preferences, hardware and / or software state of the user device, sensor information collected before, during, or immediately after the user request, and previous interactions (e.g., dialogs) between the digital assistant and the user. As described herein, contextual information is dynamic in some embodiments, changing with time, location, dialog content, and other factors.

[0171] In some embodiments, natural language processing is based on, for example, ontology 760. Ontology 760 is a hierarchical structure that includes multiple nodes, each of which represents an "actionable intent" or represents an "attribute" or other "properties" related to one or more of the "actionable intents." As described above, an "actionable intent" represents a task that a digital assistant can perform, i.e., the task is "actionable" or can be performed. An "attribute" represents a parameter associated with an actionable intent or a sub-aspect of another attribute. Links between actionable intent nodes and attribute nodes in ontology 760 define how the parameter represented by the attribute node contributes to the task represented by the actionable intent node.

[0172] In some embodiments, ontology 760 is composed of actionable intent nodes and attribute nodes. Within ontology 760, each actionable intent node is linked to one or more attribute nodes, either directly or through one or more intermediate attribute nodes. Similarly, each attribute node is linked to one or more actionable intent nodes, either directly or through one or more intermediate attribute nodes. For example, as shown in FIG. 7C , ontology 760 includes a “restaurant reservation” node (i.e., an actionable intent node). The attribute nodes “restaurant,” “date / time” (for reservation), and “number of participants” are each directly linked to an actionable intent node (i.e., the “restaurant reservation” node).

[0173] Furthermore, the attribute nodes “cuisine,” “price range,” “phone number,” and “location” are subordinate nodes of the attribute node “restaurant,” and each is linked to the “reservation restaurant” node (i.e., an actionable intention node) via the intermediate attribute node “restaurant.” As another example, as shown in FIG. 7C , ontology 760 also includes a “set reminder” node (i.e., another actionable intention node). The attribute nodes “date / time” (for setting a reminder) and “theme” (for a reminder) are each linked to the “set reminder” node. Because the attribute node “date / time” is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node “date / time” is linked to both the “reservation restaurant” node and the “set reminder” node in ontology 760.

[0174] An actionable intent node, along with its linked attribute nodes, is described as a "domain." In this discussion, each domain is associated with a distinct actionable intent and refers to the group of nodes (and the relationships between those nodes) associated with that particular actionable intent. For example, ontology 760 shown in FIG. 7C includes an example restaurant reservation domain 762 and an example reminder domain 764 within ontology 760. The restaurant reservation domain includes the actionable intent node "reservation," the attribute nodes "restaurant," "date / time," and "number of participants," and the sub-attribute nodes "cuisine," "price range," "phone number," and "location." The reminder domain 764 includes the actionable intent node "reminder setting," and the attribute nodes "theme" and "date / time." In some embodiments, ontology 760 is composed of multiple domains. Each domain shares one or more attribute nodes with one or more other domains. For example, the "date / time" attribute node is associated with a number of different domains (eg, a scheduling domain, a travel booking domain, a movie ticket domain, etc.) in addition to a restaurant reservation domain 762 and a reminder domain 764.

[0175] 7C illustrates two example domains within ontology 760, although other domains include, for example, "find a movie," "make a phone call," "find directions," "schedule a meeting," "send a message," "provide an answer to a question," "read a list," "provide navigation instructions," and "provide task instructions." The "send message" domain is associated with the actionable intent of "send a message" and further includes attribute nodes such as "recipient(s)," "message type," and "message body." The attribute node "recipient" is further defined by sub-attribute nodes such as "recipient name" and "message address."

[0176] In some implementations, ontology 760 includes all domains (and therefore actionable intents) that the digital assistant can understand and perform. In some implementations, ontology 760 is modified by adding or removing domains or entire nodes, or by modifying relationships between nodes in ontology 760, etc.

[0177] In some embodiments, nodes associated with related actionable intents are clustered under a "super domain" in ontology 760. For example, the "Travel" super domain includes a cluster of travel-related attribute nodes and actionable intent nodes. Actionable intent nodes related to travel include "book an airline ticket," "book a hotel," "rent a car," "get directions," "find points of interest," etc. Actionable intent nodes under the same super domain (e.g., the "Travel" super domain) share many attribute nodes. For example, the actionable intent nodes for "book an airline ticket," "book a hotel," "rent a car," "get directions," and "find points of interest" share one or more of the attribute nodes "departure location," "destination," "departure date / time," "arrival date / time," and "number of participants."

[0178] In some implementations, each node in ontology 760 is associated with a set of words and / or phrases related to the attribute or actionable intent represented by that node. The set of distinct words and / or phrases associated with each node is what is called the "vocabulary" associated with that node. The set of distinct words and / or phrases associated with each node, in association with the attribute or actionable intent represented by that node, is stored in vocabulary index 744. For example, returning to FIG. 7B , vocabulary associated with a node related to the attribute of "restaurant" may include words such as "food," "drink," "dish," "hungry," "eat," "pizza," "fast food," and "meal." As another example, vocabulary associated with a node related to the actionable intent of "initiate a phone call" may include words and phrases such as "call," "phone," "dial," "ring," "call this number," and "make a call to." Vocabulary index 744 optionally includes words and phrases in different languages.

[0179] Natural language processing module 732 receives candidate text representations (e.g., character string(s) or token string(s)) from STT processing module 730 and, for each candidate representation, determines which nodes are implied by the words in the candidate text representation. In some examples, if a word or phrase in the candidate text representation is found to be associated with one or more nodes in ontology 760 (via vocabulary index 744), the word or phrase "trigger" or "activate" those nodes. Based on the amount and / or relative importance of activated nodes, natural language processing module 732 selects one of those actionable intents as the task the user intends the digital assistant to perform. In some examples, the domain with the most "triggered" nodes is selected. In some examples, the domain with the highest confidence value (e.g., based on the relative importance of the various triggered nodes) is selected. In some examples, the domain is selected based on a combination of the number and importance of triggered nodes. In some embodiments, additional factors are also considered when selecting a node, such as whether the digital assistant has previously correctly interpreted a similar request from the user.

[0180] User data 748 includes user-specific information, such as user-specific vocabulary, user preferences, user addresses, the user's default and secondary languages, the user's contact list, and other short-term or long-term information about each user. In some implementations, natural language processing module 732 uses this user-specific information to supplement the information contained in the user input and further define the user intent. For example, for a user request "invite my friends to my birthday party," natural language processing module 732 can access user data 748 to determine who the "friends" are and when and where the "birthday party" will be held, without requiring the user to explicitly provide such information in the user request.

[0181] It should be appreciated that in some examples, natural language processing module 732 is implemented using one or more machine learning mechanisms (e.g., neural networks). Specifically, the one or more machine learning mechanisms are configured to receive candidate textual representations and contextual information associated with the candidate textual representations. Based on the candidate textual representations and the associated contextual information, the one or more machine learning mechanisms are configured to determine an intent confidence score across a set of actionable intent candidates. Natural language processing module 732 can select one or more actionable intent candidates from the set of actionable intent candidates based on the determined intent confidence scores. In some examples, an ontology (e.g., ontology 760) is also used to select one or more actionable intent candidates from the set of actionable intent candidates.

[0182] Further details of searching ontologies based on token strings are described in U.S. Utility Patent Application No. 12 / 341,743, filed December 22, 2008, entitled "Method and Apparatus for Searching Using an Active Ontology," the entire disclosure of which is incorporated herein by reference.

[0183] In some examples, once the natural language processing module 732 identifies an actionable intent (or domain) based on the user request, the natural language processing module 732 generates a structured query to represent the identified actionable intent. In some examples, the structured query includes parameters for one or more nodes in the domain related to the actionable intent, at least some of which are populated with specific information and requirements specified in the user request. For example, a user may say, "Make me a dinner reservation at a sushi place at 7." In this case, the natural language processing module 732 may accurately identify the actionable intent as "restaurant reservation" based on the user input. According to the ontology, a structured query for the "restaurant reservation" domain may include parameters such as {cuisine}, {time}, {date}, and {number of participants}. In some implementations, based on the speech input and text derived from the speech input using the STT processing module 730, the natural language processing module 732 generates a partially structured query for the restaurant reservation domain, where the partially structured query includes the parameter {cuisine="sushi"} and the parameter {time="7:00 PM"}. However, in this implementation, the information included in the user's utterance is insufficient to complete a structured query associated with the domain. Therefore, other required parameters, such as {number of participants} and {date}, are not specified in the structured query based on the currently available information. In some implementations, the natural language processing module 732 populates some parameters of the structured query with received context information. For example, in some implementations, if the user requests a sushi restaurant "nearby," the natural language processing module 732 populates the {location} parameter in the structured query with GPS coordinates from the user device.

[0184] In some implementations, the natural language processing module 732 identifies multiple actionable intent candidates for each text expression candidate received from the STT processing module 730. Furthermore, in some implementations, a separate (partial or complete) structured query is generated for each identified actionable intent candidate. The natural language processing module 732 determines an intent confidence score for each actionable intent candidate and ranks the actionable intent candidates based on the intent confidence score. In some implementations, the natural language processing module 732 passes the generated structured query(s), including any input parameters, to the task flow processing module 736 (“task flow processor”). In some implementations, the structured query(s) for the m best (e.g., m highest-ranked) actionable intent candidates are provided to the task flow processing module 736 (where m is a predetermined integer greater than zero). In some implementations, the structured query(s) for the m best feasible intent candidates are provided to task flow processing module 736 along with the corresponding textual representation candidate(s).

[0185] Further details of inferring user intent based on multiple possible intent candidates determined from multiple candidate text representations of a speech input are described in U.S. Utility Patent Application No. 14 / 298,725, filed June 6, 2014, for "System and Method for Inferring User Intent From Speech Inputs," the entire disclosure of which is incorporated herein by reference.

[0186] Task flow processing module 736 is configured to receive the structured query(s) from natural language processing module 732, complete the structured query as needed, and perform the actions required to "complete" the user's ultimate request. In some implementations, the various steps required to complete these tasks are provided in task flow model 754. In some implementations, task flow model 754 includes steps for obtaining additional information from the user and task flows for performing actions associated with the actionable intent.

[0187] As described above, completing a structured query may require task flow processing module 736 to initiate additional dialogue with the user to obtain additional information and / or disambiguate potentially ambiguous utterances. If such dialogue is necessary, task flow processing module 736 invokes dialog flow processing module 734 to engage in a dialogue with the user. In some embodiments, dialog flow processing module 734 determines how (and / or when) to request additional information from the user and receives and processes the user response. Questions are provided to the user and answers are received from the user via I / O processing module 728. In some embodiments, dialog flow processing module 734 presents dialog output to the user via audio and / or visual output and receives input from the user via verbal or physical (e.g., click) responses. Continuing with the above example, when task flow processing module 736 invokes dialog flow processing module 734 to determine the "number of people" and "date" information for a structured query associated with the domain "restaurant reservation," dialog flow processing module 734 generates and passes questions to the user, such as "For how many people?" and "On which day?" Once answers are received from the user, dialog flow processing module 734 then either enters additional missing information into the structured query or passes the missing information to task flow processing module 736 to complete the missing information from the structured query.

[0188] Once the taskflow processing module 736 completes a structured query for an actionable intent, the taskflow processing module 736 proceeds to execute the final task associated with the actionable intent. Thus, the taskflow processing module 736 executes steps and instructions in a taskflow model according to specific parameters included in the structured query. For example, a taskflow model for an actionable intent of "restaurant reservation" includes steps and instructions for contacting a restaurant and actually requesting a reservation for a specific number of guests at a specific time. For example, using a structured query such as {Restaurant Reservation, Restaurant=ABC Cafe, Date=3 / 12 / 2012, Time=7 PM, Number of Guests=5}, the taskflow processing module 736 executes the following steps: (1) logging on to ABC Cafe's server or a restaurant reservation system such as OPENTABLE®; (2) entering date, time, and number of guests information into a form on the website; (3) submitting the form; and (4) entering a calendar item for the reservation into the user's calendar.

[0189] In some embodiments, task flow processing module 736 employs the assistance of service processing module 738 ("service processing module") to complete a task requested in the user input or to provide an answer to information requested in the user input. For example, service processing module 738 performs functions on behalf of task flow processing module 736, such as placing phone calls, setting calendar entries, invoking map searches, invoking or interacting with other user applications installed on the user device, and invoking or interacting with third-party services (e.g., restaurant reservation portals, social networking websites, banking portals, etc.). In some embodiments, the protocols and application programming interfaces (APIs) required by each service are specified by individual service models in service models 756. Service processing module 738 accesses the appropriate service model for a service and generates requests for that service in accordance with the protocols and APIs required by that service according to the service model.

[0190] For example, if a restaurant supports an online reservation service, the restaurant submits a service model that specifies the parameters required to make a reservation and an API for communicating the values of the required parameters to the online reservation service. Upon request by task flow processing module 736, service processing module 738 establishes a network connection with the online reservation service using the web address stored in the service model and transmits the required reservation parameters (e.g., time, date, number of participants) to the online reservation interface in a format that complies with the online reservation service's API.

[0191] In some examples, the natural language processing module 732, the dialog flow processing module 734, and the task flow processing module 736 are used collectively and iteratively to infer and define a user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response (i.e., output to the user or completion of a task) to satisfy the user's intent. The generated response is a dialog response to the speech input that at least partially satisfies the user's intent. Furthermore, in some examples, the generated response is output as speech output. In these examples, the generated response is sent to the speech synthesis processing module 740 (e.g., a speech synthesizer), where it may be processed to synthesize a dialog response in the form of a speech. In still other examples, the generated response is data content related to satisfying the user request in the speech input.

[0192] In embodiments in which task flow processing module 736 receives multiple structured queries from natural language processing module 732, task flow processing module 736 first processes a first structured query among the received structured queries and attempts to complete the first structured query and / or perform one or more tasks or actions represented by the first structured query. In some embodiments, the first structured query corresponds to the highest-ranked actionable intent. In other embodiments, the first structured query is selected from the received structured queries based on a combination of the corresponding speech recognition confidence score and the corresponding intent confidence score. In some embodiments, if task flow processing module 736 encounters an error while processing the first structured query (e.g., due to a required parameter being undeterminable), task flow processing module 736 can proceed to select and process a second structured query among the received structured queries that corresponds to a lower-ranked actionable intent. This second structured query is selected based on, for example, the speech recognition confidence scores of the corresponding text representation candidates, the intent confidence scores of the corresponding actionable intent candidates, the missing required parameters in the first structured query, or any combination thereof.

[0193] The speech synthesis processing module 740 is configured to synthesize speech output for presentation to the user. The speech synthesis processing module 740 synthesizes the speech output based on text provided by the digital assistant. For example, the generated dialog response is in the form of a text string. The speech synthesis processing module 740 converts the text string into an audible speech output. The speech synthesis processing module 740 uses any suitable speech synthesis technique to generate the speech output from the text, including, but not limited to, concatenative synthesis, unit selection synthesis, diphone synthesis, domain-specific synthesis, formant synthesis, articulatory synthesis, synthesis based on a hidden Markov model (HMM), and sinusoidal synthesis. In some embodiments, the speech synthesis processing module 740 is configured to synthesize individual words based on a phoneme sequence corresponding to the word. For example, a phoneme sequence is associated with a word in the generated dialog response. The phoneme sequence is stored in metadata associated with the word. The speech synthesis processing module 740 is configured to process the phoneme sequences in the metadata directly to synthesize words in phonetic form.

[0194] In some embodiments, instead of (or in addition to) using speech synthesis processing module 740, speech synthesis is performed on a remote device (e.g., server system 108), and the synthesized speech is sent to the user device for output to the user. For example, this can be done in some implementations where output to the digital assistant is generated on a server system. Also, because server systems generally have more processing power or resources than user devices, it is possible to obtain higher quality speech output than would be practical with client-side synthesis.

[0195] More information regarding digital assistants can be found in U.S. Utility Application No. 12 / 987,982, filed January 10, 2011, entitled "Intelligent Automated Assistant," and U.S. Utility Application No. 13 / 251,088, filed September 30, 2011, entitled "Generating and Processing Task Items That Represent Tasks to Perform," the disclosures of which are incorporated herein by reference in their entireties. 4. Digital Assistant Interaction in a Communication Session

[0196] 8A-8H illustrate systems and techniques for digital assistant (DA) interaction in a communication session, according to various embodiments.

[0197] 8A shows a block diagram of a device 800 and an external device 900. As described below, each of devices 800 and 900 is configured to handle DA interactions in a communication session.

[0198] Device 800 is implemented as, for example, device 200, 400, or 600. Similarly, external device 900 is implemented as, for example, another instance of device 200, 400, or 600. Although Figures 8B-8H below show device 800 and external device 900 each implemented as a smartphone device, devices 800 and 900 may each be implemented as another type of device, such as a smart watch, a smart speaker, an information appliance, a laptop computer, a desktop computer, a tablet device, a head-mounted device, etc.

[0199] The device 800 implements in memory 802 (e.g., as computer-executable instructions) a first DA 804, an audio control module 806, and a display control module 808. The first DA 804 is configured to provide DA services to the device 800, as described with respect to Figures 7A-7C. In some embodiments, the first DA 804 includes, at least in part, the components of the DA module 726 described above with respect to Figures 7A-7B.

[0200] The audio control module 806 is configured to manage the audio experience of the device 800 when the device 800 is engaged in a communication session. For example, in cooperation with other components of the device 800 configured to provide audio communications (e.g., the communications module 228, the RF circuitry 208, the audio circuitry 210, the speaker 211, the peripherals interface 218), the audio control module 806 is configured to cause the device 800 to transmit audio output generated by the first DA 804 to the other device(s) in the communication session and / or to cause the device 800 to output received audio generated by the other DA(s) respectively operating on the other device(s).

[0201] The display control module 808 is configured to manage the displayed experience of the device 800 when the device 800 is engaged in a communication session. For example, in cooperation with other components of the device 800 (e.g., the communications module 228, the RF circuitry 208, the display controller 256, the touch-sensitive display system 212, the peripherals interface 218), the display control module 808 is configured to cause the device 800 to display DA information (e.g., display output generated by the first DA 804 and a display indication of the status of the first DA 804) and information about the participants in the communication session (e.g., information identifying the participants in the communication session), to transmit instructions to the other device(s) to display the DA information, to receive instructions from the other device(s) to display the DA information, and / or to display the DA information based on the received instructions.

[0202] Similarly, the external device 900 implements in memory 902 a second DA 904, an audio control module 906, and a display control module 908. The second DA 904, the audio control module 906, and the display control module 908 are each similar to or substantially identical to the first DA 804, the audio control module 806, and the display control module 808, respectively. For example, the second DA 904 includes at least some components of the DA module 726. The audio control module 906 is configured to manage the audio experience of the external device 900 when the external device 900 is engaged in a communication session, e.g., as described with respect to the audio control module 806. The display control module 908 is configured to manage the displayed experience of the external device 900 when the external device 900 is engaged in a communication session, e.g., as described with respect to the display control module 808.

[0203] Although the description herein describes separate DAs participating in a communication session, in some examples, the present disclosure provides each participant in the communication session with a user experience consistent with a single DA participating in the communication session. For example, the first DA 804 and the second DA 904 may each generate a respective response in the communication session, but the first DA 804 and the second DA 904 may provide the response using the same voice characteristics. Furthermore, each device in the communication session may display only a single DA representation (e.g., the DA indicator 812 described below), and the display manner of the DA representation (e.g., indicating whether the DA is speaking) may be synchronized across the devices. Thus, it may appear that a single DA processes user requests and generates / provides responses to user requests. In other examples, the present disclosure provides each participant in the communication session with a user experience consistent with multiple DAs participating in the communication session. For example, the respective responses generated by the DAs 804 and 904 may have different voice characteristics, and / or each device may indicate which DA provides each response.

[0204] 8B shows device 800 engaged in a communication session with external device 900. While Figures 8B-8H below show device 800 engaged in a communication session with a single external device 900, in other examples, device 800 is engaged in communication sessions with multiple external devices. It will be understood that the techniques described below apply when device 800 is engaged in a communication session with one or more external devices (e.g., each implemented as a separate instance of external device 900).

[0205] A communication session is provided by multiple electronic devices and allows participants in the session to share communications, e.g., text, audio, and / or video communications. For example, a communication session may correspond to an audio communication session (e.g., a phone call), a video communication session (e.g., a video conference), a text communication session (e.g., a group text messaging session), and / or a virtual or mixed reality communication session. For example, in a virtual or mixed reality communication session, each of the participants' devices provides an audiovisual experience to simulate each participant (or their individual avatar) being simultaneously present at a shared location. For example, a virtual or mixed reality communication session may simulate each participant being present in a physical or virtual room in a house. In some embodiments, the communication session includes different types of communication experiences (e.g., audio, video, text, virtual or mixed reality) provided by each of the participants' devices. For example, in a communication session, a first device may provide a virtual or mixed reality communication experience (e.g., by displaying virtual representation(s) of the other participant(s) in a virtual setting), and a second device may provide a video communication experience (e.g., by displaying video of the other participant(s)). Thus, a communication session may be provided by multiple devices with different capabilities, e.g., by devices with virtual reality capabilities and devices with limited or no virtual reality capabilities, by devices with video capabilities and devices without video capabilities.

[0206] In Figure 8B, device 800 displays representation 810 indicating that a user of external device 900 (e.g., Rae) is participating in a communication session. Similarly, external device 900 displays representation 910 indicating that a user of device 800 (e.g., Tim) is participating in a communication session. While Figure 8B shows that representations 810 and 910 each show the name of an individual user, in other examples, representations 810 and 910 each include live video of the individual user and / or each include an individual physical representation (e.g., an avatar) of the individual user.

[0207] While the device 800 is engaged in a communication session with the external device 900, the device 800 receives an input from Tim to invoke the first DA 804. In some embodiments, some types of input invoke the DA, while other types of input indicate whether the natural language input is directed to the DA without invoking the DA, as described below. In some embodiments, the input to invoke the DA includes a verbal trigger input, such as a spoken input including a predetermined word or phrase for invoking the DA, e.g., “Hey Siri,” “Siri,” “Assistant,” “Wake Up,” etc. In some embodiments, the input to invoke the DA includes a selection of a button on the corresponding device (e.g., the device 800), such as a selection of a physical button on the device or a selection of an affordance displayed by the device. In some embodiments, the input to invoke the DA includes a detected user gaze input, e.g., an input indicating that the user's gaze has been directed toward a particular displayed affordance for a predetermined duration. In some embodiments, the device determines that the user gaze input is an input to invoke the DA based on the timing of the natural language input relative to the user gaze input. For example, the device determines that the user's gaze input should invoke the DA if the user's gaze is directed at the affordance at the start time of the natural language input and / or the end time of the natural language input. In some examples, the device interprets the user's gaze input as input to invoke the DA if the communication session does not include a currently invoked DA. When the communication session includes a currently invoked DA (e.g., the DA was invoked on any device in the communication session without being dismissed), the device interprets the user's gaze input to indicate that the natural language input is intended for the DA, but does not interpret the user's gaze input as input to invoke the DA.

[0208] 8B, Tim provides the verbal trigger input "Hey Siri" to device 800. In some implementations, device 800 transmits the verbal trigger input to external device 900, which outputs the verbal trigger input.

[0209] In some embodiments, in response to receiving an input to invoke the first DA 804, the device 800 invokes the first DA 804. For example, the device 800 displays a DA indicator 812 to indicate the invoked DA and initiates execution of a particular process and / or thread corresponding to the first DA 804. In some embodiments, invoking the first DA 804 in a communication session includes the device 800 causing the external device 900 (and any other device(s) in the communication session) to display the DA indicator 812. For example, using the display control module 808, the device 800 causes the external device 900 to display the DA indicator 812 simultaneously with the device 800.

[0210] In some embodiments, the device 800 synchronizes (e.g., using the display control module 808) the display state of the DA indicator 812 across each device in the communication session. Thus, each device in the communication session simultaneously displays the DA indicator 812 in the same state. For example, the device 800 displays the DA indicator 812 in different states corresponding to different visualizations to indicate different states of the first DA 804. For example, a first display state of the DA indicator 812 indicates that the first DA 804 is ready to accept (e.g., listen to) natural language input, a second display state indicates that the first DA 804 is currently receiving natural language input, a third display state indicates that the first DA 804 is currently processing natural language input, and a fourth display state indicates that the first DA 804 is currently responding to natural language input.

[0211] While device 800 is engaged in a communication session with external device 900, device 800 receives natural language input from Tim corresponding to a task. For example, in FIG. 8B , after Tim provides the verbal trigger input “Hey Siri,” Tim provides the spoken input “Send message ‘Hello’” intended to cause the first DA 804 to send a message. In some embodiments, device 800 transmits the natural language input to external device 900 (and any other external device(s) in the communication session), and the external device(s) each output the natural language input. For example, external device 900 outputs “Send message ‘Hello’,” thereby notifying Rae that Tim intends to send a message using the first DA 804.

[0212] In some embodiments, pursuant to invoking the first DA 804, the first DA 804 generates a prompt for further user input related to the task. For example, the first DA 804 processes natural language input (e.g., as described with respect to FIGS. 7A-7C ) to generate the prompt. In the example of FIG. 8B , in response to the natural language input “Send a message saying ‘Hello’,” the first DA 804 generates the prompt “Who do you want to send a message to?” In some embodiments, the device 800 outputs the prompt and transmits the prompt to the external device 900 (and any other device(s) in the communication session), e.g., using the audio control module 806. Upon receiving the transmitted prompt, the external device 900 outputs “Who do you want to send a message to?” to notify Rae of the first DA 804’s generated prompt.

[0213] The first DA 804 can allow any participant in the communication session to respond to the prompt. Thus, operating the first DA 804 in such a manner can improve the flexibility and efficiency with which the DA can fulfill requested tasks. For example, by not limiting responses to prompts to specific participants, a participant with the correct and / or optimal response to the prompt may provide a response. Operating the DA in such a manner can also provide more intuitive DA participation in a communication session, e.g., allowing any participant to respond to a question or prompt from another participant, much like in a communication session between human participants.

[0214] 8C , after transmitting the prompt for further user input, device 800 receives a response to the prompt from external device 900. In some embodiments, the response is received at external device 900 without external device 900 receiving input to invoke second DA 904 after outputting the prompt. In FIG. 8C , Rae responds to the prompt "Who would you like to send a message to?" by speaking "Qingwei." External device 900 receives the response "Qingwei" and transmits the response to device 800 (and any other device(s) in the communication session). In some embodiments, device 800 outputs a response to the received prompt.

[0215] Upon receiving a response to the prompt, the first DA 804 initiates a task based on the response and information stored on the device 800 and corresponding to the user of the device 800. The stored information may include, for example, information about Tim's contacts (e.g., contact identities, phone numbers, and email addresses stored on Tim's device), calendar information, videos, photos, notes, documents, text and email messages, health information, financial information, appliance information (e.g., on or off, locked or unlocked, temperature settings, etc.), applications installed on the device 800, device 800 settings, etc. In some embodiments, the first DA 804 initiates the task by processing the response using the user's information, as described with respect to FIGS. 7A-7C. In FIG. 8C, for example, the first DA 804 initiates a task to send a "Hello" message to Qingwei if Qingwei is Tim's (but not Rae's) contact.

[0216] In some embodiments, the first DA 804 generates an output indicating the initiated task, and the device 800 transmits the output to the external device 900 (and to any other device(s) in the communication session). In some embodiments, the output indicates the user whose information was used to initiate the task. For example, the first DA 804 generates the output “Would you like to send the message ‘Hello’ to Tim’s contact, Qingwei?” where “Tim’s contact, Qinqwei” indicates using Tim’s contact information to initiate the task. Upon receiving the output, the external device 900 provides the output. As another example, consistent with the techniques described above, if Tim requests the first DA 804 to “Create a new note,” and the first DA 804 generates the prompt “What do you want to say in the note?” and Rae responds with “Grocery list,” the first DA 804 will generate the output “OK, Tim’s note says ‘Grocery list’.”

[0217] Thus, the first DA 804 can use the information of the user who most recently invoked the DA (e.g., the current caller, Tim) to initiate a task, even though another user (e.g., Rae) responds to the first DA 804's prompt for further input. Operating the first DA 804 in such a manner can provide a more consistent and efficient user DA experience by avoiding confusion about which information the first DA 804 uses to initiate a task. Operating the first DA 804 in such a manner can further provide improved user feedback by indicating whose information is used to initiate a task, thereby confirming that the first DA 804 initiated the correct task (using the correct user's information) or allowing the user to correct the task (if it was initiated using the wrong user's information), which further improves the accuracy and efficiency of user DA interaction.

[0218] In some examples, initiating the task includes displaying an affordance 814 corresponding to the task on the device 800. In some examples, the first DA 804 generates the affordance 814 based on information corresponding to a user of the device 800. For example, in FIG. 8C , initiating the task of sending the message "Hello" to Qingwei includes displaying an affordance 814 indicating that the composed message to Qingwei is ready to be sent.

[0219] In some embodiments, device 800 synchronizes the display of affordances corresponding to the initiated task across each device in the communication session. For example, device 800 further causes (e.g., using display control module 808) external device 900 (and any other device(s) in the communication session) to each display affordance 814, e.g., simultaneously with device 800.

[0220] 8D , in some embodiments, device 800 also synchronizes any modifications to the displayed affordance across each device in the communication session. For example, device 800 receives user input corresponding to the selection of affordance 814 and, in response, displays affordance 814 in a modified state. In FIG. 8D , Tim selects affordance 814 and modifies the message to read, "Are you coming to the viewing party?" Thus, in some embodiments, the modified state corresponds to the modified textual content of affordance 814. In other examples, the modified state corresponds to the modified appearance (e.g., size, shape, color, brightness, animation) of affordance 814. In response to modifying the state of affordance 814, device 800 (e.g., using display control module 808) causes external device(s) to each display affordance 814 in the modified state (e.g., simultaneously with device 800). For example, in FIG. 8D, the external device 900 also displays a modified affordance 814 showing the updated message "Are you coming to the viewing party?"

[0221] In some embodiments, modifications to affordance 814 made by individual user(s) of external device(s) in the communication session are also synchronized across each device in the communication session. For example, at external device 900, Rae provides an input corresponding to selecting affordance 814, e.g., an input to magnify affordance 814. In response to receiving the input, external device 900 displays affordance 814 in a second modified state, e.g., a magnified state. In response to receiving the input, external device 900 further transmits to device 800, e.g., using display control module 908, an instruction to display affordance 814 in the second modified state. In response to receiving the instruction, device 800 displays affordance 814 in the second modified state. Thus, in response to magnifying affordance 814 at external device 900, each device in the communication session simultaneously displays the magnified affordance 814.

[0222] In some embodiments, each DA operating on a device of a participant in a communication session is configured to operate in a different language. For example, the speech recognition and natural language processing capabilities of each DA (e.g., provided by the STT processing module 730, the phonetic alphabet conversion module 731, the vocabulary 744, and the natural language processing module 732) are configured to operate in a different language. For example, a first DA 804 operating on device 800 is configured to operate in a first language (e.g., English), while a second DA 904 operating on external device 900 is configured to operate in a different second language (e.g., Chinese).

[0223] In some embodiments, the first DA 804 determines whether it can interpret the natural language input (e.g., a response to a prompt for further user input) in its respective language. In some embodiments, the first DA 804 determines whether it can interpret the natural language input based on an STT confidence score corresponding to the natural language input (e.g., determined by the STT processing module 730), an intent confidence score corresponding to the natural language input (e.g., determined by the natural language processing module 732), and / or whether the first DA 804 can successfully initiate a task based on the natural language input (e.g., using the taskflow processing module 736). For example, the first DA 804 determines that it can interpret the natural language input if the STT confidence score is above a threshold, if the intent confidence score is above a threshold, and / or if the first DA 804 can successfully initiate the task.

[0224] In the example of FIG. 8B, the natural language input "Send message 'Hello'" is in a first language (English), and the first DA 804 is configured to operate in the first language. Thus, the first DA 804 is able to interpret the natural language input in the first language. However, in FIG. 8C, assume that the response to the prompt "Who would you like to send a message to?" is in a second language (e.g., Chinese) that is different from the first language. For example, because the second DA 904 is configured to operate in Chinese, Rae is accustomed to interacting with DAs in Chinese. Thus, Rae may [Table 10] A Chinese response showing a request to send a message to a contact named [Table 11] The first DA 804 determines whether it can interpret the response in the first language. In this example, the first DA 804 determines that it cannot interpret the response in the first language (e.g., because the response's respective STT confidence score and / or intent confidence score are below respective thresholds and / or the first DA 804 is unable to initiate a task based on the response). Following a determination that the first DA 804 is unable to interpret the response in an output, the first DA 804 generates an error in the first language. For example, the first DA 804 generates an output that reads, "Unfortunately, we were unable to find a contact named Song How Young," where "Song How Young" is the Chinese response attempted by the first DA 804. [Table 12] The device 800 further transmits an output indicating the error to the external device(s) in the communication session. [Table 13] In response to the prompt, the external device 900 outputs, "Unfortunately, we were unable to find a contact named Song How Young." In this manner, the first DA 804 can advantageously indicate the correct language for interacting with the first DA 804 (e.g., via an English response), thereby facilitating efficient and accurate task initiation.

[0225] 8E-8F, in some embodiments, a user of an external device (different from device 800) in a communication session calls a DA. The called DA initiates the requested task using information corresponding to the user stored on the external device. In FIG. 8E, after device 800 transmits (and external device 900 outputs) the output "Would you like to send a message 'hello' to Tim's contact, Qingwei?", external device 900 receives an input from Rae to call a second DA 904. For example, Rae may send a verbal trigger input [Table 14] Upon receiving input to call the second DA 904, the DA call is forwarded from device 800 of FIG. 8D to external device 900 of FIG. 8E, meaning that Rae is the current (e.g., most recent) caller. Once called on external device 900, second DA 904 initiates tasks based on the information corresponding to Rae, for example, until the second DA 904 disappears (e.g., DA indicator 812 is no longer displayed and / or the particular process / thread corresponding to the second DA 904 is no longer running on external device 900) or until another device receives user input to call a DA on the other device.

[0226] In FIG. 8E, the external device 900 also receives, from Rae, a second natural language input corresponding to a second task. For example, Rae speaks in Chinese, "Watch Harry Potter together" ("Let's watch Harry Potter together"). In FIG. 8E, the external device 900 transmits the second natural language input to the device 800 (and to any other device(s) within the communication session). Upon receiving the second natural language input, the device 800 outputs the second natural language input.

[0227] In accordance with receiving an input to invoke the second DA 904, the second DA 904 generates a second prompt for user input regarding the second task, e.g., [Table 15] ("Which Harry Potter movie is it?") and transmits the second prompt to the device 800 (and to any other device(s) within the communication session). Upon receiving the second prompt, the device 800 outputs the second prompt.

[0228] In FIG. 8F, Tim provides a response to the second prompt. For example, the device 800 receives, from Tim, a response "Harry Potter and the Philosopher's Stone" ("Harry Potter and the Sorcerer's Stone"). In some embodiments, the device 800 receives the response without providing an input to invoke the first DA 804 after the device 800 receives the second prompt. The device 800 transmits the response to the external device 900 (and to any other device(s) within the communication session). Upon receiving the response, the second DA 904 initiates the second task based on the response and the information corresponding to Rae stored on the external device 900. For example, the second DA 904 initiates a task to watch "Harry Potter and the Philosopher's Stone", where the movie is stored on the external device 900 or associated with Rae's media subscription account. The second DA 904 outputs an indication of the initiated second task, e.g., [Table 16] ("Start watching together"). External device 900 transmits the output to device 800 (and any other device(s) in the communication session). Upon receiving the output, device 800 provides the output.

[0229] The above examples of FIGS. 8B-8F allow users different from the current DA caller to respond to the DA's prompts for user input. However, sometimes it may be desirable to allow only the current caller to successfully respond to the DA's prompts for further user input related to the requested task. In particular, for certain predetermined types of tasks, it may be undesirable for other users to instruct the current caller's DA to initiate the task using the current caller's data. Examples of such predetermined types of tasks include payment tasks (e.g., because it may be undesirable for other users to authorize payments to the current caller) and secure tasks such as home appliance control tasks (e.g., because it may be undesirable for other users to control the current caller's home appliances (e.g., unlock doors)). Accordingly, the following FIGS. 8G-8H illustrate techniques for processing DA requests related to such predetermined types of tasks in a communication session.

[0230] In FIG. 8G , while the device 800 is engaged in a communication session with the external device 900, and while the first DA 804 is calling on the device 800 (and monitoring the natural language input), the device 800 receives natural language input corresponding to a task. For example, the current caller, Tim, provides the natural language input, "Pay Larry $100." The device 800 transmits the natural language input to the external device 900 (and to any other device(s) in the communication session). The external device 900 outputs the received natural language input. The first DA 804 further generates a prompt for further user input related to the task, e.g., "OK, I'll pay Larry $100. Is that OK?" The device 800 outputs the prompt and transmits it to the external device 900 (and to any other device(s) in the communication session).

[0231] In some embodiments, the first DA 804 determines whether the task corresponds to a predetermined type of task. For example, the first DA 804 determines whether the natural language input corresponds to a predetermined type of domain. Exemplary predetermined type domains include a payment domain (e.g., associated with an actionable intent to make a payment and / or access payment information) and a home control domain (e.g., associated with an actionable intent to control the state and / or settings of a user's home appliances). In this example, the first DA 804 determines that the task of paying Larry $100 corresponds to a predetermined type of task.

[0232] 8H , Rae provides a response to the prompt, e.g., “Yes,” on the external device 900. The device 800 receives and outputs the prompt. However, in accordance with the first DA 804 determining that the task corresponds to a predetermined type of task and in accordance with receiving a response from the external device 900, the first DA 804 generates an output indicating that it cannot complete the task. For example, because the first DA 804 determines that the task corresponds to a predetermined type and determines that a response has been received from the external device, the first DA 804 generates “Sorry, I can’t do that. Tim, would you like to confirm?” In some embodiments, the output indicating that the first DA 804 cannot complete the task indicates an authorized user (e.g., Tim) to have the first DA 804 complete the task. The device 800 provides and transmits the output to the external device 900 (and any other device(s) in the communication session).

[0233] Continuing with the example, the device 800 receives from Tim a response (e.g., "Yes") to the prompt for further input. The device 800 further transmits the response to the external device 900 (and any other device(s) in the communication session). In accordance with determining that the task corresponds to a predetermined type of task and in accordance with determining that a response is received from the current caller (e.g., Tim), the first DA 804 initiates the task based on information corresponding to Tim. For example, because the task is of a predetermined type and the current caller confirms the task, the first DA 804 initiates a task to pay Larry $100, and Larry is one of Tim's contacts. The first DA 804 further generates output indicating the initiated task, e.g., "OK, paid Larry $100." The device 800 provides an output and transmits the output to the external device 900 (and any other device(s) in the communication session). In this manner, device security may be improved by allowing only authorized users (eg, the current caller) to allow secure tasks to be completed.

[0234] In some examples, if the task corresponds to a task of a predetermined type, the displayed response (e.g., affordance) generated by the DA is not synchronized across the external device(s) in the communication session. Thus, in some examples, causing the external device(s) to display the affordance 814 of FIGS. 8C-8D, respectively, is performed pursuant to a determination that the task (e.g., sending a message) does not correspond to a task of a predetermined type. In FIGS. 8G-8H, in response to receiving the natural language input "Pay Larry $100," the first DA 804 generates an affordance 816 corresponding to the payment task. Notably, the affordance 816 is displayed on the device 800 without the device 800 (e.g., using the display control module 808) causing the external device(s) to display the affordance 816, respectively. For example, pursuant to the first DA 804 determining that the task corresponds to a predetermined type, the display control module 808 does not instruct the other external device(s) to display the generated affordance 816. Furthermore, in some embodiments, modifications that Tim makes to the state of affordance 816 are also not synchronized across the external device(s) in the communication session.

[0235] In other examples, even if the task corresponds to a task of a predetermined type, the displayed responses generated by the DAs are synchronized across the external device(s) in the communication session. For example, if the first DA 804 determines that the task corresponds to a predetermined type, the first DA 804 causes (e.g., using the display control module 808) the external device(s) to respectively display the affordances in a non-interactive state. In some examples, the affordances displayed in the non-interactive state have a different display scheme (e.g., displayed in grayscale instead of color, have a smaller display size) than the affordances displayed in the normal (e.g., interactive) state. In some examples, when the device displays the affordances in a non-interactive state, the device does not allow user input to select the affordance to perform the corresponding task and / or does not allow user input to modify the content of the affordance. In some examples, when the device displays the affordances in a non-interactive state, the device does not display the content of the affordances (e.g., text, images). For example, device 800 displays affordance 816 of Figures 8G-8H in an interactive state, thereby allowing Tim to select the "Confirm" button to pay Larry $100. Device 900 (and any other device(s) in the communication session) can simultaneously display affordance 816 with device 800, but each display affordance 816 is in a non-interactive state. For example, device 900 displays affordance 816 in grayscale and does not allow Rae to select the "Confirm" button to pay Larry $100.

[0236] In some implementations, modifications made (e.g., by Tim) to the display of affordance 816 are synchronized across devices in the communication session, but affordance 816 remains displayed in a non-interactive state on the other device(s). For example, if Tim modifies the text content of affordance 816 (e.g., to pay Larry $50 instead of $100), the other device(s) may display affordance 816 to show the modified text content, but may still display affordance 816 in grayscale and not allow the individual user(s) to select the "Confirm" button.

[0237] 9A-9F and 10A-10D below illustrate techniques for DA interaction in a communication session when the DAs of the participants in the communication session are configured to operate in different languages, according to various embodiments. Sometimes, participants in a communication session may speak different languages, and thus each participant may be accustomed to interacting with the DA in a particular language. However, a DA called in a communication session may be configured to operate in a single language (e.g., the language of the current caller) that is different from the language that another participant habitually uses to interact with the DA. Therefore, it may be desirable to define the correct language for interacting with the DA in a communication session.

[0238] In FIG. 9A , while device 800 is engaged in a communication session with external device 900 (and optionally other external device(s)), device 800 receives an input to invoke a first DA 804. For example, Tim provides a verbal trigger input, "Hey Siri." In the examples of FIGS. 9A-9F and 10A-10D, the first DA 804 is configured to operate in a first language (e.g., English) and the second DA 904 is configured to operate in a second language (e.g., Spanish). Further, in FIGS. 9A-9E below, Tim is the current (most recent) DA invoker, meaning that external device 900 (and any other device(s) in the communication session) have not received an input to invoke their respective DAs during the communication session.

[0239] The device 800 receives a first natural language input in a first language. For example, Tim says, "What time is it now?", requesting the first DA 804 to provide the current time. The device 800 transmits the first natural language input to the external device 900 (and any other external device(s) in the communication session), and the external device(s) each output the first natural language input.

[0240] In response to invoking the first DA 804, the first DA 804 generates a first response in the first language to the first natural language input. The device 800 outputs the first response and transmits the first response to the external device 900 (and to any other device(s) in the communication session), and the device(s) each output the first response. For example, in FIG. 9A , the first DA 804 generates the first response, "It's 9:00 AM in Cupertino."

[0241] 9B , in some embodiments, if a participant intends to continue interacting with the DA in the communication session, the participant issues a second natural language input (a follow-up request) in the language in which their DA is configured to operate. For example, device 800 transmits the first response, "It's 9 AM in Cupertino," and after external device 900 outputs the first response, external device 900 receives a follow-up request in the second language. For example, in FIG. 9B , Rae issues a follow-up request in Spanish to the second DA 904. [Table 17] (How's Paris?)" Device 800 (and any other device(s) in the communication session) receives the follow-up request from external device 900. In some embodiments, device 800 further outputs the follow-up request.

[0242] In some embodiments, the follow-up request is received without the external device 900 receiving any input to invoke the second DA 904 after receiving the first response. For example, Rae may call the second DA 904 without providing a verbal trigger input (e.g., "Hola Siri") and without pressing any button on the external device 900. [Table 18] In some embodiments, a follow-up request does not include a response to a DA-generated prompt (e.g., a text or voice prompt generated by the first DA 804 or the second DA 904) for user input (e.g., regarding a previously requested task). Thus, a follow-up request may be an unprompted request to the second DA 904, as opposed to the prompted responses described with respect to FIGS. 8B-8H.

[0243] After device 800 outputs a first response (e.g., "It's 9 AM in Cupertino"), each DA in the communication session monitors for follow-up requests, e.g., for a predetermined duration after the first response is output. In some embodiments, monitoring for follow-up requests at the device includes determining that the follow-up request is intended for a DA operating on the device. Thus, a DA in the communication session can distinguish between natural language input intended as a conversation between participants (not intended for the DA) and a follow-up request intended for the DA.

[0244] In some examples, determining that the follow-up request is intended for the DA includes determining that the DA received a follow-up request from a user of the DA within a predetermined duration (e.g., 5 seconds, 10 seconds) after the first response was output. Thus, the DA determines that any natural language input received directly from a user of the DA (not transmitted from another device) and received within the predetermined duration is intended for itself. For example, the second DA 904 may determine that the follow-up request is intended for itself because the second DA 904 receives a follow-up request from Rae within a predetermined duration after the external device 900 outputs "It's 9 AM in Cupertino." [Table 19] The device 800 also receives a follow-up request, but the first DA 804 determines that the follow-up request is not intended for itself because the follow-up request was received from the external device 900 and not from Tim.

[0245] In some embodiments, determining that a follow-up request is intended for a DA is based on a gaze direction of a user of the DA. For example, each device detects user gaze data (e.g., via a device camera(s)) and analyzes the gaze data to determine whether a follow-up request received from a user of the device is intended for an individual DA. In some embodiments, each device detects user gaze data within a predetermined duration after the first response is output. In some embodiments, the device determines that a follow-up request is intended for an individual DA if the user gaze is directed at a gaze target displayed within a predetermined duration before the start time of the follow-up request (e.g., in the DA indicator 812 for the predetermined period), if the user gaze is directed at the gaze target at the start time, and / or if the user gaze is directed at the gaze target at the end time of the follow-up request. For example, the second DA 904 may determine that a follow-up request is intended for an individual DA based on the gaze data detected by the external device 900. [Table 20] The device 800 also determines that the follow-up request is intended for itself. [Table 21] However, the device 800 determines that the follow-up request is not intended for the first DA 804 because the follow-up request is received from an external device 900 (not Tim) and / or because the gaze data detected by the device 800 does not indicate that the follow-up request is intended for the first DA 804.

[0246] In some embodiments, following a determination that the follow-up request is directed to the DA, the DA generates a second response to the follow-up request in the individual language. In some embodiments, the DA generates the second response based on information corresponding to the user of the individual device (e.g., the user's contact information, calendar information, videos, photos, notes, etc.), where the information is stored on the individual device. In some embodiments, the second DA 904 generates the second response after the external device 900 outputs the first response, e.g., "It's 9 AM in Cupertino," without the external device 900 receiving an input to call the second DA 904. In FIG. 9B , the second DA 904 generates the second response based on information corresponding to the user of the individual device (e.g., the user's contact information, calendar information, videos, photos, notes, etc.), where the information is stored on the individual device. [Table 22] is intended for itself, the second DA 904 generates a second response in Spanish, "en Paris, son las 6 pm (It's 6 pm in Paris)." The external device 900 then transmits the second response to the device 800 (and any other device(s) in the communication session). In some embodiments, upon receiving the second response, the device 800 outputs the second response.

[0247] In some embodiments, the device 800 transmits context information associated with the first natural language input to the external device 900 (and any other device(s) in the communication session). In some embodiments, the context information includes a transcription of the communication session determined by the STT processing module 730, e.g., a transcription of the participants' conversation and the DA-generated response. In some embodiments, the context information includes the context information described above with respect to FIG. 7B, e.g., information collected by sensors of the device 800. In some embodiments, the context information indicates a conversational context associated with the first natural language input, e.g., a determined domain corresponding to the first natural language input. For example, the first DA 804 determines that the first natural language input, "What time is it?" corresponds to a time domain (e.g., associated with an actionable intent to provide time data) and causes the device 800 to transmit the context information indicating the time domain to the external device 900. In some embodiments, the second DA 904 (or another DA responsive to the follow-up request) generates a second response to the follow-up request based on the received context information. For example, using the context information, the second DA 904 may issue a follow-up request [Table 23] to mean asking about the time in Paris to produce the second response "en Paris,son las 6pm".

[0248] In some embodiments, when a DA in a communication session generates a response based on context information, the response indicates the context information used. Operating the DA in such a manner can avoid confusion among participants in the communication session about what the response refers to. Thus, in some embodiments, in accordance with a determination that the DA uses context information to interpret the natural language input and a determination that individual devices are involved in the communication session, the DA generates a response that indicates the context information. For example, the first response in FIG. 9A , “It's 9:00 AM in Cupertino,” indicates context information for Cupertino, California (e.g., Tim's current location) that is used by the first DA 804 to respond to “What time is it now?”

[0249] 9C illustrates an example in which Tim provides a follow-up request in a first language of a first DA 804. In FIG. 9C, continuing from FIG. 9A, after the device 800 outputs and transmits the first response, "It's 9:00 AM in Cupertino," the device 800 receives a follow-up request from Tim, "What about Paris?" The follow-up request is received without the device 800 receiving an input to invoke the first DA 804 after outputting and / or transmitting the first response. In some embodiments, the follow-up request does not include a response to a DA-generated prompt for further user input (e.g., regarding a previously requested task).

[0250] 9C , the first DA 804 determines that the follow-up request “How about Paris?” is targeted to itself, e.g., so that Tim utters the follow-up request within a predetermined duration after the device 800 outputs “It’s 9 AM in Cupertino.” Pursuant to receiving the follow-up request (and optionally, pursuant to determining that the follow-up request is targeted to the first DA 804), the first DA 804 generates a second response to the follow-up request in the first language. The second response indicates an initiated task corresponding to the follow-up request (e.g., getting the current time in Paris). In some embodiments, the first DA 804 generates the second response after the device 800 outputs the first response, e.g., “It’s 9 AM in Cupertino,” without the device 800 receiving input to call the first DA 804. For example, the first DA 804 generates a second response "It's 6 pm in Paris" and the device 800 transmits the second response to the external device 900 (and to any other external device(s) in the communication session).

[0251] 9D illustrates an example in which Tim attempts to provide a follow-up request in a language different from the language of the first DA 804. In FIG. 9D, continuing from FIG. 9A, after the device 800 outputs and transmits the first response, "It's 9 AM in Cupertino," the device 800 receives a follow-up request from Tim in a language different from the first language. For example, Tim may provide a follow-up request in Spanish. [Table 24] The follow-up request is received without the device 800 receiving input to invoke the first DA 804 after outputting and / or transmitting the first response. In some embodiments, the follow-up request does not include a response to a DA-generated prompt for further user input (e.g., regarding a previously requested task).

[0252] 9D , the first DA 804 determines that the follow-up request is intended for itself. In response to receiving the follow-up request (and optionally, in response to determining that the follow-up request is intended for the first DA 804), the first DA 804 generates a second response to the follow-up request in the first language. In some embodiments, the first DA 804 generates the second response after the device 800 outputs the first response, e.g., "It's 9 AM in Cupertino," without the device 800 receiving an input to call the first DA 804. In this example, the second response indicates that the first DA 804 cannot interpret the follow-up request. For example, because the first DA 804 is configured to operate in English, the first DA 804 may generate a second response to the follow-up request in Spanish. [Table 25] and therefore generates a second response, "Sorry, I don't understand." Device 800 transmits the second response to external device 900 (and to any other device(s) in the communication session).

[0253] FIG. 9E illustrates an example in which Rae attempts to provide a follow-up request in the language of the currently calling DA. Continuing from FIG. 9A , in FIG. 9E , after device 800 transmits the first response, “It's 9 AM in Cupertino,” to external device 900, and after external device 900 outputs the first response, external device 900 receives a follow-up request in the first language. For example, Rae utters, “How about Paris?” in English. External device 900 transmits the follow-up request to device 800 (and any other device(s) in the communication session), and device 800 receives the follow-up request. The follow-up request is received at external device 900 after external device 900 receives the transmitted first response, without receiving input to call the second DA 904. In some embodiments, the follow-up request does not include a response to the DA-generated prompt for further user input.

[0254] 9E , the second DA 904 determines that the follow-up request is intended for itself. In response to receiving the follow-up request (and optionally, in response to determining that the follow-up request is intended for the second DA 904), the second DA 904 generates a second response to the follow-up request in a second language. The second response indicates that the second DA 904 cannot interpret the follow-up request. In some embodiments, the second DA 904 generates the second response after the external device 900 outputs a first response, e.g., "It's 9 AM in Cupertino," without the external device 900 receiving an input to call the second DA 904. For example, because the second DA 904 is configured to operate in Spanish, the second DA 904 cannot interpret the English follow-up request, "How's Paris?" Therefore, the second DA 904 generates the second response, "no entiendo (I don't understand)." The external device 900 transmits a second response to the device 800 (and to any other device(s) in the communication session), and the device 800 receives the response.

[0255] 9F illustrates an example of Rae successfully interacting with the second DA 904 in the second language by providing input to invoke the second DA 904. In FIG. 9F, following FIG. 9A, after the device 800 transmits the first response "It's 9 AM in Cupertino" to the external device 900, and after the external device 900 outputs the first response, the external device 900 receives natural language input in the second language from Rae. For example, Rae [Table 26] The external device 900 transmits the natural language input to the device 800 (and to any other device(s) in the communication session). The device 800 receives and outputs the natural language input.

[0256] The external device 900 further receives an input from Rae to invoke the second DA 904. For example, after Rae provides a verbal trigger input "Hola Siri" or after the second DA 904 is activated in response to a button press on the external device 900, [Table 27] Thus, by providing input to call the second DA 904, the DA call is forwarded from the device 800 of Figure 9A to the external device 900 of Figure 9F, which means that Rae is the current caller.

[0257] In accordance with the external device 900 receiving the input to call the second DA 904, the second DA 904 generates a response to the natural language input in the second language. This response indicates that the second DA 904 will initiate a task corresponding to the natural language input. For example, the second DA 904 may generate the response "en Paris,son las 6pm" to indicate that it will initiate a task providing the current time in Paris. The external device 900 transmits the response to the device 800 (and any other device(s) in the communication session), and the device 800 receives the response.

[0258] The examples of FIGS. 9A-9F show that to successfully issue a follow-up request (e.g., without providing input to invoke the DA), each participant speaks the follow-up request in the language of their own DA. In other examples, a participant may issue a follow-up request in the language in which any DA in the communication session is configured to operate. For example, after an external device(s) receives and outputs the first response of FIG. 9A , “It’s 9 AM in Cupertino,” a participant may speak a follow-up request in the language of any DA in the communication session. In some examples, the follow-up request is received by the participant’s device without the device receiving input to invoke the DA after outputting the first response. In some examples, the follow-up request does not include a response to the DA-generated prompt for further user input.

[0259] In some embodiments, the DAs in a communication session collaborate with each other to determine the correct DA to process the follow-up request, e.g., the DA configured to operate in the language of the follow-up request. For example, in response to outputting the first response “It’s 9 AM in Cupertino” at their respective devices, each DA monitors follow-up requests. Upon receiving a follow-up request (e.g., from the user of the DA or from another device), each DA determines whether the follow-up request is intended for itself. For example, each DA performs STT processing on the follow-up request (e.g., using the STT processing module 730) to determine a speech-recognition confidence score. Each DA then causes its respective device to transmit the speech-recognition confidence score to each of the other devices. Each DA compares its determined speech-recognition confidence score with the speech-recognition confidence scores received from the other device(s). Thus, the DA with the highest speech-recognition confidence score (e.g., the DA that determines its generated confidence score is higher than all received confidence scores) determines that the follow-up request is intended for itself.

[0260] In some embodiments, upon determining that the follow-up request is intended for a particular DA, the DA causes its respective device to transmit instructions to the other device(s) indicating that the follow-up request is not intended for the other DA(s). The instructions instruct the other DA(s) to cease any ongoing processing of the follow-up request and not to generate any output in response to the follow-up request. In some embodiments, upon determining that the follow-up request is intended for a particular DA, the particular DA generates a response to the follow-up request in its respective language.

[0261] For example, continuing from FIG. 9A , assume that after the external device 900 outputs the first response, “It's 9:00 AM in Cupertino,” Rae utters a follow-up request, “What about Paris?” According to the techniques described above, the first DA 804 determines that the follow-up request is intended for itself (because the follow-up request is in English and the first DA 804 is configured to operate in English). Thus, the first DA 804 generates the response, “It's 6:00 PM in Paris,” which the device 800 transmits to the external device 900 (and any other device(s) in the communication session). In this way, Rae can receive the intended response to her follow-up request, which she utters in English, even though her second DA 904 is configured to operate in Spanish.

[0262] 10A-10D below show an example where a participant issues a follow-up request in the language that the currently invoked DA is configured to operate in to successfully issue the follow-up request.

[0263] In FIG. 10A, while device 800 is engaged in a communication session with external device 900 (and optionally other external device(s)), device 800 receives an input to call a first DA 804. For example, Tim provides the verbal trigger input "Hey Siri." In FIGS. 10A, 10B, and 10D below, Tim is the current caller, meaning that external device 900 (and any other device(s) in the communication session) have not received an input to call their respective DAs during the communication session.

[0264] Device 800 also receives a first natural language input in the first language from Tim. For example, after Tim says "Hey Siri," Tim utters "What time is it?" Device 800 transmits the first natural language input to external device 900 (and any other device(s) in the communication session).

[0265] In response to invoking the first DA 804, the first DA 804 generates a first response in the first language to the first natural language speech input. For example, the first DA 804 generates the first response "It's 9 AM in Cupertino." The device 800 outputs the first response. The device 800 further transmits the first response to the external device 900 (and any other device(s) in the communication session).

[0266] 10B , in some embodiments, if a participant intends to continue interacting with the DA in the communication session, the participant issues a second natural language input (a follow-up request) in the language of the currently calling DA. For example, after device 800 transmits the first response, "It's 9 AM in Cupertino," and external device 900 outputs the first response, external device 900 receives a follow-up request in the first language. In FIG. 9B , Rae asks in English, "How's Paris?" Device 800 (and any other device(s) in the communication session) receives the follow-up request from external device 900. In some embodiments, device 800 further outputs the follow-up request.

[0267] In some embodiments, the follow-up request is received without the external device 900 receiving input to invoke the second DA 904 after receiving the transmitted first response. For example, Rae utters "How's Paris?" without providing a verbal trigger input (e.g., "Hola Siri") and without pressing a button on the external device 900. In some embodiments, the follow-up request does not include a response to a DA-generated prompt for user input (e.g., regarding a previously requested task).

[0268] In response to the device 800 outputting the first response (e.g., "It's 9 AM in Cupertino"), the first DA 804 determines whether any follow-up requests received from Tim are intended for the currently called first DA 804. In response to the device 800 outputting the first response, the second DA 904 (and any other DA(s) in the communication session) determines whether any follow-up requests received from individual users of the DAs are intended for the currently called first DA 804. This is in contrast to the example described with respect to FIGS. 9A-9E, in which each DA determines whether follow-up requests received from individual users are intended for itself. Thus, the currently called first DA 804 may be the only DA in the communication session that can respond to follow-up requests.

[0269] In some examples, determining that the follow-up request is intended for the first DA 804 includes determining that the follow-up request was received within a predetermined duration after the respective device output the first response. For example, the second DA 904 determines that Rae's follow-up request "How's Paris?" is intended for the first DA 804 because the follow-up request was received within a predetermined duration after the external device 900 output "9:00 AM in Cupertino."

[0270] In some examples, determining that the follow-up request is intended for the first DA 804 is based on the user's gaze direction. For example, the device detects user gaze data (via the device camera(s)) and analyzes the gaze data to determine whether a follow-up request received from a user of the device is intended for the first DA 804. In some examples, the device detects user gaze data within a predetermined duration after a first response is output by, for example, the device 800. In some examples, the device determines that the follow-up request is intended for the first DA 804 if the user's gaze is directed at a gaze target displayed within a predetermined duration before the start time of the follow-up request (e.g., in the DA indicator 812 for a predetermined period of time), if the user's gaze is directed at the gaze target at the start time, and / or if the user's gaze is directed at the gaze target at the end time of the follow-up request. For example, the second DA 904 determines that Rae's follow-up request "How's Paris?" is intended for the first DA 804 based on detecting Rae's gaze at the DA indicator 812.

[0271] In some embodiments, in response to a DA (other than the first DA 804) determining that a follow-up request received from an individual user is intended for the first DA 804, the individual device transmits to the device 800 an indication that the follow-up request is intended for the first DA 804. For example, the external device 900 transmits to the device 800 an indication that the follow-up request "How's Paris?" is intended for the first DA 804. If the first DA 804 determines that the follow-up request received from that individual user (e.g., Tim) is intended for itself, the first DA 804 generates a second response to the follow request.

[0272] 10B , the first DA 804 generates a second response to the follow-up request in the first language in accordance with receiving the follow-up request from the external device 900. In some embodiments, the first DA 804 generates the second response further in accordance with the device 800 receiving an indication that the follow-up request is intended for the first DA 804. In some embodiments, the first DA 804 generates the second response without receiving further input to call the first DA 804 after the device 800 outputs the first response and / or without receiving input to call the second DA 904 after the external device 900 outputs the first response. For example, in accordance with the device 800 receiving an indication from the external device 900 that the follow-up request "How's Paris?" is intended for the first DA 804, the first DA 804 generates the second response "It's 6 PM in Paris." In some implementations, the device 800 transmits the second response to the external device 900 (and to any other device(s) in the communication session).

[0273] In some embodiments, the first DA 804 generates the second response based on context information (e.g., conversational context information) associated with the first natural language input and the first response. For example, the first natural language input "What time is it now?" and the first response "It's 9:00 AM in Cupertino" are associated with context information indicating a time domain. Thus, the first DA 804 generates the second response "It's 6:00 PM in Paris" by interpreting "How's Paris?" and referencing the time in Paris.

[0274] In some embodiments, the first DA 804 determines whether the received follow-up request corresponds to a personal request. A personal request generally describes a user request, the response of which depends on the particular user who provided the request. For example, a personal request corresponds to a personal domain, e.g., a domain associated with an actionable intent requiring retrieval / modification of personal data. Exemplary personal data includes a user's contact data, email data, message data, calendar data, reminder data, photos, videos, health information, financial information, web search history, media data (e.g., songs and audiobooks), information about the user's home (e.g., the status of the user's household appliances and home security system, home security system access information), and any other confidential and / or private information that the user may not want to disclose to other users or devices. Exemplary personal requests include "call my mom" (as users may have different mothers), "how many calories did I burn today?", "how much did I spend this month?", "show me the last photo I took," "turn off the porch light," "lock the front door," etc. In contrast, a non-personal request may have a response unrelated to the user who provided the non-personal request. Examples of non-personal requests include, "What is Taylor Swift's age?", "What's the weather in Palo Alto?", and "What's the score in the Patritas game?" In some embodiments, the first DA 804 determines whether a request corresponds to a personal request by determining whether the request corresponds to a personal domain. Further techniques for determining whether a request corresponds to a personal request are described in U.S. Patent Application No. 17 / 376,991, filed July 15, 2021, entitled "PERSONAL REQUEST CLASSIFIER," the contents of which are incorporated herein by reference in their entirety.

[0275] In some embodiments, in accordance with a determination that the follow-up request received from the external device corresponds to a personal request, the first DA 804 generates a second response to the follow-up request in the first language. The second response indicates that the first DA 804 cannot fulfill the user request included in the follow-up request. In some embodiments, the first DA 804 generates the second response without receiving further input to call the first DA 804 after the device 800 outputs the first response and / or without receiving input to call the second DA 904 after the external device 900 outputs the first response. For example, assume that Rae instead provides the follow-up request "Read my message" instead of "How's Paris?" in FIG. 10B. The first DA 804 determines that the follow-up request corresponds to a personal request and therefore generates the second response "Sorry, I can't do that." Because the first DA 804 running on the device 800 does not have access to Rae's personal data (e.g., messages) stored on the external device 900, the first DA 804 generates a second response. The device 800 then transmits the second response to the external device 900 (and any other device(s) in the communication session).

[0276] In this way, the first DA 804 provides feedback to the participants of the communication session that the first DA 804 cannot fulfill the personal request from the external device, for example, because the first DA 804 cannot access personal data stored on the external device for user privacy reasons. Thus, in some embodiments, generating the second response "It's 6 PM in Paris" in Figure 10B is performed pursuant to the first DA 804 determining that the follow-up request "How's Paris?" does not correspond to the personal request.

[0277] 10C shows an example in which Rae provides input to invoke the second DA 904 and successfully process the personal request. In FIG. 10C, Rae provides input to invoke the second DA 904. For example, the external device 900 receives the spoken trigger input "Hola Siri." Rae also provides natural language input in a second language corresponding to the personal request. For example, the external device 900 receives the Spanish phrase "lee mis mensajes" ("read my messages") from Rae. The external device 900 transmits the natural language input to the device 800 (and to any other device(s) in the communication session).

[0278] In accordance with invoking the second DA 904, the second DA 904 generates a response to the natural language input in the second language. The response indicates that the second DA 904 will initiate a task corresponding to the natural language input. For example, the second DA 904 generates the response "Anthony dijo'hola'" ("Anthony says 'hello'"), indicating that the second DA 904 will initiate the task of reading Rae's message. The external device 900 further transmits the response to the device 800 (and any other device(s) in the communication session). In this manner, Rae can successfully issue a personal request in the communication session by providing input to invoke the second DA 904. By calling the second DA 904, the DA call is forwarded from device 800 in FIG. 10A to external device 900 in FIG. 10C, which means that Rae is the current DA caller and the second DA 904 can access Rae's personal data and successfully process Rae's personal request (recall that the first DA 804 can only process Tim's personal request because it only has access to Tim's personal data).

[0279] 10D shows an example in which Rae attempts to provide a follow-up request in a second language on a second DA 904. In FIG. 10D, continuing the example of FIG. 10A, after the device 800 transmits the first response, "It's 9 AM in Cupertino," to the external device 900, and after the external device 900 outputs the first response, the external device 900 receives a follow-up request in a second language. For example, Rae may provide a follow-up request in Spanish in a second language. [Table 28] The external device 900 transmits a follow-up request to the device 800 (and any other device(s) in the communication session). The follow-up request is received at the external device 900 after the external device 900 receives the transmitted first response, without receiving any input to invoke the second DA 904. In some embodiments, the follow-up request does not include a response to a DA-generated prompt for further user input.

[0280] 10D , the second DA 904 determines that the follow-up request is intended for the first DA 804. In response to receiving the follow-up request (and optionally, in response to determining that the follow-up request is intended for the first DA 804), the first DA 804 generates a second response to the follow-up request in the first language. The second response indicates that the first DA 804 cannot interpret the follow-up request. In some embodiments, the first DA 804 generates the second response without receiving further input to call the first DA 804 after the device 800 outputs the first response and / or without receiving further input to call the second DA 904 after the external device 900 outputs the first response. For example, the first DA 804 generates the second response to the Spanish follow-up request. [Table 29] , and therefore generates the second response, "Sorry, I don't understand." The device 800 further transmits the second response to the external device 900 (and any other device(s) in the communication session). In this manner, the first DA 804 can indicate the correct language for interacting with the first DA 804 (e.g., via the second response in English), thereby facilitating accurate and efficient DA interaction. If Rae wishes to successfully interact with a Spanish DA, she can provide input at the external device 900 to invoke the second DA 904, for example, as described with respect to FIG. 10C. For example, if Rae tells the external device 900, " [Table 30] ", the DA call is forwarded from device 800 to external device 900, and Rae is the current caller. Therefore, the second DA 904 can generate the response "en Paris,son las 6pm" and process Rae's request successfully. 5. Process for Digital Assistant Interaction in a Communication Session

[0281] 11A-11C illustrate a process 1100 for DA interaction in a communication session, according to various embodiments. Process 1100 may be implemented, for example, using one or more electronic devices implementing a DA. In some embodiments, process 1100 is performed using a client-server system (e.g., system 100), with blocks of process 1100 divided in any manner between a server (e.g., DA server 106) and a client device. In other embodiments, blocks of process 1100 are divided between a server and multiple client devices (e.g., a mobile phone and a smartwatch). Thus, while portions of process 1100 are described herein as being performed by a particular device in a client-server system, it will be understood that process 1100 is not so limited. In other examples, process 1100 is implemented using only a client device (e.g., device 800) or multiple client devices. In process 1100, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some embodiments, additional steps may be performed in combination with process 1100.

[0282] In block 1102, while the electronic device (e.g., device 800) is engaged in a communication session with one or more external devices (e.g., external device 900), an input to invoke a first digital assistant running on the electronic device (e.g., first DA 804) is received from a first user (e.g., Tim) of the electronic device. In some embodiments, the input to invoke the first digital assistant includes a verbal trigger input or a button selection on the electronic device.

[0283] At block 1104, a natural language input corresponding to the task (eg, "Send a message saying hi" in FIG. 8B) is received from the first user.

[0284] In block 1106, following the invocation of the first digital assistant, a prompt for further user input regarding the task (e.g., "Who do you want to send a message to?" in FIG. 8B) is generated by the first digital assistant.

[0285] At block 1108, a prompt for further user input regarding the task is transmitted (eg, by device 800) to one or more external devices.

[0286] In block 1110, after transmitting the prompt for further user input, a response to the prompt for further user input (e.g., “Qingwei” in FIG. 8C) is received from an external device (e.g., external device 900) of the one or more external devices.

[0287] In block 1112, the task is initiated by the first digital assistant based on the response and information corresponding to the first user stored on the electronic device. In some implementations, as shown in block 1114, initiating the task includes displaying an affordance corresponding to the task on the electronic device (e.g., affordance 814 of FIG. 8C), where the first digital assistant generates the affordance based on the information corresponding to the first user.

[0288] In block 1116, an output indicating the initiated task (e.g., "Send the message 'Hello' to Tim's contact, Qingwei?" in FIG. 8C) is transmitted to one or more external devices (e.g., using audio control module 806).

[0289] In some examples, at block 1118, the one or more external devices are caused to respectively display the affordance (e.g., using the display control module 808). In some examples, causing the one or more external devices to respectively display the affordance is performed in accordance with a determination that the task does not correspond to a second predetermined type of task (e.g., a secure task). In some examples, displaying the affordance on the electronic device includes displaying the affordance without causing the one or more external devices to respectively display the affordance in accordance with a determination that the task corresponds to a second predetermined type of task.

[0290] In some examples, a first user input corresponding to a selection of the displayed affordance is received at block 1120. In some examples, in response to receiving the first user input, an affordance is displayed in a modified state (e.g., affordance 814 of FIG. 8D ) at block 1122. In some examples, at block 1124, one or more external devices are caused to respectively display the affordance in the modified state.

[0291] In some examples, at block 1126, an instruction to display the affordance in a second modified state is received from a third external device of the one or more external devices, and the affordance is displayed in the second modified state on the third external device in response to receiving, at the third external device, a second user input corresponding to a selection of the affordance. In some examples, at block 1128, in response to receiving the instruction, the affordance is displayed on the electronic device in the second modified state (e.g., using the display control module 808).

[0292] In some embodiments, in block 1130, it is determined (e.g., by the first DA 804) whether the task corresponds to a task of a predetermined type. In some embodiments, in block 1132, in accordance with a determination that the task corresponds to a task of a predetermined type and in accordance with receiving a response from the external device, a third output indicating that the first digital assistant cannot complete the task (e.g., "Sorry, I can't do that. Tim, would you like to confirm?" in FIG. 8H) is generated by the first digital assistant. In some embodiments, starting the task is performed in accordance with a determination that the task does not correspond to a task of a predetermined type.

[0293] In some embodiments, at block 1134, a third response to the prompt for further user input is received from the first user (e.g., "Yes" from Tim in FIG. 8H). In some embodiments, at block 1136, in accordance with a determination that the task corresponds to a predetermined type of task and in accordance with a determination that a third response has been received from the first user, the task is initiated by the first digital assistant based on the third response and information corresponding to the first user. In some embodiments, at block 1138, a fourth output indicating the initiated task (e.g., "OK, you paid Larry $100" in FIG. 8H) is transmitted to one or more external devices.

[0294] In some embodiments, the natural language input is the first language, and the first digital assistant is configured to operate in the first language. In some embodiments, at block 1140, it is determined (e.g., by the first DA 804) whether the first digital assistant can interpret responses in the first language, and starting the task is performed in accordance with the determination that the first digital assistant can interpret responses in the first language. In some embodiments, at block 1142, in accordance with the determination that the first digital assistant cannot interpret responses in the first language, a fifth output indicating an error is transmitted to one or more external devices. In some embodiments, a third digital assistant (e.g., the second DA 904) operating on an external device is configured to operate in a second language different from the first language.

[0295] In some embodiments, after transmitting an output indicating the started task to one or more external devices, a second natural language input corresponding to a second task (e.g., "Watch Harry Potter Together" in FIG. 8E) is received from a second external device among the one or more external devices, and the second external device receives the second natural language input from a second user of the second external device. In some embodiments, after receiving the second natural language input, a second prompt (e.g., JPEG2025114525000032.jpg7170) in FIG. 8E for user input regarding the second task is received from the second external device, and the second prompt is generated by a second digital assistant (e.g., the second DA 904) operating on the second external device. In some embodiments, a second response (e.g., "Harry Potter and the Philosopher's Stone" in FIG. 8F) to the second prompt for user input is received from the first user. In some embodiments, the second response is transmitted to one or more external devices. In some embodiments, the second output (e.g., JPEG2025114525000033.jpg7170) is received from the second external device, and the second digital assistant starts a second task based on the second response and information corresponding to the second user stored on the second external device, and the second output indicates the started second task. In some embodiments, the second prompt is received in accordance with the second external device receiving a second input from the second user to invoke the second digital assistant.

[0296] The operations described above with reference to Figures 11A-11C are optionally implemented by the components shown in Figures 1-4, 6A-6B, 7A-7C, and 8A. For example, the operations of process 1100 may be implemented by devices 800 and / or 900. It will be apparent to those skilled in the art how other processes may be implemented based on the components shown in Figures 1-4, 6A-6B, and 7A-7C.

[0297] 12A-12B illustrate a process 1200 for DA interaction in a communication session, according to various embodiments. Process 1200 may be implemented, for example, using one or more electronic devices implementing a DA. In some embodiments, process 1200 is performed using a client-server system (e.g., system 100), with blocks of process 1200 divided in any manner between a server (e.g., DA server 106) and a client device. In other embodiments, blocks of process 1200 are divided between a server and multiple client devices (e.g., a mobile phone and a smartwatch). Thus, while portions of process 1200 are described herein as being performed by a particular device in a client-server system, it will be understood that process 1200 is not so limited. In other examples, process 1200 is implemented using only a client device (e.g., device 800) or multiple client devices. In process 1200, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some embodiments, additional steps may be performed in combination with process 1200.

[0298] In block 1202, while the electronic device (e.g., device 800) is engaged in a communication session with one or more external devices (e.g., external device 900), an input to invoke a first digital assistant running on the electronic device (e.g., first DA 804) is received from a first user of the electronic device. The first digital assistant is configured to operate in a first language.

[0299] At block 1204, a first natural language input in a first language (eg, "What time is it?" in FIG. 9A) is received from a first user.

[0300] In block 1206, in response to the invocation of the first digital assistant, a first response in the first language to the first natural language input (e.g., "It's 9 AM in Cupertino" in FIG. 9A) is generated by the first digital assistant.

[0301] At block 1208, the first response is transmitted to one or more external devices (eg, using the audio control module 806).

[0302] In some implementations, at block 1210, contextual information associated with the first natural language input is transmitted to an external device.

[0303] At block 1212, after transmitting the first response, a second natural language input in a second language (e.g., [Table 31] ) is received from an external device of the one or more external devices. A second natural language input is received after receiving the transmitted first response, without the external device receiving a second input for invoking a second digital assistant (e.g., second DA904) running on the external device. The second digital assistant is configured to operate in the second language. In some examples, the second natural language input does not include a response to a prompt for user input generated by the first digital assistant or the second digital assistant. In some examples, the input for invoking the first digital assistant includes a verbal trigger input or a button selection on the electronic device, and the second input for invoking the second digital assistant includes a verbal trigger input or a button selection on the external device.

[0304] At block 1214, a second response in the second language to the second natural language input (e.g., "en Paris, son las 6 pm" in FIG. 9B) is received from the external device, and the second response is generated by the second digital assistant. In some examples, the second digital assistant generates the second response based on the context information. In some examples, the external device determines that the second natural language input is intended for the second digital assistant based on a gaze direction of a third user of the external device. In some examples, the second digital assistant generates the second response in accordance with a determination that the second natural language input is intended for the second digital assistant. In some examples, the second digital assistant generates the second response based on information corresponding to the second user of the external device, the information being stored on the external device.

[0305] In some examples, at block 1216, after transmitting the first response, a third natural language input in the first language (e.g., "How's Paris?" in FIG. 9E) is received from the external device. The third natural language input is received after the external device receives the transmitted first response without receiving a third input for invoking the second digital assistant. In some examples, at block 1218, a third response in the second language to the third natural language input (e.g., "no entiendo" in FIG. 9E) is received from the external device. The third response is generated by the second digital assistant and indicates that the second digital assistant cannot interpret the third natural language input.

[0306] In some implementations, at block 1220, after transmitting the first response, a fourth natural language input in the second language (e.g., [Table 32] ) is received from the external device. In some examples, at block 1222, a fourth response in the second language to the fourth natural language input (e.g., "en Paris, son las 6 pm" in FIG. 9F) is received from the external device. The fourth response is generated by the second digital assistant. The fourth response is received in accordance with the external device receiving a fourth input to invoke the second digital assistant (e.g., "Hola Siri" in FIG. 9F). The fourth response indicates that the second digital assistant has started a task corresponding to the fourth natural language input.

[0307] In some examples, at block 1224, after transmitting the first response, a fifth natural language input in the first language (e.g., "How's Paris?" in FIG. 9C ) is received from the first user. The fifth natural language input is received without the electronic device receiving a fifth input for invoking the first digital assistant after transmitting the first response. In some examples, in accordance with receiving the fifth natural language input at block 1126, a fifth response in the first language to the fifth natural language input (e.g., "It's 6 PM in Paris" in FIG. 9C ) is generated by the first digital assistant. The fifth response indicates a started task corresponding to the fifth natural language input. In some examples, at block 1228, the fifth response is transmitted to one or more external devices (e.g., using the audio control module 806).

[0308] In some implementations, at block 1230, after transmitting the first response, a sixth natural language input in a third language different from the first language (e.g., [Table 33] ) is received from the first user. A sixth natural language input is received by the electronic device after transmitting the first response, without receiving a sixth input for invoking the first digital assistant. In some examples, in block 1232, in accordance with receiving the sixth natural language input, a sixth response in the first language to the sixth natural language input (e.g., "Sorry, I don't understand" in FIG. 9D ) is generated by the first digital assistant. The sixth response indicates that the first digital assistant cannot interpret the sixth natural language input. In some examples, in block 1234, the sixth response is transmitted to one or more external devices (e.g., using the audio control module 806).

[0309] The operations described above with reference to Figures 12A-12B are optionally implemented by the components shown in Figures 1-4, 6A-6B, 7A-7C, and 8A. For example, the operations of process 1200 may be implemented by devices 800 and / or 900. It will be apparent to those skilled in the art how other processes may be implemented based on the components shown in Figures 1-4, 6A-6B, and 7A-7C.

[0310] 13A-13B illustrate a process 1300 for DA interaction in a communication session, according to various embodiments. Process 1300 may be implemented, for example, using one or more electronic devices implementing a DA. In some embodiments, process 1300 is performed using a client-server system (e.g., system 100), with blocks of process 1300 divided in any manner between a server (e.g., DA server 106) and a client device. In other embodiments, blocks of process 1300 are divided between a server and multiple client devices (e.g., a mobile phone and a smartwatch). Thus, while portions of process 1300 are described herein as being performed by a particular device in a client-server system, it will be understood that process 1300 is not so limited. In other examples, process 1300 is implemented using only a client device (e.g., device 800) or multiple client devices. In process 1300, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some embodiments, additional steps may be performed in combination with process 1300.

[0311] In block 1302, while the electronic device (e.g., device 800) is engaged in a communication session with one or more external devices (e.g., external device 900), an input to invoke a first digital assistant (e.g., first DA 804) running on the electronic device is received from a first user of the electronic device. The first digital assistant is configured to operate in a first language.

[0312] At block 1304, a first natural language input in a first language (eg, "What time is it?" in FIG. 10A) is received from a first user.

[0313] In block 1306, a first response in the first language to the first natural language input (e.g., "It's 9 AM in Cupertino") is generated by the first digital assistant pursuant to the invocation of the first digital assistant.

[0314] At block 1308, the first response is transmitted to one or more external devices (eg, using the audio control module 806).

[0315] In block 1310, after transmitting the first response, a second natural language input in the first language (e.g., "How's Paris?" in FIG. 10B) is received from an external device (e.g., external device 900) of the one or more external devices. The second natural language input is received after the external device receives the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device (e.g., second DA 904). The second digital assistant is configured to operate in the second language. In some examples, the second natural language input does not include a response to a prompt for user input generated by the first digital assistant or the second digital assistant. In some examples, the input for invoking the first digital assistant includes a verbal trigger input or a button selection on the electronic device, and the second input for invoking the second digital assistant includes a verbal trigger input or a button selection on the external device.

[0316] In block 1312, a second response in the first language to the second natural language input (e.g., "It's 6 PM in Paris" in FIG. 10B) is generated by the first digital assistant. In some examples, the first digital assistant generates the second response based on context information associated with the first natural language input and the first response. In some examples, the external device determines that the second natural language input is intended for the first digital assistant based on a gaze direction of a second user of the external device. In some examples, the first digital assistant generates the second response in accordance with a determination that the second natural language input is intended for the first digital assistant.

[0317] At block 1314, the second response is transmitted to one or more external devices (eg, using the audio control module 806).

[0318] In some implementations, at block 1316, a third natural language input in a second language (e.g., [Table 34] ) is received from the external device. A third natural language input is received by the external device after receiving the transmitted first response, without receiving a third input for invoking the second digital assistant. In some examples, in block 1318, in accordance with receiving the third natural language input, a third response in the first language to the third natural language input (e.g., "Sorry, I don't understand" in FIG. 10D) is generated by the first digital assistant. The third response indicates that the first digital assistant cannot interpret the third natural language input. In some examples, in block 1320, the third response is transmitted to one or more external devices (e.g., using the audio control module 806).

[0319] In some examples, at block 1322, a fourth natural language input in the second language is received from the external device. In some examples, at block 1324, a fourth response in the second language to the fourth natural language input is received from the external device. The fourth response is generated by the second digital assistant. The fourth response is received in accordance with the external device receiving a fourth input to invoke the second digital assistant. The fourth response indicates that the second digital assistant will start a task corresponding to the fourth natural language input.

[0320] In some embodiments, at block 1326, it is determined (e.g., by the first DA 804) whether the second natural language input corresponds to a personal request. In some embodiments, generating the second response is performed in accordance with a determination that the second natural language input does not correspond to a personal request. In some embodiments, at block 1328, a fifth response in the first language to the second natural language input is generated by the first digital assistant in accordance with a determination that the second natural language input corresponds to a personal request. The fifth response indicates that the first digital assistant cannot fulfill the user request included in the second natural language input. In some embodiments, at block 1330, the fifth response is transmitted to one or more external devices (e.g., using the audio control module 806).

[0321] In some examples, at block 1332, a sixth natural language input in the second language (e.g., "less mis mensaje" in FIG. 10C ) is received from the external device. The sixth natural language input corresponds to a personal request. In some examples, at block 1334, a sixth response in the second language to the sixth natural language input (e.g., "Anthony dijo "hola"" in FIG. 10C ) is received from the external device. The sixth response is generated by the second digital assistant. The sixth response is received in accordance with the external device receiving a sixth input to invoke the second digital assistant. The sixth response indicates that the second digital assistant has started a task corresponding to the sixth natural language input.

[0322] The operations described above with reference to Figures 13A-13B are optionally implemented by the components shown in Figures 1-4, 6A-6B, 7A-7C, and 8A. For example, the operations of process 1300 may be implemented by devices 800 and / or 900. It will be apparent to one skilled in the art how other processes may be implemented based on the components shown in Figures 1-4, 6A-6B, and 7A-7C.

[0323] According to some implementations, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) is provided that stores one or more programs for execution by one or more processors of an electronic device, the one or more programs including instructions for performing any of the methods or processes described herein.

[0324] According to some implementations, there is provided an electronic device (eg, a portable electronic device) comprising means for performing any of the methods or processes described herein.

[0325] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes a processing unit configured to perform any of the methods or processes described herein.

[0326] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes one or more processors and a memory that stores one or more programs for execution by the one or more processors, the one or more programs including instructions for performing any of the methods or processes described herein.

[0327] The foregoing has been described with reference to specific embodiments for purposes of explanation. However, the exemplary discussion above is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described to best explain the principles of the technology and its practical applications so that others skilled in the art can best utilize the technology and various embodiments with various modifications as suited to the particular applications intended.

[0328] Although the present disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art, and such changes and modifications are to be understood as being included within the scope of the present disclosure and examples, as defined by the claims.

[0329] As explained above, one aspect of the present technology is the collection and use of available data from various sources to generate a DA response in a communication session. This disclosure contemplates that, in some cases, this collected data may include personal information data that uniquely identifies a particular person or that can be used to contact or locate a particular person. Such personal information data may include demographic data, location-based data, phone numbers, email addresses, Twitter IDs, home addresses, data or records regarding a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.

[0330] This disclosure recognizes that the use of such personal information data in the present technology can be for the benefit of the user. For example, the personal information data can be used to fulfill a user's request for a digital assistant. Additionally, other uses of the personal information data that benefit the user are contemplated by this disclosure. For example, health and fitness data can be used to provide insight into the user's overall wellness, or can be used as proactive feedback to individuals using the technology in pursuit of wellness goals.

[0331] This disclosure contemplates that entities involved in the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will adhere to robust privacy policies and / or privacy practices. Specifically, such entities should implement and consistently use privacy policies and practices that are generally recognized as meeting or exceeding industry or government requirements for maintaining the strict confidentiality of personal information data. Such policies should be easily accessible to users and should be updated as data collection and / or use changes. Personal information from users should be collected for the entity's lawful and legitimate use and should not be shared or sold except for those lawful uses. Furthermore, such collection / sharing should be carried out after the user's informed consent is obtained. Furthermore, such entities should consider taking all necessary measures to protect and secure access to such personal information data and to ensure that others with access to the personal information data adhere to their privacy policies and procedures. Furthermore, such entities may be able to undergo third-party assessments to demonstrate their adherence to widely accepted privacy policies and practices. Furthermore, policies and practices should be tailored to the specific types of personal data collected and / or accessed and should comply with applicable laws and standards, including jurisdiction-specific considerations. For example, in the United States, collection of or access to certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA). Meanwhile, health data in other countries may be subject to other regulations and policies and should be addressed accordingly. Therefore, different privacy practices should be maintained in each country with respect to different types of personal data.

[0332] Notwithstanding the foregoing, the present disclosure also contemplates embodiments in which a user selectively blocks use of or access to personal information data. That is, the present disclosure contemplates that hardware and / or software elements may be provided to prevent or block access to such personal information data. For example, when providing a DA response in a communication session, the present disclosure may be configured to allow a user to select "opt-in" or "opt-out" of participating in the collection of personal information data during service registration or any time thereafter. In another example, a user may choose not to allow a DA to access personal information data (e.g., during a communication session). In yet another example, a user may choose to limit the amount of time the DA can access personal information data. In addition to providing "opt-in" and "opt-out" options, the present disclosure contemplates providing notifications regarding access or use of personal information. For example, a user may be notified upon downloading an app that will access the user's personal information data, and then again immediately before the app accesses the user's personal information data.

[0333] Furthermore, it is the intent of this disclosure that personal information data should be managed and handled in a manner that minimizes the risk of unintentional or unauthorized access or use. Risk can be minimized by limiting data collection and deleting data when it is no longer needed. Additionally, where applicable in certain health-related applications, data anonymization can be used to protect user privacy. Anonymization can be facilitated, as needed, by removing certain identifiers (e.g., date of birth, etc.), controlling the amount or specificity of data stored (e.g., collecting location data at the city level rather than the address level), controlling how data is stored (e.g., aggregating data across users), and / or other methods.

[0334] Thus, while this disclosure broadly encompasses the use of personal information data to implement one or more various disclosed embodiments, this disclosure also contemplates that the various embodiments may be implemented without requiring access to such personal information data. That is, various embodiments of the present technology are not rendered inoperable by the absence of all or part of such personal information data. For example, a DA may fulfill a user request based on non-personal information data or a minimal amount of personal information, such as content requested by a device associated with the user, other non-personal information available to the DA, or publicly available information.

Claims

1. 1. A method comprising: In an electronic device having one or more processors and a memory, While the electronic device is engaged in a communication session with one or more external devices, Receiving an input from a first user of the electronic device to invoke a first digital assistant running on the electronic device; receiving a natural language input from the first user corresponding to a task; In response to invoking the first digital assistant, generating, by the first digital assistant, a prompt for further user input regarding the task; transmitting the prompt for further user input regarding the task to the one or more external devices; receiving, after transmitting the prompt for further user input, a response to the prompt for further user input from an external device of the one or more external devices; Initiating the task by the first digital assistant based on the response and information corresponding to the first user stored on the electronic device; transmitting an output indicating the initiated task to the one or more external devices.

2. In the electronic device, after transmitting the output indicating the initiated task to the one or more external devices; receiving a second natural language input corresponding to a second task from a second external device of the one or more external devices, the second external device receiving the second natural language input from a second user of the second external device; After receiving the second natural language input, receiving a second prompt for user input related to the second task from the second external device, the second prompt being generated by a second digital assistant running on the second external device; receiving a second response to the second prompt for user input from the first user; transmitting the second response to the one or more external devices; receiving a second output from the second external device, The second digital assistant starts the second task based on the second response and information corresponding to the second user stored on the second external device; The method of claim 1 , wherein the second output indicates the started second task.

3. 3. The method of claim 2, wherein the second prompt is received in accordance with the second external device receiving a second input from the second user to invoke the second digital assistant.

4. determining whether the task corresponds to a predetermined type of task; in response to a determination that the task corresponds to a task of the predetermined type; The method of any one of claims 1 to 3, further comprising: generating, by the first digital assistant, a third output indicating that the first digital assistant is unable to complete the task in accordance with receiving the response from the external device; wherein starting the task is performed in accordance with a determination that the task does not correspond to the predetermined type of task.

5. receiving a third response to the prompt for further user input from the first user; In accordance with a determination that the task corresponds to the predetermined type of task and in accordance with a determination that the third response has been received from the first user, starting the task by the first digital assistant based on the third response and the information corresponding to the first user; The method of claim 4 , further comprising: transmitting a fourth output indicating the initiated task to the one or more external devices.

6. The natural language input is in a first language, and the first digital assistant is configured to operate in the first language, and the method includes: Determining whether the first digital assistant can interpret the response in the first language, wherein starting the task is performed according to a determination that the first digital assistant can interpret the response in the first language; 6. The method of claim 1, further comprising: transmitting a fifth output indicating an error to the one or more external devices in accordance with a determination that the first digital assistant cannot interpret the response in the first language.

7. The method of claim 6 , wherein the third digital assistant operating on the external device is configured to operate in a second language different from the first language.

8. Initiating the task includes displaying an affordance corresponding to the task on the electronic device, and the first digital assistant generates the affordance based on the information corresponding to the first user, and the method further comprises: The method of claim 1 , further comprising causing the one or more external devices to respectively display the affordances.

9. causing the one or more external devices to respectively display the affordances is performed in accordance with a determination that the task does not correspond to a second predetermined type of task; Displaying the affordance on the electronic device comprises:

9. The method of claim 8, further comprising: displaying the affordances of the one or more external devices without causing the affordances to be displayed respectively in accordance with a determination that the task corresponds to the second predetermined type of task.

10. receiving a first user input corresponding to a selection of the displayed affordance; displaying the affordance in a modified state in response to receiving the first user input; The method of claim 8 or 9, further comprising causing the one or more external devices to respectively display the affordance in the modified state.

11. receiving, from a third external device among the one or more external devices, an instruction to display the affordance in a second modified state, the affordance being displayed in the second modified state on the third external device in response to receiving, at the third external device, a second user input corresponding to a selection of the affordance; 11. The method of claim 8, further comprising: in response to receiving the command, displaying the affordance in the second modified state on the electronic device.

12. 12. The method of claim 1, wherein the input for invoking the first digital assistant comprises a verbal trigger input or a button selection on the electronic device.

13. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions While the electronic device is engaged in a communication session with one or more external devices, Receive an input from a first user of the electronic device to invoke a first digital assistant running on the electronic device; receiving a natural language input from the first user corresponding to a task; In response to invoking the first digital assistant, generating, by the first digital assistant, a prompt for further user input regarding the task; transmitting the prompt for further user input regarding the task to the one or more external devices; receiving a response to the prompt for further user input from an external device of the one or more external devices after transmitting the prompt for further user input; Initiating, by the first digital assistant, the task based on the response and information corresponding to the first user stored on the electronic device; The electronic device transmits an output indicating the initiated task to the one or more external devices.

14. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a first electronic device, cause the first electronic device to: While the electronic device is engaged in a communication session with one or more external devices, Receive an input from a first user of the electronic device to invoke a first digital assistant running on the electronic device; receiving a natural language input from the first user corresponding to a task; In response to invoking the first digital assistant, generating a prompt by the first digital assistant for further user input regarding the task; transmitting the prompt for further user input regarding the task to the one or more external devices; receiving, after transmitting the prompt for further user input, a response to the prompt for further user input from an external device of the one or more external devices; Initiating the task by the first digital assistant based on the response and information corresponding to the first user stored on the electronic device; a non-transitory computer-readable storage medium that causes an output indicative of the initiated task to be transmitted to the one or more external devices;

15. 1. An electronic device comprising: While the electronic device is engaged in a communication session with one or more external devices, Receive an input from a first user of the electronic device to invoke a first digital assistant running on the electronic device; receiving a natural language input from the first user corresponding to a task; In response to invoking the first digital assistant, generating, by the first digital assistant, a prompt for further user input regarding the task; transmitting the prompt for further user input regarding the task to the one or more external devices; receiving a response to the prompt for further user input from an external device of the one or more external devices after transmitting the prompt for further user input; Initiating, by the first digital assistant, the task based on the response and information corresponding to the first user stored on the electronic device; means for transmitting an output indicative of the initiated task to the one or more external devices.

16. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 1 to 12.

17. 13. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 1 to 12.

18. 1. An electronic device comprising: An electronic device comprising means for carrying out the method according to any one of claims 1 to 12.

19. 1. A method comprising: In an electronic device having one or more processors and a memory, While the electronic device is engaged in a communication session with one or more external devices, Receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; In response to invoking the first digital assistant, generating, by the first digital assistant, a first response in the first language to the first natural language input; transmitting the first response to the one or more external devices; after transmitting the first response, a second natural language input in a second language from an external device of the one or more external devices; The second natural language input is received after the external device receives the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device; The second digital assistant is configured to operate in the second language, and receives a second natural language input; receiving from the external device a second response in the second language to the second natural language input, wherein the second response is generated by the second digital assistant.

20. 20. The method of claim 19, further comprising transmitting contextual information associated with the first natural language input to the external device, wherein the second digital assistant generates the second response based on the contextual information.

21. After transmitting the first response, from the external device, a third natural language input in the first language, the third natural language input being received without the external device receiving a third input for invoking the second digital assistant after receiving the transmitted first response; 21. The method of claim 19 or 20, further comprising receiving a third response in the second language from the external device to the third natural language input, the third response being generated by the second digital assistant and indicating that the second digital assistant is unable to interpret the third natural language input.

22. receiving a fourth natural language input in the second language from the external device after transmitting the first response; receiving from the external device a fourth response in the second language to the fourth natural language input, the fourth response is generated by the second digital assistant; The fourth response is received in accordance with the external device receiving a fourth input for calling the second digital assistant; 22. The method of any one of claims 19 to 21, wherein the fourth response indicates that the second digital assistant has initiated a task corresponding to the fourth natural language input.

23. 23. The method of any one of claims 19 to 22, wherein the second digital assistant generates the second response based on information corresponding to a second user of the external device, the information being stored on the external device.

24. After transmitting the first response, receive a fifth natural language input in the first language from the first user, the fifth natural language input being received without the electronic device receiving a fifth input for invoking the first digital assistant after transmitting the first response; In response to receiving the fifth natural language input, generating, by the first digital assistant, a fifth response in the first language to the fifth natural language input, the fifth response indicating a started task corresponding to the fifth natural language input; 24. The method of claim 19, further comprising: transmitting the fifth response to the one or more external devices.

25. After transmitting the first response, receiving a sixth natural language input from the first user in a third language different from the first language, the sixth natural language input being received without the electronic device receiving a sixth input for invoking the first digital assistant after transmitting the first response; In response to receiving the sixth natural language input, generating, by the first digital assistant, a sixth response in the first language to the sixth natural language input, the sixth response indicating that the first digital assistant cannot interpret the sixth natural language input; 25. The method of claim 19, further comprising: transmitting the sixth response to the one or more external devices.

26. The external device determines that the second natural language input is intended for the second digital assistant based on a gaze direction of a third user of the external device; 26. The method of any one of claims 19 to 25, wherein the second digital assistant generates the second response in accordance with a determination that the second natural language input is intended for the second digital assistant.

27. 27. The method of any one of claims 19 to 26, wherein the second natural language input does not include a response to a prompt for user input generated by the first digital assistant or the second digital assistant.

28. The input for invoking the first digital assistant includes a verbal trigger input or a button selection on the electronic device; 28. The method of claim 19, wherein the second input for invoking the second digital assistant includes the verbal trigger input or the selection of a button on the external device.

29. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions While the electronic device is engaged in a communication session with one or more external devices, receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; generating, by the first digital assistant in response to the first natural language input, a first response in the first language; transmitting the first response to the one or more external devices; after transmitting the first response, a second natural language input in a second language from an external device of the one or more external devices; The second natural language input is received by the external device after receiving the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device; The second digital assistant receives a second natural language input, the second digital assistant being configured to operate in the second language; and receiving from the external device a second response in the second language to the second natural language input, the second response being generated by the second digital assistant.

30. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a first electronic device, cause the first electronic device to: While the electronic device is engaged in a communication session with one or more external devices, receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; In response to invoking the first digital assistant, causing the first digital assistant to generate a first response in the first language to the first natural language input; transmitting the first response to the one or more external devices; after transmitting the first response, a second natural language input in a second language from an external device of the one or more external devices; The second natural language input is received after the external device receives the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device; The second digital assistant receives a second natural language input, the second digital assistant being configured to operate in the second language; A non-transitory computer-readable storage medium that receives from the external device a second response in the second language to the second natural language input, the second response being generated by the second digital assistant.

31. 1. An electronic device comprising: While the electronic device is engaged in a communication session with one or more external devices, receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; generating, by the first digital assistant in response to the first natural language input, a first response in the first language; transmitting the first response to the one or more external devices; after transmitting the first response, a second natural language input in a second language from an external device of the one or more external devices; The second natural language input is received after the external device receives the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device; The second digital assistant receives a second natural language input, the second digital assistant being configured to operate in the second language; An electronic device comprising: means for receiving from the external device a second response in the second language to the second natural language input, the second response being generated by the second digital assistant.

32. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 19 to 28.

33. 29. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 19 to 28.

34. 1. An electronic device comprising:

29. An electronic device comprising means for carrying out the method of any one of claims 19 to 28.

35. 1. A method comprising: In an electronic device having one or more processors and a memory, While the electronic device is engaged in a communication session with one or more external devices, Receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; In response to invoking the first digital assistant, generating, by the first digital assistant, a first response in the first language to the first natural language input; transmitting the first response to one or more of the external devices; after transmitting the first response, a second natural language input in the first language from an external device of the one or more external devices; The second natural language input is received after the external device receives the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device; The second digital assistant is configured to operate in a second language, and receives a second natural language input; generating, by the first digital assistant, a second response in the first language to the second natural language input; transmitting the second response to the one or more external devices.

36. 36. The method of claim 35, wherein the first digital assistant generates the second response based on contextual information associated with the first natural language input and the first response.

37. Receiving a third natural language input in the second language from the external device, the third natural language input being received by the external device after receiving the transmitted first response without receiving a third input for invoking the second digital assistant; In response to receiving the third natural language input, generating, by the first digital assistant, a third response in the first language to the third natural language input, the third response indicating that the first digital assistant cannot interpret the third natural language input; 37. The method of claim 35 or 36, further comprising: transmitting the third response to the one or more external devices.

38. receiving a fourth natural language input in the second language from the external device; receiving from the external device a fourth response in the second language to the fourth natural language input, the fourth response is generated by the second digital assistant; The fourth response is received in accordance with the external device receiving a fourth input for calling the second digital assistant; 38. The method of any one of claims 35 to 37, wherein the fourth response indicates that the second digital assistant initiates a task corresponding to the fourth natural language input.

39. 39. The method of any one of claims 35 to 38, further comprising determining whether the second natural language input corresponds to a personal request, and wherein generating the second response is performed in accordance with a determination that the second natural language input does not correspond to a personal request.

40. 1. A method comprising: In accordance with a determination that the second natural language input corresponds to a personal request, generating, by the first digital assistant, a fifth response in the first language to the second natural language input, the fifth response indicating that the first digital assistant cannot fulfill the user request included in the second natural language input; 40. The method of claim 39, further comprising: transmitting the fifth response to the one or more external devices.

41. receiving a sixth natural language input in the second language from the external device, the sixth natural language input corresponding to a personal request; receiving, from the external device, a sixth response in the second language to the sixth natural language input, The sixth response is generated by the second digital assistant; The sixth response is received in accordance with the external device receiving a sixth input for calling the second digital assistant; 41. The method of any one of claims 35 to 40, wherein the sixth response indicates that the second digital assistant has started a task corresponding to the sixth natural language input.

42. The external device determines that the second natural language input is intended for the first digital assistant based on a gaze direction of a second user of the external device; 42. The method of any one of claims 35 to 41, wherein the first digital assistant generates the second response in accordance with a determination that the second natural language input is intended for the first digital assistant.

43. 43. The method of any one of claims 35 to 42, wherein the second natural language input does not include a response to a prompt for user input generated by the first digital assistant or the second digital assistant.

44. The input for invoking the first digital assistant includes a verbal trigger input or a button selection on the electronic device; 44. The method of any one of claims 35 to 43, wherein the second input for invoking the second digital assistant includes the verbal trigger input or the selection of a button on the external device.

45. 1. An electronic device comprising: one or more processors; Memory and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions, the instructions While the electronic device is engaged in a communication session with one or more external devices, receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; generating, by the first digital assistant in response to the first natural language input, a first response in the first language; transmitting the first response to the one or more external devices; after transmitting the first response, a second natural language input in the first language from an external device of the one or more external devices; The second natural language input is received by the external device after receiving the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device; The second digital assistant receives a second natural language input, the second digital assistant being configured to operate in a second language; generating, by the first digital assistant, a second response in the first language to the second natural language input; The electronic device transmits the second response to the one or more external devices.

46. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a first electronic device, cause the first electronic device to: While the electronic device is engaged in a communication session with one or more external devices, receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; In response to invoking the first digital assistant, causing the first digital assistant to generate a first response in the first language to the first natural language input; transmitting the first response to the one or more external devices; after transmitting the first response, a second natural language input in the first language from an external device of the one or more external devices; The second natural language input is received after the external device receives the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device; The second digital assistant is configured to operate in a second language and receives a second natural language input; causing the first digital assistant to generate a second response in the first language to the second natural language input; A non-transitory computer-readable storage medium that causes the second response to be transmitted to the one or more external devices.

47. 1. An electronic device comprising: While the electronic device is engaged in a communication session with one or more external devices, receiving an input from a first user of the electronic device to invoke a first digital assistant operating on the electronic device, the first digital assistant being configured to operate in a first language; receiving a first natural language input in the first language from the first user; generating, by the first digital assistant in response to the first natural language input, a first response in the first language; transmitting the first response to the one or more external devices; after transmitting the first response, a second natural language input in the first language from an external device of the one or more external devices; The second natural language input is received after the external device receives the transmitted first response without receiving a second input for invoking a second digital assistant running on the external device; The second digital assistant receives a second natural language input, the second digital assistant being configured to operate in a second language; generating, by the first digital assistant, a second response in the first language to the second natural language input; means for transmitting the second response to the one or more external devices.

48. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 35 to 44.

49. 45. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 35 to 44.

50. 1. An electronic device comprising:

45. An electronic device comprising means for carrying out the method of any one of claims 35 to 44.

Citation Information

Patent Citations

  • Memory system and operating method thereof

    KR1020200114354A

  • Virtual assistant in a communication session

    US20160335532A1

  • Digital assistant interaction in a video communication session environment

    US20210249009A1