Digital Assistant User Interface and Response Mode
By integrating a digital assistant interface with the underlying user interface and adapting response modes based on context, the method addresses the issues of visual clutter and response format incongruence, enhancing usability and efficiency.
Patent Information
- Application Number
- JP2024012099
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-09
- Filing Date
- 2024-01-30
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2041-04-16
AI Technical Summary
Digital assistants often obscure other displayed elements on electronic devices and provide responses in undesirable formats, leading to user inconvenience and inefficiency.
A digital assistant user interface is displayed on a portion of the screen with a digital assistant indicator and response affordance on another portion, allowing the underlying user interface to remain visible, enabling simultaneous interaction with both interfaces and adapting response modes based on user context.
This approach enhances usability by reducing visual clutter, improving interaction efficiency, and optimizing power usage by allowing quicker and more efficient device operation.
Smart Images

Figure 0007809149000009 
Figure 0007809149000010 
Figure 0007809149000011
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 17 / 227,012, filed April 9, 2021, entitled "DIGITAL ASSISTANT USER INTERFACES AND RESPONSE MODES," U.S. Provisional Patent Application No. 63 / 028,821, filed May 22, 2020, entitled "DIGITAL ASSISTANT USER INTERFACES AND RESPONSE MODES," Danish Patent Application No. PA 2020 70547, filed August 24, 2020, entitled "DIGITAL ASSISTANT USER INTERFACES AND RESPONSE MODES," and Danish Patent Application No. PA 2020 70548, filed August 24, 2020, entitled "DIGITAL ASSISTANT USER INTERFACES AND RESPONSE MODES." The entire contents of each of these applications are incorporated herein by reference in their entirety.
[0002] This relates generally to intelligent automated assistants, and more specifically to user interfaces for intelligent automated assistants and ways in which intelligent automated assistants can respond to user requests. [Background technology]
[0003] Intelligent automated assistants (or digital assistants) can provide a useful interface between human users and electronic devices. Such assistants can enable users to interact with devices or systems using natural language in spoken and / or textual form. For example, a user can provide speech input containing a user request to a digital assistant running on an electronic device. The digital assistant can interpret the user's intent from the speech input and activate the user's intent into a task. The task can then be performed by executing one or more services on the electronic device and can return an associated output response to the user request to the user.
[0004] The digital assistant's displayed user interface may obscure other displayed elements that the user may be interested in. Furthermore, the digital assistant may provide responses in a format that is undesirable for the user's current situation. For example, the digital assistant may provide display output when the user does not want to (or cannot) view the device display. Summary of the Invention
[0005] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having a display and a touch-sensitive surface, receiving user input while displaying a user interface different from the digital assistant user interface, and displaying a digital assistant user interface on the user interface according to a determination that the user input satisfies criteria for starting the digital assistant, wherein the digital assistant user interface includes a digital assistant indicator displayed on a first portion of the display and a response affordance displayed on a second portion of the display, and a portion of the user interface remains visible on a third portion of the display, the third portion being between the first and second portions.
[0006] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. When executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the one or more programs cause the electronic device to receive user input while displaying a user interface different from the digital assistant user interface, and, in accordance with a determination that the user input satisfies the criteria for starting the digital assistant, display a digital assistant user interface on the user interface, wherein the digital assistant user interface includes a digital assistant indicator displayed on a first portion of the display and a response affordance displayed on a second portion of the display, and a portion of the user interface remains visible on a third portion of the display, the third portion being between the first and second portions. The digital assistant user interface is displayed.
[0007] An exemplary electronic device is disclosed herein. The exemplary electronic device includes a display, a touch-sensitive surface, one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors. While the one or more programs display a user interface different from the digital assistant user interface, the one or more programs receive user input, and in accordance with a determination that the user input satisfies criteria for starting the digital assistant, display a digital assistant user interface on the user interface, wherein the digital assistant user interface includes a digital assistant indicator displayed on a first portion of the display and a response affordance displayed on a second portion of the display, and a portion of the user interface remains visible on a third portion of the display, the third portion being between the first and second portions. The electronic device includes instructions for performing the following.
[0008] An exemplary electronic device receives user input while displaying a user interface different from the digital assistant user interface, and displays a digital assistant user interface on the user interface according to a determination that the user input satisfies criteria for starting the digital assistant, wherein the digital assistant user interface includes a digital assistant indicator displayed on a first portion of the display and a responsive affordance displayed on a second portion of the display, and a portion of the user interface remains visible on a third portion of the display, the third portion being between the first and second portions.
[0009] Displaying a digital assistant user interface on a user interface where a portion of the user interface remains visible on a portion of the display can improve the usability of the digital assistant and reduce the visual clutter of the digital assistant to user-device interactions. For example, information contained in the underlying visible user interface can allow a user to better formulate requests to the digital assistant. As another example, displaying a user interface in such a manner can facilitate interaction between elements of the digital assistant user interface and the underlying user interface (e.g., including digital assistant responses in messages of the underlying messaging user interface). Furthermore, having both the digital assistant user interface and the underlying user interface coexist on the display allows simultaneous user interaction with both user interfaces, thereby better integrating the digital assistant with both the user and the device. In this way, the user-device interface can be more efficient (e.g., by allowing the digital assistant to more accurately and efficiently perform user-requested tasks, by reducing the visual clutter of the digital assistant to what the user sees, by reducing the number of user inputs required to operate the device as desired), which further reduces power usage and improves the device's battery life by allowing users to use the device more quickly and efficiently.
[0010] An exemplary method is disclosed herein. The exemplary method is an electronic device having a display and a touch-sensitive surface. Displaying a digital assistant user interface on a user interface, wherein the digital assistant user interface includes a digital assistant indicator displayed on a first portion of the display and a response affordance displayed on a second portion of the display, and receiving a user input corresponding to a selection of a third portion of the display while displaying the digital assistant user interface on the user interface, wherein the third portion displays a portion of the user interface, and determining that the user input corresponds to a first type of input. Stopping the display of the digital assistant indicator and the response affordance, and determining that the user input corresponds to a second type of input different from the first type of input. While displaying the response affordance on the second portion, updating the display of the user interface in the third portion according to the user input.
[0011] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions, which, when executed by one or more processors of an electronic device having a display and a touch-sensitive surface, cause the electronic device to display a digital assistant user interface on a user interface, including a digital assistant indicator displayed on a first portion of the display and a response affordance displayed on a second portion of the display. While displaying the digital assistant user interface on the user interface, receive a user input corresponding to a selection of a third portion of the display, wherein the third portion displays a portion of the user interface. In accordance with a determination that the user input corresponds to a first type of input, stop displaying the digital assistant indicator and the response affordance, and in accordance with a determination that the user input corresponds to a second type of input different from the first type of input, stop displaying the digital assistant indicator and the response affordance on the second portion. Update the display of the user interface in the third portion according to the user input.
[0012] An exemplary electronic device is disclosed herein. The exemplary electronic device includes a display, a touch-sensitive surface, one or more processors, a memory, and one or more programs stored in the memory and configured to be executed by the one or more processors, wherein the one or more programs display a digital assistant user interface on a user interface, wherein the digital assistant user interface includes a digital assistant indicator displayed on a first portion of the display and a response affordance displayed on a second portion of the display. While displaying the digital assistant user interface on the user interface, receiving a user input corresponding to a selection of a third portion of the display, wherein the third portion displays a portion of the user interface, and determining that the user input corresponds to a first type of input, stopping the display of the digital assistant indicator and the response affordance, and determining that the user input corresponds to a second type of input different from the first type of input, while displaying the response affordance on the second portion, updating the display of the user interface in the third portion according to the user input.
[0013] An exemplary electronic device includes a means for displaying a digital assistant user interface on a user interface, the digital assistant user interface including a digital assistant indicator displayed on a first portion of the display and a response affordance displayed on a second portion of the display; and receiving a user input corresponding to a selection of a third portion of the display while displaying the digital assistant user interface on the user interface, wherein the third portion displays a portion of the user interface; and, in accordance with a determination that the user input corresponds to a first type of input, stopping the display of the digital assistant indicator and the response affordance; and, in accordance with a determination that the user input corresponds to a second type of input different from the first type of input, while displaying the response affordance on the second portion, updating the display of the user interface in the third portion according to the user input.
[0014] Stopping the display of the digital assistant indicator and the response affordance in accordance with a determination that the user input corresponds to a first type of input can provide an intuitive and efficient way to close the digital assistant. For example, a user can simply provide an input that selects the underlying user interface to close the digital assistant user interface, thereby reducing the digital assistant's confusion about the user's interaction with the device. Updating the display of the user interface in the third portion according to the user input while displaying the response affordance in the second portion provides an intuitive way for the digital assistant user interface to coexist with the underlying user interface. For example, a user can provide an input that selects the underlying user interface to cause the underlying user interface to be updated as if the digital assistant user interface were not displayed. Furthermore, preserving the digital assistant user interface (which may include user interest information) while allowing user interaction with the underlying user interface can reduce the digital assistant's confusion about the underlying user interface. In this way, the user device interface can be more efficient (e.g., by allowing user input to interact with the underlying user interface while the digital assistant user interface is displayed, by reducing the visual clutter of the digital assistant to what the user is looking at, by reducing the number of user inputs required to operate the device as desired), which further reduces power usage and improves the device's battery life by allowing the user to use the device more quickly and efficiently.
[0015] An exemplary method is disclosed herein. The exemplary method includes, in an electronic device having one or more processors, memory, and a display, receiving a natural language input, starting a digital assistant, obtaining a response package in response to the natural language input in accordance with starting the digital assistant, selecting a first response mode for the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device after receiving the natural language input, and presenting the response package by the digital assistant according to the first response mode in response to selecting the first response mode.
[0016] An exemplary non-transitory computer-readable medium is disclosed herein. The exemplary non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive a natural language input, start a digital assistant, and, in response to starting the digital assistant, obtain a response package in response to the natural language input. After receiving the natural language input, select a first response mode for the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device. In response to selecting the first response mode, present a response package according to the first response mode.
[0017] An exemplary electronic device is disclosed herein. The exemplary electronic device includes a display, one or more processors, a memory, and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including: receiving natural language input and starting a digital assistant; obtaining a response package in response to the natural language input in accordance with starting the digital assistant; selecting a first response mode for the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device after receiving the natural language input; and presenting the response package by the digital assistant according to the first response mode in response to selecting the first response mode.
[0018] An exemplary electronic device includes means for receiving natural language input; starting a digital assistant; obtaining a response package in response to the natural language input in accordance with starting the digital assistant; selecting a first response mode for the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device after receiving the natural language input; and presenting the response package by the digital assistant in accordance with the first response mode in response to selecting the first response mode.
[0019] Presenting the response package according to the first response mode by the digital assistant can enable the presentation of the digital assistant response in a helpful manner appropriate to the user's current context. For example, the digital assistant can present the response in audio format when the user's current context indicates that visual user device interaction is undesirable (or impossible). As another example, the digital assistant can present the response in visual format when the user's current context indicates that audible user device interaction is undesirable. As yet another example, the digital assistant can present a response having a visual component and a brief audio component when the user's current context indicates that both audible and visual user device interaction are desired, thereby reducing the length of the digital assistant's audio output. Furthermore, selecting the first response mode after receiving natural language input (and before presenting the response package) can enable a more accurate determination of the user's current context (and thus a more accurate determination of the appropriate response mode). In this way, the user device interface can be made more efficient and secure (e.g., by reducing visual clutter on the digital assistant, by efficiently presenting responses in an informative manner, by intelligently adapting the manner of responses based on the user's current context), and further reduce power usage and improve the device's battery life by allowing the user to use the device more quickly and efficiently. [Brief explanation of the drawings]
[0020] [Figure 1] FIG. 1 is a block diagram illustrating a system and environment for implementing a digital assistant, according to various embodiments.
[0021] [Figure 2A] FIG. 1 is a block diagram illustrating a portable multifunction device that implements a client-side portion of a digital assistant, according to various embodiments.
[0022] [Figure 2B] FIG. 2 is a block diagram illustrating example components for event processing, in accordance with various embodiments.
[0023] [Figure 3] FIG. 1 illustrates a portable multifunction device implementing a client-side portion of a digital assistant, according to various embodiments.
[0024] [Figure 4] FIG. 1 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface in accordance with various embodiments.
[0025] [Figure 5A] 1A-1C illustrate exemplary user interfaces for a menu of applications on a portable multifunction device in accordance with various embodiments.
[0026] [Figure 5B] 1A-1C illustrate exemplary user interfaces for a multifunction device having a touch-sensitive surface separate from the display, in accordance with various embodiments.
[0027] [Figure 6A] FIG. 1 illustrates a personal electronic device according to various embodiments.
[0028] [Figure 6B] FIG. 1 is a block diagram illustrating a personal electronic device according to various embodiments.
[0029] [Figure 7A] FIG. 1 is a block diagram illustrating a digital assistant system or a server portion thereof, according to various embodiments.
[0030] [Figure 7B] 7B illustrates the functionality of the digital assistant shown in FIG. 7A, according to various embodiments.
[0031] [Figure 7C] FIG. 2 illustrates a portion of an ontology, according to various embodiments.
[0032] [Figure 8A] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8B] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8C] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8D] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8E] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8F] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8G] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8H] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8I] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8J] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8K] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8L] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8M]1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8N] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8O] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8P] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8Q] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8R] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8S] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8T] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8U] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8V] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8W] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8X] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8Y] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8Z] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AA] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AB] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AC] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AD] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AE] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AF] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AG] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AH] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AI] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AJ] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AK] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AL] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AM] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AN] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AO] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AP] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AQ] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AR] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AS] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AT] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AU] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AV] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AW] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AX] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AY] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8AZ] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BA] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BB]1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BC] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BD] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BE] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BF] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BG] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BH] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BI] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BJ] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BK] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BL] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BM] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BN] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BO] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BP] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BQ] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BR] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BS] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BT] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BU] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BV] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BW] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BX] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BY] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8BZ] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CA] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CB] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CC] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CD] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CE] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CF] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CG] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CH] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CI] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CJ] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CK] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CL] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CM] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CN] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CO] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CP] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CQ]1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CR] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CS] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 8CT] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments.
[0033] [Figure 9A] 1 illustrates multiple devices determining which device should respond to speech input, according to various embodiments. [Figure 9B] 1 illustrates multiple devices determining which device should respond to speech input, according to various embodiments. [Figure 9C] 1 illustrates multiple devices determining which device should respond to speech input, according to various embodiments.
[0034] [Figure 10A] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10B] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10C] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10D] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10E] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10F] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10G] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10H] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10I] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10J] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10K] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10L] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10M] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10N] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10O] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10P] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10Q] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10R] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10S] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10T] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10U] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments. [Figure 10V] 1 illustrates a user interface and a digital assistant user interface, according to various embodiments.
[0035] [Figure 11] 1 illustrates a system for selecting a digital assistant response mode and presenting a response according to the selected digital assistant response mode, according to various embodiments.
[0036] [Figure 12] 1 illustrates a device that receives a natural language input and presents responses according to different digital assistant response modes, according to various embodiments.
[0037] [Figure 13] 1 illustrates an example process implemented to select a digital assistant response mode, according to various embodiments.
[0038] [Figure 14] 1 illustrates a device that presents responses according to an audio response mode when a user is determined to be in (e.g., driving) a vehicle, according to various embodiments.
[0039] [Figure 15] 1 illustrates a device presenting responses according to a voice response mode when the device is running a navigation application, according to various embodiments.
[0040] [Figure 16] 10 illustrates response mode variation over the course of a multi-turn DA interaction, according to various embodiments.
[0041] [Figure 17A] 1 illustrates a process for operating a digital assistant, according to various embodiments. [Figure 17B] 1 illustrates a process for operating a digital assistant, according to various embodiments. [Figure 17C] 1 illustrates a process for operating a digital assistant, according to various embodiments. [Figure 17D] 1 illustrates a process for operating a digital assistant, according to various embodiments. [Figure 17E] 1 illustrates a process for operating a digital assistant, according to various embodiments. [Figure 17F] 1 illustrates a process for operating a digital assistant, according to various embodiments.
[0042] [Figure 18A] 1 illustrates a process for operating a digital assistant, according to various embodiments. [Figure 18B] 1 illustrates a process for operating a digital assistant, according to various embodiments.
[0043] [Figure 19A] 1 illustrates a process for selecting a digital assistant response mode, according to various embodiments. [Figure 19B] 1 illustrates a process for selecting a digital assistant response mode, according to various embodiments. [Figure 19C] 1 illustrates a process for selecting a digital assistant response mode, according to various embodiments. [Figure 19D] 1 illustrates a process for selecting a digital assistant response mode, according to various embodiments. [Figure 19E] 1 illustrates a process for selecting a digital assistant response mode, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0044] In the following description of the embodiments, reference is made to the accompanying drawings, which show, by way of illustration, specific embodiments that may be practiced. It is to be understood that other embodiments may be utilized and structural changes may be made without departing from the scope of the various embodiments.
[0045] In the following description, terms such as "first" and "second" are used to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first input can be referred to as a second input, and similarly, a second input can be referred to as a first input, without departing from the scope of various embodiments described. The first input and the second input are both inputs, and in some cases, are separate and distinct inputs.
[0046] The terminology used in the description of the various embodiments set forth herein is for the purpose of describing particular embodiments only and is not intended to be limiting. When used in the description of the various embodiments set forth and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" should be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes," "including," "comprises," and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0047] The term "if" can be interpreted to mean "when" or "upon," or "in response to determining," or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "upon determining," or "in response to determining," or "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]," depending on the context. 1. System and Environment
[0048] FIG. 1 illustrates a block diagram of a system 100 according to various embodiments. In some embodiments, the system 100 implements a digital assistant. The terms “digital assistant,” “virtual assistant,” “intelligent automated assistant,” or “automated digital assistant” refer to any information processing system that infers user intent by interpreting natural language input in spoken and / or textual form and performs actions based on the inferred user intent. For example, to act based on the inferred user intent, the system performs one or more of the following: identifies a task flow having steps and parameters designed to fulfill the inferred user intent, inputs specific requirements from the inferred user intent into the task flow, executes the task flow by invoking programs, methods, services, APIs, or the like, and generates an output response to the user in an audible (e.g., spoken) and / or visual form.
[0049] Specifically, a digital assistant can accept user requests, at least in part, in the form of natural language commands, requests, opinions, discourse, and / or inquiries. Typically, a user request seeks either an informational answer or task performance by the digital assistant. A satisfactory response to a user request includes providing the requested informational answer, performing the requested task, or a combination of the two. For example, a user asks a digital assistant a question such as, "Where am I right now?" Based on the user's current location, the digital assistant responds, "You're in Central Park near the West Gate." The user also requests a task to be performed, such as, "Please invite my friends to my girlfriend's birthday party next week." In response, the digital assistant can acknowledge the request by stating, "Yes, right now," and then send appropriate calendar invitations on behalf of the user to each of the user's friends listed in the user's electronic address book. During the performance of a requested task, the digital assistant may interact with the user in a continuous conversation involving multiple information exchanges over an extended period of time. There are many other ways to interact with a digital assistant to request information or to perform various tasks. In addition to providing verbal responses and taking programmed actions, digital assistants also provide responses in other visual or audio formats, such as text, alerts, music, videos, animations, etc.
[0050] 1 , in some embodiments, the digital assistant is implemented according to a client-server model. The digital assistant includes a client-side portion 102 (hereinafter, "DA client 102") that runs on a user device 104 and a server-side portion 106 (hereinafter, "DA server 106") that runs on a server system 108. The DA client 102 communicates with the DA server 106 over one or more networks 110. The DA client 102 provides client-side functionality, such as user-responsive input and output processing and communication with the DA server 106. The DA server 106 provides server-side functionality to any number of DA clients 102, each resident on a respective user device 104.
[0051] In some embodiments, the DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, and an I / O interface to external services 118. The client-facing I / O interface 112 facilitates client-facing input and output processing of the DA server 106. The one or more processing modules 114 utilize the data and models 116 to process speech input and determine user intent based on natural language input. Furthermore, the one or more processing modules 114 perform task execution based on the inferred user intent. In some embodiments, the DA server 106 communicates with external services 120 over network(s) 110 to complete tasks or obtain information. The I / O interface to external services 118 facilitates such communication.
[0052] User device 104 can be any suitable electronic device. In some examples, user device 104 is a portable multifunction device (e.g., device 200 described below in connection with FIG. 2A), a multifunction device (e.g., device 400 described below in connection with FIG. 4), or a personal electronic device (e.g., device 600 described below in connection with FIGS. 6A-6B). A portable multifunction device is, for example, a mobile phone that also includes other functions, such as PDA and / or music player functionality. Specific examples of portable multifunction devices include the Apple Watch®, iPhone®, iPod Touch®, and iPad® devices by Apple Inc. (Cupertino, California). Other examples of portable multifunction devices include, but are not limited to, earphones / headphones, speakers, and laptop or tablet computers. Furthermore, in some examples, user device 104 is a non-portable multifunction device. Specifically, user device 104 is a desktop computer, a game console, a speaker, a television, or a television set-top box. In some examples, user device 104 includes a touch-sensitive surface (e.g., a touchscreen display and / or a touchpad). Additionally, user device 104 optionally includes one or more other physical user interface devices, such as a physical keyboard, a mouse, and / or a joystick. Various examples of electronic devices, such as multifunction devices, are described in further detail below.
[0053] Examples of communication network(s) 110 include a local area network (LAN) and a wide area network (WAN), such as the Internet. Communication network(s) 110 may be implemented using any known network protocol, including various wired or wireless protocols, such as, for example, Ethernet, Universal Serial Bus (USB), FIREWIRE®, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth®, Wi-Fi®, voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.
[0054] The server system 108 may be implemented on one or more standalone data processing devices or on a distributed computer network. In some embodiments, the server system 108 may also employ various virtual devices and / or the services of third-party service providers (e.g., third-party cloud service providers) to provide the underlying computing and / or infrastructure resources of the server system 108.
[0055] In some embodiments, the user device 104 communicates with the DA server 106 via a second user device 122. The second user device 122 is similar to or identical to the user device 104. For example, the second user device 122 is similar to devices 200, 400, or 600 described below in connection with FIGS. 2A, 4, and 6A-6B. The user device 104 is configured to be communicatively coupled to the second user device 122 via a direct communication connection, such as Bluetooth, NFC, or BTLE, or via a wired or wireless network, such as a local Wi-Fi network. In some embodiments, the second user device 122 is configured to act as a proxy between the user device 104 and the DA server 106. For example, the DA client 102 of the user device 104 is configured to send information (e.g., a user request received at the user device 104) to the DA server 106 via the second user device 122. The DA server 106 processes the information and returns relevant data (eg, data content responsive to the user request) to the user device 104 via the second user device 122.
[0056] In some embodiments, the user device 104 is configured to reduce the amount of information transmitted from the user device 104 by communicating with the second user device 122 via an abbreviated request for data. The second user device 122 is configured to determine supplemental information to add to the abbreviated request and generate a complete request to transmit to the DA server 106. This system architecture can advantageously allow a user device 104 (e.g., a watch or similar small electronic device) with limited communication capabilities and / or limited battery power to access services provided by the DA server 106 by using a second user device 122 (e.g., a mobile phone, laptop computer, tablet computer, etc.) with greater communication capabilities and / or battery power as a proxy to the DA server 106. While only two user devices 104 and 122 are shown in FIG. 1 , it should be understood that the system 100, in some embodiments, includes any number and type of user devices configured to communicate with the DA server system 106 in this proxy configuration.
[0057] 1 includes both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106), but in some examples, the digital assistant's functionality is implemented as an independent application installed on a user device. Furthermore, the allocation of functionality between the client and server portions of the digital assistant may vary depending on the implementation. For example, in some embodiments, the DA client is a thin client that provides only user-facing input and output processing functionality and delegates all other digital assistant functionality to a back-end server. 2. Electronic Devices
[0058] Attention now turns to embodiments of electronic devices for implementing the client-side portion of a digital assistant. FIG. 2A is a block diagram illustrating portable multifunction device 200 with touch-sensitive display system 212, according to some embodiments. Touch-sensitive display 212 may conveniently be referred to as a "touch screen" and may also be known or referred to as a "touch-sensitive display system." Device 200 includes memory 202 (optionally including one or more computer-readable storage media), a memory controller 222, one or more processing units (CPUs) 220, a peripherals interface 218, RF circuitry 208, audio circuitry 210, a speaker 211, a microphone 213, an input / output (I / O) subsystem 206, other input control devices 216, and an external port 224. Device 200 optionally includes one or more light sensors 264. Device 200 optionally includes one or more contact intensity sensors 265 that detect the intensity of a contact on device 200 (e.g., a touch-sensitive surface such as touch-sensitive display system 212 of device 200). Device 200 optionally includes one or more tactile output generators 267 that generate a tactile output on device 200 (e.g., generate a tactile output on a touch-sensitive surface such as touch-sensitive display system 212 of device 200 or touchpad 455 of device 400). These components optionally communicate via one or more communication buses or signal lines 203.
[0059] As used herein and in the claims, the term “intensity” of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or a proxy for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four distinct values and more typically includes hundreds (e.g., at least 256) distinct values. The intensity of a contact is optionally determined (or measured) using various techniques and various sensors or combinations of sensors. For example, one or more force sensors under or adjacent to the touch-sensitive surface are optionally used to measure force at various points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine an estimated force of the contact. Similarly, a pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change in the contact area detected on the touch-sensitive surface, the capacitance and / or change in the capacitance of the touch-sensitive surface proximate the contact, and / or the resistance and / or change in the capacitance of the touch-sensitive surface proximate the contact are optionally used as a surrogate for the force or pressure of the contact on the touch-sensitive surface. In some implementations, the surrogate measure of the force or pressure of the contact is used directly to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measure). In some implementations, the surrogate measure of the contact force or pressure is converted to an estimate of the force or pressure, and the estimate of the force or pressure is used to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using contact intensity as an attribute of user input allows users to access additional device functionality (e.g., on a touch-sensitive display) and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls such as knobs or buttons) that may not otherwise be accessible to users on devices of reduced size that have limited footprint for displaying affordances.
[0060] As used herein and in the claims, the term “tactile output” refers to a physical displacement of a device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device, that will be detected by a user with the user's sense of touch. For example, in a situation where a device or a component of a device is in contact with a touch-sensitive surface of a user (e.g., the fingers, palm, or other part of the user's hand), the tactile output produced by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in a physical property of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is optionally interpreted by the user as a “downclick” or “upclick” of a physical actuator button. In some cases, a user feels a tactile sensation such as a “downclick” or “upclick” even when there is no movement of a physical actuator button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's action. As another example, movement of a touch-sensitive surface is optionally interpreted or perceived by a user as "roughness" of the touch-sensitive surface, even when there is no change in the smoothness of the touch-sensitive surface. While such user interpretation of touch depends on the user's personal sensory perception, there are many sensory perceptions of touch that are common to the majority of users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., "upclick," "downclick," "roughness"), unless otherwise specified, the generated tactile output corresponds to a physical displacement of the device, or a component of the device, that produces the described sensory perception for a typical (or average) user.
[0061] It should be understood that device 200 is only one example of a portable multifunction device, and that device 200 optionally has more or fewer components than those shown, optionally combines two or more components, or optionally has a different configuration or arrangement of its components. The various components shown in Figure 2A are implemented as hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.
[0062] Memory 202 includes one or more computer-readable storage media that are, for example, tangible and non-transitory. Memory 202 includes high-speed random-access memory and also includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 222 controls access to memory 202 by other components of device 200.
[0063] In some embodiments, the non-transitory computer-readable storage medium of memory 202 is used to store instructions (e.g., to perform aspects of the processes described below) for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other system capable of fetching instructions from the instruction execution system, apparatus, or device and executing the instructions. In other examples, the instructions (e.g., to perform aspects of the processes described below) are stored in a non-transitory computer-readable storage medium (not shown) of server system 108 or are split between the non-transitory computer-readable storage medium of memory 202 and the non-transitory computer-readable storage medium of server system 108.
[0064] Peripheral interface 218 is used to couple input and output peripherals of the device to CPU 220 and memory 202. One or more processors 220 operate or execute various software programs and / or instruction sets stored in memory 202 to perform various functions and process data for device 200. In some embodiments, peripheral interface 218, CPU 220, and memory controller 222 are implemented on a single chip, such as chip 204. In some other embodiments, they are implemented on separate chips.
[0065] RF (radio frequency) circuitry 208 transmits and receives RF signals, also called electromagnetic signals. RF circuitry 208 converts electrical signals to electromagnetic signals and vice versa, and communicates with communication networks and other communication devices via electromagnetic signals. RF circuitry 208 optionally includes well-known circuitry for performing these functions, including, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, etc. RF circuitry 208 optionally communicates via wireless communication with networks, such as the Internet, also known as the World Wide Web (WWW), an intranet, and / or wireless networks, such as cellular telephone networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs), and with other devices. RF circuitry 208 optionally includes known circuitry for detecting near field communication (NFC) fields, such as by short-range radios. Wireless communication optionally includes, but is not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPA), Long Term Evolution (LTE), and other technologies.Wireless technology includes, but is not limited to, wireless technology such as LTE evolution, near field communications (NFC), wideband code division multiple access (W-CDMA), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet message access protocol (IMAP) and / or post office protocol (POP)), instant messaging (e.g., extensible messaging and presence protocol), and the like. The present invention may use any of a number of communication standards, protocols, and technologies, including the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (XMPP), the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), the Instant Messaging and Presence Service (IMPS), and / or the Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the filing date of this application.
[0066] The audio circuit 210, speaker 211, and microphone 213 provide an audio interface between a user and device 200. Audio circuit 210 receives audio data from peripherals interface 218, converts the audio data into electrical signals, and transmits the electrical signals to speaker 211. Speaker 211 converts the electrical signals into sound waves audible to humans. Audio circuit 210 also receives electrical signals converted from sound waves by microphone 213. Audio circuit 210 converts the electrical signals into audio data and transmits the audio data to peripherals interface 218 for processing. The audio data is retrieved from and / or transmitted to memory 202 and / or RF circuit 208 by peripherals interface 218. In some embodiments, audio circuit 210 also includes a headset jack (e.g., 312 in FIG. 3 ). The headset jack provides an interface between audio circuitry 210 and a detachable audio input / output peripheral, such as an output-only headphone or a headset with both an output (e.g., single or double ear headphones) and an input (e.g., a microphone).
[0067] I / O subsystem 206 couples input / output peripherals on device 200, such as touchscreen 212 and other input control devices 216, to peripheral interface 218. I / O subsystem 206 optionally includes one or more input controllers 260 for display controller 256, light sensor controller 258, intensity sensor controller 259, haptic feedback controller 261, and other input or control devices. One or more input controllers 260 receive / send electrical signals from / to other input control devices 216. Other input control devices 216 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some alternative embodiments, input controller 260 is optionally coupled to any (or none) of a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. The one or more buttons (e.g., 308 in FIG. 3) optionally include up / down buttons for volume control of speaker 211 and / or microphone 213. The one or more buttons optionally include a push button (e.g., 306 in FIG. 3).
[0068] A quick press of a push button unlocks the touch screen 212 or initiates the process of using gestures on the touch screen to unlock the device, as described in U.S. Patent Application No. 11 / 322,549, filed December 23, 2005, and U.S. Patent No. 7,657,849, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," both of which are incorporated herein by reference in their entireties. A long press of a push button (e.g., 306) powers the device 200 on or off. The user can customize the functionality of one or more buttons. The touch screen 212 can be used to implement virtual or soft buttons and one or more soft keyboards.
[0069] The touch-sensitive display 212 provides an input and output interface between the device and a user. The display controller 256 receives and / or sends electrical signals to and from the touchscreen 212. The touchscreen 212 displays visual output to the user. The visual output includes graphics, text, icons, animation, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output corresponds to user interface objects.
[0070] Touchscreen 212 has a touch-sensitive surface, sensor, or set of sensors that accepts input from a user based on haptic and / or tactile contact. Touchscreen 212 and display controller 256 (along with any associated modules and / or instruction sets in memory 202) detects contacts (and any movement or cessation of contact) on touchscreen 212 and translates the detected contacts into interactions with user interface objects (e.g., one or more softkeys, icons, web pages, or images) displayed on touchscreen 212. In an exemplary embodiment, the point of contact between touchscreen 212 and the user corresponds to the user's finger.
[0071] Touchscreen 212 uses LCD (liquid crystal display), LPD (light emitting polymer display), or LED (light emitting diode) technology, although other display technologies may be used in other embodiments. Touchscreen 212 and display controller 256 detect contact and its movement or breaking using any of several now known or later developed touch sensing technologies, including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements that determine one or more points of contact using touchscreen 212. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that found in the iPhone® and iPod Touch® from Apple Inc. of Cupertino, California.
[0072] The touch-sensitive display of some embodiments of touchscreen 212 is similar to the multi-touch-sensing touchpad described in U.S. Patents 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman) and / or U.S. Patent Publication 2002 / 0015024 (A1), each of which is incorporated by reference in its entirety. However, touchscreen 212 displays visual output from device 200, whereas touch-sensitive touchpads do not provide visual output.
[0073] The touch-sensitive display in some embodiments of touchscreen 212 is described in the following applications: (1) U.S. patent application Ser. No. 11 / 381,313, filed May 2, 2006, entitled "Multipoint Touch Surface Controller"; (2) U.S. patent application Ser. No. 10 / 840,862, filed May 6, 2004, entitled "Multipoint Touchscreen"; (3) U.S. patent application Ser. No. 10 / 903,964, filed July 30, 2004, entitled "Gestures For Touch Sensitive Input Devices"; (4) U.S. patent application Ser. No. 11 / 048,264, filed January 31, 2005, entitled "Gestures For Touch Sensitive Input Devices"; and (5) U.S. patent application Ser. No. 11 / 038,590, filed January 18, 2005, entitled "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices." No. 11 / 228,758, filed September 16, 2005, entitled "Virtual Input Device Placement On A Touch Screen User Interface," (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, entitled "Operation Of A Computer With A Touch Screen Interface," (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, entitled "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard," and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, entitled "Multi-Functional Hand-Held Device," all of which are incorporated herein by reference in their entireties.
[0074] The touchscreen 212 has a video resolution of, for example, greater than 100 dpi. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. A user contacts the touchscreen 212 using a suitable object or accessory, such as a stylus, finger, or the like. In some embodiments, the user interface is designed to operate primarily using finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of a finger on the touchscreen. In some embodiments, the device translates the coarse finger input into precise pointer / cursor positions or commands to perform the action desired by the user.
[0075] In some embodiments, in addition to the touchscreen, device 200 includes a touchpad (not shown) for activating or deactivating certain functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touchscreen, does not display visual output. The touchpad may be a touch-sensitive surface separate from touchscreen 212 or may be an extension of the touch-sensitive surface formed by the touchscreen.
[0076] Device 200 also includes a power system 262 that provides power to the various components. Power system 262 includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a charging system, power failure detection circuitry, power converters or inverters, power status indicators (e.g., light emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in a portable device.
[0077] Device 200 also includes one or more light sensors 264. FIG. 2A shows a light sensor coupled to light sensor controller 258 in I / O subsystem 206. Light sensor 264 includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. Light sensor 264 receives light from the environment projected through one or more lenses and converts the light into data representing an image. In conjunction with imaging module 243 (also called a camera module), light sensor 264 captures still images or video. In some embodiments, the light sensor is located on the back of device 200, opposite touchscreen display 212 on the front of the device, so that the touchscreen display is used as a viewfinder for still image and / or video capture. In some embodiments, the light sensor is located on the front of the device so that an image of the user for a video conference is captured while the user views other video conference participants on the touchscreen display. In some embodiments, the position of the light sensor 264 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), so that a single light sensor 264 is used for both video conferencing and capturing still images and / or video, along with a touchscreen display.
[0078] Device 200 also optionally includes one or more contact intensity sensors 265. FIG. 2A shows a contact intensity sensor coupled to intensity sensor controller 259 in I / O subsystem 206. Contact intensity sensor 265 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of a contact on a touch-sensitive surface). Contact intensity sensor 265 receives contact intensity information (e.g., pressure information or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed with or proximate to the touch-sensitive surface (e.g., touch-sensitive display system 212). In some embodiments, at least one contact intensity sensor is located on the back of device 200, opposite touchscreen display 212, which is located on the front of device 200.
[0079] Device 200 also includes one or more proximity sensors 266. Figure 2A shows proximity sensor 266 coupled to peripheral interface 218. Alternatively, proximity sensor 266 is coupled to input controller 260 within I / O subsystem 206. Proximity sensor 266 functions as described in U.S. patent application Ser. Nos. 11 / 241,839, "Proximity Detector In Handheld Device," 11 / 240,788, "Proximity Detector In Handheld Device," 11 / 620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output," 11 / 586,862, "Automated Response To And Sensing Of User Activity In Portable Devices," and 11 / 638,251, "Methods And Systems For Automatic Configuration Of Peripherals." In some embodiments, when the multifunction device is placed near the user's ear (eg, when the user is making a phone call), the proximity sensor turns off and disables the touchscreen 212.
[0080] Device 200 also optionally includes one or more tactile output generators 267. FIG. 2A shows tactile output generators 267 coupled to haptic feedback controller 261 in I / O subsystem 206. Tactile output generator 267 optionally includes one or more electroacoustic devices, such as speakers or other audio components, and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components that convert electrical signals into tactile output on the device). Contact intensity sensor 265 receives tactile feedback generation commands from haptic feedback module 233 and generates a tactile output on device 200 that can be sensed by a user of device 200. In some embodiments, at least one tactile output generator is juxtaposed with or proximate to a touch-sensitive surface (e.g., touch-sensitive display system 212) and, optionally, generates a tactile output by moving the touch-sensitive surface vertically (e.g., in / out of the surface of device 200) or horizontally (e.g., back and forth in the same plane as the surface of device 200). In some embodiments, at least one tactile output generator sensor is located on the back of device 200, opposite touchscreen display 212, which is located on the front of device 200.
[0081] Device 200 also includes one or more accelerometers 268. FIG. 2A shows accelerometer 268 coupled to peripherals interface 218. Alternatively, accelerometer 268 is coupled to input controller 260 within I / O subsystem 206. Accelerometer 268 operates as described, for example, in U.S. Patent Publication No. 20050190059, entitled "Acceleration-Based Theft Detection for Portable Electronic Devices," and U.S. Patent Publication No. 20060017692, entitled "Acceleration-Based Theft Detection System for Portable Electronic Devices," both of which are incorporated herein by reference in their entireties. In some embodiments, information is displayed on the touchscreen display in portrait or landscape orientation based on an analysis of data received from the one or more accelerometers. In addition to accelerometer 268, device 200 optionally includes a magnetometer (not shown) and a GPS (or GLONASS or other global navigation system) receiver (not shown) for obtaining information regarding the location and orientation (e.g., portrait or landscape) of device 200.
[0082] In some embodiments, software components stored in memory 202 include operating system 226, communication module (or instruction set) 228, touch / motion module (or instruction set) 230, graphics module (or instruction set) 232, text input module (or instruction set) 234, Global Positioning System (GPS) module (or instruction set) 235, digital assistant client module 229, and applications (or instruction sets) 236. Additionally, memory 202 stores data and models, such as user data and models 231. Additionally, in some embodiments, as shown in FIGS. 2A and 4, memory 202 (FIG. 2A) or memory 470 (FIG. 4) stores device / global internal state 257. The device / global internal state 257 includes one or more of: an active application state indicating which application, if any, is currently active; a display state indicating which applications, views, or other information occupy various areas of the touchscreen display 212; a sensor state including information obtained from the device's various sensors and input control devices 216; and position information regarding the device's position and / or orientation.
[0083] Operating system 226 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers that control and manage general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware and software components.
[0084] Communications module 228 facilitates communication with other devices via one or more external ports 224 and also includes various software components for processing data received by RF circuitry 208 and / or external port 224. External port 224 (e.g., Universal Serial Bus (USB), FIREWIRE®, etc.) is adapted to couple to other devices directly or indirectly via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as, similar to, and / or compatible with the 30-pin connector used on iPod® (trademark of Apple Inc.) devices.
[0085] Contact / motion module 230, optionally in conjunction with display controller 256, detects contact with touchscreen 212 and other touch-sensing devices (e.g., a touchpad or physical click wheel). Contact / motion module 230 includes various software components for performing various operations related to contact detection, such as determining whether contact occurs (e.g., detecting a finger-down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there is contact movement and tracking the movement across the touch-sensitive surface (e.g., detecting one or more finger-drag events), and determining whether the contact has stopped (e.g., detecting a finger-up event or an interruption of contact). Contact / motion module 230 receives contact data from the touch-sensitive surface. Determining the movement of the contact point, as represented by the series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point. These actions are optionally applied to a single contact (e.g., a single finger contact) or multiple simultaneous contacts (e.g., "multi-touch" / multiple finger contacts). In some embodiments, contact / motion module 230 and display controller 256 detect contacts on the touchpad.
[0086] In some embodiments, the contact / motion module 230 uses one or more sets of intensity thresholds to determine whether an action has been performed by the user (e.g., to determine whether the user has “clicked” on an icon). In some embodiments, at least a subset of the intensity thresholds are determined according to software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator, but can be adjusted without modifying the physical hardware of the device 200). For example, the mouse “click” threshold of a trackpad or touchscreen display can be set to any of a wide range of pre-defined thresholds without modifying the trackpad or touchscreen display hardware. Additionally, in some implementations, a user of the device is provided with a software setting to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once via a system-level click “intensity” parameter).
[0087] Contact / motion module 230 optionally detects gesture input by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different movements, timing, and / or intensities of detected contacts). Thus, gestures are optionally detected by detecting particular contact patterns. For example, detecting a finger tap gesture includes detecting a finger down event, followed by detecting a finger up (lift off) event at the same location (or substantially the same location) as the finger down event (e.g., the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger down event, followed by one or more finger drag events, followed by detecting a finger up (lift off) event.
[0088] Graphics module 232 includes various known software components that render and display graphics on touchscreen 212 or other display, including components that vary the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual characteristics) of the displayed graphics. As used herein, the term "graphics" includes any object that can be displayed to a user, including, but not limited to, text, web pages, icons (e.g., user interface objects, including softkeys), digital images, video, animation, etc.
[0089] In some embodiments, graphics module 232 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. Graphics module 232 receives one or more codes specifying the graphics to be displayed, including coordinate data and other graphic characteristic data, as needed, from an application or the like, and then generates screen image data to output to display controller 256.
[0090] The tactile feedback module 233 includes various software components for generating instructions used by the tactile output generator 267 to generate tactile outputs at one or more locations on the device 200 in response to a user's interaction with the device 200.
[0091] A text input module 234, a component of the graphics module 232, in some examples provides a soft keyboard for entering text in various applications (e.g., contacts 237, email 240, IM 241, browser 247, and any other application requiring text input).
[0092] The GPS module 235 determines the location of the device and provides this information for use within various applications (e.g., to the phone 238 for use in location-based dialing, to the camera 243 as picture / video metadata, and to applications that provide location-based services such as a weather widget, a local yellow pages widget, and a maps / navigation widget).
[0093] Digital assistant client module 229 includes various client-side digital assistant instructions for providing client-side functionality of the digital assistant. For example, digital assistant client module 229 can accept audio input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces of portable multifunction device 200 (e.g., microphone 213, accelerometer 268, touch-sensitive display system 212, light sensor(s) 264, other input control devices 216, etc.). Digital assistant client module 229 can also provide audio (e.g., speech output), visual, and / or tactile output, etc., through various output interfaces of portable multifunction device 200 (e.g., speaker 211, touch-sensitive display system 212, tactile output generator(s) 267, etc.). For example, output may be provided as voice, sound, an alert, a text message, a menu, a graphic, a video, an animation, a vibration, and / or a combination of two or more of the above. During operation, the digital assistant client module 229 communicates with the DA server 106 using the RF circuitry 208.
[0094] User data and models 231 includes various data associated with a user (e.g., user-specific vocabulary data, user preference data, user-specified name pronunciations, data from the user's electronic address book, to-do lists, shopping lists, etc.) for providing the client-side functionality of the digital assistant. Additionally, user data and models 231 includes various models (e.g., speech recognition models, statistical language models, natural language processing models, ontologies, task flow models, service models, etc.) for processing user input and determining user intent.
[0095] In some embodiments, digital assistant client module 229 establishes a context associated with the user, the current user interaction, and / or the current user input by utilizing various sensors, subsystems, and peripherals of portable multifunction device 200 to gather additional information from the environment surrounding portable multifunction device 200. In some embodiments, digital assistant client module 229 provides context information, or a subset thereof, along with the user input to DA server 106 to assist in inferring the user's intent. In some embodiments, the digital assistant also uses the context information to determine how to prepare and deliver output to the user. The context information is referred to as context data.
[0096] In some embodiments, the context information accompanying the user input includes sensor information, such as lighting, ambient noise, ambient temperature, images or videos of the surrounding environment, etc. In some embodiments, the context information may also include the physical state of the device, such as device orientation, device location, device temperature, power level, speed, acceleration, motion patterns, cellular signal strength, etc. In some embodiments, information about the software state of DA server 106, such as running processes, installed programs, past and present network activity, background services, error logs, resource usage, etc., as well as information about the software state of portable multifunction device 200, is provided to DA server 106 as context information associated with the user input.
[0097] In some embodiments, digital assistant client module 229 selectively provides information stored on portable multifunction device 200 (e.g., user data 231) in response to a request from DA server 106. In some embodiments, digital assistant client module 229 also elicits additional input from the user via a natural language dialog or other user interface in response to a request by DA server 106. Digital assistant client module 229 passes the additional input to DA server 106 to assist DA server 106 in intent inference and / or fulfillment of the user's intent expressed in the user request.
[0098] A more detailed description of the digital assistant is provided below with reference to Figures 7A-7C. It should be appreciated that the digital assistant client module 229 can include any number of sub-modules of the digital assistant module 726 described below.
[0099] The application 236 includes the following modules (or sets of instructions), or a subset or superset thereof: a contacts module 237 (sometimes called an address book or contact list); Telephone module 238, ·Videoconferencing module 239, an email client module 240; Instant messaging (IM) module 241, · Training Support Module 242, camera module 243 for still images and / or video; Image management module 244, Video player module, Music player module, Browser module 247, · Calendar module 248, a widget module 249, which in some embodiments includes one or more of a weather widget 249-1, a stocks widget 249-2, a calculator widget 249-3, an alarm clock widget 249-4, a dictionary widget 249-5, and a user-created widget 249-6; a widget creator module 250 for creating user-created widgets 249-6; · Search module 251, A video and music player module 252 that integrates a video player module and a music player module; · Memo module 253, a map module 254, and / or ·255 online video modules.
[0100] Examples of other applications 236 stored in memory 202 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice duplication.
[0101] The contacts module 237, in conjunction with the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, and the text input module 234, is used to manage an address book or contact list (e.g., stored in the application internal state 292 of the contacts module 237 in the memory 202 or the memory 470), including adding names to the address book, removing names from the address book, associating phone numbers, email addresses, physical addresses, or other information with names, associating pictures with names, categorizing and sorting names, providing phone numbers or email addresses to initiate and / or facilitate communication via telephone 238, videoconferencing module 239, email 240, or IM 241, and the like.
[0102] The telephone module 238, in conjunction with the RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, is used to input character strings corresponding to telephone numbers, access one or more telephone numbers in contact module 237, modify input telephone numbers, dial each telephone number, conduct a conversation, and disconnect or hang up when the conversation is completed. Thus, wireless communication uses any of a number of communication standards, protocols, and technologies.
[0103] Videoconferencing module 239 includes executable instructions to cooperate with RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touchscreen 212, display controller 256, light sensor 264, light sensor controller 258, contact / motion module 230, graphics module 232, text input module 234, contact module 237, and telephone module 238 to initiate, conduct, and end a videoconference between a user and one or more other participants according to the user's commands.
[0104] Email client module 240, in conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, contains executable instructions for composing, sending, receiving, and managing emails in response to user commands. In conjunction with image management module 244, email client module 240 greatly facilitates the creation and sending of emails with still or video images captured by camera module 243.
[0105] Instant messaging module 241, in conjunction with RF circuitry 208, touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, includes executable instructions for entering character sequences corresponding to instant messages, modifying previously entered characters, sending individual instant messages (e.g., using Short Message Service (SMS) or Multimedia Message Service (MMS) protocols for telephony-based instant messaging, or XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, sent and / or received instant messages include graphics, photos, audio files, video files, and / or other attachments supported by MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both telephony-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0106] The training support module 242 includes executable instructions to cooperate with the RF circuitry 208, the touchscreen 212, the display controller 256, the contact / motion module 230, the graphics module 232, the text input module 234, the GPS module 235, the map module 254, and the music player module to create workouts (e.g., with time, distance, and / or calorie burn goals), communicate with training sensors (sports devices), receive training sensor data, calibrate sensors used to monitor workouts, select and play music for workouts, and display, store, and transmit workout data.
[0107] Camera module 243, in conjunction with touchscreen 212, display controller 256, light sensor 264, light sensor controller 258, contact / motion module 230, graphics module 232, and image management module 244, contains executable instructions for capturing and storing still images or video (including video streams) in memory 202, modifying the characteristics of the still images or video, or deleting the still images or video from memory 202.
[0108] Image management module 244 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slideshow or album), and storing still and / or video images in conjunction with touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and camera module 243.
[0109] Browser module 247, in conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, contains executable instructions for browsing the Internet according to user commands, including retrieving, linking to, receiving, and displaying web pages or portions thereof, as well as attachments and other files linked to web pages.
[0110] The calendar module 248 includes executable instructions to cooperate with the RF circuitry 208, the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, the text input module 234, the email client module 240, and the browser module 247 to create, display, modify, and store calendars and data associated with the calendars (e.g., calendar items, to-do lists, etc.) according to user instructions.
[0111] In conjunction with RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, and browser module 247, widget modules 249 are mini-applications that can be downloaded and used by a user (e.g., weather widget 249-1, stock quotes widget 249-2, calculator widget 249-3, alarm clock widget 249-4, and dictionary widget 249-5) or created by a user (e.g., user-created widget 249-6). In some embodiments, a widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, a widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! Widgets).
[0112] In conjunction with the RF circuitry 208, touch screen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, and browser module 247, widget creation module 250 is used by a user to create widgets (e.g., turn user-specified portions of a web page into widgets).
[0113] The search module 251 includes executable instructions for working in conjunction with the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, and the text input module 234 to search for text, music, sound, images, video, and / or other files in the memory 202 that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with a user's commands.
[0114] Video and music player module 252 includes executable instructions that, in conjunction with touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, and browser module 247, enable a user to download and play pre-recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing videos (e.g., on touchscreen 212 or on an external display connected via external port 224). In some embodiments, device 200 optionally includes the functionality of an MP3 player, such as an iPod (a trademark of Apple Inc.).
[0115] The notes module 253 includes executable instructions for cooperating with the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, and the text input module 234 to create and manage notes, to-do lists, and the like according to user commands.
[0116] In conjunction with RF circuitry 208, touch screen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, GPS module 235, and browser module 247, map module 254 is used to receive, display, modify, and store maps and data associated with maps (e.g., driving directions, data about stores, other locations at or near a particular location, and other location-based data) in accordance with user instructions.
[0117] Online video module 255, in conjunction with touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, text input module 234, email client module 240, and browser module 247, contains instructions that enable a user to access, browse for, receive (e.g., by streaming and / or downloading), and play (e.g., on the touchscreen or on an external display connected via external port 224) particular online videos, send emails with links to particular online videos, and otherwise manage online videos in one or more file formats, such as H.264. In some embodiments, instant messaging module 241 is used to send links to particular online videos, rather than email client module 240. For additional description of online video applications, see U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," the contents of which are incorporated herein by reference in their entireties.
[0118] Each of the above-identified modules and applications corresponds to a set of executable instructions that perform one or more of the functions and methods described herein (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules; various embodiments may combine or otherwise rearrange various subsets of these modules. For example, a video player module may be combined with a music player module into a single module (e.g., video and music player module 252, FIG. 2A). In some embodiments, memory 202 stores a subset of the above-identified modules and data structures. Additionally, memory 202 stores additional modules and data structures not described above.
[0119] In some embodiments, device 200 is a device in which operation of a predefined set of functions on the device is performed exclusively via a touchscreen and / or touchpad. By using the touchscreen and / or touchpad as the primary input control device for operation of device 200, the number of physical input control devices (push buttons, dials, etc.) on device 200 is reduced.
[0120] The set of predefined functions performed only through the touchscreen and / or touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 200 to a main menu, home menu, or root menu from any user interface displayed on device 200. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is a physical push button or other physical input control device rather than a touchpad.
[0121] 2B is a block diagram illustrating exemplary components for event processing, according to some embodiments. In some embodiments, memory 202 (FIG. 2A) or memory 470 (FIG. 4) includes event sorter 270 (e.g., in operating system 226) and corresponding application 236-1 (e.g., any of applications 237-251, 255, 480-490 described above).
[0122] Event sorter 270 receives the event information and determines which application 236-1 to deliver the event information to and application view 291 for application 236-1. Event sorter 270 includes event monitor 271 and event dispatcher module 274. In some embodiments, application 236-1 includes application internal state 292 that indicates the current application view that is displayed on touch-sensitive display 212 when the application is active or running. In some embodiments, device / global internal state 257 is used by event sorter 270 to determine which application(s) is currently active, and application internal state 292 is used by event sorter 270 to determine which application(s) is / are delivered to.
[0123] In some embodiments, application internal state 292 includes additional information such as one or more of resume information to be used when application 236-1 resumes execution, user interface state information indicating or ready to display information being displayed by application 236-1, state cues that allow the user to return to a previous state or view of application 236-1, and redo / undo cues of previous actions taken by the user.
[0124] Event monitor 271 receives event information from peripherals interface 218. The event information includes information about sub-events (e.g., a user touch as part of a multi-touch gesture on touch-sensitive display 212). Peripherals interface 218 transmits information it receives from I / O subsystem 206 or sensors such as proximity sensor 266, accelerometer(s) 268, and / or microphone 213 (via audio circuitry 210). Information that peripherals interface 218 receives from I / O subsystem 206 includes information from touch-sensitive display 212 or a touch-sensitive surface.
[0125] In some embodiments, event monitor 271 sends requests to peripherals interface 218 at predetermined intervals. In response, peripherals interface 218 transmits event information. In other embodiments, peripherals interface 218 transmits event information only when there is a significant event (e.g., receipt of an input above a predetermined noise threshold and / or for more than a predetermined period of time).
[0126] In some embodiments, event sorter 270 also includes a hit view determination module 272 and / or an active event recognizer determination module 273 .
[0127] Hit view determination module 272 provides a software procedure that determines where a sub-event occurred within one or more views when touch-sensitive display 212 is displaying more than one view. A view consists of the controls and other elements that a user can see on the display.
[0128] Another aspect of a user interface associated with an application is the set of views, sometimes referred to herein as application views or user interface windows, within which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected corresponds to a programmatic level within the application's program or view hierarchy. For example, the lowest-level view in which a touch is detected is called a hit view, and the set of events that are recognized as valid inputs is determined based at least in part on the hit view of the initial touch that initiates a touch-based gesture.
[0129] Hit view determination module 272 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized hierarchically, hit view determination module 272 identifies the hit view as the lowest view in the hierarchy that should process the sub-events. In most situations, the hit view is the lowest-level view in which an initiating sub-event occurs (e.g., the first sub-event in a sequence of sub-events that form an event or potential event). Once a hit view is identified by hit view determination module 272, the hit view typically receives all sub-events related to the same touch or input source as the touch or input source identified as the hit view.
[0130] The active event recognizer determination module 273 determines which view(s) in the view hierarchy should receive the particular sequence of sub-events. In some embodiments, the active event recognizer determination module 273 determines that only the hit view should receive the particular sequence of sub-events. In other embodiments, the active event recognizer determination module 273 determines that all views that contain the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive the particular sequence of sub-events. In other embodiments, even if the touch sub-event is completely confined to the area associated with one particular view, views higher in the hierarchy still remain actively participating views.
[0131] Event dispatcher module 274 dispatches event information to event recognizers (e.g., event recognizer 280). In embodiments that include active event recognizer determination module 273, event dispatcher module 274 delivers event information to the event recognizers determined by active event recognizer determination module 273. In some embodiments, event dispatcher module 274 stores event information obtained by individual event receivers 282 in an event queue.
[0132] In some embodiments, operating system 226 includes event sorter 270. Alternatively, application 236-1 includes event sorter 270. In still other embodiments, event sorter 270 is a stand-alone module or is part of another module stored in memory 202, such as contact / motion module 230.
[0133] In some embodiments, application 236-1 includes multiple event handlers 290 and one or more application views 291, each containing instructions for processing touch events that occur within a respective view of the application's user interface. Each application view 291 of application 236-1 includes one or more event recognizers 280. Typically, an individual application view 291 includes multiple event recognizers 280. In other embodiments, one or more of the event recognizers 280 are part of a separate module, such as a user interface kit (not shown) or a higher-level object from which application 236-1 inherits methods and other properties. In some embodiments, an individual event handler 290 includes one or more of a data updater 276, an object updater 277, a GUI updater 278, and / or event data 279 received from event sorter 270. Event handler 290 utilizes or calls data updater 276, object updater 277, or GUI updater 278 to update application internal state 292. Alternatively, one or more of the application views 291 include one or more separate event handlers 290. Also, in some embodiments, one or more of the data updater 276, object updater 277, and GUI updater 278 are included in the separate application views 291.
[0134] Individual event recognizer 280 receives event information (e.g., event data 279) from event sorter 270 and identifies events from the event information. Event recognizer 280 includes event receiver 282 and event comparator 284. In some embodiments, event recognizer 280 includes at least a subset of metadata 283 and event delivery instructions 288 (including sub-event delivery instructions).
[0135] Event receiver 282 receives event information from event sorter 270. The event information includes information about a sub-event, e.g., a touch or a movement of a touch. Depending on the sub-event, the event information also includes additional information, such as the position of the sub-event. If the sub-event is related to a movement of a touch, the event information also includes the speed and direction of the sub-event. In some embodiments, the event includes a rotation of the device from one orientation to another (e.g., from portrait to landscape or vice versa), and the event information includes corresponding information about the current orientation of the device (also called the device's attitude).
[0136] The event comparator 284 compares the event information to predefined event or sub-event definitions and determines the event or sub-event, or determines or updates the state of the event or sub-event, based on the comparison. In some embodiments, the event comparator 284 includes an event definition 286. The event definition 286 includes definitions of events (e.g., a sequence of predefined sub-events), such as Event 1 (287-1) and Event 2 (287-2). In some embodiments, sub-events within Event 1 (287-1) include, for example, touch start, touch end, touch movement, touch cancellation, and multiple touches. In one example, the definition for Event 1 (287-1) is a double tap on a displayed object. A double tap includes, for example, a first touch on a displayed object relative to a predetermined phase (touch start), a first lift-off (touch end) relative to the predetermined phase, a second touch on a displayed object relative to the predetermined phase (touch start), and a second lift-off (touch end) relative to the predetermined phase. In another example, a definition of event 2 (287-2) is a drag on a displayed object. Drag includes, for example, a touch (or contact) on the displayed object to a predetermined stage, a movement of the touch across the touch-sensitive display 212, and a lift-off of the touch (touch end). In some embodiments, the event also includes information about one or more associated event handlers 290.
[0137] In some embodiments, event definition 287 includes definitions of events for individual user interface objects. In some embodiments, event comparator 284 performs a hit test to determine which user interface objects are associated with the sub-event. For example, if a touch is detected on touch-sensitive display 212 in an application view in which three user interface objects are displayed on touch-sensitive display 212, event comparator 284 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a separate event handler 290, event comparator 284 uses the results of the hit test to determine which event handler 290 to activate. For example, event comparator 284 selects the event handler associated with the sub-event and object that triggers the hit test.
[0138] In some embodiments, the definition of an individual event 287 also includes a delay action that delays transmission of the event information until it is determined whether the sequence of sub-events corresponds to the event type of the event recognizer.
[0139] If the individual event recognizer 280 determines that the sequence of sub-events does not match any of the events in the event definition 286, the individual event recognizer 280 enters an event disabled, event failed, or event finished state, after which it ignores the next sub-event of the touch-based gesture. In this situation, any other event recognizers that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.
[0140] In some embodiments, individual event recognizers 280 include metadata 283 with configurable properties, flags, and / or lists that indicate to actively participating event recognizers how the event delivery system should perform sub-event delivery. In some embodiments, metadata 283 includes configurable properties, flags, and / or lists that indicate how event recognizers interact with each other or how event recognizers are allowed to interact with each other. In some embodiments, metadata 283 includes configurable properties, flags, and / or lists that indicate how sub-events are delivered to various levels in a view or programmatic hierarchy.
[0141] In some embodiments, an individual event recognizer 280 activates an event handler 290 associated with an event when one or more specific sub-events of the event are recognized. In some embodiments, the individual event recognizer 280 delivers event information associated with the event to the event handler 290. Activating the event handler 290 is separate from sending (and postponing sending) sub-events to the individual hit view. In some embodiments, the event recognizer 280 pops a flag associated with the recognized event, and the event handler 290 associated with the flag captures the flag and performs a predetermined process.
[0142] In some embodiments, the event delivery instructions 288 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver the event information to an event handler associated with a set of sub-events or to an actively participating view. The event handler associated with the set of sub-events or the actively participating view receives the event information and performs a predetermined process.
[0143] In some embodiments, data updater 276 creates and updates data used by application 236-1. For example, data updater 276 updates phone numbers used by contacts module 237 or stores video files used by video player module. In some embodiments, object updater 277 creates and updates objects used by application 236-1. For example, object updater 277 creates new user interface objects or updates the positions of user interface objects. GUI updater 278 updates the GUI. For example, GUI updater 278 prepares display information and sends the display information to graphics module 232 for display on the touch-sensitive display.
[0144] In some embodiments, event handler(s) 290 include or have access to data updater 276, object updater 277, and GUI updater 278. In some embodiments, data updater 276, object updater 277, and GUI updater 278 are included in a single module of an individual application 236-1 or application view 291. In other embodiments, they are included in two or more software modules.
[0145] It should be understood that the foregoing description of event processing of a user's touch on a touch-sensitive display also applies to other forms of user input for operating multifunction device 200 using input devices, although not all of them are initiated on a touchscreen. For example, mouse movements and mouse button presses, contact movements such as tapping, dragging, scrolling on a touchpad, optionally coordinated with single or multiple keyboard presses or holds, pen stylus input, device movement, verbal commands, detected eye movements, biometric input, and / or any combination thereof, optionally utilize as inputs corresponding to sub-events that define the recognized event.
[0146] 3 illustrates portable multifunction device 200 having touchscreen 212, according to some embodiments. The touchscreen optionally displays one or more graphics within user interface (UI) 300. In this embodiment, as well as other embodiments described below, a user may select one or more of the graphics by performing a gesture on the graphics, for example, using one or more fingers 302 (not drawn to scale) or one or more styluses 303 (not drawn to scale). In some embodiments, selection of one or more graphics is performed when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (left to right, right to left, upward and / or downward), and / or rolling (right to left, left to right, upward and / or downward) of a finger in contact with device 200. In some implementations or situations, accidental contact with a graphic does not select the graphic, for example, if the gesture corresponding to selection is a tap, a swipe gesture sweeping over an application icon optionally does not select the corresponding application.
[0147] Device 200 also includes one or more physical buttons, such as a "home" or menu button 304. As described above, menu button 304 is used to navigate to any application 236 within a set of applications running on device 200. Alternatively, in some embodiments, the menu button is implemented as a soft key within a GUI displayed on touchscreen 212.
[0148] In one embodiment, device 200 includes touchscreen 212, menu button 304, pushbutton 306 for powering the device on / off and locking the device, volume control button(s) 308, subscriber identity module (SIM) card slot 310, headset jack 312, and external docking / charging port 224. Pushbutton 306 is optionally used to power the device on / off by pressing and holding the button down for a predetermined period of time, to lock the device by pressing and releasing the button before the predetermined time has elapsed, and / or to unlock the device or initiate the unlocking process. In an alternative embodiment, device 200 also accepts verbal input via microphone 213 for activating or deactivating certain functions. Device 200 also optionally includes one or more contact intensity sensors 265 for detecting the intensity of a contact on touchscreen 212 and / or one or more tactile output generators 267 for generating a tactile output for a user of device 200.
[0149] 4 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface, according to some embodiments. Device 400 need not be portable. In some embodiments, device 400 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a child's learning toy), a gaming system, or a control device (e.g., a home or commercial controller). Device 400 typically includes one or more processing units (CPUs) 410, one or more network or other communication interfaces 460, memory 470, and one or more communication buses 420 interconnecting these components. Communication bus 420 optionally includes circuitry (sometimes called a chipset) that interconnects and controls communication between system components. Device 400 includes input / output (I / O) interface 430, including display 440, which is typically a touchscreen display. I / O interface 430 also optionally includes a keyboard and / or mouse (or other pointing device) 450, as well as a touchpad 455, a tactile output generator 457 (e.g., similar to tactile output generator(s) 267 described above with reference to FIG. 2A ) for generating tactile output on device 400, sensors 459 (e.g., optical sensors, acceleration sensors, proximity sensors, touch-sensitive sensors, and / or contact intensity sensors similar to contact intensity sensor(s) 265 described above with reference to FIG. 2A ). Memory 470 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 470 optionally includes one or more storage devices located remotely from CPU(s) 410.In some embodiments, memory 470 stores programs, modules, and data structures similar to, or a subset of, programs, modules, and data structures stored in memory 202 of portable multifunction device 200 (FIG. 2A). Additionally, memory 470 optionally stores additional programs, modules, and data structures not present in memory 202 of portable multifunction device 200. For example, memory 470 of device 400 optionally stores drawing module 480, presentation module 482, word processing module 484, website creation module 486, disc authoring module 488, and / or spreadsheet module 490, while memory 202 of portable multifunction device 200 (FIG. 2A) optionally does not store these modules.
[0150] Each of the above-identified elements in FIG. 4 may, in some examples, be stored in any one or more of the memory devices mentioned above. Each of the above-identified modules corresponds to a set of instructions that perform the functions described above. The above-identified modules or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules; thus, various subsets of these modules may be combined or otherwise reconfigured in various embodiments. In some embodiments, memory 470 stores a subset of the above-identified modules and data structures. Additionally, memory 470 stores additional modules and data structures not described above.
[0151] Attention is now directed to user interface embodiments that may be implemented in portable multifunction device 200, for example.
[0152] 5A shows an exemplary user interface for a menu of applications on portable multifunction device 200, according to some embodiments. A similar user interface is implemented on device 400. In some embodiments, user interface 500 includes the following elements, or a subset or superset thereof:
[0153] signal strength indicator(s) 502 for wireless communication(s), such as cellular and Wi-Fi signals; ●Time 504, ●Bluetooth indicator 505, ● Battery status indicator 506, A tray 508 with icons of frequently used applications, such as: An icon 516 for the phone module 238, labeled "Phone," optionally including an indicator 514 of the number of missed calls or voicemail messages; An icon 518 for the email client module 240, labeled "Mail," optionally including an indicator 510 of the number of unread emails; ○ An icon 520 for the browser module 247, labeled "Browser"; and ○ An icon 522 for the video and music player module 252, also called the iPod (trademark of Apple Inc.) module 252, labeled "iPod"; and ● Icons of other applications, such as: ○ Icon 524 of IM module 241, labeled "Messages" ○ Icon 526 of the calendar module 248, labeled "Calendar" ○ Icon 528 of the image management module 244, labeled "Photos" ○ An icon 530 for the camera module 243, labeled "camera"; ○ Icon 532 of the online video module 255, labeled "Online Video"; Icon 534 of Stock Price Widget 249-2, labeled "Stock Price" ○ Icon 536 of the map module 254, labeled "Map" ○ Icon 538 of weather widget 249-1, labeled "Weather" ○ Icon 540 of alarm clock widget 249-4, labeled "Clock" ○ Icon 542 of Training Support Module 242, labeled "Training Support"; ○ An icon 544 in the Notes module 253 labeled "Notes," and A settings application or module icon 546 labeled "Settings" that provides access to settings for the device 200 and its various applications 236.
[0154] 5A are merely exemplary. For example, icon 522 for video and music player module 252 is optionally labeled "Music" or "Music Player." Other labels are optionally used for various application icons. In some embodiments, the label for an individual application icon includes the name of the application that corresponds to the individual application icon. In some embodiments, the label for a particular application icon is different from the name of the application that corresponds to that particular application icon.
[0155] 5B shows an exemplary user interface on a device (e.g., device 400 of FIG. 4) that has a touch-sensitive surface 551 (e.g., tablet or touchpad 455 of FIG. 4) that is separate from display 550 (e.g., touchscreen display 212). Device 400 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 457) that detect the intensity of a contact on touch-sensitive surface 551, and / or one or more tactile output generators 459 that generate a tactile output for a user of device 400.
[0156] Although some of the following examples are described with reference to input on touchscreen display 212 (when the touch-sensitive surface and display are combined), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, as shown in FIG. 5B . In some embodiments, this touch-sensitive surface (e.g., 551 in FIG. 5B ) has a major axis (e.g., 552 in FIG. 5B ) that corresponds to a major axis (e.g., 553 in FIG. 5B ) on the display (e.g., 550). According to these embodiments, the device detects contact with touch-sensitive surface 551 (e.g., 560 and 562 in FIG. 5B ) at locations that correspond to respective locations on the display (e.g., in FIG. 5B , 560 corresponds to 568 and 562 corresponds to 570). In this manner, when the touch-sensitive surface is separate from the display, user input (e.g., contacts 560 and 562 and their movement) detected by the device on the touch-sensitive surface (e.g., 551 in FIG. 5B ) is used by the device to operate a user interface on the display (e.g., 550 in FIG. 5B ) of the multifunction device. It should be understood that similar methods are optionally used for the other user interfaces described herein.
[0157] Additionally, while the following examples are given primarily with reference to finger input (e.g., finger contact, finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of the finger inputs are replaced with input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a contact) followed by movement of a cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is optionally replaced by a mouse click (e.g., instead of detecting a contact and then ceasing contact detection) while the cursor is located over the location of the tap gesture. Similarly, it should be understood that when multiple user inputs are detected simultaneously, multiple computer mice are optionally used simultaneously, or a mouse and finger contacts are optionally used simultaneously.
[0158] FIG. 6A shows an exemplary personal electronic device 600. Device 600 includes a main body 602. In some embodiments, device 600 includes some or all of the features described in connection with devices 200 and 400 (e.g., FIGS. 2A-4 ). In some embodiments, device 600 includes a touch-sensitive display screen 604, hereafter touchscreen 604. Instead of, or in addition to, touchscreen 604, device 600 includes a display and a touch-sensitive surface. As with devices 200 and 400, in some embodiments, touchscreen 604 (or the touch-sensitive surface) includes one or more intensity sensors that detect the intensity of an applied contact (e.g., a touch). The one or more intensity sensors in touchscreen 604 (or the touch-sensitive surface) provide output data that represents the intensity of the touch. The user interface of device 600 responds to touches based on the intensity of the touch, meaning that touches of different intensities can invoke different user interface actions on device 600.
[0159] Techniques for detecting and processing touch intensity can be found, for example, in related applications: International Patent Application No. PCT / US2013 / 040061, filed May 8, 2013, entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," and International Patent Application No. PCT / US2013 / 069483, filed November 11, 2013, entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," each of which is incorporated herein by reference in its entirety.
[0160] In some embodiments, device 600 has one or more input mechanisms 606 and 608. Input mechanisms 606 and 608, if included, are physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 600 has one or more attachment mechanisms. Such attachment mechanisms, if included, can allow device 600 to be attached to, for example, hats, eyewear, earrings, necklaces, shirts, jackets, bracelets, watch bands, chains, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow device 600 to be worn by a user.
[0161] FIG. 6B illustrates an exemplary personal electronic device 600. In some embodiments, device 600 includes some or all of the components described in connection with FIGS. 2A, 2B, and 4. Device 600 includes a bus 612 operably coupling an I / O unit 614 to one or more computer processors 616 and memory 618. The I / O unit 614 is connected to a display 604, which may have touch-sensing components 622 and, optionally, touch-intensity-sensing components 624. Additionally, I / O unit 614 is connected to a communication unit 630 that receives application and operating system data using Wi-Fi, Bluetooth, near-field communication (NFC), cellular, and / or other wireless communication technologies. Device 600 includes input mechanisms 606 and / or 608. Input mechanism 606 is, for example, a rotatable input device or a depressible and rotatable input device. Input mechanism 608 is, in some examples, a button.
[0162] The input mechanism 608 is, in some examples, a microphone. The personal electronic device 600 includes various sensors, such as, for example, a GPS sensor 632, an accelerometer 634, an orientation sensor 640 (e.g., a compass), a gyroscope 636, a motion sensor 638, and / or combinations thereof, all of which are operably connected to the I / O section 614.
[0163] The memory 618 of the personal electronic device 600 is a non-transitory computer-readable storage medium that stores computer-executable instructions that, when executed by one or more computer processors 616, for example, cause the computer processors to perform the following techniques and processes. Those computer-executable instructions may also be stored and / or transmitted in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as, for example, a computer-based system, a system including a processor, or other system capable of fetching instructions from and executing those instructions. The personal electronic device 600 is not limited to the components and configuration of FIG. 6B and may include other or additional components in multiple configurations.
[0164] As used herein, the term "affordance" refers to a user-interactive graphical user interface object displayed on a display screen of, for example, device 200, 400, 600, 800, 900, 902, or 904 (FIGS. 2A, 4, 6A-6B, 8A-8CT, 9A-9C, 10A-10V, 12, 14, 15, and 16). For example, an image (e.g., an icon), a button, and text (e.g., a hyperlink) each constitute an affordance.
[0165] As used herein, the term “focus selector” refers to an input element that indicates the current portion of the user interface with which the user is interacting. In some implementations including a cursor or other position marker, the cursor serves as a “focus selector” such that when input (e.g., a press input) is detected on a touch-sensitive surface (e.g., touchpad 455 in FIG. 4 or touch-sensitive surface 551 in FIG. 5B ) while the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted according to the detected input. In some implementations including a touchscreen display (e.g., touch-sensitive display system 212 in FIG. 2A or touchscreen 212 in FIG. 5A ) that allows direct interaction with user interface elements on the touchscreen display, a contact detected on the touchscreen serves as a “focus selector” such that when input (e.g., a press input by a contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, the particular user interface element is adjusted according to the detected input. In some implementations, focus is moved from one region of the user interface to another region of the user interface without a corresponding cursor movement or contact movement on the touchscreen display (e.g., by using the tab key or arrow keys to move focus from one button to another), and in these implementations, the focus selector moves to follow the movement of focus between various regions of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or contact on a touchscreen display) that is controlled by the user to communicate the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface through which the user intends to interact).For example, the position of a focus selector (e.g., cursor, touch, or selection box) over an individual button while a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen) indicates that the user intends to activate that individual button (and not other user interface elements shown on the device's display).
[0166] As used herein and in the claims, the term "characteristic intensity" of a contact refers to a characteristic of that contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on a plurality of intensity samples. The characteristic intensity is optionally based on a predetermined number of intensity samples, i.e., a set of intensity samples collected during a predetermined time period (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined event (e.g., after detecting the contact, before detecting lift-off of the contact, before or after detecting the start of contact movement, before detecting the end of the contact, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The characteristic intensity of the contact is optionally based on one or more of the maximum intensity of the contact, the median intensity of the contact, the average intensity of the contact, the top 10 percent of the intensity of the contact, half the maximum intensity of the contact, 90 percent of the maximum intensity of the contact, etc. In some embodiments, the duration of the contact is used in determining the characteristic intensity (e.g., when the characteristic intensity is an average of the intensity of the contact over time). In some embodiments, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether an action is performed by the user. For example, the set of one or more intensity thresholds includes a first intensity threshold and a second intensity threshold. In this example, a contact having a characteristic intensity that does not exceed the first threshold results in a first action, a contact having a characteristic intensity above the first intensity threshold but not above the second intensity threshold results in a second action, and a contact having a characteristic intensity above the second threshold results in a third action. In some embodiments, the comparison of the characteristic intensity to one or more thresholds is not used to determine whether to perform the first action or the second action, but rather to determine whether to perform one or more actions (e.g., whether to perform the respective action or to forgo performing the respective action).
[0167] In some embodiments, a portion of the gesture is identified for purposes of determining the characteristic intensity. For example, the touch-sensitive surface receives a continuous swipe contact transitioning from a start location point to an end location point of increasing contact intensity. In this example, the characteristic intensity of the contact at the end location is based on only the portion of the continuous swipe contact, rather than the entire swipe contact (e.g., only the portion of the swipe contact at the end location). In some embodiments, a smoothing algorithm is applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of an unweighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some situations, these smoothing algorithms eliminate narrow spikes or dips in the swipe contact intensity for purposes of determining the characteristic intensity.
[0168] The intensity of a contact on the touch-sensitive surface is characterized with respect to one or more intensity thresholds, such as a contact-detection intensity threshold, a light pressure intensity threshold, a deep pressure intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light pressure intensity threshold corresponds to an intensity at which the device performs an action typically associated with clicking a physical mouse button or trackpad. In some embodiments, the deep pressure intensity threshold corresponds to an intensity at which the device performs an action different from an action typically associated with clicking a physical mouse button or trackpad. In some embodiments, when a contact is detected having a characteristic intensity below the light pressure intensity threshold (e.g., and above a nominal contact-detection intensity threshold below which the contact is not detected), the device follows the movement of the contact on the touch-sensitive surface and moves the focus selector without performing an action associated with the light pressure intensity threshold or the deep pressure intensity threshold. In general, unless otherwise specified, these intensity thresholds are consistent across various sets of user interface values.
[0169] An increase in the characteristic intensity of a contact from an intensity below the light pressure intensity threshold to an intensity between the light pressure intensity threshold and the deep pressure intensity threshold may be referred to as inputting a "light press." An increase in the characteristic intensity of a contact from an intensity below the deep pressure intensity threshold to an intensity above the deep pressure intensity threshold may be referred to as inputting a "deep press." An increase in the characteristic intensity of a contact from an intensity below the contact-detection intensity threshold to an intensity between the contact-detection intensity threshold and the light pressure intensity threshold may be referred to as detecting a contact on the touch surface. A decrease in the characteristic intensity of a contact from an intensity above the contact-detection intensity threshold to an intensity below the contact-detection intensity threshold may be referred to as detecting a lift-off of the contact from the touch surface. In some embodiments, the contact-detection intensity threshold is zero. In some embodiments, the contact-detection intensity threshold is greater than zero.
[0170] In some embodiments described herein, one or more actions are performed in response to detecting a gesture including an individual pressure input or in response to detecting an individual pressure input performed by an individual contact (or multiple contacts), where the individual pressure input is detected based at least in part on detecting an increase in intensity of the contact (or multiple contacts) above a pressure input intensity threshold. In some embodiments, the individual action is performed in response to detecting an increase in intensity of the individual contact above the pressure input intensity threshold (e.g., a "downstroke" of the individual pressure input). In some embodiments, the pressure input includes an increase in intensity of the individual contact above the pressure input intensity threshold followed by a decrease in intensity of the contact below the pressure input intensity threshold, and the individual action is performed in response to detecting a subsequent decrease in intensity of the individual contact below the pressure input threshold (e.g., an "upstroke" of the individual pressure input).
[0171] In some embodiments, the device employs intensity hysteresis to avoid accidental input, sometimes referred to as “jitter,” and the device defines or selects a hysteresis intensity threshold that has a predetermined relationship to the pressure input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units below the pressure input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the pressure input intensity threshold). Thus, in some embodiments, the pressure input includes an increase in the intensity of a discrete contact above the pressure input intensity threshold followed by a decrease in the intensity of the contact below the hysteresis intensity threshold corresponding to the pressure input intensity threshold, and a discrete action is performed in response to detecting a subsequent decrease in the intensity of the discrete contact below the hysteresis intensity threshold (e.g., an “upstroke” of the discrete pressure input). Similarly, in some embodiments, a pressure input is detected only when the device detects an increase in the intensity of the contact from an intensity below the hysteresis intensity threshold to an intensity above the pressure input intensity threshold, and optionally a subsequent decrease in the intensity of the contact to an intensity below the hysteresis intensity, and a distinct action is performed in response to detecting the pressure input (e.g., an increase in the intensity of the contact or a decrease in the intensity of the contact, as the case may be).
[0172] For ease of explanation, descriptions of operations performed in response to a pressure input associated with a pressure input intensity threshold, or a gesture including a pressure input, are optionally triggered in response to detecting any of: an increase in the intensity of the contact above the pressure input intensity threshold; an increase in the intensity of the contact from an intensity below a hysteresis intensity threshold to an intensity above the pressure input intensity threshold; a decrease in the intensity of the contact below the pressure input intensity threshold; and / or a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to the pressure input intensity threshold. Further, in examples where an operation is described as being performed in response to detecting a decrease in the intensity of the contact below a pressure input intensity threshold, the operation is optionally performed in response to detecting a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to and lower than the pressure input intensity threshold. 3. Digital Assistant System
[0173] FIG. 7A shows a block diagram of a digital assistant system 700 according to various embodiments. In some embodiments, digital assistant system 700 is implemented on a standalone computer system. In some embodiments, digital assistant system 700 is distributed across multiple computers. In some embodiments, some of the modules and functionality of the digital assistant are allocated to a server portion and a client portion, where the client portion resides on one or more user devices (e.g., devices 104, 122, 200, 400, 600, 800, 900, 902, or 904), for example, as shown in FIG. 1, and communicates with the server portion (e.g., server system 108) through one or more networks. In some embodiments, digital assistant system 700 is an implementation of server system 108 (and / or DA server 106) shown in FIG. 1. It should be noted that digital assistant system 700 is only one example of a digital assistant system, and that digital assistant system 700 may have more or fewer components than those shown, may combine two or more components, or may have a different configuration or arrangement of the components. The various components shown in FIG. 7A may be implemented as hardware, including one or more signal processing circuits and / or application specific integrated circuits, software instructions executed by one or more processors, firmware, or a combination thereof.
[0174] Digital assistant system 700 includes memory 702, one or more processors 704, an input / output (I / O) interface 706, and a network communication interface 708. These components can communicate with each other via one or more communication buses or signal lines 710.
[0175] In some embodiments, memory 702 includes a non-transitory computer-readable medium, such as high-speed random access memory and / or a non-volatile computer-readable storage medium (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).
[0176] In some embodiments, I / O interface 706 couples input / output devices 716 of digital assistant system 700, such as a display, keyboard, touchscreen, and microphone, to user interface module 722. I / O interface 706 interfaces with user interface module 722 to receive user inputs (e.g., voice input, keyboard input, touch input, etc.) and process them accordingly. In some embodiments, for example, when the digital assistant is implemented on a standalone user device, digital assistant system 700 includes any of the components and I / O communication interfaces described in connection with devices 200, 400, 600, 800, 900, 902, and 904 in FIGS. 2A, 4, 6A-6B, 8A-8CT, 9A-9C, 10A-10V, 12, 14, 15, and 16, respectively. In some embodiments, digital assistant system 700 represents the server portion of a digital assistant implementation and can interact with a user through a client-side portion that resides on a user device (e.g., device 104, 200, 400, 600, 800, 900, 902, or 904).
[0177] In some embodiments, network communication interface 708 includes wired communication port(s) 712 and / or wireless transceiver circuitry 714. The wired communication port(s) transmit and receive communication signals via one or more wired interfaces, such as Ethernet, Universal Serial Bus (USB), FIREWIRE®, etc. The wireless circuitry 714 transmits and receives RF and / or optical signals to and from communication networks and other communication devices. Wireless communication uses any of a number of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. Network communication interface 708 enables communication between digital assistant system 700 and networks, such as the Internet, intranets, and / or wireless networks, such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs), and with other devices.
[0178] In some embodiments, memory 702, or the computer-readable storage medium of memory 702, stores programs, modules, instructions, and data structures, including all or a subset of an operating system 718, a communications module 720, a user interface module 722, one or more applications 724, and a digital assistant module 726. In particular, memory 702, or the computer-readable storage medium of memory 702, stores instructions for performing the processes described below. One or more processors 704 execute these programs, modules, and instructions and read / write from / to the data structures.
[0179] Operating system 718 (e.g., Darwin, RTXC, LINUX, UNIX, iOS, OS X, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitating communication between various hardware, firmware, and software components.
[0180] Communications module 720 facilitates communication between digital assistant system 700 and other devices via network communications interface 708. For example, communications module 720 communicates with RF circuitry 208 of electronic devices such as devices 200, 400, and 600 shown in FIGS. 2A, 4, and 6A-6B, respectively. Communications module 720 also includes various components for processing data received by wireless circuitry 714 and / or wired communications port 712.
[0181] The user interface module 722 receives commands and / or input from a user via the I / O interface 706 (e.g., from a keyboard, touchscreen, pointing device, controller, and / or microphone) and generates user interface objects on the display. The user interface module 722 also prepares and delivers output (e.g., speech, sound, animation, text, icons, vibration, haptic feedback, light, etc.) to the user via the I / O interface 706 (e.g., through a display, audio channel, speaker, touchpad, etc.).
[0182] Applications 724 include programs and / or modules configured to be executed by one or more processors 704. For example, if the digital assistant system is implemented on a standalone user device, applications 724 include user applications such as games, calendar applications, navigation applications, or email applications. If the digital assistant system 700 is implemented on a server, applications 724 include, for example, resource management applications, diagnostic applications, or scheduling applications.
[0183] Memory 702 also stores digital assistant module 726 (or the server portion of the digital assistant). In some embodiments, digital assistant module 726 includes the following submodules, or a subset or superset thereof: input / output processing module 728, speech-to-text (STT) processing module 730, natural language processing module 732, dialog flow processing module 734, task flow processing module 736, service processing module 738, and speech synthesis processing module 740. Each of these modules has access to one or more of the following systems or data and models of digital assistant module 726, or a subset or superset thereof: ontology 760, vocabulary index 744, user data 748, task flow model 754, service model 756, and ASR system 758.
[0184] In some examples, using the processing modules, data, and models implemented in digital assistant module 726, the digital assistant can perform at least some of the following: converting speech input to text and identifying the user's intent expressed in the natural language input received from the user; actively eliciting and obtaining the information necessary to fully infer the user's intent (e.g., by disambiguating words, games, intent, etc.); determining a task flow to satisfy the inferred intent; and executing the task flow to satisfy the inferred intent.
[0185] In some embodiments, as shown in FIG. 7B , I / O processing module 728 interacts with a user through I / O device 716 of FIG. 7A or with a user device (e.g., device 104, 200, 400, 600, or 800) through network communication interface 708 of FIG. 7A to obtain user input (e.g., speech input) and to provide responses to the user input (e.g., as speech output). I / O processing module 728 optionally obtains contextual information associated with the user input from the user device along with or immediately after receiving the user input. The contextual information includes user-specific data, vocabulary, and / or preferences related to the user input. In some embodiments, the contextual information also includes information about the software and hardware state of the user device at the time the user request is received and / or the user's ambient environment at the time the user request is received. In some embodiments, I / O processing module 728 also sends follow-up questions to the user regarding the user request and receives answers from the user. When a user request is received by the I / O processing module 728 and the user request includes speech input, the I / O processing module 728 forwards the speech input to the STT processing module 730 (or speech recognizer) for speech-to-text conversion.
[0186] The STT processing module 730 includes one or more ASR systems 758. The one or more ASR systems 758 can process speech input received via the I / O processing module 728 to generate recognition results. Each ASR system 758 includes a front-end speech preprocessor. The front-end speech preprocessor extracts representative features from the speech input. For example, the front-end speech preprocessor performs a Fourier transform on the speech input to extract spectral features that characterize the speech input as a sequence of representative multi-dimensional vectors. Furthermore, each ASR system 758 includes one or more speech recognition models (e.g., acoustic models and / or language models) and implements one or more speech recognition engines. Examples of speech recognition models include hidden Markov models, Gaussian mixture models, deep neural network models, n-gram language models, and other statistical models. Examples of speech recognition engines include dynamic time warping-based engines and weighted finite-state transducer (WFST)-based engines. One or more speech recognition models and one or more speech recognition engines are used to process the extracted representative features of the front-end speech preprocessor to generate intermediate recognition results (e.g., phonemes, phoneme strings, subwords) and ultimately text recognition results (words, word strings, sequences of tokens). In some implementations, the speech input is at least partially processed by a third-party service or on the user's device (e.g., device 104, 200, 400, 600, or 800) to generate the recognition results. Once the STT processing module 730 generates a recognition result including a text string (e.g., a word, a string of words, or a string of tokens), the recognition result is passed to the natural language processing module 732 for intent inference. In some implementations, the STT processing module 730 generates multiple candidate text representations of the speech input. Each candidate text representation is a sequence of words or tokens corresponding to the speech input. In some implementations, each candidate text representation is associated with a speech recognition confidence score.Based on the speech recognition confidence scores, the STT processing module 730 ranks the candidate text representations and provides the n best (e.g., the n highest-ranked) candidate text representation(s) to the natural language processing module 732 for intent inference, where n is a predetermined integer greater than zero. For example, in one embodiment, only the highest-ranked (n=1) candidate text representation is passed to the natural language processing module 732 for intent inference. In another embodiment, the five highest-ranked (n=5) candidate text representations are passed to the natural language processing module 732 for intent inference.
[0187] Further details regarding speech-to-text processing are described in U.S. Utility Patent Application No. 13 / 236,942, filed September 20, 2011, for "Consolidating Speech Recognition Results," the entire disclosure of which is incorporated herein by reference.
[0188] In some implementations, the STT processing module 730 includes a vocabulary of recognizable words and / or has access to that vocabulary via the phonetic alphabet conversion module 731. Each vocabulary word is associated with one or more candidate pronunciations of that word, represented in a speech recognition phonetic alphabet. Specifically, the vocabulary of recognizable words includes words associated with multiple candidate pronunciations. For example, the vocabulary may include: The example includes the word "tomato," which is associated with a pronunciation candidate of TIFF0007809149000001.tif6128. Additionally, vocabulary words are associated with custom pronunciation candidates based on previous speech input from the user. Such custom pronunciation candidates are stored within the STT processing module 730 and associated with a particular user via the user's profile on the device. In some examples, pronunciation candidates for a word are determined based on the spelling of the word and one or more linguistic and / or phonetic rules. In some examples, pronunciation candidates are generated manually, for example, based on known canonical pronunciations.
[0189] In some implementations, the pronunciation candidates are ranked based on the commonality of the pronunciation candidate. TIFF0007809149000002.tif6128 because the former is a more commonly used pronunciation (e.g., among all users, for users in a particular geographic region, or for any other suitable subset of users). In some examples, pronunciation candidates are ranked based on whether the pronunciation candidate is a custom pronunciation candidate associated with the user. For example, custom pronunciation candidates are ranked higher than regular pronunciation candidates. This can be useful for recognizing proper nouns that have unique pronunciations that deviate from the regular pronunciation. In some examples, pronunciation candidates are associated with one or more speech characteristics, such as place of origin, nationality, or ethnicity. For example, pronunciation candidate TIFF0007809149000003.tif6128 is associated with the United States, while pronunciation candidates TIFF0007809149000004.tif6128 is associated with the United Kingdom. Furthermore, the ranking of the pronunciation candidates is based on one or more characteristics of the user (e.g., place of origin, nationality, ethnicity, etc.) stored in the user's profile on the device. For example, it may be determined from the user's profile that the user is associated with the United States. Based on the user's association with the United States, the pronunciation candidates (associated with the United States) may be ranked based on the user's association with the United States. TIFF0007809149000005.tif6128 is a pronunciation candidate (associated with the UK) TIFF0007809149000006.tif6128 is ranked higher than TIFF0007809149000006.tif6128. In some implementations, one of the ranked pronunciation candidates is selected as the predicted pronunciation (e.g., the most likely pronunciation).
[0190] When speech input is received, the STT processing module 730 is used to determine the phonemes that correspond to the speech input (e.g., using an acoustic model) and then attempts to determine a word that matches the phonemes (e.g., using a language model). For example, the STT processing module 730 first determines a sequence of phonemes that correspond to a portion of the speech input. If TIFF0007809149000007.tif6128 is located, then based on lexical index 744 it can be determined that this string corresponds to the word "tomato."
[0191] In some implementations, the STT processing module 730 uses proximity matching techniques to determine words in an utterance. Thus, for example, the STT processing module 730 may use a sequence of phonemes: It determines that TIFF0007809149000008.tif6128 corresponds to the word "tomato" even though that particular sequence of phonemes is not one of the candidate sequences of phonemes for that word.
[0192] The digital assistant's natural language processing module 732 ("natural language processor") takes the n best candidate text representation(s) ("word string(s)" or "token string(s)") generated by the STT processing module 730 and attempts to associate each of those candidate text representations with one or more "actionable intents" recognized by the digital assistant. An "actionable intent" (or "user intent") represents a task that can be performed by the digital assistant and may have an associated task flow implemented in task flow model 754. This associated task flow is a sequence of programmed actions and steps that the digital assistant performs to accomplish the task. The scope of the digital assistant's capabilities is determined according to the number and variety of task flows implemented and stored in task flow model 754, or in other words, the number and variety of "actionable intents" recognized by the digital assistant. However, the effectiveness of a digital assistant is also judged according to its ability to infer the correct "actionable intent(s)" from user requests expressed in natural language.
[0193] In some embodiments, in addition to the string of words or tokens obtained from STT processing module 730, natural language processing module 732 also receives contextual information associated with the user request, e.g., from I / O processing module 728. Natural language processing module 732 optionally uses the contextual information to clarify, complement, and / or further define the information contained in the candidate text representation received from STT processing module 730. Contextual information includes, for example, user preferences, hardware and / or software state of the user device, sensor information collected before, during, or immediately after the user request, and previous interactions (e.g., dialogs) between the digital assistant and the user. As described herein, contextual information is dynamic in some embodiments, changing with time, location, dialog content, and other factors.
[0194] In some embodiments, natural language processing is based on, for example, ontology 760. Ontology 760 is a hierarchical structure that includes multiple nodes, each of which represents an "actionable intent" or represents an "attribute" or other "properties" related to one or more of the "actionable intents." As described above, an "actionable intent" represents a task that a digital assistant can perform, i.e., the task is "actionable" or can be performed. An "attribute" represents a parameter associated with an actionable intent or a sub-aspect of another attribute. Links between actionable intent nodes and attribute nodes in ontology 760 define how the parameter represented by the attribute node contributes to the task represented by the actionable intent node.
[0195] In some embodiments, ontology 760 is composed of actionable intent nodes and attribute nodes. Within ontology 760, each actionable intent node is linked to one or more attribute nodes, either directly or through one or more intermediate attribute nodes. Similarly, each attribute node is linked to one or more actionable intent nodes, either directly or through one or more intermediate attribute nodes. For example, as shown in FIG. 7C , ontology 760 includes a "Restaurant Reservation" node (i.e., an actionable intent node). The attribute nodes "Restaurant," "Date / Time" (for the reservation), and "Number of Participants" are each directly linked to an actionable intent node (i.e., the "Restaurant Reservation" node).
[0196] Furthermore, the attribute nodes “cuisine,” “price range,” “phone number,” and “location” are subordinate nodes of the attribute node “restaurant,” and are each linked to the “restaurant reservation” node (i.e., an actionable intention node) via the intermediate attribute node “restaurant.” As another example, as shown in FIG. 7C , ontology 760 also includes a “set reminder” node (i.e., another actionable intention node). The attribute nodes “date / time” (for setting a reminder) and “theme” (for a reminder) are each linked to the “set reminder” node. Because the attribute node “date / time” is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node “date / time” is linked to both the “restaurant reservation” node and the “set reminder” node in ontology 760.
[0197] An actionable intent node, along with its linked attribute nodes, is described as a "domain." In this discussion, each domain is associated with a corresponding actionable intent and refers to a group of nodes (and the relationships between those nodes) associated with that particular actionable intent. For example, ontology 760 shown in FIG. 7C includes an example of a restaurant reservation domain 762 and an example of a reminder domain 764 within ontology 760. The restaurant reservation domain includes the actionable intent node "reservation of restaurant," the attribute nodes "restaurant," "date / time," and "number of participants," and the sub-attribute nodes "cuisine," "price range," "phone number," and "location." The reminder domain 764 includes the actionable intent node "reminder setting," and the attribute nodes "theme" and "date / time." In some embodiments, ontology 760 is composed of multiple domains. Each domain shares one or more attribute nodes with one or more other domains. For example, the "date / time" attribute node is associated with a number of different domains (eg, a scheduling domain, a travel booking domain, a movie ticket domain, etc.) in addition to a restaurant reservation domain 762 and a reminder domain 764.
[0198] 7C illustrates two example domains within ontology 760, although other domains include, for example, "find a movie," "make a phone call," "find directions," "schedule a meeting," "send a message," "provide an answer to a question," "read a list," "provide navigation instructions," and "provide task instructions." The "send message" domain is associated with the actionable intent of "send a message" and further includes attribute nodes such as "recipient," "message type," and "message body." The attribute node "recipient" is further defined by sub-attribute nodes such as "recipient name" and "message address."
[0199] In some implementations, ontology 760 includes all domains (and therefore actionable intents) that the digital assistant can understand and perform. In some implementations, ontology 760 is modified by adding or removing domains or entire nodes, or by modifying relationships between nodes in ontology 760, etc.
[0200] In some embodiments, nodes associated with related actionable intents are clustered under a "super domain" in ontology 760. For example, the "Travel" super domain includes a cluster of travel-related attribute nodes and actionable intent nodes. Actionable intent nodes related to travel include "book an airline ticket," "book a hotel," "rent a car," "get directions," "find points of interest," etc. Actionable intent nodes under the same super domain (e.g., the "Travel" super domain) share many attribute nodes. For example, the actionable intent nodes for "book an airline ticket," "book a hotel," "rent a car," "get directions," and "find points of interest" share one or more of the attribute nodes "origin," "destination," "departure date / time," "arrival date / time," and "number of participants."
[0201] In some embodiments, each node in ontology 760 is associated with a set of words and / or phrases related to the attribute or actionable intent represented by that node. The set of corresponding words and / or phrases associated with each node is the so-called "vocabulary" associated with that node. The set of corresponding words and / or phrases associated with each node, in association with the attribute or actionable intent represented by that node, is stored in vocabulary index 744. For example, returning to FIG. 7B , vocabulary associated with a node related to the attribute of "restaurant" includes words such as "food," "drink," "dish," "hungry," "eat," "pizza," "fast food," and "meal." As another example, vocabulary associated with a node related to the actionable intent of "initiate a phone call" includes words and phrases such as "call," "phone," "dial," "ring," "call this number," and "make a call to." Vocabulary index 744 optionally includes words and phrases in different languages.
[0202] Natural language processing module 732 receives candidate text representations (e.g., character strings or token strings) from STT processing module 730 and, for each candidate representation, determines which nodes are implied by the words in the candidate text representation. In some examples, if a word or phrase in the candidate text representation is found to be associated with one or more nodes in ontology 760 (via vocabulary index 744), the word or phrase "trigger" or "activate" those nodes. Based on the amount and / or relative importance of activated nodes, natural language processing module 732 selects one of those actionable intents as the task the user intends the digital assistant to perform. In some examples, the domain with the most "triggered" nodes is selected. In some examples, the domain with the highest confidence value (e.g., based on the relative importance of the various triggered nodes) is selected. In some examples, the domain is selected based on a combination of the number and importance of triggered nodes. In some examples, additional factors are also considered when selecting a node, such as whether the digital assistant has previously accurately interpreted a similar request from the user.
[0203] User data 748 includes user-specific information, such as user-specific vocabulary, user preferences, user addresses, the user's default and secondary languages, the user's contact list, and other short-term or long-term information about each user. In some implementations, natural language processing module 732 uses this user-specific information to supplement the information contained in the user input and further define the user intent. For example, for a user request "invite my friends to my birthday party," natural language processing module 732 can access user data 748 to determine who the "friends" are and when and where the "birthday party" will be held, without requiring the user to explicitly provide such information in the user request.
[0204] It should be appreciated that in some examples, natural language processing module 732 is implemented using one or more machine learning mechanisms (e.g., neural networks). Specifically, the one or more machine learning mechanisms are configured to receive candidate textual representations and contextual information associated with the candidate textual representations. Based on the candidate textual representations and the associated contextual information, the one or more machine learning mechanisms are configured to determine an intent confidence score across a set of actionable intent candidates. Natural language processing module 732 can select one or more actionable intent candidates from the set of actionable intent candidates based on the determined intent confidence scores. In some examples, an ontology (e.g., ontology 760) is also used to select one or more actionable intent candidates from the set of actionable intent candidates.
[0205] Further details of searching ontologies based on token strings are described in U.S. Utility Patent Application No. 12 / 341,743, filed December 22, 2008, entitled "Method and Apparatus for Searching Using an Active Ontology," the entire disclosure of which is incorporated herein by reference.
[0206] In some examples, once the natural language processing module 732 identifies an actionable intent (or domain) based on the user request, the natural language processing module 732 generates a structured query to represent the identified actionable intent. In some examples, the structured query includes parameters for one or more nodes in the domain related to the actionable intent, at least some of which are populated with specific information and requirements specified in the user request. For example, a user may say, "Make me a dinner reservation at a sushi place at 7." In this case, the natural language processing module 732 may accurately identify the actionable intent as "restaurant reservation" based on the user input. According to the ontology, a structured query for the "restaurant reservation" domain may include parameters such as {cuisine}, {time}, {date}, and {number of participants}. In some implementations, based on the speech input and text derived from the speech input using the STT processing module 730, the natural language processing module 732 generates a partially structured query for the restaurant reservation domain, where the partially structured query includes the parameter {cuisine="sushi"} and the parameter {time="7:00 PM"}. However, in this implementation, the information included in the user's utterance is insufficient to complete a structured query associated with the domain. Therefore, other required parameters, such as {number of participants} and {date}, are not specified in the structured query based on the currently available information. In some implementations, the natural language processing module 732 populates some parameters of the structured query with received context information. For example, in some implementations, if the user requests a sushi restaurant "nearby," the natural language processing module 732 populates the {location} parameter in the structured query with GPS coordinates from the user device.
[0207] In some implementations, the natural language processing module 732 identifies multiple actionable intent candidates for each text expression candidate received from the STT processing module 730. Furthermore, in some implementations, a corresponding (partial or complete) structured query is generated for each identified actionable intent candidate. The natural language processing module 732 determines an intent confidence score for each actionable intent candidate and ranks the actionable intent candidates based on the intent confidence score. In some implementations, the natural language processing module 732 passes the generated structured query(s), including any input parameters, to the task flow processing module 736 (“task flow processor”). In some implementations, the structured query(s) for the m best (e.g., m highest-ranked) actionable intent candidates are provided to the task flow processing module 736 (where m is a predetermined integer greater than zero). In some implementations, the structured query(s) for the m best feasible intent candidates are provided to task flow processing module 736 along with the corresponding textual representation candidate(s).
[0208] Further details of inferring user intent based on multiple possible intent candidates determined from multiple candidate text representations of a speech input are described in U.S. Utility Patent Application No. 14 / 298,725, filed June 6, 2014, for "System and Method for Inferring User Intent From Speech Inputs," the entire disclosure of which is incorporated herein by reference.
[0209] Task flow processing module 736 is configured to receive the structured query(s) from natural language processing module 732, complete the structured query as needed, and perform the actions required to "complete" the user's ultimate request. In some implementations, the various steps required to complete these tasks are provided in task flow model 754. In some implementations, task flow model 754 includes steps for obtaining additional information from the user and task flows for performing actions associated with the actionable intent.
[0210] As described above, completing a structured query may require task flow processing module 736 to initiate additional dialogue with the user to obtain additional information and / or disambiguate potentially ambiguous utterances. If such dialogue is necessary, task flow processing module 736 invokes dialog flow processing module 734 to engage in a dialogue with the user. In some implementations, dialog flow processing module 734 determines how (and / or when) to request additional information from the user and receives and processes the user response. Questions are provided to the user and answers are received from the user via I / O processing module 728. In some examples, dialog flow processing module 734 presents dialog output to the user via audio and / or visual output and receives input from the user via spoken or physical (e.g., click) responses. Continuing with the above example, when task flow processing module 736 invokes dialog flow processing module 734 to determine the "number of people" and "date" information for a structured query associated with the domain "restaurant reservation," dialog flow processing module 734 generates and passes questions to the user, such as "For how many people?" and "On which day?" Once answers are received from the user, dialog flow processing module 734 then either enters additional missing information into the structured query or passes the missing information to task flow processing module 736 to complete the missing information from the structured query.
[0211] Once the taskflow processing module 736 completes a structured query for an actionable intent, the taskflow processing module 736 proceeds to execute the final task associated with the actionable intent. Thus, the taskflow processing module 736 executes steps and instructions in a taskflow model according to specific parameters included in the structured query. For example, a taskflow model for an actionable intent of "restaurant reservation" includes steps and instructions for contacting a restaurant and actually requesting a reservation for a specific number of guests at a specific time. For example, using a structured query such as {Restaurant Reservation, Restaurant=ABC Cafe, Date=3 / 12 / 2012, Time=7 PM, Number of Guests=5}, the taskflow processing module 736 executes the following steps: (1) logging on to ABC Cafe's server or a restaurant reservation system such as OPENTABLE®; (2) entering date, time, and number of guests information into a form on the website; (3) submitting the form; and (4) entering a calendar item for the reservation into the user's calendar.
[0212] In some embodiments, task flow processing module 736 employs the assistance of service processing module 738 ("service processing module") to complete a task requested in the user input or to provide an answer to information requested in the user input. For example, service processing module 738 performs functions on behalf of task flow processing module 736, such as placing phone calls, setting calendar entries, invoking map searches, invoking or interacting with other user applications installed on the user device, and invoking or interacting with third-party services (e.g., restaurant reservation portals, social networking websites, banking portals, etc.). In some embodiments, the protocols and application programming interfaces (APIs) required by each service are specified by a corresponding service model in service models 756. Service processing module 738 accesses the appropriate service model for a service and generates requests for that service in accordance with the protocols and APIs required by that service according to the service model.
[0213] For example, if a restaurant supports an online reservation service, the restaurant submits a service model that specifies the parameters required to make a reservation and an API for communicating the values of the required parameters to the online reservation service. Upon request by task flow processing module 736, service processing module 738 establishes a network connection with the online reservation service using the web address stored in the service model and transmits the required reservation parameters (e.g., time, date, number of participants) to the online reservation interface in a format that complies with the online reservation service's API.
[0214] In some examples, the natural language processing module 732, the dialog flow processing module 734, and the task flow processing module 736 are used collectively and iteratively to infer and define a user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response (i.e., output to the user or completion of a task) to satisfy the user's intent. The generated response is a dialog response to the speech input that at least partially satisfies the user's intent. Furthermore, in some examples, the generated response is output as speech output. In these examples, the generated response is sent to the speech synthesis processing module 740 (e.g., a speech synthesizer), where it may be processed to synthesize a dialog response in the form of a speech. In still other examples, the generated response is data content related to satisfying the user request in the speech input.
[0215] In embodiments in which task flow processing module 736 receives multiple structured queries from natural language processing module 732, task flow processing module 736 first processes a first structured query among the received structured queries and attempts to complete the first structured query and / or perform one or more tasks or actions represented by the first structured query. In some embodiments, the first structured query corresponds to the highest-ranked actionable intent. In other embodiments, the first structured query is selected from the received structured queries based on a combination of the corresponding speech recognition confidence score and the corresponding intent confidence score. In some embodiments, if task flow processing module 736 encounters an error while processing the first structured query (e.g., due to a required parameter being undeterminable), task flow processing module 736 can proceed to select and process a second structured query among the received structured queries that corresponds to a lower-ranked actionable intent. This second structured query is selected based on, for example, the speech recognition confidence scores of the corresponding text representation candidates, the intent confidence scores of the corresponding actionable intent candidates, the missing required parameters in the first structured query, or any combination thereof.
[0216] The speech synthesis processing module 740 is configured to synthesize speech output for presentation to the user. The speech synthesis processing module 740 synthesizes the speech output based on text provided by the digital assistant. For example, the generated dialog response is in the form of a text string. The speech synthesis processing module 740 converts the text string into an audible speech output. The speech synthesis processing module 740 uses any suitable speech synthesis technique to generate the speech output from the text, including, but not limited to, concatenative synthesis, unit selection synthesis, diphone synthesis, domain-specific synthesis, formant synthesis, articulatory synthesis, synthesis based on a hidden Markov model (HMM), and sinusoidal synthesis. In some embodiments, the speech synthesis processing module 740 is configured to synthesize individual words based on a phoneme sequence corresponding to the word. For example, a phoneme sequence is associated with a word in the generated dialog response. The phoneme sequence is stored in metadata associated with the word. The speech synthesis processing module 740 is configured to process the phoneme sequences in the metadata directly to synthesize words in phonetic form.
[0217] In some embodiments, instead of (or in addition to) using speech synthesis processing module 740, speech synthesis is performed on a remote device (e.g., server system 108), and the synthesized speech is sent to the user device for output to the user. For example, this can be done in some implementations where output to the digital assistant is generated on a server system. Also, because server systems generally have more processing power or resources than user devices, it is possible to obtain higher quality speech output than would be practical with client-side synthesis.
[0218] More information regarding digital assistants can be found in U.S. Utility Application No. 12 / 987,982, filed January 10, 2011, entitled "Intelligent Automated Assistant," and U.S. Utility Application No. 13 / 251,088, filed September 30, 2011, entitled "Generating and Processing Task Items That Represent Tasks to Perform," the disclosures of which are incorporated herein by reference in their entireties. 4. Digital Assistant User Interface
[0219] 8A-8CT show user interfaces and digital assistant user interfaces according to various embodiments. 8A-8CT are used to illustrate processes described below, including the processes of FIGS. 17A-17F.
[0220] 8A illustrates electronic device 800. Device 800 may be implemented as device 104, device 122, device 200, or device 600. In some implementations, device 800 at least partially implements digital assistant system 700. In the example of FIG. 8A, device 800 is a smartphone having a display and a touch-sensitive surface. In other examples, device 800 is a different type of device, such as a wearable device (e.g., a smartwatch), a tablet device, a laptop computer, or a desktop computer.
[0221] 8A , device 800 displays on display 801 a user interface 802 that differs from a digital assistant (DA) user interface 803, discussed below. In the example of FIG. 8A , user interface 802 is a home screen user interface. In other examples, the user interface is a lock screen user interface or another type of application-specific user interface, such as a map application user interface, a weather application user interface, a messaging application user interface, a music application user interface, a video application user interface, or the like.
[0222] In some embodiments, device 800 receives user input while displaying a user interface different from DA user interface 803. Device 800 determines whether the user input meets criteria for initiating DA. Exemplary user inputs that meet the criteria for initiating DA include a predetermined type of speech input (e.g., "Hey, Siri"), an input selecting a virtual or physical button on device 800 (or an input selecting such a button for a predetermined period of time), a type of input received on an external device coupled to device 800, a type of user gesture performed on display 801 (e.g., a drag or swipe gesture from a corner of display 801 toward the center of display 801), and a type of input representing a movement of device 800 (e.g., lifting device 800 into a viewing position).
[0223] In some embodiments, following a determination that the user input meets the criteria for initiating DA, device 800 displays DA user interface 803 on a user interface. In some embodiments, displaying DA user interface 803 (or another display element) on the user interface includes replacing at least a portion of the display of the user interface with a representation of DA user interface 803 (or a representation of another graphical element). In some embodiments, following a determination that the user input does not meet the criteria for initiating DA, device 800 ceases displaying DA user interface 803 and instead performs an action in response to the user input (e.g., updates user interface 802).
[0224] FIG. 8B illustrates a DA user interface 803 displayed on user interface 802. In some embodiments, as shown in FIG. 8B, DA user interface 803 includes a DA indicator 804. In some embodiments, indicator 804 is displayed in different states to indicate the respective states of the DA. DA states include a listening state (indicating that the DA is sampling speech input), a processing state (indicating that the DA is processing natural language requests), a speaking state (indicating that the DA is providing audio and / or text output), and an idle state. In some embodiments, indicator 804 includes different visualizations to indicate different DA states. FIG. 8B illustrates indicator 804 in a listening state because the DA is ready to accept speech input after initiation based on detection of user input that meets criteria.
[0225] In some implementations, the size of the listening state indicator 804 changes based on the received natural language input. For example, the indicator 804 expands and contracts in real time according to the amplitude of the received speech input. Figure 8C shows the indicator 804 in the listening state. In Figure 8C, the device 800 receives the natural language speech input "What's the weather like today?" and the indicator 804 expands and contracts in real time according to the speech input.
[0226] Figure 8D shows a processing state indicator 804, e.g., indicating that the DA is processing the request "What's the weather today?" Figure 8E shows a speaking state indicator 804, e.g., indicating that the DA is currently providing audio output "Nice weather today" in response to the request. Figure 8F shows the indicator 804 in an idle state. In some embodiments, user input that selects the idle state indicator 804 transitions the DA (and indicator 804) to a listening state, e.g., by activating one or more microphones to sample audio input.
[0227] In some embodiments, the DA provides audio output in response to a user request while the device 800 provides other audio output. In some embodiments, the DA reduces the volume of the other audio output while simultaneously providing audio output in response to a user request and the other audio output. For example, the DA user interface 803 is displayed over a user interface that includes the currently playing media (e.g., a movie or song). When the DA provides audio output in response to a user request, the DA reduces the volume of the audio output of the playing media.
[0228] In some examples, the DA user interface 803 includes a DA response affordance. In some examples, the response affordance corresponds to a response for receiving natural language input by the DA. For example, FIG. 8E shows a device 800 displaying a response affordance 805 that includes weather information in response to received speech input.
[0229] As shown in FIGS. 8E-8F, device 800 displays indicator 804 in a first portion of display 801 and in a response affordance 805 in a second portion of display 801. In a third portion of display 801, a portion of user interface 802 where DA user interface 803 is displayed remains visible (e.g., is not visually obscured). For example, the portion of user interface 802 that remains visible was displayed in the third portion of display 801 prior to receiving the user input that initiated the digital assistant (e.g., FIG. 8A). In some examples, the first, second, and third portions of display 801 are referred to as the "indicator portion," the "response portion," and the "user interface (UI) portion," respectively.
[0230] In some examples, the UI portion is between the indicator portion (which displays indicator 804) and the response portion (which displays response affordance 805). For example, in Figure 8F, the UI portion includes a display area 8011 (e.g., a rectangular area) between the bottom of response affordance 805 and the top of indicator 804, with the side edges of display area 8011 defined by the side edges of response affordance 805 (or display 801). In some examples, the portion of user interface 802 that remains visible in the UI portion of display 801 includes one or more user-selectable graphical elements, e.g., links and / or affordances, such as the home screen application affordances of Figure 8F.
[0231] In some examples, device 800 displays responsive affordance 805 in a first state. In some examples, the first state includes a compact state, where the display size of responsive affordance 805 is small (e.g., compared to an expanded responsive affordance state described below) and / or responsive affordance 805 displays information in a compact (e.g., condensed) form (e.g., compared to an expanded responsive affordance state). In some examples, device 800 receives user input corresponding to selection of responsive affordance 805 in the first state, and in response, replaces the display of responsive affordance 805 in the first state with a display of responsive affordance 805 in a second state. In some examples, the second state is an expanded state, where the display size of responsive affordance 805 is large (e.g., compared to the compact state), and / or responsive affordance 805 displays a larger amount of information / more detailed information (e.g., compared to the compact state). In some examples, device 800 displays response affordance 805 in the first state by default, for example, such that device 800 initially displays response affordance 805 in the first state (FIGS. 8E-8G).
[0232] 8E-8G illustrate a first-state responsive affordance 805. As shown, the responsive affordance 805 provides compact weather information, for example, by providing the current temperature and conditions and omitting more detailed weather information (e.g., hourly weather information). FIG. 8G illustrates device 800 receiving user input 806 (e.g., a tap gesture) corresponding to selection of the first-state responsive affordance 805. While FIGS. 8G-8P generally illustrate that the user input corresponding to each selection of the responsive affordance is touch input, in other examples, the user input corresponding to selection of the responsive affordance is another type of input, such as voice input (e.g., "Show me more") or peripheral device input (e.g., input from a mouse or touchpad). FIG. 8H illustrates that in response to receiving user input 806, device 800 replaces the display of the first-state responsive affordance 805 with the display of the second-state responsive affordance 805. As shown, the response affordance 805 in the second state includes more detailed weather information.
[0233] In some examples, while displaying responsive affordance 805 in the second state, device 800 receives user input requesting to display responsive affordance 805 in the first state. In some examples, in response to receiving the user input, device 800 replaces the display of responsive affordance 805 in the second state with the display of responsive affordance 805 in the first state. For example, in FIG. 8H , DA user interface 803 includes a selectable element (e.g., a back button) 807. User input selecting selectable element 807 returns device 800 to the display of FIG. 8F .
[0234] In some examples, while displaying response affordance 805 in the second state, device 800 receives user input corresponding to selection of response affordance 805. In response to receiving the user input, device 800 displays a user interface of an application corresponding to response affordance 805. For example, FIG. 81 shows device 800 receiving user input 808 (e.g., a tap gesture) corresponding to selection of response affordance 805. FIG. 8J shows device 800 displaying a user interface 809 of a weather application in response to receiving user input 808.
[0235] In some embodiments, while displaying the user interface of an application, device 800 displays a selectable DA indicator. For example, FIG. 8J shows selectable DA indicator 810. In some embodiments, while displaying the user interface of an application, device 800 additionally or alternatively displays indicator 804 in a first portion of display 801, e.g., in an idle state.
[0236] In some embodiments, while displaying a user interface of an application, device 800 receives user input selecting a selectable DA indicator. In some embodiments, in response to receiving the user input, device 800 replaces the display of the application's user interface with DA user interface 803. In some embodiments, DA user interface 803 is the DA user interface displayed immediately before displaying the application's user interface. For example, FIG. 8K shows device 800 receiving user input 811 (e.g., a tap gesture) selecting DA indicator 810. FIG. 8L shows device 800 replacing the display of weather application user interface 809 with the display of DA user interface 803 in response to receiving user input 811.
[0237] User input 806 in FIG. 8G corresponds to a selection of a first portion of response affordance 805. In some embodiments, while device 800 is displaying response affordance 805 in a first state (e.g., a compact state), device 800 receives a user input corresponding to a selection of a second portion of response affordance 805. In some embodiments, the first portion (e.g., a bottom portion) of response affordance 805 includes information intended to respond to a user request. In some embodiments, the second portion (e.g., an top portion) of response affordance 805 includes a glyph indicating a category of response affordance 805 and / or associated text. Exemplary categories of response affordances include weather, stock prices, knowledge, calculator, messaging, music, maps, etc. The categories may correspond to categories of services that the DA can provide. In some embodiments, the first portion of response affordance 805 occupies a larger display area than the second portion of response affordance 805.
[0238] In some examples, in response to receiving user input corresponding to a selection of a second portion of response affordance 805, device 800 displays a user interface of an application corresponding to response affordance 805 (e.g., without displaying response affordance 805 in a second state). For example, FIG. 8M shows device 800 receiving user input 812 (e.g., a tap gesture) that selects a second portion of response affordance 805 displayed in a first state. FIG. 8N shows device 800 displaying user interface 809 of a weather application in response to receiving user input 812 (e.g., without displaying response affordance 805 in an expanded state). In this manner, a user can provide input that selects a different portion of response affordance 805 that expands response affordance 805 or causes the display of an application corresponding to response affordance 805, as shown in FIGS. 8G-8H and 8M-8N.
[0239] Figure 8N further shows that while displaying user interface 809, device 800 displays selectable DA indicator 810. User input selecting DA indicator 810 causes device 800 to return to the display of Figure 8M, e.g., similar to the examples illustrated by Figures 8K-8L. In some embodiments, while displaying user interface 809, device 800 displays DA indicator 804 in a first portion of display 801 (e.g., in an idle state).
[0240] In some examples, for some types of response affordances, user input corresponding to selecting any portion of the response affordance causes device 800 to display a user interface of the application corresponding to the response affordance. In some examples, this is because the response affordance cannot be displayed in more detail (e.g., in a second state). For example, a DA may not have additional information it can provide in response to a natural language input. Consider, for example, the natural language input of "What is 5 times 6?" FIG. 8O shows a DA user interface 803 displayed in response to the natural language input. DA user interface 803 includes a response affordance 813 displayed in a first state. Response affordance 813 includes the answer "5 x 6 = 30," but no additional information the DA can provide. FIG. 8O further shows device 800 receiving user input 814 (e.g., a tap gesture) that selects a first portion of response affordance 813. FIG. 8P shows that in response to receiving user input 814, device 800 displays a user interface 815 of an application that corresponds to response affordance 813, for example, a calculator application user interface.
[0241] In some examples, the response affordance includes a selectable element, such as selectable text indicating a link. FIG. 8Q shows a DA user interface 803 displayed in response to the natural language input "Tell me about famous bands." The DA user interface 803 includes a response affordance 816. The response affordance 816 includes information about "famous bands" and a selectable element 817 corresponding to member #1 of the "famous band." In some examples, the device 800 receives user input corresponding to the selection of the selectable element and, in response, displays an affordance (a second response affordance) corresponding to the selectable element on the response affordance. FIG. 8R shows the device 800 receiving user input 818 (e.g., a tap gesture) that selects the selectable element 817. FIG. 8S shows that, in response to receiving the user input 818, the device 800 displays a second response affordance 819 on the response affordance 816 that includes information about member #1, forming a stack of response affordances.
[0242] In some examples, device 800 visually hides the user interface in a third portion of display 801 (e.g., a portion that does not display any response affordances or indicators 804) or portion thereof while displaying the second response affordance over response affordance 816. In some examples, visually obscuring the user interface includes darkening the user interface or blurring the user interface. FIG. 8S shows device 800 visually hiding user interface 802 in a third portion of display 801 while second response affordance 819 is displayed over response affordance 816.
[0243] 8S shows that a portion of responsive affordance 816 remains visible while a second responsive affordance 819 is displayed over it. In other examples, second responsive affordance 819 replaces the display of responsive affordance 816, thereby obscuring a portion of responsive affordance 816.
[0244] Figure 8T shows device 800 receiving user input 820 (e.g., a tap gesture) that selects selectable element 821 ("Detroit") in second response affordance 819. Figure 8U shows that in response to receiving user input 820, device 800 displays third response affordance 822 in second response affordance 819. Third response affordance 822 includes information about member #1's hometown of Detroit. Figure 8U shows that user interface 802 remains visually obscured in a third portion of display 801.
[0245] 8U further illustrates that although there are three response affordances (e.g., 816, 819, and 822) in the stack of response affordances, device 800 only displays two response affordances in the stack. For example, third response affordance 822 and part of second response affordance 819 are displayed, but part of response affordance 816 is not displayed. Thus, in some examples, when three or more response affordances are stacked, device 800 only visually indicates that two response affordances are in the stack. In other examples, when response affordances are stacked, device 800 only visually indicates a single response affordance of the stack (e.g., such that the display of the next response affordance completely replaces the display of the previous response affordance).
[0246] 8V-8Y illustrate a user providing an input to return to a previous response affordance in the stack. In particular, in FIG. 8V, device 800 receives user input 823 (e.g., a swipe gesture) on third response affordance 822 that requests returning to second response affordance 819. FIG. 8W illustrates that in response to receiving user input 823, device 800 stops displaying third response affordance 822 and displays second response affordance 819 in its entirety. Device 800 further displays (e.g., reveals) a portion of response affordance 816. FIG. 8X illustrates device 800 receiving user input 824 (e.g., a swipe gesture) on second response affordance 819 that requests returning to response affordance 816. 8Y shows that in response to receiving user input 824, device 800 stops displaying second responsive affordance 819 and displays responsive affordance 816 in its entirety. In some examples, device 800 receives input to display the next responsive affordance in the stack (e.g., a swipe gesture in the opposite direction) and, in response, displays the next responsive affordance in the stack, in a manner similar to that described above. In other examples, navigating through the responsive affordances in the stack depends on other input means (e.g., user selection of a displayed "back" or "next" button), in a manner similar to that described above.
[0247] Figure 8Y further illustrates that user interface 802 is no longer visually obscured by the third portion of display 801. Thus, in some examples, as shown in Figures 8Q-8Y, user interface 802 is visually obscured when the responsive affordances are stacked and is not visually obscured when the affordances are not stacked. For example, user interface 802 is visually obscured when initial responsive affordance 816 is not displayed (or only partially displayed), and user interface 802 is visually not obscured when initial responsive affordance 816 is displayed in its entirety.
[0248] In some examples, a user interface (e.g., the user interface on which DA user interface 803 is displayed) includes an input field that occupies a fourth portion (e.g., an "input field portion") of display 801. The input field includes an area where a user can provide natural language input. In some examples, the input field corresponds to an application such as a messaging application, an email application, a note-taking application, a reminder application, a calendar application, or the like. FIG. 8Z shows a user interface 825 of a messaging application that includes an input field 826 that occupies the fourth portion of display 801.
[0249] 8AA shows DA user interface 803 displayed on user interface 825. Device 800 displays DA user interface 803 in response to the natural language input "What's this song?" DA user interface 803 includes indicator 804 in a first portion of display 801 and response affordance 827 (indicating the song identified by the DA) in a second portion of display 801.
[0250] In some examples, device 800 receives user input corresponding to a displacement of a response affordance from a first portion of display 801 to a fourth portion of display 801. In response to receiving the user input, device 800 replaces the display of the response affordance in the first portion of display 801 with a display of the response affordance in an input field. For example, FIGS. 8AB-8AD show device 800 receiving user input 828 that displaces response affordance 827 from the first portion of display 801 to input field 826. User input 828 corresponds to a drag gesture from the first portion of display 801 to the fourth portion of display 801, terminating in a lift-off event (e.g., a finger lift-off event) at the display of input field 826.
[0251] 8AB-8AD, while receiving user input 828, device 800 continuously displaces response affordance 827 from a first portion of display 801 to a fourth portion of display 801. For example, while response affordance 827 is displacing, device 800 displays response affordance 827 at a position corresponding to the current display contact position of each of user inputs 828. In some examples, while response affordance 827 is displacing, the displayed size of response affordance 827 decreases, for example, so that response affordance 827 contracts under the user's finger (or other input device) as it displaces. FIGS. 8AB-8AD further show that indicator 804 ceases to be displayed while continuously displacing response affordance 827.
[0252] FIG. 8AD now shows that reply affordance 827 is displayed in an input field 826 of a messaging application. FIG. 8AE shows that electronic device 800 receives user input 830 (e.g., a tap gesture) corresponding to selection of send message affordance 829. FIG. 8AF shows that in response to receiving user input 830, device 800 sends reply affordance 827 as a message. In this manner, a user can send a reply affordance in a communication (e.g., a text message, an email) by providing input (e.g., drag and drop) to displace the reply affordance into the appropriate input field. In other examples, a user can include a reply affordance in a note, a calendar entry, a word processing document, a reminder entry, etc. in a similar manner.
[0253] In some examples, a user input corresponding to a displacement of the response affordance from a first portion of the display 801 (displaying an input field) to a fourth portion of the display 801 corresponds to a selection of an affordance. In some examples, the affordance is either a share affordance (e.g., for sharing the response affordance during communication) or a save affordance (e.g., for saving the affordance in a note or reminder entry). For example, when the device 800 displays the DA user interface 803 over a user interface that includes input fields, the response affordance includes either a share affordance or a save affordance, depending on the type of user interface. For example, if the user interface corresponds to a communication application (e.g., messaging or email), the response affordance includes a share affordance, and if the user interface corresponds to another type of application that has input fields (e.g., word processing, reminders, calendar, notes), the response affordance includes a save affordance. User input selecting the share or save affordance causes device 800 to replace the display of the response affordance in the first portion of display 801 with a display of the response affordance in the input field, similar to that described above. For example, when the response affordance is displayed in the input field, device 800 stops displaying indicator 804.
[0254] In some examples, a user interface (e.g., the user interface on which DA user interface 803 is displayed) includes a widget area that occupies a fifth portion (e.g., a "widget portion") of display 801. In the example of FIG. 8AG, device 800 is a tablet device. Device 800 displays user interface 831 on display 801, including widget area 832 that occupies a fifth portion of display 801. Device 800 also displays DA user interface 803 on user interface 831. DA user interface 803 is displayed in response to the natural language input "track flight 23." DA user interface 803 includes indicator 804 that is displayed on a first portion of display 801 and responsive affordance 833 (including information about flight 23) that is displayed on a second portion of display 801.
[0255] In some examples, device 800 receives user input corresponding to the displacement of a responsive affordance from a first portion of display 801 to a fifth portion of display 801. In some examples, in response to receiving the user input, device 800 replaces the display of the responsive affordance in the first portion of the display with the display of the responsive affordance in the widget region. For example, FIGS. 8AH-8AJ show device 800 receiving user input 834 that displaces responsive affordance 833 from the first portion of display 801 to widget region 832. User input 834 corresponds to a drag gesture from the first portion of display 801 to the fifth portion of display 801, terminating in a lift-off event on the display of widget region 832. In some examples, the displacement of responsive affordance 833 from the first portion of display 801 to the fifth portion of display 801 is performed in a manner similar to the displacement of responsive affordance 827 described above. For example, indicator 804 ceases to be displayed while responsive affordance 833 is displaced.
[0256] 8AJ shows that responsive affordance 833 is currently displayed within widget region 832 with displayed calendar and music widgets. In this manner, the user can provide input (e.g., drag and drop) that displaces responsive affordance 833 into widget region 832, adding responsive affordance 833 as a widget.
[0257] In some examples, user input corresponding to a displacement of the responsive affordance from the first portion of display 801 to the fifth portion of display 801 corresponds to a selection of the affordance. In some examples, the affordance is a "display in widget" affordance. For example, when device 800 displays DA user interface 803 over a user interface that includes a widget region, the responsive affordance includes the "display in widget" affordance. User input selecting the "display in widget" affordance causes device 800 to replace the display of the responsive affordance on the first portion of display 801 with the display of the responsive affordance in the widget region in a manner similar to that described above.
[0258] In some examples, the responsive affordance corresponds to an event, and device 800 determines completion of the event. In some examples, in response to determining completion of the event, device 800 stops displaying the responsive affordance in the widget region (e.g., for a predetermined period of time after determining completion). For example, responsive affordance 833 corresponds to a flight, and in response to determining that the flight has completed (e.g., landed), device 800 stops displaying responsive affordance 833 in widget region 832. As another example, the responsive affordance corresponds to a sports game, and in response to determining that the sports game has ended, device 800 stops displaying the responsive affordance in the widget region.
[0259] 8AK-8AN illustrate various exemplary types of response affordances. In particular, FIG. 8AK illustrates a compact response affordance 835 displayed in response to the natural language request "How old is Celebrity X?" The compact response affordance 835 includes a direct answer to the request (e.g., "30 years old") without further information (e.g., additional information about Celebrity X). In some examples, all compact response affordances have the same maximum size, allowing the compact response affordances to occupy only a (relatively small) area of the display 801. FIG. 8AL illustrates a detailed response affordance 836 displayed in response to the natural language request "What are Team #1's stats?" The detailed response affordance 836 includes detailed information about Team #1 (e.g., various statistics) and has a larger display size than the compact response affordance 835. FIG. 8AM illustrates a list response affordance 837 displayed in response to the natural language request "Show me a list of nearby restaurants." List response affordance 837 includes a list of options (e.g., restaurants) and has a larger display size than compact response affordance 835. FIG. 8AN shows disambiguation response affordance 838 displayed in response to the natural language request "Call Neal." The disambiguation response affordance includes selectable disambiguation options: (1) Neal Ellis, (2) Neal Smith, and (3) Neal Johnson. Device 800 further provides a speech output asking, "Which Neal?"
[0260] As FIGS. 8AK-8AN show, the type of response affordance displayed (e.g., compact, detailed, list, disambiguated) depends on the content of the natural language input and / or the DA's interpretation of the natural language input. In some examples, affordance authoring rules specify a particular type of response affordance to display a particular type of natural language input. In some examples, the authoring rules specify that device 800 attempts to display a compact response affordance by default, for example, so that device 800 displays a compact response affordance in response to natural language input that can be fully answered by the compact response affordance. In some examples, if a response affordance can be displayed in different states (e.g., a first compact state and a second expanded (detailed) state), the authoring rules specify to initially display the response affordance as a compact affordance. As discussed with respect to FIGS. 8G-8H, a detailed version of the compact affordance may be available for display in response to receiving appropriate user input. It will be appreciated that some natural language inputs (e.g., "Give me the statistics on Team #1," "Show me a list of restaurants near me") cannot be adequately answered with compact affordances (or it may not be desirable to answer the input). Thus, authoring rules can specify particular kinds of affordances to display (e.g., in detail, a list) for such inputs.
[0261] In some examples, the DA determines a plurality of results corresponding to the received natural language input. In some examples, the device 800 displays a response affordance that includes a single result from the plurality of results. In some examples, while displaying the response affordance, other results from the plurality of results are not displayed. For example, consider the natural language input "nearest coffee." The DA determines a plurality of results (multiple nearby coffee shops) corresponding to the input. FIG. 8AO shows a response affordance 839 (e.g., a compact affordance) that is displayed in response to the input. The response affordance 839 includes a single result from the plurality of results (the coffee shop closest to the location of the device 800). The device 800 further provides a spoken output of "This is the nearest coffee shop." In this way, to accommodate natural language requests that imply multiple results, the DA can initially provide a single result, e.g., the most relevant result.
[0262] In some examples, after providing a single result (e.g., displaying response affordance 839), the DA provides the next result of the multiple results. For example, in FIG. 8AP, device 800 replaces response affordance 839 with response affordance 840 including the second-closest coffee shop. Device 800 further provides the speech output "This is the second-closest coffee shop." In some examples, device 800 transitions from FIG. 8AO to FIG. 8AP in response to receiving user input rejecting the single result (e.g., "I don't want that") or user input indicating that the next result should be provided. In some examples, device 800 transitions from FIG. 8AO to FIG. 8AP for a predetermined period after displaying affordance 839 and / or providing the speech output "This is the closest coffee shop," e.g., if user input selecting affordance 839 is not received. In this manner, device 800 can sequentially provide results of a natural language input implying multiple results.
[0263] In some examples, response affordances include one or more task affordances. User input (e.g., a tap gesture) that selects a task affordance causes device 800 to perform the corresponding task. For example, in FIG. 8AN , response affordance 838 includes task affordances 841, 842, and 843, where user selection of task affordance 841 causes device 800 to initiate a call to Neal Ellis, user selection of task affordance 842 causes device 800 to initiate a call to Neal Smith, and so on. As another example, response affordance 839 includes task affordance 844, and response affordance 840 includes task affordance 845. User selection of task affordance 844 causes device 800 to launch a maps application that displays directions to the nearest coffee shop, while user selection of task affordance 845 causes device 800 to launch a maps application that displays directions to the second-closest coffee shop.
[0264] In some examples, device 800 simultaneously displays multiple response affordances responsive to the natural language input. In some examples, each of the multiple response affordances corresponds to a different possible domain for the natural language input. In some examples, device 800 displays multiple response affordances when the natural language input is determined to be ambiguous, e.g., corresponds to multiple domains.
[0265] For example, consider the natural language input "Beyoncé." FIG. 8AQ shows response affordances 846, 847, and 848 displayed simultaneously in response to the natural language input. Response affordances 846, 847, and 848 correspond to the news domain (e.g., a user has requested news about Beyoncé), the music domain (e.g., a user has requested to play music by Beyoncé), and the knowledge domain (e.g., a user has requested information about Beyoncé), respectively. In some embodiments, each user input corresponding to a selection of response affordances 846, 847, and 848 causes device 800 to perform a corresponding action. For example, selection of response affordance 846 causes the display of a detailed response affordance including news about Beyoncé, selection of response affordance 847 causes device 800 to launch a music application including songs by Beyoncé, and selection of response affordance 848 causes the display of a detailed response affordance including information about Beyoncé.
[0266] In some implementations, the response affordance includes an editable text field, an editable text field containing text determined from the natural language input. For example, FIG. 8AR shows response affordance 849 displayed in response to the natural language utterance input "Text, Mom, I'm home." Response affordance 849 includes an editable text field 850 containing the text "I'm hole," for example, because the DA incorrectly recognized "I'm home" as "I'm hole." The response affordance further includes task affordance 851. User input selecting task affordance 851 causes device 800 to send a text message.
[0267] In some examples, device 800 receives user input corresponding to the selection of an editable text field and, in response, displays a keyboard while displaying a response affordance. For example, FIG. 8AS shows device 800 receiving user input 852 (e.g., a tap gesture) that selects editable text field 850. FIG. 8AT shows device 800 displaying keyboard 853 while displaying response affordance 849 in response to receiving user input 852. As shown, device 800 displays keyboard 853 on user interface 802 (e.g., user interface DA user interface 803 is displayed on top). While FIGS. 8AT-8AV show that portions of user interface 802 are not visually obscured while the response affordance and keyboard are displayed on user interface 802, in other examples, at least a portion of user interface 802 is visually obscured (e.g., in portions of display 801 that do not display the keyboard or response affordance 849).
[0268] In some implementations, device 800 receives one or more keyboard inputs and, in response, updates the text in an editable text field according to the one or more keyboard inputs. For example, FIG. 8AU shows device 800 receiving a keyboard input correcting "hole" to "home." Device 800 displays the corrected text in editable text field 850 of response affordance 849.
[0269] In other examples, device 800 receives a spoken input requesting that the text displayed in the editable text field be edited. In response to receiving the spoken input, device 800 updates the text in the editable text field according to the spoken input. For example, in FIG. 8AR, the user provides the spoken input "No, I said I'm home," causing device 800 to update the text in editable text field 850 accordingly.
[0270] In some examples, after updating the text in the editable text field, device 800 receives user input requesting that it perform a task associated with the affordance. In response to receiving the user input, device 800 performs the requested task based on the updated text. For example, FIG. 8AV shows device 800 receiving user input 854 (e.g., a tap gesture) corresponding to selection of task affordance 851 after editing "hole" to "home." FIG. 8AW shows device 800 sending the message "I'm back" to the user's mother in response to receiving user input 854. Device 800 further displays glyph 855 indicating completion of the task. FIG. 8AW further shows that in response to receiving user input 854, device 800 stops displaying keyboard 853 and displays (e.g., reveals) a portion of user interface 802, and device 800 displays indicator 804.
[0271] In this way, the user can edit the text included in the response affordance (e.g., if the DA misrecognizes the user's spoken input) and have the DA perform an action using the corrected text. While Figures 8AR-8AW show an example of editing and sending a text message, in other embodiments, the user can edit and save (or send) a note, a calendar entry, a reminder entry, an email entry, etc. in a similar manner.
[0272] In some examples, device 800 receives user input to close the DA. In some examples, closing the DA includes ceasing to display the DA user interface 803. Closing a DA is discussed in more detail with respect to FIGS. 10A-10V below. In some examples, after closing the DA, device 800 receives user input (e.g., user input that meets criteria for starting the DA) to resume the DA. In some examples, pursuant to receiving user input to resume the DA, device 800 displays a DA user interface including the same response affordances, e.g., the response affordances that were displayed before the DA was closed.
[0273] In some implementations, device 800 displays the same response affordance pursuant to determining that the same response affordance corresponds to a response to receiving the natural language input (e.g., an input intended to resume the DA). For example, FIG. 8AX shows DA user interface 803 including response affordance 856. Device 800 displays DA user interface 803 in response to the natural language input "What's the weather?". FIG. 8AY shows device 800 receiving user input 857 to close the DA, e.g., a tap gesture corresponding to selection of user interface 802. FIG. 8AZ shows device 800 closing the DA, e.g., ceasing to display DA user interface 803, in response to receiving user input 857. FIG. 8BA shows device 800 receiving input to resume the DA and now receiving the natural language input "Is it windy?". FIG. 8BB shows device 800 displaying DA user interface 803 including the same response affordance 856 and providing the speech output "Yes, it's windy." For example, the DA has determined that the same response affordance 856 corresponds to the natural language inputs "What's the weather?" and "Is it windy?" In this way, if the previous response affordance is relevant to the current natural language request, the previous response affordance can be included in a subsequently launched DA user interface.
[0274] In some examples, device 800 displays the same response affordance pursuant to a determination that user input to resume the DA is received within a predetermined period of closing the DA. For example, FIG. 8BC shows DA user interface 803 displayed in response to the natural language input "What is 3 times 5?" DA user interface 803 includes response affordance 858. FIG. 8BD shows that the DA was closed a first time. FIG. 8BE shows that device 800 received user input to resume the DA within a predetermined period (e.g., 5 seconds) of the first time. For example, device 800 received any one of the above types of input that meets the criteria for starting the DA, but has not received another natural language input including a different request to the DA. Thus, in FIG. 8BE, device 800 displays DA user interface 803 including the same response affordance 858 and indicator 804 that is in a listening state. In this way, the previous response affordance can be included in a subsequently started DA user interface if, for example, the user quickly resumes the DA because the user previously closed the DA by mistake.
[0275] FIG. 8BF shows device 800 in landscape orientation. In some embodiments, because device 800 is in landscape orientation, device 800 displays a user interface in landscape mode. For example, FIG. 8BF shows messaging application user interface 859 displayed in landscape mode. In some embodiments, device 800 displays DA user interface 803 in landscape mode via the landscape mode user interface. For example, FIG. 8BG shows DA user interface 803 in landscape mode displayed on user interface 859. It will be understood that a user can provide one or more inputs to interact with DA user interface 803 in landscape mode in a manner consistent with the techniques discussed herein.
[0276] In some implementations, some user interfaces do not have a landscape mode. For example, the display of a user interface is the same regardless of whether device 800 is in landscape or portrait orientation. Examples of user interfaces without a landscape mode include a home screen user interface and a lock screen user interface. Figure 8BH shows a home screen user interface 860 (without landscape mode) that is displayed when device 800 is in landscape orientation.
[0277] In some embodiments, when device 800 is in landscape orientation, device 800 displays DA user interface 803 over a user interface without a landscape mode. In some embodiments, when displaying DA user interface 803 (in landscape mode) over a user interface without a landscape mode, device 800 visually obscures the user interface, e.g., visually obscures portions of the user interface where DA user interface 803 is not displayed. For example, FIG. 8BI shows device 800 in landscape orientation and displays DA user interface 803 in landscape mode over home screen user interface 860. Home screen user interface 860 is displayed in portrait mode (even though device 800 is in landscape orientation) because home screen user interface 860 does not have a landscape mode. As shown, device 800 visually obscures home screen user interface 860. In this way, device 800 avoids simultaneously displaying landscape mode DA user interface 803 and a visually unobscured portrait mode user interface (e.g., home screen user interface 860), which may disrupt the user's visual experience.
[0278] In some embodiments, when the device 800 displays the DA user interface 803 over a predetermined type of user interface, the device 800 visually hides the predetermined type of user interface. An exemplary predetermined type of user interface includes a lock screen user interface. FIG. 8BJ shows the device 800 displaying an exemplary lock screen user interface 861. FIG. 8BK shows the device 800 displaying the DA user interface 803 over the lock screen user interface 861. As shown, the device 800 visually obscures the lock screen user interface 861 with the portion of the lock screen user interface 861 where the DA user interface 803 is not displayed.
[0279] In some examples, the DA user interface 803 includes a dialog affordance. In some examples, the dialog affordance includes a dialog generated by the DA in response to received natural language input. In some examples, the dialog affordance is displayed in a sixth portion (e.g., a “conversation portion”) of the display 801, the sixth portion of the display 801 being between a first portion of the display 801 (displaying the DA indicator 804) and a second portion of the display 801 (displaying the response affordance). For example, FIG. 8BL shows a dialog affordance 862 including a dialog generated by the DA in response to natural language input, as described further below. FIG. 8BM shows a dialog affordance 863 including a dialog generated by the DA in response to the natural language input “delete meeting #1,” as described further below. FIG. 8BM further shows that the device 800 displays the dialog affordance 863 in the sixth portion of the display 801, the sixth portion being between the display of the indicator 804 and the display of the response affordance 864.
[0280] In some examples, the DA determines multiple selectable disambiguation options for the received natural language input. In some examples, a dialog of a dialog affordance includes multiple selectable disambiguation options. The multiple disambiguation options are determined in some examples according to the DA determining that the natural language input is ambiguous. An ambiguous natural language input corresponds to multiple possible actionable intents, e.g., each having a relatively high (and / or equal) confidence score. For example, in FIG. 8BL, consider the natural language input "Play Frozen." The DA determines two selectable disambiguation options: option 865 "Play video" (e.g., the user intended to play the video "Frozen") and option 866 "Play music" (e.g., the user intended to play music from the video "Frozen"). Dialog affordance 862 includes options 865 and 866, where user selection of option 865 causes device 800 to play the video "Frozen," and user selection of option 866 causes device 800 to play music from the video "Frozen." As another example, consider the natural language input "Delete Meeting #1" in FIG. 8BM, where "Meeting #1" is a recurring meeting. The DA determines two selectable disambiguation options: option 867 "Delete Single" (e.g., the user intended to delete a single instance of Meeting #1) and option 868 "Delete All" (e.g., the user intended to delete all instances of Meeting #1). Dialog affordance 863 includes options 867 and 868, along with a cancel option 869.
[0281] In some embodiments, the DA determines that additional information is needed to perform a task based on the received natural language input. In some embodiments, the dialog of the dialog affordance includes one or more selectable options suggested by the DA for the needed additional information. For example, the DA may determine a domain for the received natural language input, but is unable to determine the parameters needed to complete a task associated with the domain. Consider, for example, the natural language input "phone call." The DA determines a phone call domain (e.g., a domain associated with an actionable intent for a phone call) for the natural language input, but is unable to determine the parameters of the call. Thus, in some embodiments, the DA determines one or more selectable options as parameter suggestions. For example, device 800 displays selectable options in the dialog affordance that correspond to the user's most frequently called contacts. User selection of any one of the selectable options causes device 800 to call the respective contact.
[0282] In some examples, the DA determines a primary user intent based on the received natural language input and alternative user intents based on the received natural language input. In some examples, the primary intent is the highest-ranked actionable intent, and the alternative user intent is the second-highest-ranked actionable intent. In some examples, the displayed response affordance corresponds to the primary user intent, and the simultaneously displayed dialog of the dialog affordance includes selectable options corresponding to the alternative user intents. For example, FIG. 8BN shows a DA user interface 803 displayed in response to the natural language input “Phil’s directions,” in which the DA determines a primary user intent where the user intends to get directions to “Phil’s Coffee” and an alternative user intent where the user intends to get directions to the home of a contact named “Phil.” The DA user interface 803 includes a response affordance 870 corresponding to the primary user intent and a dialog affordance 871. A dialog 872 of the dialog affordance 871 corresponds to the secondary user intent. The dialog-based user input selection 872 causes the device 800 to get directions to the home of a contact named "Phil," while the user input selection response affordance 870 causes the device 800 to get directions to "Phil's Coffee."
[0283] In some examples, the dialog affordance is displayed in a first state. In some examples, the first state is an initial state, e.g., describing how the dialog affordance is initially displayed before receiving user input to interact with the dialog affordance. FIG. 8BO shows a DA user interface 803 including a dialog affordance 873 displayed in an initial state. The device 800 displays the DA user interface 803 in response to the natural language input "What's the weather?" The dialog affordance 873 includes at least a portion of a dialog generated by the DA in response to the input, e.g., "It's currently 70 degrees and windy...". Further discussion regarding whether to display a dialog generated by the DA is discussed below with respect to FIGS. 11-16.
[0284] In some examples, device 800 receives user input corresponding to selection of a dialog affordance displayed in a first state. In response to receiving the user input, device 800 replaces the display of the dialog affordance in the first state with a display of the dialog affordance in a second state. In some examples, the second state is an expanded state, where the display size of the dialog affordance in the expanded state is larger than the display size of the dialog affordance in the initial state, and / or the dialog affordance in the expanded state displays more content than the dialog affordance in the initial state. Figure 8BP shows device 800 receiving user input 874 (e.g., a drag gesture) corresponding to selection of dialog affordance 873 displayed in the initial state. Figure 8BQ shows device 800 replacing the display of dialog affordance 873 displayed in the initial state with a display of dialog affordance 873 in the expanded state in response to receiving user input 874 (or a portion thereof). As shown, dialog affordance 873 in Figure 8BQ has a larger display size and includes more text than the dialog affordance in Figure 8BP.
[0285] In some examples, the display size of the dialog affordance (in the second state) is proportional to the length of the user input that causes the dialog affordance to be displayed in the second state. For example, in FIGS. 8BP-8BQ, the display size of dialog affordance 873 increases proportionally to the length (e.g., physical distance) of drag gesture 874. In this manner, the user can provide successive drag gestures to extend response affordance 873 according to the drag length of the drag gesture. Furthermore, FIGS. 8BO-8BQ show that device 800 may first display dialog affordance 873 as shown in FIG. 8BO and then extend the dialog affordance in FIG. 8BQ, while in other examples, device 800 may first display dialog affordance 873 as shown in FIG. 8BQ. Thus, in some examples, device 800 initially displays the dialog affordance so that it displays the maximum amount of content, e.g., without obscuring (covering) simultaneously displayed response affordances.
[0286] In some examples, the display of the dialog affordance obscures the display of a simultaneously displayed response affordance. Specifically, in some examples, the display of the dialog affordance in the second (e.g., expanded) state occupies at least a portion of a second portion of display 801 (displaying the response affordance). In some examples, displaying the dialog affordance in the second state further includes displaying the dialog affordance over at least a portion of the response affordance. For example, FIG. 8BQ shows that drag gesture 874 continues. FIG. 8BR shows that in response to receiving the continued drag gesture 874, device 800 extends the display of dialog affordance 873 over the display of response affordance 875.
[0287] In some examples, before receiving user input, the dialog affordance is displayed in a second state (e.g., to expand), and the response affordance is displayed in its original state. In some examples, the original state represents the state of the response affordance before the dialog affordance (or a portion thereof) is displayed over the response affordance. For example, FIGS. 8BO-8BQ show response affordance 875 displayed in its original state. In some examples, displaying the dialog affordance in a second (e.g., expanded) state over at least a portion of the response affordance includes replacing the display of the response affordance in its original state with a display of the response affordance in a covered state. FIG. 8BR shows response affordance 875 displayed in a covered state. In some examples, when displayed in the covered state, the response affordance is reduced in display size (e.g., relative to its original state) and / or dimmed (e.g., displayed less prominently than in its original state). In some examples, the degree to which the response affordance shrinks and / or dims is proportional to the amount of the dialog affordance displayed over the response affordance.
[0288] In some examples, the dialog affordance has a maximum display size, and a second (e.g., expanded) state of the dialog affordance corresponds to the maximum display size. In some examples, a dialog affordance displayed at the maximum display size cannot be further expanded in response to a user input, such as a drag gesture. In some examples, a dialog affordance displayed at the maximum display size displays the entire content of the dialog affordance. In other examples, a dialog affordance displayed at the maximum display size does not display the entire content of the dialog affordance. Thus, in some examples, while the device 800 displays the dialog affordance in the second state (having the maximum display size), the device 800 allows user input (e.g., a drag gesture / swipe gesture) to scroll the content of the dialog affordance. FIG. 8BS shows a dialog affordance 873 displayed at the maximum display size. Notably, in FIG. 8BR, the drag gesture 874 continues. In response to receiving the continued drag gesture 874, device 800 displays (e.g., expands) dialog affordance 873 to its maximum display size of FIG. 8BS. Dialog affordance 873 includes scroll indicator 876, indicating that the user can provide input to scroll the content of dialog affordance 873.
[0289] In some examples, a portion of the response affordance remains visible when the dialog affordance is displayed in the second state (and at its maximum size). Thus, in some examples, device 800 constrains the maximum size of the dialog affordance displayed above the response affordance so that the dialog affordance does not completely cover the response affordance. In some examples, the portion of the response affordance that is visible is the second portion of the response affordance described above with respect to FIG. 8M. For example, the portion is the top portion of the response affordance that includes a glyph indicating the category of the response affordance and / or associated text. FIG. 8BS shows that the top portion of response affordance 875 remains visible when device 800 displays dialog affordance 873 at its maximum size above response affordance 875.
[0290] In some examples, device 800 receives user input corresponding to a selection of a portion of the response affordance that remains visible (when the dialog affordance is displayed in the second state over the response affordance). In response to receiving the user input, the device displays the response affordance on a first portion of display 801, e.g., displaying the response affordance in its original state. In response to receiving the user input, device 800 further replaces the display of the dialog affordance in the second (e.g., expanded) state with a display of the dialog affordance in a third state. In some examples, the third state is a state in which the dialog affordance in the third state has a smaller display size (than the dialog affordance in the initial or expanded state) and / or the dialog affordance includes a smaller amount of content (than the dialog affordance in the initial or expanded state). In other examples, the third state is the first state (e.g., the initial state). FIG. 8BT shows device 800 receiving user input 877 (e.g., a tap gesture) that selects an upper portion of response affordance 875. 8BU shows that in response to receiving user input 877, device 800 replaces the display of dialog affordance 873 in the expanded state (FIG. 8BT) with a display of dialog affordance 873 in the collapsed state. Device 800 further displays response affordance 875 in its original state.
[0291] In some examples, device 800 receives user input corresponding to a selection of the dialog affordance displayed in the third state. In response to receiving the user input, device 800 replaces the display of the response affordance in the third state with the display of the dialog affordance in the first state. For example, in FIG. 8BU, a user can provide input (e.g., a tap gesture) that selects dialog affordance 873 displayed in the collapsed state. In response to receiving the input, device 800 displays the dialog affordance in the initial state, e.g., returning to the display of FIG. 8BO.
[0292] In some examples, while a dialog affordance is displayed in a first or second state (e.g., an initial or expanded state), device 800 receives user input corresponding to a selection of a simultaneously displayed response affordance. In response to receiving the user input, device 800 replaces the display of the dialog affordance in the first or second state with a display of the dialog affordance in a third (e.g., collapsed) state. For example, FIG. 8BV shows a DA user interface 803 displayed in response to the natural language input "Show me the roster for team #1." DA user interface 803 includes a detail response affordance 878 and a dialog affordance 879 displayed in the initial state. FIG. 8BV further shows device 800 receiving user input 880 (e.g., a drag gesture) that selects response affordance 878. FIG. 8BW shows device 800 replacing the display of dialog affordance 879 with a display of dialog affordance 879 in the collapsed state in response to receiving user input 880.
[0293] In some examples, while displaying a dialog affordance in a first or second state (e.g., an initial or expanded state), device 800 receives user input corresponding to a selection of the dialog affordance. In response to receiving the user input, device 800 replaces the display of the dialog affordance in the first or second state with a display of the dialog affordance in a third (e.g., collapsed) state. For example, FIG. 8BX shows DA user interface 803 displayed in response to the natural language input "What music do you have for me?" DA user interface 803 includes response affordance 881 and dialog affordance 882 displayed in the initial state. FIG. 8BX further shows device 800 receiving user input 883 (e.g., a downward drag or swipe gesture) selecting dialog affordance 882. Figure 8BY shows that in response to receiving user input 883, device 800 replaces the display of dialog affordance 882 in its initial state with a display of dialog affordance 882 in its collapsed state. While Figures 8BX-8BY show that the user input corresponding to the selection of a dialog affordance is a drag or swipe gesture, in other examples, the user input is a selection of a displayed affordance included in the dialog affordance. For example, user input (e.g., a tap gesture) that selects the "collapse" affordance within a dialog affordance displayed in the first or second state causes device 800 to replace the display of the dialog affordance in the first or second state with a display of the dialog affordance in a third state.
[0294] In some implementations, device 800 displays a transcription of the received natural language speech input in a dialog affordance. This transcription is obtained by performing automatic speech recognition (ASR) on the natural language speech input. FIG. 8BZ shows a DA user interface 803 displayed in response to the natural language speech input "What's the weather like?" The DA user interface includes a response affordance 884 and a dialog affordance 885. Dialog affordance 885 includes a transcription 886 of the speech input and a dialog 887 generated by the DA in response to the speech input.
[0295] In some examples, device 800 does not display a transcription of a received natural language speech input by default. In some examples, device 800 includes a setting that, when activated, causes device 800 to always display a transcription of a received natural language speech input. Various other examples are now described in which device 800 may display a transcription of a received natural language speech input.
[0296] In some embodiments, the natural language speech input (with the displayed transcription) is consecutive to a second natural language speech input received before the natural language speech input. In some embodiments, displaying the transcription is performed pursuant to a determination that the DA was unable to determine a user intent for the natural language speech input and was unable to determine a second user intent for the second natural language speech input. Thus, in some embodiments, the device 800 displays a transcription for a natural language input if the DA was unable to determine a feasible intent for two consecutive natural language speech inputs.
[0297] For example, FIG. 8CA shows that device 800 receives the speech input "How far is Dish n' Dash?" and the DA is unable to determine the user's intent for the natural language input. For example, device 800 provides the speech output "I'm not sure, can you say that again?" unless I understand. Thus, the user repeats the speech input. For example, FIG. 8CB shows device 800 receiving the continuous speech input "How far is Dish n' Dash?". FIG. 8CC shows that the DA is unable to determine the user's intent for the continuous speech input. For example, device 800 provides the speech output "I'm not sure I understand." Thus, device 800 further displays dialog affordance 888, which includes transcription 889 of the continuous speech input, "How far is Rish and Rash?". In this example, transcription 889 reveals that the DA incorrectly recognized "How far is Dish n' Dash?" twice as "How far is Rish and Rash?" Since "Rish and Rash" may not be an actual location, the DA was unable to determine the user intent of both speech inputs.
[0298] In some examples, displaying a representation of the received natural language speech input is performed pursuant to a determination that the natural language speech input repeats a previous natural language speech input. For example, FIG. 8CD shows a DA user interface 803 displayed in response to the speech input (previous speech input) "Where's Starbucks?" The DA displays a response affordance 890 including "Star Mall" because it incorrectly recognized the speech input as "Where's Star Mall?" Because the DA incorrectly understood the speech input, the user repeats the speech input. For example, FIG. 8CE shows device 800 receiving a repetition (e.g., successive repetitions) of the previous speech input "Where's Starbucks?" The DA determines that the speech input repeats the previous speech input. FIG. 8CF shows device 800 displaying a dialog affordance 891 including a transcription 892 pursuant to such a determination. The transcription 892 reveals that the DA incorrectly recognized (e.g., twice) "Where's Starbucks?" as "Where's Star Mall?"
[0299] In some examples, after receiving the natural language speech input (e.g., a transcription is displayed), the device receives a second natural language speech input that follows the natural language speech input. In some examples, displaying the transcription is performed pursuant to a determination that the second natural language speech input indicates a speech recognition error. Thus, in some examples, the device 800 displays a transcription of the previous speech input if the DA indicates that it incorrectly recognized the previous speech input. For example, FIG. 8CG shows a DA user interface 803 displayed in response to the speech input “set timer for 15 minutes.” The DA incorrectly recognized “15 minutes” as “50 minutes.” Thus, the DA user interface 803 includes a response affordance 893 indicating that the timer is set to 50 minutes. Because the DA incorrectly recognized the speech input, the user provides a second speech input indicating a speech recognition error (e.g., “that's not what I said,” “you heard me wrong,” “that's wrong,” etc.). For example, FIG. 8CH shows device 800 receiving a second speech input, "That's not what I said." The DA determines that the second speech input indicates a speech recognition error. FIG. 8CI shows that, pursuant to such a determination, device 800 displays dialog affordance 894, which includes transcription 895. Transcription 895 reveals that the DA incorrectly recognized "15 minutes" as "50 minutes."
[0300] In some embodiments, device 800 receives user input corresponding to selection of the displayed transcription. In response to receiving the user input, device 800 simultaneously displays a keyboard and an editable text field including the transcription, e.g., displayed a keyboard and an editable text field on user interface 803. In some embodiments, device 800 further visually hides at least a portion of the user interface (e.g., a portion of display 801 that does not display the keyboard or editable text field). Continuing with the example of FIG. 8CI, FIG. 8CJ shows device 800 receiving user input 896 (e.g., a tap gesture) that selects transcription 895. FIG. 8CK shows that in response to receiving user input 896, device 800 displays a keyboard 897 and an editable text field 898 including transcription 895. FIG. 8CK further shows device 800 visually hiding a portion of user interface 802.
[0301] FIG. 8CL shows that device 800 receives one or more keyboard inputs and has an edited transcription 895, e.g., from "set timer for 50 minutes" to "set timer for 15 minutes," according to the one or more keyboard inputs. FIG. 8CL further shows device 800 receiving user input 899 (e.g., a tap gesture) corresponding to selection of done key 8001 on keyboard 897. FIG. 8CM shows that in response to receiving user input 899, the DA performs a task based on the current (e.g., edited) transcription 895. For example, device 800 displays DA user interface 803 including response affordance 8002 indicating that the timer is set for 15 minutes. Device 800 further provides the speech output "Ok, set timer for 15 minutes." In this manner, the user can manually correct the incorrect transcription (e.g., using keyboard input) to trigger performance of the correct task.
[0302] In some embodiments, while displaying the keyboard and editable text field, device 800 receives user input corresponding to selection of the visually obscured user interface. In some embodiments, in response to receiving the user input, device 800 ceases displaying the keyboard and editable text field. In some embodiments, device 800 additionally or alternatively ceases displaying DA user interface 803. For example, in FIGS. 8CK-8CL, user input (e.g., a tap gesture) that selects visually obscured user interface 802 may cause device 800 to revert to the display of FIG. 8CI, or, as shown in FIG. 8A, may cause device 800 to not display DA user interface 803 and display user interface 802 in its entirety.
[0303] In some examples, device 800 presents a digital assistant result (e.g., a response affordance and / or a speech output) at a first time. In some examples, in accordance with determining that the digital assistant result corresponds to a predetermined type of digital assistant result, device 800 automatically stops displaying DA user interface 803 for a predetermined period of time after the first time. Thus, in some examples, device 800 can quickly (e.g., within 5 seconds) close DA user interface 803 after providing a predetermined type of result. Exemplary predetermined type results correspond to completed tasks for which no further user input is required (or no further user interaction is desired). For example, such results include results that confirm a timer has been set, a message has been sent, or a household appliance (e.g., a light) has changed state. Examples of results that do not correspond to a predetermined type include results in which the DA asks for further user input and results for which the DA provides information (e.g., news, Wikipedia articles, locations) in response to a user's information request.
[0304] For example, Figure 8CM shows device 800 presenting a result at a first time and terminating by providing, e.g., the speech output "Ok, set a 15 minute timer." Because the result corresponds to a predetermined type, Figure 8CN shows device 800 automatically (e.g., without further user input) closing the DA for a predetermined period of time (e.g., 5 seconds) after the first time.
[0305] 8CO-8CT show examples of DA user interface 803 and exemplary user interfaces when device 800 is a tablet device. It will be understood that any of the techniques discussed herein with respect to device 800 being a tablet device are equally applicable when device 800 is another type of device (and vice versa).
[0306] FIG. 8CO shows device 800 displaying user interface 8003. User interface 8003 includes dock region 8004. In FIG. 8CO, device 800 displays DA user interface 803 on user interface 8003. DA user interface 803 includes indicator 804 displayed on a first portion of display 801 and responsive affordance 8005 displayed on a second portion of display 801. As shown, a portion of user interface 8003 remains visible (e.g., is not visually obscured) on a third portion of display 801. In some examples, the third portion is between the first portion of display 801 and the second portion of display 801. In some examples, as shown in FIG. 8CO, the display of DA user interface 803 does not visually obscure dock region 8004, e.g., a portion of DA user interface 803 is not displayed on dock region 8004.
[0307] 8CP shows device 800 displaying a DA user interface 803 that includes a dialog affordance 8006. As shown, dialog affordance 8006 is displayed in a portion of display 801 between a first portion of display 801 (displaying indicator 804) and a second portion of the display (displaying response affordance 8005). Displaying dialog affordance 8006 also causes response affordance 8005 to be displaced (from FIG. 8CO) toward the top of display 801.
[0308] 8CQ shows device 800 displaying user interface 8003 including media panel 8007 showing currently playing media. FIG. 8CR shows device 800 displaying DA user interface 803 in user interface 8003. DA user interface 803 includes response affordance 8008 and indicator 804. As shown, displaying DA user interface 803 does not visually obscure media panel 8007. For example, as shown, displaying elements of DA user interface 803 (e.g., indicator 804, response affordance 8008, dialog affordance) displaces media panel 8007 toward the top of display 801.
[0309] For example, Figure 8CS shows device 800 displaying user interface 8009 that includes keyboard 8010. Figure 8CT shows DA user interface 803 displayed on user interface 8009. Figure 8CT shows that, in some embodiments, displaying DA user interface 803 on user interface 8009 that includes keyboard 8010 causes device 800 to visually hide (e.g., gray out) the keys of keyboard 8010.
[0310] 9A-9C illustrate multiple devices that determine which device should respond to speech input, according to various embodiments. In particular, FIG. 9A illustrates devices 900, 902, and 904. Devices 900, 902, and 904 are each implemented as device 104, device 122, device 200, or device 600. In some embodiments, devices 900, 902, and 904 each at least partially implement DA system 700.
[0311] 9A , the displays of each of devices 900, 902, and 904 are silent when a user provides speech input that includes a trigger phrase (e.g., “Hey Siri”) to initiate DA, e.g., “Hey Siri, what's the weather?” In some examples, the displays of each of at least one of devices 900, 902, and 904 display a user interface (e.g., a home screen user interface, an application-specific user interface) when the user provides speech input. FIG. 9B shows that in response to receiving speech input that includes the trigger phrase, devices 900, 902, and 904 each display indicator 804. In some examples, each indicator 804 is displayed in a listening state, e.g., to indicate that the respective device is sampling audio input.
[0312] Devices 900, 902, and 904 in FIG. 9B cooperate among themselves (or via a fourth device) to determine which device should respond to a user request. Exemplary techniques for coordinating devices to determine which device should respond to a user request are described in U.S. Patent No. 10,089,072, entitled "INTELLIGENT DEVICE ARBITRATION AND CONTROL," filed October 2, 2018, and U.S. Patent Application No. 63 / 022,942, entitled "DIGITAL ASSISTANT HARDWARE ABSTRACTION," filed May 11, 2020, the contents of which are incorporated herein by reference in their entireties. As shown in FIG. 9B, each device determines whether to respond to the user request, and each device displays only indicator 804. For example, the respective portions of the displays of devices 900, 902, and 904 that do not display indicator 804 are not displayed. In some embodiments, when at least one of devices 900, 902, and 904 displays a user interface (previous user interface) when a user provides speech input, the at least one device determines whether to respond to the user request, and the at least one device displays only indicator 804 on the previous user interface.
[0313] 9C illustrates that device 902 is determined to be the device that will respond to the user request. As shown, in response to a determination that another device (e.g., device 902) should respond to the user request, the displays of devices 900 and 904 cease displaying (or cease displaying indicator 804 and fully display the previous user interface). As further illustrated, in response to a determination that device 902 should respond to the user request, device 902 displays a user interface (e.g., a lock screen user interface) 906 and a DA user interface 803 on user interface 906. DA user interface 803 includes a response to the user request. In this manner, visual clutter is minimized when determining which of multiple devices should respond to the speech input. For example, in FIG. 9B, the display of the device determined not to respond to the user request displays only indicator 804, as opposed to, for example, displaying a user interface on the entire display.
[0314] Determining which device of a plurality of devices should respond to the speech input in the above manner provides feedback to the user that the speech input is being received and processed. Furthermore, providing feedback in such a manner can advantageously reduce unnecessary visual or audible clutter when responding to the speech input. For example, the user is not required to turn off display and / or audible output on non-selected devices, and visual clutter on the user interfaces of non-selected devices is minimized (e.g., if the user was previously interacting with the user interface of the non-selected device). Providing improved visual feedback to the user enhances device usability (e.g., by helping the user avoid providing unnecessary input) and makes the user-device interface more efficient, which in turn reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.
[0315] 10A-10V show user interfaces and digital assistant user interfaces according to various embodiments. 10A-10V are used to illustrate processes described below, including the processes of FIGS. 18A-18B.
[0316] Figure 10A shows a device 800. The device 800 displays a DA user interface 803 on a user interface on a display 801. In Figure 10A, the device 800 displays the DA user interface 803 on a home screen user interface 1001. In other examples, the user interface is another type of user interface, such as a lock screen user interface or an application-specific user interface.
[0317] In some examples, DA user interface 803 includes indicator 804 displayed in a first portion (e.g., "indicator portion") of display 801 and response affordances displayed in a second portion (e.g., "response portion") of display 801. A third portion (e.g., "UI portion") of display 801 displays a portion of the user interface (with DA user interface 803 displayed on top). For example, in FIG. 10A , the first portion of display 801 displays indicator 804, the second portion of display 801 displays response affordance 1002, and the third portion of display 801 displays a portion of home screen user interface 1001.
[0318] In some embodiments, while displaying the DA user interface 803 on the user interface, the device 800 receives user input corresponding to a selection of a third portion of the display 801. The device 800 determines whether the user input corresponds to a first type of input or a second type of input. In some embodiments, the first type of user input includes a tap gesture and the second type of user input includes a drag or swipe gesture.
[0319] In some examples, in accordance with determining that the user input corresponds to the first type of input, device 800 stops displaying DA user interface 803. Stopping displaying DA user interface 803 includes stopping displaying any portion of DA user interface 803, such as indicator 804, response affordances, and dialog affordances (if included). In some examples, stopping displaying DA user interface 803 includes replacing the display of elements of DA user interface 803 in their respective portions of display 801 with a display of the user interface in the respective portions. For example, device 800 replaces the display of indicator 804 with a display of a first portion of the user interface in a first portion of display 801 and replaces the display of the response affordances with a display of a second portion of the user interface in a second portion of display 801.
[0320] 10B shows device 800 receiving user input 1003 (e.g., a tap gesture) corresponding to a selection of a third portion of display 801. Device 800 determines that user input 1003 corresponds to a first type of input. FIG. 10C shows that, following such a determination, device 800 ceases displaying DA user interface 803 and displays user interface 1001 in its entirety.
[0321] In this manner, a user can dismiss DA user interface 803 by providing input that selects a portion of display 801 that does not display any portion of DA user interface 803. For example, in Figures 8S-8X above, a tap gesture that selects a portion of display 801 that displays visually hidden home screen user interface 802 causes device 800 to revert to the display of Figure 8A.
[0322] In some examples, the user input corresponds to a selection of a selectable element displayed in a third portion of display 801. In some examples, in accordance with a determination that the user input corresponds to the first type of input, device 800 displays a user interface corresponding to the selectable element. For example, device 800 replaces the display of indicator 804 with a display of the portion of the user interface (displayed in the third portion of display 801), a display of the responsive affordance, and a display of the user interface corresponding to the selectable element.
[0323] In some examples, the user interface is a home screen user interface 1001, the selectable element is an application affordance of the home screen user interface 1001, and the user interface corresponding to the selectable element is a user interface corresponding to the application affordance. For example, FIG. 10D shows a DA user interface 803 displayed on the home screen user interface 1001. The display 801 displays an indicator 804 in a first portion, a response affordance 1004 in a second portion, and a portion of the user interface 1001 in a third portion. FIG. 10E shows the device 800 receiving a user input 1005 (e.g., a tap gesture) that selects a health application affordance 1006 displayed in the third portion. FIG. 10F shows the device 800 ceasing to display the indicator 804, the response affordance 1004, and the portion of the user interface 1001 in accordance with determining that the user input 1005 corresponds to a first type of input. The device 800 further displays a user interface 1007 corresponding to a health application.
[0324] In some examples, the selectable element is a link, and the user interface corresponding to the selectable element is a user interface corresponding to the link. For example, FIG. 10G shows a DA user interface 803 displayed on a web browsing application user interface 1008. The display 801 displays an indicator 804 in a first portion, a response affordance 1009 in a second portion, and a portion of the user interface 1008 in a third portion. FIG. 10G further shows the device 800 receiving user input 1010 (e.g., a tap gesture) that selects a link 1011 (e.g., a web page link) displayed in the third portion. FIG. 10H shows the device 800 ceasing to display the indicator 804, the response affordance 1009, and the portion of the user interface 1008 in accordance with the device 800 determining that the user input 1010 corresponds to a first type of input. The device 800 further displays a user interface 1012 corresponding to the web page link 1011.
[0325] In this manner, a user input selecting a third portion of the display 801 can close the DA user interface 803 and trigger the performance of further actions (e.g., updating the display 801) according to the user selection.
[0326] In some embodiments, in accordance with determining that the user input corresponds to a second type of input (e.g., a drag or swipe gesture), device 800 updates the display of the user interface in the third portion of display 801 in accordance with the user input. In some embodiments, while device 800 updates the display of the user interface in the third portion of display 801, device 800 continues to display at least some of the elements of DA user interface 803 in the respective display portions. For example, device 800 displays (e.g., continues to display) a response affordance in the second portion of display 801. In some embodiments, device 800 further displays (e.g., continues to display) indicator 804 in the first portion of display 801. In some embodiments, updating the display of the user interface in the third portion includes scrolling content of the user interface.
[0327] For example, FIG. 10I shows DA user interface 803 displayed on a web browser application user interface 1013 displaying a web page. Display 801 displays indicator 804 in a first portion, responsive affordance 1014 in a second portion, and a portion of user interface 1013 in a third portion. FIG. 10I further shows device 800 receiving user input 1015 (e.g., a drag gesture) that selects the third portion. FIG. 10J shows device 800 determining that user input 1015 corresponds to a second type of input, and updating (e.g., scrolling) the content of user interface 1013 in accordance with user input 1015, e.g., scrolling the content of a web page. FIGS. 10I-10J show device 800 continuing to display indicator 804 in the first portion of display 801 and responsive affordance 1014 in the second portion of display 801 while updating user interface 1013 (in the third portion of display 801).
[0328] As another example, FIG. 10K shows DA user interface 803 displayed on home screen user interface 1001. Display 801 displays indicator 804 in a first portion, response affordance 1016 in a second portion, and a portion of user interface 1001 in a third portion. FIG. 10K further shows device 800 receiving user input 1017 (e.g., a swipe gesture) that selects the third portion. FIG. 10L shows device 800 determining that user input 1017 corresponds to a second type of input, and device 800 updating the content of user interface 1001 according to user input 1017. For example, as shown, device 800 updates user interface 1001 to display secondary home screen user interface 1018 that includes one or more application affordances that differ from those of home screen user interface 1001. 10K-10L show that while updating user interface 1001, device 800 continues to display indicator 804 on a first portion of display 801 and responsive affordance 1016 on a second portion of display 801.
[0329] In this way, the user can provide input that causes DA user interface 803 to update the user interface displayed without the input stalling DA user interface 803 .
[0330] In some embodiments, updating the display of the user interface in the third portion of display 801 is performed pursuant to a determination that the DA is in a listening state. Thus, device 800 may allow a drag or swipe gesture to update the user interface (DA user interface 803 displayed above) only when the DA is in a listening state. In such an example, if the DA is not in a listening state, in response to receiving user input corresponding to the second type (and corresponding to a selection of the third portion of display 801), device 800 does not update display 801 in response to the user input or cease displaying DA user interface 803. In some embodiments, while updating the display of the user interface while the DA is in a listening state, the display size of indicator 804 changes based on the magnitude of the received speech input, as described above.
[0331] In some embodiments, while device 800 is displaying DA user interface 803 on its user interface, device 800 receives a second user input. In some embodiments, in accordance with a determination that the second user input corresponds to a third type of input, device 800 stops displaying DA user interface 803. In some embodiments, the third type of input includes a swipe gesture occurring from the bottom of display 801 toward the top of display 801. The third type of input may be considered a "home swipe" upon receiving such input when device 800 displays a user interface different from the home screen user interface (and does not display DA user interface 803) and causes device 800 to return to displaying the home screen user interface.
[0332] Figure 10M shows device 800 displaying DA user interface 803 on home screen user interface 1001. DA user interface 803 includes response affordance 1020 and indicator 804. Figure 10M further shows device 800 receiving user input 1019, which is a swipe gesture from the bottom of display 801 toward the top of display 801. Figure 10N shows device 800 determining that user input 1019 corresponds to a third type of input, and device 800 ceasing to display response affordance 1020 and indicator 804.
[0333] In some embodiments, the user interface (on which DA user interface 803 is displayed) is an application-specific user interface. In some embodiments, device 800 displays DA user interface 803 on the application-specific user interface, while device 800 receives a second user input. In some embodiments, in accordance with a determination that the second user input corresponds to a third type of input, the device ceases displaying DA user interface 803 and further displays a home screen user interface. For example, FIG. 10O shows device 800 displaying DA user interface 803 on a health application user interface 1022. DA user interface 803 includes response affordance 1021 and indicator 804. FIG. 10O further shows device 800 receiving user input 1023, which is a swipe gesture from the bottom of display 801 toward the top of display 801. FIG. 10P shows device 800 determining that user input 1023 corresponds to a third type of input, and device 800 displays home screen user interface 1001. For example, as shown, device 800 replaces the display of indicator 804, response affordance 1021, and messaging application user interface 1022 with the display of home screen user interface 1001.
[0334] In some examples, while device 800 is displaying DA user interface 803 on a user interface, device 800 receives a third user input corresponding to a selection of a response affordance. In response to receiving the third user input, device 800 stops displaying DA user interface 803. For example, FIG. 10Q shows DA user interface 803 displayed on home screen user interface 1001. DA user interface 803 includes response affordance 1024, dialog affordance 1025, and indicator 804. FIG. 10Q further shows device 800 receiving user input 1026 (e.g., an upward swipe or drag gesture) that selects response affordance 1024. FIG. 10R shows device 800 stopping displaying DA user interface 803 in response to receiving user input 1026.
[0335] In some embodiments, while device 800 is displaying user interface 803 over a user interface, device 800 receives a fourth user input corresponding to a displacement of indicator 804 from a first portion of display 801 to an edge of display 801. In response to receiving the fourth user input, device 800 stops displaying DA user interface 803. For example, FIG. 10S shows DA user interface 803 displayed over home screen user interface 1001. In FIG. 10S, device 800 receives user input 1027 (e.g., a drag or swipe gesture) that displaces the indicator from a first portion of display 801 to an edge of display 801. FIGS. 10S-10V show that in response to receiving user input 1027 (e.g., in response to indicator 804 reaching an edge of display 801), device 800 stops displaying DA user interface 803. 5. Digital Assistant Response Mode
[0336] 11 illustrates a system 1100 for selecting a DA response mode and presenting a response according to the selected DA response mode, according to various examples. In some embodiments, system 1100 is implemented on a standalone computer system (e.g., device 104, 122, 200, 400, 600, 800, 900, 902, or 904). System 1100 is implemented using hardware, software, or a combination of hardware and software to carry out the principles discussed herein. In some embodiments, the modules and functionality of system 1100 are implemented in a DA system, as described above with respect to FIGS. 7A-7C.
[0337] System 1100 is exemplary, and thus system 1100 may have more or fewer components than shown, may combine two or more components, or may have a different configuration or arrangement of components. Furthermore, while the following discussion describes functions performed by a single component of system 1100, it should be understood that such functions may be performed by other components of system 1100, and that such functions may be performed by two or more components of system 1100.
[0338] 12 illustrates device 800 presenting a response to a received natural language input according to different DA response modes, according to various examples. In FIG. 12, for each example of device 800, device 800 initiates a DA and presents a response to the spoken input "What's the weather?" according to a silent response mode, a mixed response mode, or an audio response mode, as discussed below. Device 800 implementing system 1100 selects a DA response mode and presents a response according to the selected response mode using techniques discussed below.
[0339] The system 1100 includes an acquisition module 1102. The acquisition module 1102 acquires a response package in response to a natural language input. The response package includes content (e.g., speakable text) intended as a response to the natural language input. In some examples, the response package includes first text (content text) associated with a digital assistant response affordance (e.g., response affordance 1202) and second text (caption text) associated with the response affordance. In some examples, the caption text is less detailed (e.g., includes fewer words) than the content text. The content text can provide a complete response to a user's request, while the caption text can provide an abbreviated (e.g., incomplete) response to the request. For a complete response to a request, the device 800 can, for example, present the caption text simultaneously with the response affordance, whereas presentation of the content text may not require presentation of the response affordance for a complete response.
[0340] For example, consider the natural language input "What's the weather?" in Figure 12. The content text is "It's currently 70 degrees and sunny with no chance of rain today. The high today is 75 degrees and the low is 60 degrees." The caption text is simply "Nice weather today." As shown, the caption text is intended for presentation with a response affordance 1202 that visually indicates the information in the content text. Thus, presentation of the content text alone can fully answer the request while presenting both the caption text and the response affordance.
[0341] In some embodiments, the acquisition module 1102 acquires the response package locally, for example, by the device 800 that processes the natural language input, as described with respect to Figures 7A-7C. In some embodiments, the acquisition module 1102 acquires the response package from an external device, such as the DA server 106. In such an example, the DA server 106 processes the natural language input and determines the response package, as described with respect to Figures 7A-7C. In some embodiments, the acquisition module 1102 acquires one portion of the response package locally and another portion of the response package from the external device.
[0342] The system 1100 includes a mode selection module 1104. The selection module 1104 selects a DA response mode from a plurality of DA response modes based on context information associated with the device 800. The DA response mode specifies how (e.g., in what format) the DA presents a response to a natural language input (e.g., a response package).
[0343] In some embodiments, the selection module 1104 selects the DA response mode based on current context information obtained after the device 800 receives the natural language input, e.g., after receiving the natural language input. In some embodiments, the selection module 1104 selects the DA response mode based on current context information obtained after the acquisition module 1102 acquires the response package, e.g., after acquiring the response package. The current context information describes the time the selection module 1104 uses the context information to select the DA response mode. In some embodiments, the time is after receiving the natural language input and before presenting a response to the natural language input. In some embodiments, the multiple DA response modes include a silent response mode, a mixed response mode, and a voice response mode, which are discussed further below.
[0344] The system 1100 includes a formatting module 1106. In response to the selection module 1104 selecting a DA response mode, the formatting module 1106 causes the DA to present a response package (e.g., in a format consistent with the selected DA response mode). In some embodiments, the selected DA response mode is a silent response mode. In some embodiments, presenting the response package according to the silent response mode includes displaying a response affordance and displaying caption text without providing (e.g., speaking) audio output representing the caption text (without providing content text). In some embodiments, the selected DA response mode is a mixed response mode. In some embodiments, presenting the response package according to the mixed response mode includes displaying a response affordance and speaking caption text without displaying the caption text (without providing context text). In some embodiments, the selected DA response mode is an audio response mode. In some embodiments, presenting the response package according to the audio response mode includes, for example, speaking content text without presenting caption text and / or without displaying a response affordance.
[0345] 12 , presenting a response package according to a silent response mode includes displaying response affordance 1202 and displaying the caption text "It's a nice day today" within dialog affordance 1204 without speaking the caption text. Presenting a response package according to a mixed response mode includes displaying response affordance 1202 and speaking the caption text "It's a nice day today" without displaying the caption text. Presenting a response package according to a voice response mode includes displaying content text "It's currently 70 degrees and there's no chance of rain today. The high today is 75 degrees and the low is 60 degrees." Although FIG. 12 shows device 800 displaying response affordance 1202 when presenting a response package according to a voice response mode, in other examples, the response affordance is not displayed when presenting a response package according to a voice response mode.
[0346] In some examples, when the DA presents a response according to the silent response mode, the device 800 displays the response affordance without displaying a dialog affordance (e.g., including text). In some examples, the device 800 refrains from providing text according to determining that the response affordance includes a direct answer to the natural language request. For example, the device 800 determines that the caption text and the response affordance each include respective matching text that responds to the user request (thus making the caption text redundant). For example, for the natural language request "What's the temperature?", if the response affordance includes the current temperature, in silent mode the device 800 does not display any caption text because the caption text including the current temperature is redundant with the response affordance. In contrast, consider the exemplary natural language request "Is it cold?" The response affordance for the request may include the current temperature and weather conditions but may not include a direct (e.g., explicit) answer to the request, such as "yes" or "no." Thus, for such natural language input, in silent mode, device 800 displays both a response affordance and caption text containing a direct answer to the request, e.g., "No, it's not cold."
[0347] 12 illustrates that in some embodiments, selecting a DA response mode includes determining whether to (1) display caption text without speaking the caption text, or (2) speak the caption text without displaying the caption text. In some embodiments, selecting a response mode includes determining whether to speak content text.
[0348] In general, the silent response mode may be preferred when the user wishes to view the display and does not wish to have audio output. The mixed response mode may be preferred when the user wishes to view the display and desires audio output. The audio response mode may be preferred when the user does not wish to view (or is unable to see) the display. Various techniques and context information selection module 1104 used to select the DA response mode are now described.
[0349] 13 shows an example process 1300 implemented by the selection module 1104 to select a DA response mode, according to various embodiments. In some embodiments, the selection module 1104 implements the process 1300 as computer-executable instructions stored in memory of the device 800, for example.
[0350] At block 1302, the selection module 1104 obtains (e.g., determines) current context information. At block 1304, the module 1104 determines whether to select voice mode based on the current context information. If the module 1104 determines to select voice mode, the module 1104 selects the voice mode at block 1306. If the module 1104 determines not to select voice mode, the process 1300 proceeds to block 1308. At block 1308, the module 1104 selects between silent mode and mixed mode. If the module 1104 determines to select silent mode, the module 1104 selects silent mode at block 1310. If the module 1104 determines to select mixed mode, the module 1104 selects mixed mode at block 1312.
[0351] In some embodiments, blocks 1304 and 1308 are implemented using a rule-based system. For example, in block 1304, module 1104 determines whether the current context information meets certain conditions for selecting the voice mode. If the certain conditions are met, module 1104 selects the voice mode. If the certain conditions are not met (meaning the current context information meets the conditions for selecting the mixed mode or the voice mode), module 1104 proceeds to block 1308. Similarly, in block 1308, module 1104 determines whether the current context information meets certain conditions for selecting the silent mode or the mixed mode and selects the silent mode or the mixed mode accordingly.
[0352] In some embodiments, blocks 1304 and 1308 are implemented using a probabilistic (e.g., machine learning) system. For example, in block 1304, module 1104 determines the probability of selecting voice mode and the probability of not selecting voice mode (e.g., the probability of selecting silent mode or mixed mode) based on current context information, and selects the branch with the highest probability. In block 1308, module 1104 determines the probability of selecting mixed mode and the probability of selecting silent mode based on current context information, and selects the mode with the highest probability. In some embodiments, the voice mode, mixed mode, and silent mode probabilities sum to one.
[0353] Various types of current context information that may be used in the determinations of blocks 1304 and / or 1308 will now be discussed.
[0354] In some embodiments, the context information includes whether the device 800 has a display. In a rule-based system, a determination that the device 800 does not have a display satisfies a condition for selecting a voice mode. In a probabilistic system, a determination that the device 800 does not have a display increases the probability of a voice mode and / or decreases the probability of a mixed mode, which decreases the probability of a silent mode.
[0355] In some embodiments, the context information includes whether device 800 detects a voice input (e.g., "Hey Siri") to initiate DA. In a rule-based system, detecting a voice input that initiates DA satisfies a condition for selecting voice mode. In a rule-based system, not detecting a voice input that initiates DA does not satisfy a condition for selecting voice mode (and thus satisfies a condition for selecting mixed mode or silent mode). In a probabilistic system, in some embodiments, detecting a voice input that initiates DA increases the probability of voice mode and / or decreases the probability of mixed mode, and decreases the probability of silent mode. In a probabilistic system, in some embodiments, not detecting a voice input that initiates DA decreases the probability of voice mode and / or increases the probability of mixed mode, and increases the probability of silent mode.
[0356] In some embodiments, the context information includes whether the device 800 detects physical contact with the device 800 to initiate DA. In a rule-based system, not detecting physical contact satisfies a condition for selecting voice mode. In a rule-based system, detecting physical contact does not satisfy a condition for selecting voice mode. In a probabilistic system, in some embodiments, not detecting physical contact increases the probability of voice mode and / or decreases the probability of mixed mode, and decreases the probability of silent mode. In a probabilistic system, in some embodiments, detecting physical contact decreases the probability of voice mode and / or increases the probability of mixed mode, and increases the probability of silent mode.
[0357] In some embodiments, the context information includes whether the device 800 is in a locked state. In a rule-based system, a determination that the device 800 is in a locked state satisfies a condition for selecting a voice mode. In a rule-based system, a determination that the device 800 is not in a locked state does not satisfy a condition for selecting a voice mode. In a probabilistic system, in some embodiments, a determination that the device 800 is in a locked state increases the probability of a voice mode and / or decreases the probability of a mixed mode, and decreases the probability of a silent mode. In a probabilistic system, in some embodiments, a determination that the device 800 is not in a locked state decreases the probability of a voice mode and / or increases the probability of a mixed mode, and increases the probability of a silent mode.
[0358] In some embodiments, the context information includes whether a display of device 800 was displayed before initiating DA. In a rule-based system, a determination that a display was not displayed before initiating DA satisfies a condition for selecting voice mode. In a rule-based system, a determination that a display was displayed before initiating DA does not satisfy a condition for selecting voice mode. In a probabilistic system, in some embodiments, a determination that a display was not displayed before initiating DA increases the probability of voice mode and / or decreases the probability of mixed mode and decreases the probability of silent mode. In a probabilistic system, in some embodiments, a determination that a display was displayed before initiating DA decreases the probability of voice mode and / or increases the probability of mixed mode and increases the probability of silent mode.
[0359] In some embodiments, the context information includes a display orientation of device 800. In a rule-based system, a determination that the display is face-down satisfies a condition for selecting a voice mode. In a rule-based system, a determination that the display is face-up does not satisfy a condition for selecting a voice mode. In a probabilistic system, in some embodiments, a determination that the display is face-down increases the probability of a voice mode and / or decreases the probability of a mixed mode and decreases the probability of a silent mode. In a probabilistic system, in some embodiments, a determination that the display is face-up decreases the probability of a voice mode and / or increases the probability of a mixed mode and increases the probability of a silent mode.
[0360] In some embodiments, the context information includes whether the display of device 800 is obscured. For example, device 800 may use one or more sensors (e.g., a light sensor, a microphone, a proximity sensor) to determine whether the user cannot see the display. For example, the display may be in an at least partially enclosed space (e.g., a pocket, a bag, or a drawer) or may be covered by an object. In a rule-based system, a determination that the display is obscured satisfies a condition for selecting a voice mode. In a rule-based system, a determination that the display is not obscured does not satisfy a condition for selecting a voice mode. In a probabilistic system, in some embodiments, a determination that the display is obscured increases the probability of a voice mode and / or decreases the probability of a mixed mode and decreases the probability of a silent mode. In a probabilistic system, in some embodiments, a determination that the display is not obscured decreases the probability of a voice mode and / or increases the probability of a mixed mode and increases the probability of a silent mode.
[0361] In some embodiments, the context information includes whether device 800 is coupled to an external audio output device (e.g., headphones, Bluetooth device, speaker). In a rule-based system, a determination that device 800 is coupled to an external device satisfies a condition for selecting audio mode. In a rule-based system, a determination that device 800 is not coupled to an external device does not satisfy a condition for selecting audio mode. In a probabilistic system, in some embodiments, a determination that device 800 is coupled to an external device increases the probability of audio mode and / or decreases the probability of mixed mode and decreases the probability of silent mode. In a probabilistic system, in some embodiments, a determination that device 800 is not coupled to an external device decreases the probability of audio mode and / or increases the probability of mixed mode and increases the probability of silent mode.
[0362] In some embodiments, the context information includes whether the direction of the user's gaze is directed toward the device 800. In a rule-based system, a determination that the direction of the user's gaze is not directed toward the device 800 satisfies a condition for selecting a voice mode. In a rule-based system, a determination that the direction of the user's gaze is directed toward the device 800 does not satisfy a condition for selecting a voice mode. In a probabilistic system, in some embodiments, a determination that the direction of the user's gaze is not directed toward the device 800 increases the probability of a voice mode and / or decreases the probability of a mixed mode and decreases the probability of a silent mode. In a probabilistic system, in some embodiments, a determination that the direction of the user's gaze is directed toward the device 800 decreases the probability of a voice mode and / or increases the probability of a mixed mode and increases the probability of a silent mode.
[0363] In some embodiments, the context information includes whether a predetermined type of gesture was detected on the device 800 within a predetermined period of time prior to selecting the responsive mode. The predetermined type of gesture may include, for example, a lift and / or rotate gesture to turn on the display on the device 800. In a rule-based system, not detecting a predetermined type of gesture within a predetermined period of time satisfies a condition for selecting the voice mode. In a rule-based system, detecting a predetermined type of gesture within a predetermined period of time does not satisfy a condition for selecting the voice mode. In a probabilistic system, in some embodiments, not detecting a predetermined type of gesture within a predetermined period of time increases the probability of the voice mode and / or decreases the probability of the mixed mode, and decreases the probability of the silent mode. In a probabilistic system, in some embodiments, detecting a predetermined type of gesture within a predetermined period of time decreases the probability of the voice mode and / or increases the probability of the mixed mode, and increases the probability of the silent mode.
[0364] In some embodiments, the context information includes a direction of the natural language input. In a rule-based system, a determination that the direction of the natural language input is not directed toward the device 800 satisfies a condition for selecting a voice mode. In a rule-based system, a determination that the direction of the natural language input is directed toward the device 800 does not satisfy a condition for selecting a voice mode. In a probabilistic system, in some embodiments, a determination that the direction of the natural language input is not directed toward the device 800 increases the probability of a voice mode and / or decreases the probability of a mixed mode and decreases the probability of a silent mode. In a probabilistic system, in some embodiments, a determination that the direction of the natural language input is directed toward the device 800 decreases the probability of a voice mode and / or increases the probability of a mixed mode and increases the probability of a silent mode.
[0365] In some embodiments, the context information includes whether device 800 detected a touch (e.g., user input selecting a responsive affordance) performed on device 800 within a predetermined period of time prior to selecting the responsive mode. In a rule-based system, not detecting a touch within a predetermined period of time satisfies a condition for selecting voice mode. In a rule-based system, detecting a touch within a predetermined period of time does not satisfy a condition for selecting voice mode. In a probabilistic system, in some embodiments, not detecting a touch within a predetermined period of time increases the probability of voice mode and / or decreases the probability of mixed mode and decreases the probability of silent mode. In a probabilistic system, in some embodiments, detecting a touch within a predetermined period of time decreases the probability of voice mode and / or increases the probability of mixed mode and increases the probability of silent mode.
[0366] In some embodiments, the conte...
Claims
1. 1. A method of operating a digital assistant, comprising: In an electronic device having one or more processors, a memory, and a display, receiving a natural language input; starting the digital assistant; Obtaining a response package responsive to the natural language input in accordance with initiating the digital assistant; After receiving the natural language input, selecting a first response mode of the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device, wherein the context information includes whether a user's gaze direction is directed toward or not directed toward the electronic device, and the context information includes a physical state of the electronic device; In response to selecting the first response mode, presenting the response package by the digital assistant according to the first response mode, wherein presenting the response package according to the first response mode includes: Displaying first text associated with a digital assistant response affordance, the first text being a complete response to the natural language input; Displaying second text associated with the digital assistant response affordance, the second text being an abbreviated response to the natural language input; and A method comprising:
2. The method of claim 1 , wherein the second text has fewer words than the first text.
3. Selecting the first response mode comprises: displaying the second text without providing an audio output representing the second text; or providing the audio output representing the second text without displaying the second text; The method of any one of claims 1 to 2, comprising determining whether to:
4. The method of any one of claims 1 to 3, wherein selecting the first response mode comprises determining whether to provide a speech output representing the first text.
5. the first response mode is a silent response mode; presenting, by the digital assistant, the response package according to the first response mode; Displaying the digital assistant response affordance; and displaying the second text without providing a second audio output representing the second text; The method according to any one of claims 1 to 4, comprising:
6. the context information includes digital assistant voice feedback settings; Selecting the first response mode is based on determining that the digital assistant voice feedback setting indicates not to provide voice feedback; The method of claim 5.
7. the context information includes detecting a physical contact of the electronic device to initiate the digital assistant; selecting the first response mode based on the detection of the physical contact. The method according to any one of claims 5 to 6.
8. the context information includes whether the electronic device is in a locked state; selecting the first response mode is based on determining that the electronic device is not in the locked state; The method according to any one of claims 5 to 7.
9. The context information includes whether the display of the electronic device was displaying before starting the digital assistant; Selecting the first response mode is based on determining that the display was displaying before starting the digital assistant; The method according to any one of claims 5 to 8.
10. the contextual information includes detecting a touch performed on the electronic device within a predetermined period of time before selecting the first response mode; selecting the first response mode based on the detection of the touch; The method according to any one of claims 5 to 9.
11. the context information includes detecting a predetermined gesture of the electronic device within a second predetermined period of time prior to selecting the first response mode; selecting the first response mode is based on the detection of the predetermined gesture. The method according to any one of claims 5 to 10.
12. the first response mode is a mixed response mode; Presenting the response package according to the first response mode by the digital assistant includes displaying the digital assistant response affordance and providing a second speech output representing the second text without displaying the second text. The method according to any one of claims 5 to 11.
13. the context information includes digital assistant voice feedback settings; Selecting the first response mode is based on determining that the digital assistant voice feedback setting indicates providing voice feedback; The method of claim 12.
14. the context information includes detecting a physical contact of the electronic device to initiate the digital assistant; selecting the first response mode is based on the detection of the physical contact. The method according to any one of claims 12 to 13.
15. the context information includes whether the electronic device is in a locked state; selecting the first response mode is based on determining that the electronic device is not in the locked state; The method according to any one of claims 12 to 14.
16. The context information includes whether the display of the electronic device was displaying before starting the digital assistant; Selecting the first response mode is based on determining that the display was displaying before starting the digital assistant; The method according to any one of claims 12 to 15.
17. the contextual information includes detecting a touch performed on the electronic device within a predetermined period of time before selecting the first response mode; selecting the first response mode based on the detection of the touch; The method according to any one of claims 12 to 16.
18. the context information includes detecting a predetermined gesture of the electronic device within a second predetermined period of time prior to selecting the first response mode; selecting the first response mode is based on the detection of the predetermined gesture. The method according to any one of claims 12 to 17.
19. the first response mode is a voice response mode; presenting, by the digital assistant, the response package according to the first response mode includes providing a speech output representing the first text. The method according to any one of claims 1 to 4.
20. the context information includes a determination that the electronic device is in a vehicle; selecting the first response mode is based on the determination that the electronic device is within the vehicle.
20. The method of claim 19.
21. the contextual information includes a determination that the electronic device is coupled to an external audio output device; selecting the first response mode is based on the determination that the electronic device is coupled to the external audio output device. The method according to any one of claims 19 to 20.
22. the context information includes detecting a voice input to initiate the digital assistant; selecting the first response mode is based on the detection of the voice input. The method according to any one of claims 19 to 21.
23. the context information includes whether the electronic device is in a locked state; selecting the first response mode is based on determining that the electronic device is in the locked state. The method according to any one of claims 19 to 22.
24. The context information includes whether the display of the electronic device was displaying before starting the digital assistant; Selecting the first response mode is based on determining that the display of the electronic device was not displaying before starting the digital assistant; The method according to any one of claims 19 to 23.
25. After presenting the response package with the digital assistant, receiving a second natural language input in response to the presentation of the response package; obtaining a second response package responsive to the second natural language input; After receiving the second natural language input, selecting a second response mode of the digital assistant from the plurality of digital assistant response modes, wherein the second response mode is different from the first response mode; In response to selecting the second response mode, presenting the second response package by the digital assistant according to the second response mode; The method of any one of claims 1 to 24, further comprising:
26. The method of any one of claims 1 to 25, wherein selecting the first response mode of the digital assistant is performed after obtaining the response package.
27. The display and a touch-sensitive surface; and one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for performing the method of any one of claims 1 to 26.
28. 27. A computer program comprising instructions that, when executed by one or more processors of an electronic device comprising a display and a touch-sensitive surface, cause the electronic device to perform the method of any one of claims 1 to 26.
29. An electronic device comprising means for carrying out the method according to any one of claims 1 to 26.
Citation Information
Patent Citations
Changing device behavior based on eye gaze
JP2014516181A
Information output device, information output method, and program
JP2016181047A
Activation of Virtual Assistant
JP2018505491A
Adjustment of Digital Personal Assistant Agent Across Devices
JP2018509014A
Context-Sensitive Handling of Interruptions by Intelligent Digital Assistant
US20140074483A1