Accelerated task execution

By displaying task-related suggestions on electronic devices and responding to user input, the identification limitations of digital assistants when performing tasks are solved, the user experience and device operability are improved, and more efficient task execution and battery life are achieved.

CN113342156BActive Publication Date: 2025-05-30APPLE INC
View PDF 26 Cites 0 Cited by

Patent Information

Application Number
CN202110612652.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-06-03
Filing Date
2018-09-27
Publication Date
2025-05-30
Estimated Expiration
2038-09-27

AI Technical Summary

Technical Problem

Existing digital assistants are limited by the way they recognize tasks when performing tasks, making it difficult for users to directly indicate tasks through natural language voice input, and digital assistants cannot adjust according to user behavior to optimize user experience.

Method used

By displaying a user interface with suggested indications associated with a task on an electronic device, detecting user input and responding, performing a task or displaying a confirmation interface according to the task type.

Benefits of technology

It provides an easy to identify and intuitive way to enable users to perform tasks, reducing the amount of user input, improving device operability and user experience, thereby reducing power usage and extending device battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113342156B_ABST
    Figure CN113342156B_ABST
Patent Text Reader

Abstract

This application relates to accelerated task execution. The present invention provides systems and processes for accelerating task execution. Example methods include, at an electronic device including a display and one or more input devices, displaying, on the display, a user interface including a suggested affordance representation associated with a task; detecting, via the one or more input devices, a first user input corresponding to a selection of the suggested affordance representation, and in response to detecting the first user input: executing the task based on determining that the task is a first type of task; and displaying a confirmation interface including a confirmation affordance representation based on determining that the task is a second type of task different from the first type.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application is a divisional application of the Chinese patent application with the application number 201811134435.X, the filing date of September 27, 2018, and the invention title of "Accelerated Task Execution". Technical Field

[0003] The present invention generally relates to digital assistants, and more particularly, to using digital assistants to accelerate task execution. Background Art

[0004] Intelligent automation assistants (or digital assistants) can provide a beneficial interface between human users and electronic devices. Such assistants can allow users to interact with the device or system in natural language in voice form and / or text form. For example, a user can provide voice input containing a user request to a digital assistant running on an electronic device. The digital assistant can interpret the user's intention from the voice input, operationalize the user's intention into a task, and execute the task. In some systems, performing tasks in this way can be restricted in the way the tasks are recognized. However, in some cases, the user may be restricted to a specific set of commands, such that the user cannot easily instruct the digital assistant to perform tasks using natural language voice input. In addition, in many cases, the digital assistant cannot be adjusted based on previous user behavior, and thus lacks the desired optimization of the user experience. Summary of the Invention

[0005] Example methods are described herein. The example methods include, on an electronic device having a display and a touch-sensitive surface, displaying a user interface on the display that includes a suggested affordance representation associated with a task, detecting a first user input corresponding to a selection of the suggested affordance representation via one or more input devices; in response to detecting the first user input: performing the task according to a determination that the task is a first type of task; and displaying a confirmation interface that includes a confirmation affordance representation according to a determination that the task is a second type of task different from the first type.

[0006] Example electronic devices are described herein. The example electronic devices include one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for the following operations: displaying a user interface on the display that includes a suggested affordance representation associated with a task; detecting a first user input corresponding to a selection of the suggested affordance representation via one or more input devices; in response to detecting the first user input: performing the task according to a determination that the task is a first type of task; and displaying a confirmation interface that includes a confirmation affordance representation according to a determination that the task is a second type of task different from the first type.

[0007] An example electronic device includes: means for displaying on a display a user interface including a suggested affordance representation associated with a task; means for detecting, via one or more input devices, a first user input corresponding to a selection of the suggested affordance representation; and means for, in response to detecting the first user input, performing the following operations: performing the task according to a determination that the task is a first type of task; and displaying a confirmation interface including a confirmation affordance representation according to a determination that the task is a second type of task different from the first type.

[0008] The present disclosure provides example non-transitory computer-readable media. The example non-transitory computer-readable media stores one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: display on a display a user interface including a suggested affordance representation associated with a task; detect, via one or more input devices, a first user input corresponding to a selection of the suggested affordance representation; and in response to detecting the first user input: perform the task according to a determination that the task is a first type of task; and display a confirmation interface including a confirmation affordance representation according to a determination that the task is a second type of task different from the first type.

[0009] Displaying a user interface including a suggested affordance representation and, in response to a selection of the suggested affordance representation, selectively requesting confirmation to perform a task provides an easy-to-recognize and intuitive method for a user to perform tasks on an electronic device, thereby reducing the amount of user input otherwise required to perform such tasks. Thus, displaying the user interface in this way enhances the operability of the device and makes the user device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn further reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0010] An example method includes, at an electronic device having one or more processors: displaying a plurality of candidate task affordance representations including a candidate task affordance representation associated with a task; detecting a set of inputs including a first user input corresponding to a selection of the candidate task affordance representation associated with the task; in response to detecting the set of user inputs, displaying a first interface for generating a voice shortcut associated with the task, and while displaying the first interface: receiving natural language voice input via an audio input device and displaying candidate phrases in the first interface, where the candidate phrases are based on the natural language voice input; after displaying the candidate phrases, detecting a second user input via a touch-sensitive surface; and in response to detecting the second user input, associating the candidate phrase with the task.

[0011] An example non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to display a plurality of candidate task enabling representations including candidate task enabling representations associated with a task; detect a set of inputs including a first user input that corresponds to a selection of a candidate task enabling representation associated with the task; in response to detecting the set of user inputs, display a first interface for generating a voice shortcut associated with the task, and when displaying the first interface: receive natural language voice input via an audio input device and display candidate phrases in the first interface, where the candidate phrases are based on the natural language voice input; after displaying the candidate phrases, detect a second user input via a touch-sensitive surface; and in response to detecting the second user input, associate the candidate phrase with the task.

[0012] An example electronic device includes one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for: displaying a plurality of candidate task enabling representations including candidate task enabling representations associated with a task; detecting a set of inputs including a first user input that corresponds to a selection of a candidate task enabling representation associated with the task; in response to detecting the set of user inputs, displaying a first interface for generating a voice shortcut associated with the task, and when displaying the first interface: receiving natural language voice input via an audio input device and displaying candidate phrases in the first interface, where the candidate phrases are based on the natural language voice input; after displaying the candidate phrases, detecting a second user input via a touch-sensitive surface; and in response to detecting the second user input, associating the candidate phrase with the task.

[0013] An example electronic device includes: means for displaying a plurality of candidate task enabling representations including candidate task enabling representations associated with a task; means for detecting a set of inputs including a first user input that corresponds to a selection of a candidate task enabling representation associated with the task; means for, in response to detecting the set of user inputs, displaying a first interface for generating a voice shortcut associated with the task; means for, when displaying the first interface, performing the following operations: receiving natural language voice input via an audio input device and displaying candidate phrases in the first interface, where the candidate phrases are based on the natural language voice input; means for, after displaying the candidate phrases, detecting a second user input via a touch-sensitive surface; and means for, in response to detecting the second user input, associating the candidate phrase with the task.

[0014] Providing candidate phrases based on natural language voice input and associating the candidate phrases with corresponding tasks allows a user to accurately and efficiently generate user-specific voice shortcuts that can be used to perform tasks on an electronic device. For example, allowing the user to associate voice shortcuts with tasks in this way allows the user to visually confirm that the desired voice shortcut has been selected and assigned to the correct task, thereby reducing the likelihood of incorrect or unwanted associations. Thus, providing candidate phrases in the described manner makes the use of the electronic device more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn further reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0015] An example method includes: at an electronic device having one or more processors, receiving context data associated with the electronic device; determining a task probability of a task based on the context data; determining a parameter probability of a parameter based on the context data, where the parameter is associated with the task; determining whether the task meets a recommendation criterion based on the task probability and the parameter probability; in accordance with determining that the task meets the recommendation criterion, displaying a recommended affordance representation corresponding to the task and the parameter on a display; and in accordance with determining that the task does not meet the recommendation criterion, refraining from displaying the recommended affordance representation.

[0016] An example non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive context data associated with the electronic device; determine a task probability of a task based on the context data; determine a parameter probability of a parameter based on the context data, where the parameter is associated with the task; determine whether the task meets a recommendation criterion based on the task probability and the parameter probability; in accordance with determining that the task meets the recommendation criterion, display a recommended affordance representation corresponding to the task and the parameter on a display; and in accordance with determining that the task does not meet the recommendation criterion, refrain from displaying the recommended affordance representation.

[0017] An example electronic device includes one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for: receiving context data associated with the electronic device; determining a task probability of a task based on the context data; determining a parameter probability of a parameter based on the context data, where the parameter is associated with the task; determining whether the task meets a recommendation criterion based on the task probability and the parameter probability; in accordance with determining that the task meets the recommendation criterion, displaying a recommended affordance representation corresponding to the task and the parameter on a display; and in accordance with determining that the task does not meet the recommendation criterion, refrain from displaying the recommended affordance representation.

[0018] An example electronic device includes means for receiving context data associated with the electronic device; means for determining a task probability of a task based on the context data; means for determining a parameter probability of a parameter based on the context data, wherein the parameter is associated with the task; means for determining whether the task meets a recommendation criterion based on the task probability and the parameter probability; means for displaying, on a display, a recommended affordance representation corresponding to the task and the parameter according to determining that the task meets the recommendation criterion; and means for refraining from displaying the recommended affordance representation according to determining that the task does not meet the recommendation criterion.

[0019] Selectively providing a recommended affordance representation associated with a task as described herein allows a user to effectively and conveniently perform user-related tasks on an electronic device. As an example, the recommended affordance representation displayed by the electronic device may correspond to a task identified based on context data of the electronic device (such as context data indicating a user's previous use of the electronic device). Thus, selectively providing recommendations in this manner reduces the amount of input and time required for a user to operate the electronic device (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device.

[0020] An example method includes: at an electronic device having one or more processors, receiving a natural language voice input; determining whether the natural language voice input meets a voice shortcut criterion; according to determining that the natural language voice input meets the voice shortcut criterion: identifying a task associated with the voice shortcut, and performing the task associated with the voice shortcut; and according to determining that the natural language voice input does not meet the voice shortcut criterion: identifying a task associated with the natural language voice input, and performing the task associated with the natural language voice input.

[0021] An example non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive a natural language voice input; determine whether the natural language voice input meets a voice shortcut criterion; according to determining that the natural language voice input meets the voice shortcut criterion: identify a task associated with the voice shortcut, and perform the task associated with the voice shortcut; and according to determining that the natural language voice input does not meet the voice shortcut criterion: identify a task associated with the natural language voice input, and perform the task associated with the natural language voice input.

[0022] The example electronic device includes one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors. The one or more programs include instructions for the following operations: receiving a natural language voice input; determining whether the natural language voice input meets the voice shortcut criteria; based on determining that the natural language voice input meets the voice shortcut criteria: identifying a task associated with the voice shortcut and executing the task associated with the voice shortcut; and based on determining that the natural language voice input does not meet the voice shortcut criteria: identifying a task associated with the natural language voice input and executing the task associated with the natural language voice input.

[0023] The example electronic device includes means for receiving a natural language voice input; means for determining whether the natural language voice input meets the voice shortcut criteria; means for performing the following operations based on determining that the natural language voice input meets the voice shortcut criteria: identifying a task associated with the voice shortcut and executing the task associated with the voice shortcut; and means for performing the following operations based on determining that the natural language voice input does not meet the voice shortcut criteria: identifying a task associated with the natural language voice input and executing the task associated with the natural language voice input.

[0024] As described herein, performing tasks in response to natural language voice inputs (e.g., voice shortcuts) provides an intuitive and efficient way to perform tasks on an electronic device. As an example, one or more tasks can be performed in response to a natural language voice input without any additional input from the user. Thus, performing tasks in this manner in response to natural language voice inputs reduces the amount of input and time required for the user to operate the electronic device (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device.

[0025] The example method includes: at an electronic device having one or more processors, receiving a natural language voice input using a digital assistant; determining a voice shortcut associated with the natural language voice input; determining a task corresponding to the voice shortcut; causing an application to initiate execution of the task; receiving a response from the application, where the response is associated with the task; determining whether the task was successfully executed based on the response, and providing an output indicating whether the task was successfully executed.

[0026] An example non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device, cause the electronic device to receive natural language speech input using a digital assistant; determine a voice shortcut associated with the natural language speech input; determine a task corresponding to the voice shortcut; cause an application to initiate execution of the task; receive a response from the application, where the response is associated with the task; determine whether the task was successfully executed based on the response, and provide an output indicating whether the task was successfully executed.

[0027] An example electronic device includes one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors. The one or more programs include instructions for: receiving natural language speech input using a digital assistant; determining a voice shortcut associated with the natural language speech input; determining a task corresponding to the voice shortcut; causing an application to initiate execution of the task; receiving a response from the application, where the response is associated with the task; determining whether the task was successfully executed based on the response, and providing an output indicating whether the task was successfully executed.

[0028] An example electronic device includes means for receiving natural language speech input using a digital assistant; means for determining a voice shortcut associated with the natural language speech input; means for determining a task corresponding to the voice shortcut; means for causing an application to initiate execution of the task; means for receiving a response from the application, where the response is associated with the task; means for determining whether the task was successfully executed based on the response; and means for providing an output indicating whether the task was successfully executed.

[0029] As described herein, providing the output allows the digital assistant to provide feedback and / or other information from the application in an intuitive and flexible manner, for example, during a conversation (e.g., a conversational dialogue) between the user and the digital assistant. As an example, the digital assistant can provide (e.g., relay) a natural language expression from the application to the user such that the user can interact with the application without having to open or otherwise directly access the application. Thus, providing natural language output in this manner reduces the amount of input and time required for the user to operate the electronic device (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device.

[0030] Example methods include receiving, from an application, a plurality of media items; receiving context data associated with an electronic device; determining a task based on the plurality of media items and the context data; determining whether the task meets a recommendation criterion; in accordance with determining that the task meets the recommendation criterion, displaying, on a display, a recommended affordance representation corresponding to the task; and in accordance with determining that the task does not meet the recommendation criterion, refraining from displaying the recommended affordance representation.

[0031] Example electronic devices include one or more processors, memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: receiving, from an application, a plurality of media items; receiving context data associated with the electronic device; determining a task based on the plurality of media items and the context data; determining whether the task meets a recommendation criterion; in accordance with determining that the task meets the recommendation criterion, displaying, on a display, a recommended affordance representation corresponding to the task; and in accordance with determining that the task does not meet the recommendation criterion, refraining from displaying the recommended affordance representation.

[0032] Example electronic devices include means for receiving, from an application, a plurality of media items; means for receiving context data associated with the electronic device; means for determining a task based on the plurality of media items and the context data; means for determining whether the task meets a recommendation criterion; means for, in accordance with determining that the task meets the recommendation criterion, displaying, on a display, a recommended affordance representation corresponding to the task; and means for, in accordance with determining that the task does not meet the recommendation criterion, refraining from displaying the recommended affordance representation.

[0033] Exemplary non-transitory computer-readable media store one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: receive, from an application, a plurality of media items; receive context data associated with the electronic device; determine a task based on the plurality of media items and the context data; determine whether the task meets a recommendation criterion; in accordance with determining that the task meets the recommendation criterion, display, on a display, a recommended affordance representation corresponding to the task; and in accordance with determining that the task does not meet the recommendation criterion, refrain from displaying the recommended affordance representation.

[0034] Selectively providing a recommended affordance representation corresponding to a task as described herein allows a user to efficiently and conveniently perform user-related tasks on an electronic device. As an example, the recommended affordance representation displayed by the electronic device may correspond to a task identified based on media consumption and / or a determined media preference of the user. Thus, selectively providing recommendations in this manner reduces the amount of input and time required by the user to operate the electronic device (e.g., by assisting the user in providing appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 FIG. is a block diagram showing a system and environment for implementing a digital assistant according to various examples.

[0036] Figure 2A FIG. is a block diagram of a portable multifunctional device showing a client-side portion for implementing a digital assistant according to various examples.

[0037] Figure 2B FIG. is a block diagram showing exemplary components for event handling according to various examples.

[0038] Figure 3 FIG. shows a portable multifunctional device implementing a client-side portion of a digital assistant according to various examples.

[0039] Figure 4 FIG. is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface according to various examples.

[0040] Figure 5A FIG. shows an exemplary user interface of a menu of an application on a portable multifunctional device according to various examples.

[0041] Figure 5B FIG. shows an exemplary user interface of a multifunctional device having a touch-sensitive surface separate from the display according to various examples.

[0042] Figure 6A FIG. shows a personal electronic device according to various examples.

[0043] Figure 6B FIG. is a block diagram showing a personal electronic device according to various examples.

[0044] Figures 6C - 6D FIG. shows exemplary components of a personal electronic device having a touch-sensitive display and an intensity sensor according to some embodiments.

[0045] Figures 6E - 6H FIG. shows exemplary components and a user interface of a personal electronic device according to some embodiments.

[0046] Figure 7A FIG. is a block diagram showing a digital assistant system or its server portion according to various examples.

[0047] Figure 7B FIG. shows, according to various examples, the functions of the digital assistant shown in Figure 7A

[0048] Figure 7C FIG. shows a portion of a knowledge ontology according to various examples.

[0049] Figures 8A - 8AFShows an exemplary user interface for providing suggestions according to various examples.

[0050] Figures 9A - 9B Is a flowchart showing a method for providing suggestions according to various examples.

[0051] Figures 10A - 10AJ Shows an exemplary user interface for providing voice shortcuts according to various examples.

[0052] Figures 11A - 11B Is a flowchart showing a method for providing voice shortcuts according to various examples.

[0053] Figure 12 Is a block diagram of a task recommendation system according to various examples.

[0054] Figure 13 Is a flowchart showing a method for providing suggestions according to various examples.

[0055] Figure 14 Shows an exemplary operation sequence for performing a task in a privacy - protected manner according to various examples.

[0056] Figure 15 Is a flowchart showing a method for performing a task according to various examples.

[0057] Figures 16A - 16S Shows an exemplary user interface for performing a task using a digital assistant according to various examples.

[0058] Figure 17 Is a flowchart showing a method for performing a task using a digital assistant according to various examples.

[0059] Figures 18A - 18D Shows an exemplary user interface for providing media item suggestions according to various examples.

[0060] Figure 19 Is a flowchart showing a method for providing media item suggestions according to various examples. Detailed Description

[0061] In the following description of the examples, reference will be made to the accompanying drawings, in which specific examples that can be implemented are shown by way of illustration. It should be understood that other examples can be used and structural changes can be made without departing from the scope of the various examples.

[0062] Although the following description uses terms such as "first" and "second" to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various examples, the first input may be referred to as the second input, and similarly, the second input may be referred to as the first input. The first input and the second input are both inputs and, in some cases, are independent and different inputs.

[0063] The terms used in the description of the various examples herein are for the purpose of describing particular examples only and are not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms "includes", "including", "comprises", and / or "comprising", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0064] Depending on the context, the term "if" can be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined..." or "if detected [the stated condition or event]" can be interpreted to mean "when determining..." or "in response to determining..." or "when detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]".

[0065] 1. System and environment

[0066] Figure 1FIG. 0 shows a block diagram of a system 100 according to various examples. In some examples, the system 100 implements a digital assistant. The terms “digital assistant,” “virtual assistant,” “intelligent automation assistant,” or “automated digital assistant” refer to any information processing system that interprets natural language input in oral form and / or text form to infer a user's intent and performs an action based on the inferred user intent. For example, to act on the inferred user intent, the system performs one or more of the following steps: identifying a task flow having steps and parameters designed to implement the inferred user intent, entering specific requirements into the task flow based on the inferred user intent; executing the task flow by invoking a program, method, service, API, etc.; and generating an output response to the user in an audible (e.g., speech) and / or visual form.

[0067] Specifically, the digital assistant is capable of accepting user requests that are at least partially in the form of natural language commands, requests, statements, utterances, and / or queries. Generally, a user request either seeks an informative answer from the digital assistant or seeks the digital assistant to perform a task. A satisfactory response to a user request includes providing the requested informative answer, performing the requested task, or a combination of both. For example, a user asks the digital assistant a question such as “Where am I now?” Based on the user's current location, the digital assistant answers “You are near the west gate of Central Park.” The user also requests to perform a task, such as “Please invite my friends to my girlfriend's birthday party next week.” In response, the digital assistant can confirm the request by saying “Okay, right away” and then send an appropriate calendar invitation to each of the user's friends listed in the user's electronic address book on behalf of the user. During the execution of the requested task, the digital assistant sometimes interacts with the user in an ongoing conversation involving multiple information exchanges over a long period of time. There are many other ways to interact with the digital assistant to request information or perform various tasks. In addition to providing verbal responses and taking programmed actions, the digital assistant also provides responses in other video or audio forms, such as text, alerts, music, video, animation, etc.

[0068] As Figure 1 shown, in some examples, the digital assistant is implemented according to a client-server model. The digital assistant includes a client-side portion 102 (hereinafter referred to as “DA client 102”) that executes on the user device 104 and a server-side portion 106 (hereinafter referred to as “DA server 106”) that executes on the server system 108. The DA client 102 communicates with the DA server 106 via one or more networks 110. The DA client 102 provides client-side functions, such as user-facing input and output processing, and communication with the DA server 106. The DA server 106 provides server-side functions for any number of DA clients 102 located on respective user devices 104.

[0069] In some examples, the DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, and an I / O interface 118 to external services. The client-facing I / O interface 112 facilitates client-facing input and output processing of the DA server 106. One or more processing modules 114 utilize the data and models 116 to process voice input and determine user intent based on the natural language input. Additionally, one or more processing modules 114 perform task execution based on the inferred user intent. In some examples, the DA server 106 communicates with an external service 120 via one or more networks 110 to complete a task or gather information. The I / O interface 118 to external services enables such communication.

[0070] The user device 104 can be any suitable electronic device. In some examples, the user device 104 is a portable multifunctional device (e.g., the device 200 described below with reference to Figure 2A ), a multifunctional device (e.g., the device 400 described below with reference to Figure 4 ), or a personal electronic device (e.g., the device 600 described below with reference to Figures 6A - 6B ). A portable multifunctional device is, for example, a mobile phone that also includes other functions such as PDA and / or music player functionality. Specific examples of portable multifunctional devices include the Apple iPod and devices from Apple Inc. (Cupertino, California). Other examples of portable multifunctional devices include, but are not limited to, earbuds / headphones, speakers, and laptop or tablet computers. Additionally, in some examples, the user device 104 is a non-portable multifunctional device. Specifically, the user device 104 is a desktop computer, gaming console, speaker, television, or set-top box. In some examples, the user device 104 includes a touch-sensitive surface (e.g., a touchscreen display and / or a touchpad). Additionally, the user device 104 optionally includes one or more other physical user interface devices such as a physical keyboard, mouse, and / or joystick. Various examples of electronic devices such as multifunctional devices are described in more detail below.

[0071] Examples of one or more communication networks 110 include local area networks (LANs) and wide area networks (WANs), such as the Internet. The one or more communication networks 110 are implemented using any known network protocols, including various wired or wireless protocols such as Ethernet, Universal Serial Bus (USB), FireWire, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi, Voice over Internet Protocol (VoIP), WiMAX, or any other suitable communication protocol.

[0072] The server system 108 is implemented on one or more stand-alone data processing devices or a distributed computer network. In some examples, the server system 108 also employs various virtual devices and / or services of a third-party service provider (e.g., a third-party cloud service provider) to provide potential computing resources and / or infrastructure resources of the server system 108.

[0073] In some examples, the user device 104 communicates with the DA server 106 via a second user device 122. The second user device 122 is similar or identical to the user device 104. For example, the second user device 122 is similar to the devices 200, 400, or 600 described below with reference to Figure 2A , Figure 4 and Figures 6A - 6B The user device 104 is configured to be communicatively coupled to the second user device 122 via a direct communication connection such as Bluetooth, NFC, BTLE, etc., or via a wired or wireless network such as a local Wi-Fi network. In some examples, the second user device 122 is configured to act as a proxy between the user device 104 and the DA server 106. For example, the DA client 102 of the user device 104 is configured to transmit information (e.g., a user request received at the user device 104) to the DA server 106 via the second user device 122. The DA server 106 processes the information and returns relevant data (e.g., data content in response to the user request) to the user device 104 via the second user device 122.

[0074] In some examples, the user device 104 may be configured to send a thumbnail request for data to a second user device 122 to reduce the amount of information transmitted from the user device 104. The second user device 122 is configured to determine supplementary information to add to the thumbnail request to generate a complete request for transmission to the DA server 106. This system architecture may advantageously allow a user device 104 with limited communication capabilities and / or limited battery power (e.g., a watch or similar compact electronic device) to access the services provided by the DA server 106 by using a second user device 122 (e.g., a mobile phone, laptop computer, tablet, etc.) with stronger communication capabilities and / or battery power as a proxy to the DA server 106. Although Figure 1 only two user devices 104 and 122 are shown, it should be understood that in some examples, the system 100 may include any number and type of user devices configured to communicate with the DA server system 106 in this proxy configuration.

[0075] Although Figure 1 the digital assistant shown includes both a client-side portion (e.g., DA client 102) and a server-side portion (e.g., DA server 106), in some examples, the functionality of the digital assistant is implemented as a stand-alone application installed on the user device. Additionally, the functional division between the client portion and the server portion of the digital assistant may vary in different implementations. For example, in some examples, the DA client is a thin client that provides only user-facing input and output processing functionality and delegates all other functions of the digital assistant to a backend server.

[0076] 2. Electronic device

[0077] Attention is now turned to an embodiment of an electronic device for implementing the client-side portion of the digital assistant. Figure 2Ais a block diagram showing a portable multifunctional device 200 having a touch-sensitive display system 212 according to some embodiments. The touch-sensitive display 212 is sometimes called a "touch screen" for convenience and may sometimes be referred to as or called a "touch-sensitive display system". The device 200 includes a memory 202 (which optionally includes one or more computer-readable storage media), a memory controller 222, one or more processing units (CPUs) 220, a peripheral device interface 218, an RF circuit 208, an audio circuit 210, a speaker 211, a microphone 213, an input / output (I / O) subsystem 206, other input control devices 216, and an external port 224. The device 200 optionally includes one or more optical sensors 264. The device 200 optionally includes one or more contact intensity sensors 265 for detecting the intensity of contacts on the device 200 (e.g., a touch-sensitive surface such as the touch-sensitive display system 212 of the device 200). The device 200 optionally includes one or more tactile output generators 267 for generating tactile output on the device 200 (e.g., generating tactile output on a touch-sensitive surface such as the touch-sensitive display system 212 of the device 200 or the touchpad 455 of the device 400). These components optionally communicate via one or more communication buses or signal lines 203.

[0078] As used in this specification and the claims, the "intensity" of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or to a surrogate for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four different values and more typically includes hundreds of different values (e.g., at least 256). The intensity of a contact is optionally determined (or measured) using a variety of methods and a variety of sensors or combinations of sensors. For example, one or more force sensors beneath or adjacent to the touch-sensitive surface are optionally used to measure the force at different points on the touch-sensitive surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change thereof of the contact area detected on the touch-sensitive surface, the capacitance and / or change thereof of the touch-sensitive surface near the contact, and / or the resistance and / or change thereof of the touch-sensitive surface near the contact are optionally used as a surrogate for the force or pressure of a contact on the touch-sensitive surface. In some embodiments, the surrogate measurement of the contact force or pressure is directly used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measurement). In some embodiments, the surrogate measurement of the contact force or pressure is converted into an estimated force or pressure, and the estimated force or pressure is used to determine whether the intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of a contact as an attribute of a user input allows a user to access additional device functions that would otherwise be inaccessible on a smaller device with a limited footprint, the smaller device being used to (e.g., on a touch-sensitive display) display affordances and / or receive user input (e.g., via a touch-sensitive display, a touch-sensitive surface, or physical / mechanical controls such as a knob or button).

[0079] As used in this specification and the claims, the term "haptic output" refers to a physical displacement of the device relative to a previous position of the device detected by a user using the user's sense of touch, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., the housing), or a displacement of the component relative to the center of mass of the device. For example, in the case of contact between the device or a component of the device and a surface of the user sensitive to touch (e.g., a finger, palm, or other part of the user's hand), a haptic output generated by the physical displacement will be interpreted by the user as a sense of touch corresponding to a perceived change in a physical characteristic of the device or the component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or a touchpad) may optionally be interpreted by the user as a "press click" or "release click" of a physical actuation button. In some cases, the user will sense a haptic sensation, such as a "press click" or "release click", even when the physical actuation button associated with the touch-sensitive surface that is physically depressed (e.g., displaced) by the user's movement does not move. As another example, even when there is no change in the smoothness of the touch-sensitive surface, movement of the touch-sensitive surface may optionally be interpreted or sensed by the user as "roughness" of the touch-sensitive surface. Although such interpretations of touch by the user will be limited by the user's individual sensory perception, many sensory perceptions of touch are common to most users. Thus, when a haptic output is described as corresponding to a particular sensory perception of the user (e.g., "press click", "release click", "roughness"), unless otherwise stated, the generated haptic output corresponds to a physical displacement of the device or its component that would generate the described sensory perception of a typical (or ordinary) user.

[0080] It should be understood that device 200 is merely an example of a portable multifunctional device, and device 200 optionally has more or fewer components than those shown, optionally combines two or more components, or optionally has different configurations or arrangements of these components. Figure 2A The various components shown are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.

[0081] Memory 202 includes one or more computer-readable storage media. These computer-readable storage media are tangible and non-transitory, for example. Memory 202 includes high-speed random access memory and also includes non-volatile memory such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. A memory controller 222 controls access to memory 202 by other components of device 200.

[0082] In some examples, the non-transitory computer-readable storage medium of the memory 202 is used to store instructions (e.g., for performing aspects of the processes described below) for use by or in conjunction with an instruction execution system, apparatus, or device such as a computer-based system, a system including a processor, or other systems that can retrieve and execute instructions from the instruction execution system, apparatus, or device. In other examples, the instructions (e.g., for performing aspects of the processes described below) are stored on the non-transitory computer-readable storage medium of the server system 108, or are divided between the non-transitory computer-readable storage medium of the memory 202 and the non-transitory computer-readable storage medium of the server system 108.

[0083] The peripheral device interface 218 is used to couple the input and output peripheral devices of the device to the CPU 220 and the memory 202. The one or more processors 220 run or execute various software programs and / or instruction sets stored in the memory 202 to perform the various functions of the device 200 and process data. In some embodiments, the peripheral device interface 218, the CPU 220, and the memory controller 222 are implemented on a single chip such as chip 204. In some other embodiments, they are implemented on separate chips.

[0084] The RF (Radio Frequency) circuit 208 receives and transmits RF signals, which are also referred to as electromagnetic signals. The RF circuit 208 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with a communication network and other communication devices via electromagnetic signals. The RF circuit 208 optionally includes well-known circuits for performing these functions, including but not limited to an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a codec chipset, a subscriber identity module (SIM) card, a memory, and so on. The RF circuit 208 optionally communicates with the network and other devices via wireless communication, and these networks are such as the Internet (also known as the World Wide Web (WWW)), an intranet, and / or a wireless network (such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN)). The RF circuit 208 optionally includes well-known circuits for detecting a near field communication (NFC) field, such as detecting via a short-range communication radio component. The wireless communication optionally uses any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), Evolution-Data Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPDA), Long-Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wi-Fi (e.g., IEEE802.11a, IEEE 802.11b, IEEE 802.11g, IEEE802.11n, and / or IEEE802.11ac), Voice over Internet Protocol (VoIP), WiMAX, email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the date of submission of this document.

[0085] The audio circuit 210, speaker 211, and microphone 213 provide an audio interface between the user and the device 200. The audio circuit 210 receives audio data from the peripheral interface 218, converts the audio data into an electrical signal, and transmits the electrical signal to the speaker 211. The speaker 211 converts the electrical signal into sound waves audible to humans. The audio circuit 210 also receives the electrical signal converted from sound waves by the microphone 213. The audio circuit 210 converts the electrical signal into audio data and transmits the audio data to the peripheral interface 218 for processing. The audio data is retrieved from and / or transmitted to the memory 202 and / or the RF circuit 208 via the peripheral interface 218. In some embodiments, the audio circuit 210 also includes an earphone jack (e.g., Figure 3 312 in

[0086] ). The earphone jack provides an interface between the audio circuit 210 and a removable audio input / output peripheral device, such as an output-only headset or an earphone with both output (e.g., a mono or stereo headset) and input (e.g., a microphone). Figure 3 One or more buttons (e.g., Figure 3 308 in

[0087] ) optionally include increase / decrease buttons for volume control of the speaker 211 and / or the microphone 213. One or more buttons optionally include a push button (e.g., Figure 3 306 in

[0087] ).

[0087] Rapidly pressing the depress button optionally unlocks the touch screen 212 or optionally begins the process of unlocking the device using gestures on the touch screen, as described in U.S. Patent Application No. 11 / 322,549, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," filed on December 23, 2005 (i.e., U.S. Patent No. 7,657,849), which is hereby incorporated by reference in its entirety. Long pressing the depress button (e.g., 306) optionally powers on or powers off the device 200. The functions of one or more of the buttons can optionally be customized by the user. The touch screen 212 is used to implement virtual buttons or soft buttons and one or more soft keyboards.

[0088] The touch-sensitive display 212 provides an input interface and an output interface between the device and the user. The display controller 256 receives electrical signals from the touch screen 212 and / or sends electrical signals to the touch screen 112. The touch screen 212 displays visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.

[0089] The touch screen 212 has a touch-sensitive surface, sensor, or group of sensors that accepts input from the user based on haptic and / or tactile contact. The touch screen 212 and the display controller 256 (along with any associated modules and / or instruction sets in the memory 202) detect contact (and any movement or interruption of the contact) on the touch screen 212 and convert the detected contact into an interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 212. In an exemplary embodiment, the point of contact between the touch screen 212 and the user corresponds to the user's finger.

[0090] The touch screen 212 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, but uses other display technologies in other embodiments. The touch screen 212 and the display controller 256 optionally use any of a variety of touch sensing technologies now known or later developed, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 212 to detect contact and any movement or interruption thereof, the variety of touch sensing technologies including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as in and iPod The technology used in

[0091] In some embodiments of the touchscreen 212, the touch-sensitive display is optionally similar to the multi-touch sensitive touchpad described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.) and / or 6,677,932 (Westerman et al.) and / or U.S. Patent Publication 2002 / 0015024A1, each of which is hereby incorporated by reference in its entirety. However, the touchscreen 212 displays the visual output from the device 200, while the touch-sensitive touchpad does not provide a visual output.

[0092] In some embodiments, the touch-sensitive display of the touch screen 212 is as described in the following patent applications: (1) U.S. Patent Application No. 11 / 381,313, entitled "Multipoint Touch Surface Controller," filed on May 2, 2006; (2) U.S. Patent Application No. 10 / 840,862, entitled "Multipoint Touchscreen," filed on May 6, 2004; (3) U.S. Patent Application No. 10 / 903,964, entitled "Gestures For TouchSensitive Input Devices," filed on July 30, 2004; (4) U.S. Patent Application No. 11 / 048,264, entitled "Gestures For Touch Sensitive Input Devices," filed on January 31, 2005; (5) U.S. Patent Application No. 11 / 038,590, entitled "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices," filed on January 18, 2005; (6) U.S. Patent Application No. 11 / 228,758, entitled "Virtual Input Device Placement On A Touch Screen UserInterface," filed on September 16, 2005; (7) U.S. Patent Application No. 11 / 228,700, entitled "Operation Of AComputer With A Touch ScreenInterface," filed on September 16, 2005; (8) U.S. Patent Application No. 11 / 228,737, entitled "Activating Virtual Keys Of A Touch-Screen VirtualKeyboard," filed on September 16, 2005; and (9) U.S. Patent Application No. 11 / 367,749, entitled "Multi-Functional Hand-Held Device," filed on March 3, 2006. All of these applications are hereby incorporated by reference in their entirety.

[0093] The touch screen 212 optionally has a video resolution of more than 100 dpi. In some embodiments, the touch screen has a video resolution of about 160 dpi. The user optionally uses any suitable object or attachment such as a stylus, finger, etc. to contact the touch screen 212. In some embodiments, the user interface is designed to work primarily through finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts the rough finger-based input into an accurate pointer / cursor position or command for performing the action desired by the user.

[0094] In some embodiments, in addition to the touch screen, the device 200 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device, which, unlike the touch screen, does not display a visual output. The touchpad is optionally a touch-sensitive surface separate from the touch screen 212 or an extension of the touch-sensitive surface formed by the touch screen.

[0095] The device 200 also includes a power system 262 for powering various components. The power system 262 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., a light-emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.

[0096] The device 200 optionally also includes one or more optical sensors 264. Figure 2AAn optical sensor coupled to an optical sensor controller 258 in the I / O subsystem 206 is shown. The optical sensor 264 includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. The optical sensor 264 optionally receives light projected through one or more lenses from the environment and converts the light into data representative of an image. In conjunction with the imaging module 243 (also referred to as the camera module), the optical sensor 264 captures still images or video. In some embodiments, the optical sensor is located on the rear of the device 200, opposite the touch screen display 212 on the front of the device, such that the touch screen display can be used as a viewfinder for still image and / or video image capture. In some embodiments, the optical sensor is located on the front of the device such that an image of the user can optionally be acquired while the user views other video conference participants on the touch screen display, for video conferencing. In some embodiments, the position of the optical sensor 264 can be changed by the user (e.g., by rotating the lens and sensor in the device housing), such that a single optical sensor 264 can be used with the touch screen display for both video conferencing and still image and / or video image capture.

[0097] The device 200 optionally further includes one or more depth camera sensors 275. Figure 2A A depth camera sensor coupled to a depth camera controller 269 in the I / O subsystem 206 is shown. The depth camera sensor 275 receives data from the environment to create a three-dimensional model of an object (e.g., a face) within a scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with the imaging module 243 (also referred to as the camera module), the depth camera sensor 275 is optionally used to determine depth maps of different portions of an image captured by the imaging module 243. In some embodiments, the depth camera sensor is located on the front of the device 200 such that an image of the user with depth information can optionally be acquired while the user views other video conference participants on the touch screen display, for video conferencing, and a selfie with depth map data can be captured. In some embodiments, the depth camera sensor 275 is located on the back of the device, or on both the back and the front of the device 200. In some embodiments, the position of the depth camera sensor 275 can be changed by the user (e.g., by rotating the lens and sensor in the device housing), such that the depth camera sensor 275 can be used with the touch screen display for both video conferencing and still image and / or video image capture.

[0098] In some embodiments, a depth map (e.g., a depth map image) contains information (e.g., values) related to the distance of objects in a scene from a viewpoint (e.g., a camera, an optical sensor, a depth camera sensor). In one embodiment of a depth map, each depth pixel defines the position in the Z-axis of the viewpoint where its corresponding two-dimensional pixel is located. In some embodiments, a depth map is composed of pixels, where each pixel is defined by a value (e.g., from 0 to 255). For example, a value of "0" represents a pixel that is farthest from the viewpoint (e.g., a camera, an optical sensor, a depth camera sensor) in a "three-dimensional" scene, and a value of "255" represents a pixel that is closest to the viewpoint in the "three-dimensional" scene. In other embodiments, a depth map represents the distance between an object in a scene and a plane of the viewpoint. In some embodiments, a depth map includes information about the relative depth of various features of an object of interest in the field of view of a depth camera (e.g., the relative depth of the eyes, nose, mouth, ears of a user's face). In some embodiments, a depth map includes information that enables a device to determine the profile of an object of interest in the z-direction.

[0099] Device 200 optionally further includes one or more contact intensity sensors 265. Figure 2A Shown is a contact intensity sensor coupled to an intensity sensor controller 259 in I / O subsystem 206. Contact intensity sensor 265 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electric power sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). Contact intensity sensor 265 receives contact intensity information (e.g., pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 212). In some embodiments, at least one contact intensity sensor is located on the rear of device 200 that is opposite to the touchscreen display 212 located on the front of device 200.

[0100] Device 200 optionally further includes one or more proximity sensors 266. Figure 2AA proximity sensor 266 is shown coupled to the peripheral device interface 218. Alternatively, the proximity sensor 266 is optionally coupled to the input controller 260 in the I / O subsystem 206. The proximity sensor 266 optionally operates as described in the following U.S. patent applications: 11 / 241,839, titled "Proximity Detector In Handheld Device"; No. 11 / 240,788, titled "Proximity Detector In Handheld Device"; No. 11 / 620,702, titled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; No. 11 / 586,862, titled "Automated Response To And Sensing Of User Activity In Portable Devices"; and No. 11 / 638,251, titled "Methods And Systems For Automatic Configuration Of Peripherals", which U.S. patent applications are hereby incorporated by reference in their entirety. In some embodiments, when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call), the proximity sensor turns off and disables the touch screen 212.

[0101] Device 200 optionally further includes one or more haptic output generators 267. Figure 2AIllustrated is a haptic output generator coupled to a haptic feedback controller 261 in an I / O subsystem 206. The haptic output generator 267 optionally includes one or more electroacoustic devices, such as speakers or other audio components; and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components that convert an electrical signal into a haptic output on the device). A contact intensity sensor 265 receives haptic feedback generation instructions from a haptic feedback module 233 and generates a haptic output on the device 200 that can be felt by a user of the device 200. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a touch-sensitive surface (e.g., a touch-sensitive display system 212), and optionally generates a haptic output by moving the touch-sensitive surface vertically (e.g., into / out of the surface of the device 200) or laterally (e.g., backward and forward in the same plane as the surface of the device 200). In some embodiments, at least one haptic output generator sensor is located on the rear of the device 200 opposite the touch screen display 212 located on the front of the device 200.

[0102] The device 200 optionally further includes one or more accelerometers 268. Figure 2A Illustrated is an accelerometer 268 coupled to a peripheral device interface 218. Alternatively, the accelerometer 268 is optionally coupled to an input controller 260 in the I / O subsystem 206. The accelerometer 268 optionally operates as described in the following U.S. Patent Publications: U.S. Patent Publication 20050190059, entitled "Acceleration-based Theft Detection System for Portable Electronic Devices" and U.S. Patent Publication 20060017692, entitled "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer", both of which are incorporated herein by reference in their entireties. In some embodiments, information is displayed in a portrait or landscape view on the touch screen display based on an analysis of data received from one or more accelerometers. In addition to the accelerometer 268, the device 200 optionally further includes a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for obtaining information about the location and orientation (e.g., portrait or landscape) of the device 200.

[0103] In some embodiments, the software components stored in the memory 202 include an operating system 226, a communication module (or instruction set) 228, a touch / motion module (or instruction set) 230, a graphics module (or instruction set) 232, a text input module (or instruction set) 234, a Global Positioning System (GPS) module (or instruction set) 235, a digital assistant client module 229, and application programs (or instruction sets) 236. In addition, the memory 202 stores data and models, such as user data and models 231. Further, in some embodiments, the memory 202 ( Figure 2A ) or 470 ( Figure 4 ) stores a device / global internal state 257, as Figure 2A and Figure 4 shown in. The device / global internal state 257 includes one or more of the following: an active application state, which indicates which applications (if any) are currently active; a display state, which indicates what applications, views, or other information occupy the various regions of the touch screen display 212; a sensor state, including information obtained from the various sensors and input control devices 216 of the device; and location information regarding the location and / or orientation of the device.

[0104] The operating system 226 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.), and facilitates communication between the various hardware components and software components.

[0105] The communication module 228 facilitates communication with other devices via one or more external ports 224, and also includes various software components for processing data received by the RF circuit 208 and / or the external ports 224. The external ports 224 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices or indirectly coupled via a network (e.g., the Internet, a wireless LAN, etc.). In some embodiments, the external port is the same as or similar to and / or compatible with the 30-pin connector used on (a trademark of Apple Inc.) devices and / or a multi-pin (e.g., 30-pin) connector.

[0106] The contact / motion module 230 optionally detects contact with the touchscreen 212 (in conjunction with the display controller 256) and other touch-sensitive devices (e.g., a touchpad or a physical click wheel). The contact / motion module 230 includes various software components for performing various operations related to contact detection, such as determining whether contact has occurred (e.g., detecting a finger press event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there is movement of the contact and tracking the movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or contact break). The contact / motion module 230 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of the contact point being represented by a series of contact data. These operations are optionally applied to single-point contact (e.g., single-finger contact) or multi-point simultaneous contact (e.g., "multi-touch" / multiple finger contact). In some embodiments, the contact / motion module 230 and the display controller 256 detect contact on the touchpad.

[0107] In some embodiments, the contact / motion module 230 uses a set of one or more intensity thresholds to determine whether the user has performed an operation (e.g., determining whether the user has "clicked" on an icon). In some embodiments, at least a subset of the intensity thresholds is determined based on software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator and can be adjusted without changing the physical hardware of the device 200). For example, the mouse "click" threshold of the touchpad or touchscreen can be set to any one of a wide range of predefined thresholds without changing the touchpad or touchscreen display hardware. Additionally, in some implementations, software settings are provided to the user of the device for adjusting one or more of the intensity thresholds in a set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by using a system-level click on an "intensity" parameter to adjust multiple intensity thresholds at once).

[0108] The contact / motion module 230 optionally detects gesture inputs made by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timing, and / or intensities of the detected contacts). Thus, gestures are optionally detected by detecting a particular contact pattern. For example, detecting a finger tap gesture includes detecting a finger press event and then detecting a finger lift (lift-off) event at the same location (or substantially the same location) as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift (lift-off) event.

[0109] The graphics module 232 includes various known software components for presenting and displaying graphics on the touch screen 212 or other display, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual characteristics). As used herein, the term "graphics" includes any object that can be displayed to a user, including, without limitation, text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.

[0110] In some embodiments, the graphics module 232 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 232 receives one or more codes specifying the graphics to be displayed from an application, etc., and also receives coordinate data and other graphic attribute data as necessary, and then generates screen image data for output to the display controller 256.

[0111] The haptic feedback module 233 includes various software components for generating instructions that are used by the haptic output generator 267 to generate haptic output at one or more locations on the device 200 in response to user interaction with the device 200.

[0112] The text input module 234, which is optionally a component of the graphics module 232, provides a soft keyboard for entering text in various applications (e.g., contacts 237, email 240, IM 241, browser 247, and any other application that requires text input).

[0113] The GPS module 235 determines the location of the device and provides this information for use in various applications (e.g., provided to the phone 238 for use in location-based dialing; provided to the camera 243 as picture / video metadata; and provided to applications that provide location-based services, such as weather widgets, local yellow pages widgets, and map / navigation widgets).

[0114] The digital assistant client module 229 includes various client - side digital assistant instructions to provide client - side functionality of the digital assistant. For example, the digital assistant client module 229 is capable of receiving voice input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces of the portable multifunctional device 200 (e.g., microphone 213, accelerometer 268, touch - sensitive display system 212, optical sensor 229, other input control devices 216, etc.). The digital assistant client module 229 is also capable of providing output in audio form (e.g., voice output), visual form, and / or tactile form through various output interfaces of the portable multifunctional device 200 (e.g., speaker 211, touch - sensitive display system 212, tactile output generator 267, etc.). For example, the output is provided as voice, sound, alert, text message, menu, graphic, video, animation, vibration, and / or a combination of two or more of the above. During operation, the digital assistant client module 229 communicates with the DA server 106 using the RF circuit 208.

[0115] The user data and model 231 includes various data associated with the user (e.g., user - specific vocabulary data, user preference data, user - specified name pronunciations, data from the user's electronic address book, to - do lists, shopping lists, etc.) to provide client - side functionality of the digital assistant. In addition, the user data and model 231 includes various models for processing user input and determining user intent (e.g., speech recognition models, statistical language models, natural language processing models, knowledge ontologies, task flow models, service models, etc.).

[0116] In some examples, the digital assistant client module 229 utilizes various sensors, subsystems, and peripherals of the portable multifunctional device 200 to collect additional information from the surrounding environment of the portable multifunctional device 200 to establish context associated with the user, the current user interaction, and / or the current user input. In some examples, the digital assistant client module 229 provides context information or a subset thereof together with the user input to the DA server 106 to assist in inferring user intent. In some examples, the digital assistant also uses the context information to determine how to prepare and deliver the output to the user. The context information is referred to as context data.

[0117] In some examples, the context information accompanying the user input includes sensor information, such as lighting, ambient noise, ambient temperature, an image or video of the surrounding environment, etc. In some examples, the context information may further include the physical state of the device, such as device orientation, device location, device temperature, power level, speed, acceleration, motion pattern, cellular signal strength, etc. In some examples, information related to the software state of the DA server 106, such as the running process of the portable multifunctional device 200, installed programs, past and current network activities, background services, error logs, resource usage, etc., is provided as context information associated with the user input to the DA server 106.

[0118] In some examples, the digital assistant client module 229 selectively provides information stored on the portable multifunctional device 200 (e.g., user data 231) in response to a request from the DA server 106. In some examples, the digital assistant client module 229 also solicits additional input from the user via natural language conversation or other user interfaces when requested by the DA server 106. The digital assistant client module 229 transmits this additional input to the DA server 106 to assist the DA server 106 in intent inference and / or fulfilling the user intent expressed in the user request.

[0119] The following Figure 7A -C provides a more detailed description of the digital assistant. It should be recognized that the digital assistant client module 229 may include any number of sub-modules of the digital assistant module 726 described below.

[0120] The application 236 optionally includes the following modules (or instruction sets) or subsets or supersets thereof:

[0121] · A contacts module 237 (sometimes referred to as an address book or contacts list);

[0122] · A phone module 238;

[0123] · A video conferencing module 239;

[0124] · An email client module 240;

[0125] · An instant messaging (IM) module 241;

[0126] · A fitness support module 242;

[0127] · A camera module 243 for still images and / or video images;

[0128] · An image management module 244;

[0129] · A video player module;

[0130] · Music player module;

[0131] · Browser module 247;

[0132] · Calendar module 248;

[0133] · Desktop widget module 249, which optionally includes one or more of the following: weather desktop widget 249-1, stock market desktop widget 249-2, calculator desktop widget 249-3, alarm clock desktop widget 249-4, dictionary desktop widget 249-5, and other desktop widgets obtained by the user, as well as user-created desktop widget 249-6;

[0134] · Desktop widget creator module 250 for forming user-created desktop widget 249-6;

[0135] · Search module 251;

[0136] · Video and music player module 252, which combines a video player module and a music player module;

[0137] · Notepad module 253;

[0138] · Map module 254; and / or

[0139] · Online video module 255.

[0140] Examples of other applications 236 that are optionally stored in the memory 202 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-supported applications, encryption, digital rights management, speech recognition, and speech replication.

[0141] In combination with the touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the contact module 237 is optionally used to manage an address book or contact list (e.g., the application internal state 292 of the contact module 237 stored in the memory 202 or memory 470), including: adding one or more names to the address book; deleting names from the address book; associating a phone number, email address, physical address, or other information with a name; associating an image with a name; categorizing and classifying names; providing a phone number or email address to initiate and / or facilitate communication via the phone 238, video conferencing module 239, email 240, or IM 241; and so on.

[0142] In combination with RF circuit 208, audio circuit 210, speaker 211, microphone 213, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the telephone module 238 is optionally used to input a character sequence corresponding to a telephone number, access one or more telephone numbers in the contact module 237, modify the entered telephone number, dial the corresponding telephone number, conduct a session, and disconnect or hang up when the session is completed. As described above, the wireless communication optionally uses any one of a variety of communication standards, protocols, and technologies.

[0143] In combination with RF circuit 208, audio circuit 210, speaker 211, microphone 213, touch screen 212, display controller 256, optical sensor 264, optical sensor controller 258, contact / motion module 230, graphics module 232, text input module 234, contact module 237, and telephone module 238, the video conferencing module 239 includes executable instructions for initiating, conducting, and terminating a video conference between the user and one or more other participants according to user instructions.

[0144] In combination with RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the email client module 240 includes executable instructions for creating, sending, receiving, and managing emails in response to user instructions. In combination with the image management module 244, the email client module 240 makes it very easy to create and send emails with static images or video images captured by the camera module 243.

[0145] In combination with RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the instant messaging module 241 includes executable instructions for: inputting a character sequence corresponding to an instant message, modifying a previously entered character, transmitting the corresponding instant message (e.g., using the Short Message Service (SMS) or Multimedia Message Service (MMS) protocol for phone-based instant messaging or using XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing the received instant messages. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant message" refers to both phone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).

[0146] In combination with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, GPS module 235, map module 254, and music player module, the fitness support module 242 includes executable instructions for: creating a fitness routine (e.g., with time, distance, and / or calorie burn goals); communicating with fitness sensors (exercise equipment); receiving fitness sensor data; calibrating sensors for monitoring fitness; selecting and playing music for fitness; and displaying, storing, and transmitting fitness data.

[0147] In combination with the touch screen 212, display controller 256, one or more optical sensors 264, optical sensor controller 258, contact / motion module 230, graphics module 232, and image management module 244, the camera module 243 includes executable instructions for: capturing still images or videos (including video streams) and storing them in the memory 202, modifying the characteristics of still images or videos, or deleting still images or videos from the memory 202.

[0148] In combination with the touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and camera module 243, the image management module 244 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, tagging, deleting, presenting (e.g., in a digital slide show or album), and storing still images and / or video images.

[0149] In combination with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the browser module 247 includes executable instructions for browsing the Internet according to user instructions (including searching, linking to, receiving, and displaying web pages or portions thereof, and linking to attachments and other files of web pages).

[0150] In combination with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, email client module 240, and browser module 247, the calendar module 248 includes executable instructions for creating, displaying, modifying, and storing calendars and data associated with the calendars (e.g., calendar entries, to-do items, etc.) according to user instructions.

[0151] In combination with the RF circuit 208, touch screen 212, display system controller 256, contact / motion module 230, graphics module 232, text input module 234, and browser module 247, the desktop widget module 249 is a mini - application (e.g., weather desktop widget 249 - 1, stock market desktop widget 249 - 2, calculator desktop widget 249 - 3, alarm clock desktop widget 249 - 4, and dictionary desktop widget 249 - 5) optionally downloaded and used by the user or a mini - application created by the user (e.g., user - created desktop widget 249 - 6). In some embodiments, the desktop widget includes an HTML (HyperText Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, the desktop widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! desktop widget).

[0152] In combination with the RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and browser module 247, the desktop widget creator module 250 is optionally used by the user to create a desktop widget (e.g., transfer a user - specified portion of a web page into a desktop widget).

[0153] In combination with the touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the search module 251 includes executable instructions for searching the memory 202 for text, music, sound, images, video, and / or other files that match one or more search criteria (e.g., one or more user - specified search terms) according to a user instruction.

[0154] In combination with the touch screen 212, display controller 256, contact / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, and browser module 247, the video and music player module 252 includes executable instructions that allow the user to download and play back recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, and executable instructions for displaying, presenting, or otherwise playing back video (e.g., on the touch screen 212 or on an external display connected via the external port 224). In some embodiments, the device 200 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0155] In combination with the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, and the text input module 234, the notepad module 253 includes executable instructions for creating and managing notepads, to-do lists, etc. according to user instructions.

[0156] In combination with the RF circuit 208, the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, the text input module 234, the GPS module 235, and the browser module 247, the map module 254 optionally receives, displays, modifies, and stores maps and data associated with the maps (e.g., driving directions, data related to stores and other points of interest at or near a particular location, and other location-based data) according to user instructions.

[0157] In combination with the touch screen 212, the display controller 256, the contact / motion module 230, the graphics module 232, the audio circuit 210, the speaker 211, the RF circuit 208, the text input module 234, the email client module 240, and the browser module 247, the online video module 255 includes instructions that allow a user to access, browse, receive (e.g., by streaming and / or downloading), play back (e.g., on the touch screen or on an external display connected via the external port 224), send an email with a link to a particular online video, and otherwise manage online videos in one or more file formats (such as, H.264). In some embodiments, the instant message module 241 is used instead of the email client module 240 to send a link to a particular online video. Other descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed on June 20, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos" and U.S. Patent Application No. 11 / 968,067, filed on December 31, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos", the contents of both of which are hereby incorporated by reference in their entirety.

[0158] Each of the above modules and applications corresponds to a set of executable instructions for performing one or more of the above functions and the methods described in this patent application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) need not be implemented as separate software programs, processes, or modules, and thus various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. For example, the video player module is optionally combined with the music player module into a single module (e.g., Figure 2A the video and music player module 252 in

[0159] In some embodiments, the device 200 is a device in which a predefined set of functions on the device is uniquely performed via a touchscreen and / or a touchpad. By using the touchscreen and / or the touchpad as the primary input control device for operating the device 200, the number of physical input control devices (e.g., push buttons, dials, etc.) on the device 200 is optionally reduced.

[0160] The predefined set of functions uniquely performed via the touchscreen and / or the touchpad optionally includes navigating between user interfaces. In some embodiments, when the user touches the touchpad, the device 200 is navigated from any user interface displayed on the device 200 to the main menu, the home menu, or the root menu. In such embodiments, the touchpad is used to implement the "menu button". In some other embodiments, the menu button is a physical push button or other physical input control device rather than the touchpad.

[0161] Figure 2B is a block diagram showing exemplary components for event processing according to some embodiments. In some embodiments, the memory 202 ( Figure 2A ) or the memory 470 ( Figure 4 ) includes an event classifier 270 (e.g., in the operating system 226) and a corresponding application 236-1 (e.g., any one of the foregoing applications 237-251, 255, 480-490).

[0162] The event classifier 270 receives event information and determines the application 236-1 to which the event information is to be delivered and the application view 291 of the application 236-1. The event classifier 270 includes an event monitor 271 and an event dispatcher module 274. In some embodiments, the application 236-1 includes an application internal state 292 that indicates one or more current application views that are displayed on the touch-sensitive display 212 when the application is active or executing. In some embodiments, the device / global internal state 257 is used by the event classifier 270 to determine which application(s) is / are currently active, and the application internal state 292 is used by the event classifier 270 to determine the application view 291 to which the event information is to be delivered.

[0163] In some embodiments, the application internal state 292 includes additional information such as one or more of the following: recovery information that will be used when the application 236-1 resumes execution, user interface state information indicating information being displayed or ready to be displayed by the application 236-1, a state queue for enabling the user to return to a previous state or view of the application 236-1, and a repeat / undo queue of previous actions taken by the user.

[0164] The event monitor 271 receives event information from the peripheral device interface 218. The event information includes information about sub-events (e.g., a user touch on the touch-sensitive display 212 as part of a multi-touch gesture). The peripheral device interface 218 transmits the information it receives from the I / O subsystem 206 or sensors such as the proximity sensor 266, the accelerometer 268, and / or the microphone 213 (via the audio circuit 210). The information that the peripheral device interface 218 receives from the I / O subsystem 206 includes information from the touch-sensitive display 212 or a touch-sensitive surface.

[0165] In some embodiments, the event monitor 271 sends requests to the peripheral device interface 218 at pre-determined intervals. In response, the peripheral device interface 218 transmits event information. In other embodiments, the peripheral device interface 218 transmits event information only when there is a significant event (e.g., an input received above a pre-determined noise threshold and / or an input received for longer than a pre-determined duration).

[0166] In some embodiments, the event classifier 270 further includes a hit view determination module 272 and / or an active event recognizer determination module 273.

[0167] When the touch-sensitive display 212 displays more than one view, the hit view determination module 272 provides a software process for determining where within one or more of the views a sub-event has occurred. A view consists of the controls and other elements that a user can see on the display.

[0168] Another aspect of the user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected optionally corresponds to a programmatic level within the programmatic or view hierarchy of the application. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events recognized as correct inputs is optionally determined at least in part based on the hit view of the initial touch that begins the touch-based gesture.

[0169] The hit view determination module 272 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized in a hierarchy, the hit view determination module 272 identifies the hit view as the lowest view in the hierarchy that should process the sub-event. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events that form an event or potential event) occurs. Once the hit view is identified by the hit view determination module 272, the hit view generally receives all sub-events related to the same touch or input source for which it was identified as the hit view.

[0170] The active event recognizer determination module 273 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, the active event recognizer determination module 273 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, the active event recognizer determination module 273 determines that all views that include the physical location of the sub-event are actively participating views and thus determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with a particular view, higher views in the hierarchy will still remain as actively participating views.

[0171] The event dispatcher module 274 distributes event information to event recognizers (e.g., event recognizer 280). In embodiments that include the active event recognizer determination module 273, the event dispatcher module 274 delivers the event information to the event recognizer determined by the active event recognizer determination module 273. In some embodiments, the event dispatcher module 274 stores the event information in an event queue, which is retrieved by the corresponding event receiver 282.

[0172] In some embodiments, the operating system 226 includes an event classifier 270. Alternatively, the application 236-1 includes an event classifier 270. In yet another embodiment, the event classifier 270 is an independent module or part of another module (such as, the contact / motion module 230) stored in the memory 202.

[0173] In some embodiments, the application 236-1 includes a plurality of event handlers 290 and one or more application views 291, where each application view includes instructions for handling touch events occurring within a corresponding view of the application's user interface. Each application view 291 of the application 236-1 includes one or more event recognizers 280. Typically, a corresponding application view 291 includes a plurality of event recognizers 280. In other embodiments, one or more of the event recognizers 280 are part of an independent module, such as a user interface toolkit or a higher-level object from which the application 236-1 inherits methods and other properties. In some embodiments, the corresponding event handlers 290 include one or more of the following: a data updater 276, an object updater 277, a GUI updater 278, and / or event data 279 received from the event classifier 270. The event handlers 290 utilize or call the data updater 276, the object updater 277, or the GUI updater 278 to update the application internal state 292. Alternatively, one or more of the application views 291 include one or more corresponding event handlers 290. Additionally, in some embodiments, one or more of the data updater 276, the object updater 277, and the GUI updater 278 are included in the corresponding application view 291.

[0174] The corresponding event recognizer 280 receives event information (e.g., event data 279) from the event classifier 270 and identifies an event from the event information. The event recognizer 280 includes an event receiver 282 and an event comparator 284. In some embodiments, the event recognizer 280 also includes at least a subset of metadata 283 and event delivery instructions 288 (which optionally include sub-event delivery instructions).

[0175] The event receiver 282 receives event information from the event classifier 270. The event information includes information about sub-events such as a touch or a touch movement. Depending on the sub-event, the event information also includes additional information such as the location of the sub-event. When the sub-event involves the movement of a touch, the event information optionally also includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a landscape orientation, or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the device pose).

[0176] The event comparator 284 compares the event information with predefined event or sub-event definitions and, based on the comparison, determines the event or sub-event or determines or updates the state of the event or sub-event. In some embodiments, the event comparator 284 includes an event definition 286. The event definition 286 contains definitions of events (e.g., predefined sequences of sub-events), such as event 1 (287-1), event 2 (287-2), and other events. In some embodiments, the sub-events in an event (287) include, for example, touch start, touch end, touch movement, touch cancellation, and multi-touch. In one example, the definition of event 1 (287-1) is a double-tap on a displayed object. For example, a double-tap includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift-off (touch end) of a predetermined duration, a second touch (touch start) of a predetermined duration on the displayed object, and a second lift-off (touch end) of a predetermined duration. In another example, the definition of event 2 (287-2) is a drag on a displayed object. For example, a drag includes a touch (or contact) of a predetermined duration on the displayed object, movement of the touch on the touch-sensitive display 212, and lift-off of the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 290.

[0177] In some embodiments, the event definition 287 includes definitions of events for corresponding user interface objects. In some embodiments, the event comparator 284 performs a hit test to determine which user interface object is associated with the sub-event. For example, in an application view that displays three user interface objects on the touch-sensitive display 212, when a touch is detected on the touch-sensitive display 212, the event comparator 284 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 290, the event comparator uses the result of the hit test to determine which event handler 290 should be activated. For example, the event comparator 284 selects the event handler associated with the sub-event and the object that triggered the hit test.

[0178] In some embodiments, the definition of a corresponding event (287) also includes a delay action that delays the delivery of event information until it has been determined whether a subsequence of sub-events does or does not correspond to the event type of an event recognizer.

[0179] When the corresponding event recognizer 280 determines that a subsequence of sub-events does not match any of the events in the event definition 286, the corresponding event recognizer 280 enters an event impossible, event failed, or event ended state, after which subsequent sub-events of a touch-based gesture are ignored. In such a case, any other event recognizers (if any) for which the hit view remains active continue to track and process sub-events of an ongoing touch-based gesture.

[0180] In some embodiments, the corresponding event recognizer 280 includes metadata 283 having configurable attributes, flags, and / or lists indicating how the event delivery system should perform sub-event delivery to active participating event recognizers. In some embodiments, the metadata 283 includes configurable attributes, flags, and / or lists indicating how event recognizers interact with or can interact with each other. In some embodiments, the metadata 283 includes configurable attributes, flags, and / or lists indicating whether sub-events are delivered to different levels in a view or a programmatic hierarchy.

[0181] In some embodiments, when one or more specific sub-events of an event are recognized, the corresponding event recognizer 280 activates an event handler 290 associated with the event. In some embodiments, the corresponding event recognizer 280 delivers event information associated with the event to the event handler 290. Activating the event handler 290 is different from sending (and deferring the sending of) sub-events to the corresponding hit view. In some embodiments, the event recognizer 280 throws a token associated with the recognized event, and the event handler 290 associated with the token retrieves the token and executes a predefined process.

[0182] In some embodiments, the event delivery instruction 288 includes a sub-event delivery instruction that delivers event information about a sub-event without activating the event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with the subsequence of sub-events or to an actively participating view. The event handler associated with the subsequence of sub-events or with the actively participating view receives the event information and executes a predetermined process.

[0183] In some embodiments, the data updater 276 creates and updates data used in the application 236-1. For example, the data updater 276 updates the phone numbers used in the contacts module 237, or stores the video files used in the video player module. In some embodiments, the object updater 277 creates and updates objects used in the application 236-1. For example, the object updater 277 creates new user interface objects or updates the positions of user interface objects. The GUI updater 278 updates the GUI. For example, the GUI updater 278 prepares display information and sends it to the graphics module 232 for display on the touch-sensitive display.

[0184] In some embodiments, the event handler 290 includes or has access to the data updater 276, the object updater 277, and the GUI updater 278. In some embodiments, the data updater 276, the object updater 277, and the GUI updater 278 are included in a single module of the corresponding application 236-1 or application view 291. In other embodiments, they are included in two or more software modules.

[0185] It should be understood that the above discussion of event handling for user touches on the touch-sensitive display also applies to other forms of user input for operating the multifunctional device 200 using an input device, and not all user input is initiated on the touch screen. For example, mouse movement and mouse button presses optionally in cooperation with single or multiple keyboard presses or holds; contact movement on a touchpad, such as tapping, dragging, scrolling, etc.; stylus input; movement of the device; voice commands; detected eye movement; biometric input; and / or any combination thereof are optionally used as inputs corresponding to sub-events that define the events to be recognized.

[0186] Figure 3FIG. 0 shows a portable multifunctional device 200 having a touch screen 212, in accordance with some embodiments. The touch screen optionally displays one or more graphics within a user interface (UI) 300. In this and other embodiments described below, a user is able to select one or more of these graphics by making gestures on the graphics, for example, by using one or more fingers 302 (not drawn to scale in the figures) or one or more styli 303 (not drawn to scale in the figures). In some embodiments, selection of one or more graphics occurs when the user breaks contact with the one or more graphics. In some embodiments, the gestures optionally include one or more taps, one or more swipes (from left to right, right to left, up, and / or down), and / or rolling of a finger that has made contact with the device 200 (from right to left, left to right, up, and / or down). In some implementations or in some cases, inadvertently contacting a graphic does not select the graphic. For example, when the gesture corresponding to selection is a tap, a swipe gesture that sweeps over an application icon optionally does not select the corresponding application.

[0187] The device 200 optionally further includes one or more physical buttons, such as a "home" button, or a menu button 304. As previously described, the menu button 304 is optionally used to navigate to any of a set of applications 236 that are optionally executed on the device 200. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI that is displayed on the touch screen 212.

[0188] In some embodiments, the device 200 includes a touch screen 212, a menu button 304, a push button 306 for powering on / off the device and for locking the device, one or more volume adjustment buttons 308, a subscriber identity module (SIM) card slot 310, a headset jack 312, and a docking / charging external port 224. The push button 306 is optionally used to power on / off the device by pressing the button and holding the button in the pressed state for a predefined time interval; to lock the device by pressing the button and releasing the button before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlocking process. In an alternative embodiment, the device 200 also accepts voice input for activating or deactivating certain functions via a microphone 213. The device 200 also optionally includes one or more contact intensity sensors 265 for detecting the intensity of contact on the touch screen 212, and / or one or more tactile output generators 267 for generating tactile output for a user of the device 200.

[0189] Figure 4is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface, in accordance with some embodiments. Device 400 need not be portable. In some embodiments, device 400 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a children's learning toy), gaming system, or control device (e.g., a home or industrial controller). Device 400 generally includes one or more processing units (CPUs) 410, one or more network or other communication interfaces 460, memory 470, and one or more communication buses 420 for interconnecting these components. Communication bus 420 optionally includes circuitry (sometimes termed a chipset) for interconnecting system components and controlling the communication between them. Device 400 includes an input / output (I / O) interface 430 having a display 440, which is typically a touchscreen display. I / O interface 430 also optionally includes a keyboard and / or mouse (or other pointing device) 450 and a touchpad 455, a haptic output generator 457 for generating haptic output on device 400 (e.g., similar to one or more of the haptic output generators 267 described above with reference to Figure 2A ), sensors 459 (e.g., optical sensors, acceleration sensors, proximity sensors, touch-sensitive sensors, and / or contact intensity sensors similar to one or more of the contact intensity sensors 265 described above with reference to Figure 2A ). Memory 470 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 470 optionally includes one or more storage devices located remotely from CPU 410. In some embodiments, memory 470 stores programs, modules, and data structures similar to or a subset of the programs, modules, and data structures stored in memory 202 of portable multifunctional device 200 ( Figure 2A ). Additionally, memory 470 optionally stores additional programs, modules, and data structures not present in memory 202 of portable multifunctional device 200. For example, memory 470 of device 400 optionally stores a drawing module 480, a presentation module 482, a word processing module 484, a website creation module 486, a disk editing module 488, and / or a spreadsheet module 490, while memory 202 of portable multifunctional device 200 ( Figure 2A ) optionally does not store these modules.

[0190] Figure 4Each of the above elements in [element name] is optionally stored in one or more of the aforementioned memory devices of the memory device. Each of the above modules corresponds to an instruction set for performing the above functions. The above modules or programs (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, the memory 470 optionally stores a subset of the above modules and data structures. In addition, the memory 470 optionally stores additional modules and data structures not described above.

[0191] Attention is now turned to an embodiment of a user interface that is optionally implemented on, for example, the portable multifunctional device 200.

[0192] Figure 5A An exemplary user interface of an application menu on the portable multifunctional device 200 according to some embodiments is shown. A similar user interface is optionally implemented on the device 400. In some embodiments, the user interface 500 includes the following elements or subsets or supersets thereof:

[0193] One or more signal strength indicators 502 for one or more wireless communications (such as cellular signals and Wi-Fi signals);

[0194] · Time 504;

[0195] · Bluetooth indicator 505;

[0196] · Battery status indicator 506;

[0197] · A tray 508 with icons of common applications, such as common application icons:

[0198] ○ An icon 516 marked "Phone" of the phone module 238, which optionally includes an indicator 514 of the number of missed calls or voicemails;

[0199] ○ An icon 518 marked "Mail" of the email client module 240, which optionally includes an indicator 510 of the number of unread emails;

[0200] ○ An icon 520 marked "Browser" of the browser module 247; and

[0201] ○ An icon 522 marked "iPod" of the video and music player module 252 (also known as the iPod (trademark of Apple Inc.) module 252); and

[0202] · Icons of other applications, such as:

[0203] ○ The icon 524 of the IM module 241 marked as "Message";

[0204] ○ The icon 526 of the calendar module 248 marked as "Calendar";

[0205] ○ The icon 528 of the image management module 244 marked as "Photo";

[0206] ○ The icon 530 of the camera module 243 marked as "Camera";

[0207] ○ The icon 532 of the online video module 255 marked as "Online Video";

[0208] ○ The icon 534 of the stock market desktop applet 249-2 marked as "Stock Market";

[0209] ○ The icon 536 of the map module 254 marked as "Map";

[0210] ○ The icon 538 of the weather desktop applet 249-1 marked as "Weather";

[0211] ○ The icon 540 of the alarm clock desktop applet 249-4 marked as "Clock";

[0212] ○ The icon 542 of the fitness support module 242 marked as "Fitness Support";

[0213] ○ The icon 544 of the notepad module 253 marked as "Notepad"; and

[0214] ○ The icon 546 marked as "Settings" for setting the application or module, which provides access to the settings of the device 200 and its various applications 236.

[0215] It should be indicated that Figure 5A The icon labels shown are merely exemplary. For example, the icon 522 of the video and music player module 252 is optionally marked as "Music" or "Music Player". Other labels are optionally used for various application icons. In some embodiments, the label of the corresponding application icon includes the name of the application corresponding to the corresponding application icon. In some embodiments, the label of a specific application icon is different from the name of the application corresponding to the specific application icon.

[0216] Figure 5B A device (e.g., Figure 4 is shown having a touch-sensitive surface 551 (e.g., Figure 4An exemplary user interface on a device 400). The device 400 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 457) for detecting the intensity of contacts on the touch-sensitive surface 551, and / or one or more haptic output generators 459 for generating haptic output for a user of the device 400.

[0217] Although some examples below will be given with reference to input on a touch screen display 212 (where a touch-sensitive surface and a display are combined), in some embodiments, the device detects input on a touch-sensitive surface separate from the display, as Figure 5B shown. In some embodiments, the touch-sensitive surface (e.g., Figure 5B 551 in ) has a major axis (e.g., Figure 5B 552 in ) corresponding to the major axis (e.g., Figure 5B 553 in ) of the display (e.g., 550). According to these embodiments, the device detects contacts (e.g., Figure 5B 560 and 562 in ) with the touch-sensitive surface 551 at positions corresponding to respective positions on the display (e.g., in Figure 5B 560 corresponds to 568 and 562 corresponds to 570). Thus, when the touch-sensitive surface (e.g., Figure 5B 551 in ) is separate from the display of the multifunctional device ( Figure 5B 550 in ), user input detected by the device on the touch-sensitive surface (e.g., contacts 560 and 562 and their movements) is used by the device to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein.

[0218] Additionally, although the examples below are mainly given with reference to finger input (e.g., finger contacts, single-finger tap gestures, finger swipe gestures), it should be understood that in some embodiments, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a contact), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of a contact). As another example, a tap gesture is optionally replaced by a mouse click when the cursor is above the position of the tap gesture (e.g., instead of detecting a contact, followed by ceasing to detect the contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are optionally used simultaneously, or a mouse and a finger contact are optionally used simultaneously.

[0219] Figure 6AAn exemplary personal electronic device 600 is shown. The device 600 includes a body 602. In some embodiments, the device 600 may include some or all of the features described with respect to devices 200 and 400 (e.g., Figures 2A - 4 ). In some embodiments, the device 600 has a touch-sensitive display screen 604 hereinafter referred to as a touch screen 604. As an alternative or addition to the touch screen 604, the device 600 has a display and a touch-sensitive surface. As in the case of devices 200 and 400, in some embodiments, the touch screen 604 (or touch-sensitive surface) optionally includes one or more intensity sensors for detecting the intensity of an applied contact (e.g., a touch). One or more intensity sensors of the touch screen 604 (or touch-sensitive surface) may provide output data representative of the touch intensity. The user interface of the device 600 responds to touches based on the touch intensity, meaning that touches of different intensities may invoke different user interface operations on the device 600.

[0220] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related patent applications: International Patent Application Serial No. PCT / US2013 / 040061, filed May 8, 2013, entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application", published as WIPO Patent Publication No. WO / 2013 / 169849; and International Patent Application Serial No. PCT / US2013 / 069483, filed November 11, 2013, entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships", published as WIPO Patent Publication No. WO / 2014 / 105276, each of which is hereby incorporated by reference in its entirety.

[0221] In some embodiments, device 600 has one or more input mechanisms 606 and input mechanism 608. Input mechanisms 606 and 608 (if included) can be in physical form. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 600 has one or more attachment mechanisms. Such attachment mechanisms (if included) may allow device 600 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watchbands, bracelets, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow the user to wear device 600.

[0222] Figure 6B An exemplary personal electronic device 600 is shown. In some embodiments, device 600 optionally includes some or all of the components referred to Figure 2A , Figure 2B and Figure 4 . Device 600 has a bus 612 that operatively couples the I / O section 614 to one or more computer processors 616 and a memory 618. The I / O section 614 is optionally connected to a display 604 that may have a touch-sensitive component 622 and optionally has an intensity sensor 624 (e.g., a contact intensity sensor). Additionally, the I / O section 614 is optionally connected to a communication unit 630 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication technologies. Device 600 optionally includes input mechanism 606 and / or input mechanism 608. For example, input mechanism 606 is optionally a rotatable input device or a pressable input device and a rotatable input device. In some examples, input mechanism 608 is optionally a button.

[0223] In some examples, input mechanism 608 is optionally a microphone. Personal electronic device 600 optionally includes various sensors, such as a GPS sensor 632, an accelerometer 634, an orientation sensor 640 (e.g., a compass), a gyroscope 636, a motion sensor 638, and / or combinations thereof, all of which are optionally operatively connected to the I / O section 614.

[0224] The memory 618 of the personal electronic device 600 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 616, cause the computer processors to perform, for example, the techniques and processes described above. The computer-executable instructions are also stored and / or transmitted, for example, within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device such as a computer-based system, a system including a processor, or other systems that can obtain instructions from and execute the instructions of the instruction execution system, apparatus, or device. In some examples, the storage medium is a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium can be any type of storage device, including but not limited to magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like. The personal electronic device 600 is not limited to Figure 6B the components and configurations thereof, but may include other components or additional components in a variety of configurations.

[0225] As used herein, the term "indicative representation" refers to a user-interactive graphical user interface object optionally displayed on the display screen of devices 200, 400, 600, 800, 1000, 1600, and / or 1800 ( Figure 2A , Figure 4 , and Figures 6A - 6B , Figures 8A - 8AF , Figures 10A - 10AJ , Figures 16A - 16S , Figures 18A - 18D ). For example, an image (e.g., an icon), a button, and text (e.g., a hyperlink) each optionally constitute an indicative representation.

[0226] As used herein, the term "focus selector" refers to an input element for indicating the current portion of the user interface with which the user is interacting. In some embodiments including a cursor or other position marker, the cursor acts as the "focus selector" such that when an input (e.g., a press input) is detected on a touch-sensitive surface (e.g., Figure 4 the touchpad 455 in Figure 5B or Figure 2A the touch-sensitive surface 551 in Figure 5AIn some specific implementations of the touch screen 212), the detected contact on the touch screen acts as a "focus selector", such that when an input (e.g., a press input made by a contact) is detected at the position of a specific user interface element (e.g., a button, a window, a slider, or other user interface element) on the touch screen display, the specific user interface element is adjusted according to the detected input. In some specific implementations, the focus moves from one area of the user interface to another area of the user interface without a corresponding movement of the cursor or a movement of the contact on the touch screen display (e.g., moving the focus from one button to another button by using the tab key or arrow keys); in these specific implementations, the focus selector moves according to the movement of the focus between different areas of the user interface. Regardless of the specific form taken by the focus selector, the focus selector is generally a user interface element (or a contact on the touch screen display) that is controlled by the user to deliver the interaction with the user interface expected by the user (e.g., by indicating to the device the element of the user interface with which the user expects to interact). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or a touch screen), the position of the focus selector (e.g., a cursor, a contact, or a selection box) above the corresponding button will indicate that the user expects to activate the corresponding button (rather than other user interface elements shown on the device display).

[0227] As used in the specification and claims, the term "feature intensity" of a contact refers to a feature of the contact based on one or more intensities of the contact. In some embodiments, the feature intensity is based on a plurality of intensity samples. The feature intensity is optionally based on a predefined number or set of intensity samples collected during a predefined period of time (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after detecting a contact, before detecting a contact lift-off, before or after detecting a contact start to move, before detecting a contact end, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The feature intensity of a contact is optionally based on one or more of the following: the maximum value of the contact intensity, the mean value of the contact intensity, the average value of the contact intensity, the value at the top 10% of the contact intensity, the half-maximum value of the contact intensity, the 90% maximum value of the contact intensity, etc. In some embodiments, the duration of the contact is used in determining the feature intensity (e.g., when the feature intensity is the average value of the contact intensity over time). In some embodiments, the feature intensity is compared to a set of one or more intensity thresholds to determine whether a user has performed an operation. For example, the set of one or more intensity thresholds optionally includes a first intensity threshold and a second intensity threshold. In this example, a contact with a feature intensity not exceeding the first threshold results in a first operation, a contact with a feature intensity exceeding the first intensity threshold but not exceeding the second intensity threshold results in a second operation, and a contact with a feature intensity exceeding the second threshold results in a third operation. In some embodiments, the comparison between the feature intensity and one or more thresholds is used to determine whether to perform one or more operations (e.g., whether to perform the corresponding operation or to forgo performing the corresponding operation), rather than for determining whether to perform a first operation or a second operation.

[0228] Figure 6C Shows multiple contacts 652A to 652E detected on the touch-sensitive display screen 604 using multiple intensity sensors 624A to 624D. Figure 6C Also includes an intensity map that shows the current intensity measurements of the intensity sensors 624A to 624D relative to intensity units. In this example, the intensity measurements of both intensity sensors 624A and 624D are 9 intensity units, and the intensity measurements of both intensity sensors 624B and 624C are 7 intensity units. In some specific implementations, the cumulative intensity is the sum of the intensity measurements of the multiple intensity sensors 624A to 624D, which is 32 intensity units in this example. In some embodiments, each contact is assigned a corresponding intensity, which is a portion of the cumulative intensity. Figure 6DShows the assignment of cumulative intensity to contacts 652A to 652E based on their distance from the force center 654. In this example, each of contacts 652A, 652B, and 652E is assigned an intensity of 8 intensity units of cumulative intensity, and each of contacts 652C and 652D is assigned an intensity of 4 intensity units of cumulative intensity. More generally, in some embodiments, each contact j is assigned a corresponding intensity Ij according to a predefined mathematical function Ij = A·(Dj / ΣDi), which is a part of the cumulative intensity A, where Dj is the distance of the corresponding contact j from the force center, and ΣDi is the sum of the distances of all corresponding contacts (e.g., i = 1 to the last) from the force center. An electronic device similar to or equivalent to devices 104, 200, 400, or 600 can be used to perform the operations referred to Figures 6C - 6D above. In some embodiments, the characteristic intensity of a contact is based on one or more intensities of the contact. In some embodiments, an intensity sensor is used to determine a single characteristic intensity (e.g., a single characteristic intensity of a single contact). It should be noted that the intensity map is not part of the displayed user interface, but is included in Figures 6C - 6D to assist the reader.

[0229] In some embodiments, a part of the recognized gesture is used to determine the characteristic intensity. For example, the touch-sensitive surface optionally receives a continuous swiping contact that transitions from a starting position and reaches an ending position, at which the contact intensity increases. In this example, the characteristic intensity of the contact at the ending position is optionally based only on a part of the continuous swiping contact, rather than the entire swiping contact (e.g., only the part of the swiping contact at the ending position). In some embodiments, a smoothing algorithm is optionally applied to the intensity of the swiping contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of the following: an unweighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow spikes or dips in the intensity of the swiping contact for the purpose of determining the characteristic intensity.

[0230] Optionally, characterize the contact intensity on the touch-sensitive surface relative to one or more intensity thresholds such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device will perform an operation typically associated with clicking a button of a physical mouse or touchpad. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device will perform an operation different from an operation typically associated with clicking a button of a physical mouse or touchpad. In some embodiments, when a contact with a characteristic intensity below the light press intensity threshold (e.g., and above a nominal contact detection intensity threshold, contacts below the nominal contact detection intensity threshold are no longer detected) is detected, the device will move a focus selector based on the movement of the contact on the touch-sensitive surface without performing an operation associated with the light press intensity threshold or the deep press intensity threshold. Generally speaking, unless otherwise stated, these intensity thresholds are consistent between different sets of user interface figures.

[0231] An increase in contact characteristic intensity from an intensity below the light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in contact characteristic intensity from an intensity below the deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in contact characteristic intensity from an intensity below the contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on the touch surface. A decrease in contact characteristic intensity from an intensity above the contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting a contact being lifted from the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.

[0232] In some embodiments described herein, one or more operations are performed in response to detecting a gesture including a corresponding press input or in response to detecting a corresponding press input performed using a corresponding contact (or contacts), where the corresponding press input is detected at least in part based on the detected intensity of the contact (or contacts) increasing above a press input intensity threshold. In some embodiments, a corresponding operation is performed in response to detecting an intensity of the corresponding contact increasing above the press input intensity threshold (e.g., the "down stroke" of the corresponding press input). In some embodiments, the press input includes an increase in the intensity of the corresponding contact above the press input intensity threshold and a subsequent decrease in the intensity of the contact below the press input intensity threshold, and a corresponding operation is performed in response to detecting a subsequent decrease in the intensity of the corresponding contact below the press input threshold (e.g., the "up stroke" of the corresponding press input).

[0233] Figures 6E - 6HDetection of a gesture is shown, the gesture comprising contact 662 with an intensity from below Figure 6E The light press intensity threshold in L ”) increases to a strength higher than Figure 6H The deep compression intensity threshold in D ”) corresponding to a press input with an intensity of a deep press. On the displayed user interface 670 including application icons 672A to 672D displayed in a predefined area 674, a cursor 676 is displayed above application icon 672B corresponding to application 2, and a gesture performed using contact 662 is detected on the touch-sensitive surface 660. In some embodiments, gestures are detected on the touch-sensitive display 604. The intensity sensor detects the intensity of the contact on the touch-sensitive surface 660. The device determines that the intensity of contact 662 is within a deep press intensity threshold (for example, “IT D ”) reaches a peak above. Contact 662 is maintained on touch-sensitive surface 660. In response to detecting the gesture, and according to the intensity increasing to a deep press intensity threshold (e.g., “IT D ”) above contact 662, displays reduced-scale representations 678A-678C (e.g., thumbnails) of documents recently opened for application 2, such as Figures 6F - 6H In some embodiments, the intensity is a characteristic intensity of the contact compared to one or more intensity thresholds. It should be noted that the intensity map for contact 662 is not part of the displayed user interface, but is included in the Figures 6E - 6H To help readers.

[0234] In some embodiments, the display of representations 678A-678C includes animation. For example, representation 678A is initially displayed near application icon 672B, such as Figure 6F As the animation progresses, representation 678A moves upward and representation 678B is displayed near application icon 672B, as shown in FIG. Figure 6G Then, representation 678A moves upward, 678B moves upward toward representation 678A, and representation 678C is displayed near application icon 672B, as shown in FIG. Figure 6H 678A-678C form an array above icon 672B. In some embodiments, the animation progresses according to the intensity of contact 662, such as Figures 6F - 6G , where representations 678A-678C appear and increase as the strength of contact 662 approaches a deep press strength threshold (eg, “IT D ”) increases and moves upward. In some embodiments, the intensity according to which the animation progresses is the characteristic intensity of the contact. The reference may be performed using an electronic device similar to or equivalent to device 104, 200, 400, or 600. Figures 6E - 6H The operation described.

[0235] In some embodiments, the device employs strength hysteresis to avoid unexpected inputs sometimes referred to as “jitter,” where the device defines or selects a hysteresis strength threshold that has a predefined relationship to a press input strength threshold (e.g., the hysteresis strength threshold is X strength units lower than the press input strength threshold, or the hysteresis strength threshold is 75%, 90%, or some reasonable percentage of the press input strength threshold). Thus, in some embodiments, a press input includes the strength of a corresponding contact increasing above the press input strength threshold and the strength of that contact subsequently decreasing below the hysteresis strength threshold corresponding to the press input strength threshold, and a corresponding operation is performed in response to detecting that the strength of the corresponding contact subsequently decreases below the hysteresis strength threshold (e.g., the “upstroke” of the corresponding press input). Similarly, in some embodiments, a press input is detected only when the device detects that the contact strength increases from a strength equal to or lower than the hysteresis strength threshold to a strength equal to or higher than the press input strength threshold and optionally the contact strength subsequently decreases to a strength equal to or lower than the hysteresis strength, and a corresponding operation is performed in response to detecting the press input (e.g., depending on the context, the contact strength increases or the contact strength decreases).

[0236] For ease of explanation, optionally, a description of an operation performed in response to a press input associated with a press input strength threshold or in response to a gesture including a press input is triggered in response to detecting any one of the following various conditions: the contact strength increases above the press input strength threshold, the contact strength increases from a strength below the hysteresis strength threshold to a strength above the press input strength threshold, the contact strength decreases below the press input strength threshold, and / or the contact strength decreases below the hysteresis strength threshold corresponding to the press input strength threshold. Additionally, in an example where an operation is described as being performed in response to detecting that the strength of a contact decreases below the press input strength threshold, the operation is optionally performed in response to detecting that the strength of the contact decreases below the hysteresis strength threshold corresponding to and less than the press input strength threshold.

[0237] As used herein, an “installed application” refers to a software application that has been downloaded to an electronic device (e.g., devices 100, 200, 400, and / or 600) and is ready to be launched (e.g., made open) on the device. In some embodiments, a downloaded application becomes an installed application using an installer, and the installed application extracts program portions from the downloaded software package and integrates the extracted portions with the operating system of the computer system.

[0238] As used herein, the term "open application" or "executing application" refers to a software application that maintains state information (e.g., as part of a device / global internal state 157 and / or an application internal state 192). An open or executing application is optionally any of the following types of applications:

[0239] · An active application that is currently displayed on the display screen of the device on which the application is being used;

[0240] · A background application (or background process) that is not currently displayed but one or more processes of the application are being processed by one or more processors; and

[0241] · A paused or dormant application that is not running but has state information stored in memory (volatile and non-volatile respectively) and is available to resume execution of the application.

[0242] As used herein, the term "closed application" refers to a software application that does not maintain state information (e.g., the state information of a closed application is not stored in the memory of the device). Thus, closing an application includes stopping and / or removing the application process and removing the state information of the application from the memory of the device. Generally speaking, when in a first application, opening a second application does not close the first application. When the second application is displayed and the first application stops being displayed, the first application becomes a background application.

[0243] 3. Digital assistant system

[0244] Figure 7A A block diagram illustrating a digital assistant system 700 according to various examples. In some examples, the digital assistant system 700 is implemented on a stand-alone computer system. In some examples, the digital assistant system 700 is distributed across multiple computers. In some examples, some of the modules and functions of the digital assistant are divided into a server part and a client part, where the client part is located on one or more user devices (e.g., devices 104, 122, 200, 400, 600, 800, 1000, 1404, 1600, 1800) and communicates with the server part (e.g., server system 108) via one or more networks, e.g., as Figure 1 shown. In some examples, the digital assistant system 700 is Figure 1Specific implementations of the server system 108 (and / or the DA server 106) shown in []. It should be noted that the digital assistant system 700 is only one example of a digital assistant system, and the digital assistant system 700 has more or fewer components than shown, combines two or more components, or may have different configurations or layouts of components. Figure 7A The various components shown in [] are implemented in hardware, software instructions for execution by one or more processors, firmware (including one or more signal processing integrated circuits and / or application specific integrated circuits), or combinations thereof.

[0245] The digital assistant system 700 includes a memory 702, an input / output (I / O) interface 706, a network communication interface 708, and one or more processors 704. These components may communicate with each other via one or more communication buses or signal lines 710.

[0246] In some examples, the memory 702 includes non-transitory computer-readable media, such as high-speed random access memory and / or non-volatile computer-readable storage media (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).

[0247] In some examples, the I / O interface 706 couples input / output devices 716 of the digital assistant system 700, such as a display, keyboard, touch screen, and microphone, to a user interface module 722. The I / O interface 706, in combination with the user interface module 722, receives user input (e.g., voice input, keyboard input, touch input, etc.) and processes these inputs accordingly. In some examples, for instance, when the digital assistant is implemented on a stand-alone user device, the digital assistant system 700 includes any of the components and I / O communication interfaces described for each of the devices 200, 400, 600, 1200, and 1404 in []. In some examples, the digital assistant system 700 represents the server portion of a digital assistant implementation and may interact with users via a client-side portion located on a user device (e.g., devices 104, 200, 400, 600, 800, 1000, 1404, 1600, 1800). Figure 2A 、 Figure 4 、 Figures 6A - 6H 、 Figure 12 and Figure 14 In some examples, the digital assistant system 700 represents the server portion of a digital assistant implementation and may interact with users via a client-side portion located on a user device (e.g., devices 104, 200, 400, 600, 800, 1000, 1404, 1600, 1800).

[0248] In some examples, network communication interface 708 includes one or more wired communication ports 712, and / or wireless transmission and reception circuitry 714. The one or more wired communication ports receive and transmit communication signals via one or more wired interfaces such as Ethernet, Universal Serial Bus (USB), FireWire, etc. The wireless circuitry 714 receives RF signals and / or optical signals from communication networks and other communication devices and transmits RF signals and / or optical signals to communication networks and other communication devices. Wireless communication uses any one of a variety of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. Network communication interface 708 enables the digital assistant system 700 to communicate with other devices via a network, such as the Internet, an intranet, and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN).

[0249] In some examples, the memory 702 or the computer-readable storage medium of the memory 702 stores programs, modules, instructions, and data structures, including all or a subset of the following: an operating system 718, a communication module 720, a user interface module 722, one or more application programs 724, and a digital assistant module 726. Specifically, the memory 702 or the computer-readable storage medium of the memory 702 stores instructions for performing the above processes. The one or more processors 704 execute these programs, modules, and instructions and read data from or write data to the data structures.

[0250] The operating system 718 (e.g., Darwin, RTXC, LINUX, UNIX, iOS, OSX, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware, firmware, and software components.

[0251] The communication module 720 facilitates communication between the digital assistant system 700 and other devices via the network communication interface 708. For example, the communication module 720 communicates with the RF circuitry 208 of an electronic device such as the devices 200, 400, or 600 shown respectively in Figure 2A , Figure 4 , Figures 6A - 6B . The communication module 720 also includes various components for processing data received by the wireless circuitry 714 and / or the wired communication ports 712.

[0252] The user interface module 722 receives commands and / or inputs from a user (e.g., from a keyboard, touch screen, pointing device, controller, and / or microphone) via the I / O interface 706 and generates user interface objects on a display. The user interface module 722 also prepares outputs (e.g., speech, sound, animation, text, icons, vibration, tactile feedback, lighting, etc.) and transmits them to the user via the I / O interface 706 (e.g., through a display, audio channel, speaker, touchpad, etc.).

[0253] The application 724 includes programs and / or modules configured to be executed by the one or more processors 704. For example, if the digital assistant system is implemented on a stand-alone user device, the application 724 includes user applications such as games, calendar applications, navigation applications, or mail applications. If the digital assistant system 700 is implemented on a server, the application 724 includes, for example, a resource management application, a diagnostic application, or a scheduling application.

[0254] The memory 702 also stores a digital assistant module 726 (or the server portion of the digital assistant). In some examples, the digital assistant module 726 includes the following sub-modules or a subset or superset thereof: an input / output processing module 728, a speech-to-text (STT) processing module 730, a natural language processing module 732, a dialogue flow processing module 734, a task flow processing module 736, a service processing module 738, and a speech synthesis processing module 740. Each of these modules has access to one or more of the following systems or data and models of the digital assistant module 726 or a subset or superset thereof: a knowledge ontology 760, a lexical index 744, user data 748, a task flow model 754, a service model 756, and an ASR system 758.

[0255] In some examples, using the processing modules, data, and models implemented in the digital assistant module 726, the digital assistant can perform at least some of the following: convert speech input into text; recognize a user intention expressed in a natural language input received from the user; actively elicit and obtain information required to fully infer the user intention (e.g., by disambiguating words, names, intentions, etc.); determine a task flow for satisfying the inferred intention; and execute the task flow to satisfy the inferred intention.

[0256] In some examples, as Figure 7B shown, the I / O processing module 728 can interact with the user through the Figure 7A I / O device 716 in Figure 7AThe network communication interface 708 therein interacts with a user device (e.g., device 104, device 200, device 400, or device 600) to obtain user input (e.g., voice input) and provide a response to the user input (e.g., as a voice output). The I / O processing module 728 optionally obtains context information associated with the user input from the user device either with or shortly after receiving the user input. The context information includes user-specific data, vocabulary, and / or preferences related to the user input. In some examples, the context information further includes the software and hardware state of the user device when the user request is received, and / or information related to the user's surrounding environment when the user request is received. In some examples, the I / O processing module 728 also sends follow-up questions related to the user request to the user and receives answers from the user. When the user request is received by the I / O processing module 728 and the user request includes voice input, the I / O processing module 728 forwards the voice input to the STT processing module 730 (or speech recognizer) for speech-to-text conversion.

[0257] The STT processing module 730 includes one or more ASR systems 758. The one or more ASR systems 758 can process the speech input received through the I / O processing module 728 to generate recognition results. Each ASR system 758 can include a front-end speech pre-processor. The front-end speech pre-processor extracts representative features from the speech input. For example, the front-end speech pre-processor performs a Fourier transform on the speech input to extract spectral features characterizing the speech input as a sequence of representative multi-dimensional vectors. Additionally, each ASR system 758 includes one or more speech recognition models (e.g., acoustic models and / or language models) and implements one or more speech recognition engines. Examples of speech recognition models include hidden Markov models, Gaussian mixture models, deep neural network models, n-gram language models, and other statistical models. Examples of speech recognition engines include engines based on dynamic time warping and engines based on weighted finite state transducers (WFSTs). The one or more speech recognition models and the one or more speech recognition engines are used to process the extracted representative features of the front-end speech pre-processor to generate intermediate recognition results (e.g., phonemes, strings of phonemes, and sub-words), and ultimately generate text recognition results (e.g., words, strings of words, or sequences of symbols). In some examples, the speech input is at least partially processed by a third-party service or on the user's device (e.g., device 104, device 200, device 400, or device 600) to generate recognition results. Once the STT processing module 730 generates a recognition result that includes a text string (e.g., a word, or a sequence of words, or a sequence of symbols), the recognition result is transmitted to the natural language processing module 732 for intent inference. In some examples, the STT processing module 730 generates multiple candidate text representations of the speech input. Each candidate text representation is a sequence of words or symbols corresponding to the speech input. In some examples, each candidate text representation is associated with a speech recognition confidence score. Based on the speech recognition confidence score, the STT processing module 730 ranks the candidate text representations and provides the n best (e.g., the n highest-ranked) candidate text representations to the natural language processing module 732 for intent inference, where n is a predetermined integer greater than zero. For example, in one example, only the highest-ranked (n = 1) candidate text representation is delivered to the natural language processing module 732 for intent inference. As another example, the 5 highest-ranked (n = 5) candidate text representations are passed to the natural language processing module 732 for intent inference.

[0258] More details regarding speech-to-text processing are described in U.S. Utility Patent Application Serial No. 13 / 236,942, entitled "Consolidating Speech Recognition Results," filed on September 20, 2011, the entire disclosure of which is incorporated herein by reference.

[0259] In some examples, the STT processing module 730 includes a vocabulary of recognizable words and / or accesses the vocabulary via a phonetic conversion module 731. Each vocabulary word is associated with one or more candidate pronunciations of the word represented in a speech recognition phonetic alphabet. Specifically, the vocabulary of recognizable words includes words associated with multiple candidate pronunciations. For example, the vocabulary includes the word "tomato" associated with candidate pronunciations of and In addition, vocabulary words are associated with custom candidate pronunciations based on previous speech input from the user. Such custom candidate pronunciations are stored in the STT processing module 730 and are associated with a specific user via a user profile on the device. In some examples, candidate pronunciations of a word are determined based on the spelling of the word and one or more linguistic and / or phonetic rules. In some examples, candidate pronunciations are generated manually, e.g., based on known standard pronunciations.

[0260] In some examples, candidate pronunciations are ranked based on their prevalence. For example, candidate pronunciation is ranked higher than because the former is a more commonly used pronunciation (e.g., among all users, among users in a specific geographical region, or among any other suitable subset of users). In some examples, candidate pronunciations are ranked based on whether they are custom candidate pronunciations associated with the user. For example, custom candidate pronunciations are ranked higher than standard candidate pronunciations. This can be used to identify proper names with unique pronunciations that deviate from the norm. In some examples, candidate pronunciations are associated with one or more phonetic characteristics such as geographical origin, country, or ethnicity. For example, candidate pronunciation is associated with the United States, while candidate pronunciation is associated with the United Kingdom. Additionally, the ranking of candidate pronunciations is based on one or more characteristics of the user (e.g., geographical origin, country, ethnicity, etc.) stored in the user profile on the device. For example, it can be determined from the user profile that the user is associated with the United States. Based on the user being associated with the United States, candidate pronunciation (associated with the United States) may be ranked higher than candidate pronunciations (associated with the United Kingdom). In some examples, one of the ranked candidate pronunciations can be selected as the predicted pronunciation (e.g., the most likely pronunciation).

[0261] When a voice input is received, the STT processing module 730 is used (e.g., using a voice model) to determine the phonemes corresponding to the voice input, and then attempts to (e.g., using a language model) determine the words that match the phonemes. For example, if the STT processing module 730 first identifies a sequence of phonemes corresponding to a portion of the voice input then it may subsequently determine that the sequence corresponds to the word "tomato" based on the lexical index 744.

[0262] In some examples, the STT processing module 730 uses fuzzy matching techniques to determine the words in the utterance. Thus, for example, the STT processing module 730 determines that a sequence of phonemes corresponds to the word "tomato", even though that particular sequence of phonemes is not a candidate sequence of phonemes for that word.

[0263] The natural language processing module 732 of the digital assistant ("natural language processor") obtains the n-best candidate text representations ("sequence of words" or "sequence of symbols") generated by the STT processing module 730, and attempts to associate each candidate text representation with one or more "executable intents" recognized by the digital assistant. An "executable intent" (or "user intent") represents a task that can be performed by the digital assistant and can have an associated task flow that is implemented in the task flow model 754. The associated task flow is a series of programmed actions and steps that the digital assistant takes to perform the task. The scope of the digital assistant's capabilities depends on the number and variety of task flows that have been implemented and stored in the task flow model 754, or in other words, on the number and variety of "executable intents" recognized by the digital assistant. However, the effectiveness of the digital assistant also depends on the assistant's ability to infer the correct "one or more executable intents" from a user request expressed in natural language.

[0264] In some examples, in addition to the sequence of words or symbols obtained from the STT processing module 730, the natural language processing module 732 also receives (e.g., from the I / O processing module 728) context information associated with the user request. The natural language processing module 732 optionally uses the context information to clarify, supplement, and / or further qualify the information contained in the candidate text representation received from the STT processing module 730. Context information includes, for example, user preferences, the hardware and / or software state of the user's device, sensor information collected before, during, or shortly after the user request, previous interactions (e.g., conversations) between the digital assistant and the user, and so on. As described herein, in some examples, the context information is dynamic and varies with the time, location, content, and other factors of the conversation.

[0265] In some examples, natural language processing is based on, for example, an ontology 760. The ontology 760 is a hierarchical structure that includes many nodes, where each node represents an "executable intent" or an "attribute" related to one or more of an "executable intent" or other "attributes". As described above, an "executable intent" represents a task that a digital assistant can perform, that is, the task is "executable" or can be carried out. An "attribute" represents a parameter associated with a sub-aspect of an executable intent or another attribute. The connections between the executable intent nodes and the attribute nodes in the ontology 760 define how the parameters represented by the attribute nodes are subordinate to the tasks represented by the executable intent nodes.

[0266] In some examples, the ontology 760 consists of executable intent nodes and attribute nodes. Within the ontology 760, each executable intent node is directly connected to or connected to one or more attribute nodes through one or more intermediate attribute nodes. Similarly, each attribute node is directly connected to or connected to one or more executable intent nodes through one or more intermediate attribute nodes. For example, as Figure 7C shown, the ontology 760 includes a "restaurant reservation" node (i.e., an executable intent node). The attribute nodes "restaurant", "date / time" (for the reservation), and "number of guests" are all directly connected to the executable intent node (i.e., the "restaurant reservation" node).

[0267] In addition, the attribute nodes "cuisine", "price range", "phone number", and "location" are child nodes of the attribute node "restaurant" and are all connected to the "restaurant reservation" node (i.e., the executable intent node) through the intermediate attribute node "restaurant". Another example is, as Figure 7C shown, the ontology 760 also includes a "set reminder" node (i.e., another executable intent node). The attribute nodes "date / time" (for setting the reminder) and "subject" (for the reminder) are both connected to the "set reminder" node. Since the attribute "date / time" is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node "date / time" is connected to both the "restaurant reservation" node and the "set reminder" node in the ontology 760.

[0268] An executable intent node, together with its linked attribute nodes, is described as a "domain". In this discussion, each domain is associated with a corresponding executable intent and involves a set of nodes (and the relationships between these nodes) related to a specific executable intent. For example, Figure 7CThe ontology 760 shown includes an example of a restaurant reservation domain 762 within the ontology 760 and an example of a reminder domain 764. The restaurant reservation domain includes an executable intent node "Restaurant reservation", attribute nodes "Restaurant", "Date / Time", and "Number of companions", and sub-attribute nodes "Cuisine", "Price range", "Phone number", and "Location". The reminder domain 764 includes an executable intent node "Set reminder" and attribute nodes "Subject" and "Date / Time". In some examples, the ontology 760 consists of multiple domains. Each domain shares one or more attribute nodes with one or more other domains. For example, in addition to the restaurant reservation domain 762 and the reminder domain 764, the "Date / Time" attribute node is also associated with many different domains (e.g., itinerary domain, travel reservation domain, movie ticket domain, etc.).

[0269] Although Figure 7C Two exemplary domains within the ontology 760 are shown, but other domains include, for example, "Find movie", "Initiate phone call", "Find directions", "Schedule meeting", "Send message", and "Provide answer to question", "Reading list", "Provide navigation instructions", "Provide instructions for task", etc. The "Send message" domain is associated with the "Send message" executable intent node and further includes attribute nodes such as "One or more recipients", "Message type", and "Message body". The attribute node "Recipient" is further defined, for example, by sub-attribute nodes such as "Recipient name" and "Message address".

[0270] In some examples, the ontology 760 includes all domains (and thus executable intents) that the digital assistant can understand and act on. In some examples, the ontology 760 is modified, for example, by adding or removing entire domains or nodes, or by modifying the relationships between nodes within the ontology 760.

[0271] In some examples, nodes associated with multiple related executable intents are clustered under a "superdomain" in the ontology 760. For example, the "Travel" superdomain includes a cluster of travel-related attribute nodes and executable intent nodes. Travel-related executable intent nodes include "Flight reservation", "Hotel reservation", "Car rental", "Get route", "Find points of interest", etc. Executable intent nodes under the same superdomain (e.g., the "Travel" superdomain) have multiple shared attribute nodes. For example, the executable intent nodes for "Flight reservation", "Hotel reservation", "Car rental", "Route planning", and "Find points of interest" share one or more of the attribute nodes "Starting location", "Destination", "Departure date / time", "Arrival date / time", and "Number of companions".

[0272] In some examples, each node in the ontology 760 is associated with a set of words and / or phrases related to the attribute or executable intent represented by the node. The corresponding set of words and / or phrases associated with each node is the so-called "lexicon" associated with the node. The corresponding set of words and / or phrases associated with each node is stored in a lexicon index 744 associated with the attribute or executable intent represented by the node. For example, returning Figure 7B , the lexicon associated with the node for the "restaurant" attribute includes words such as "food", "drinks", "cuisine", "hunger", "eat", "pizza", "fast food", "meal", etc. As another example, the lexicon associated with the node for the "initiate phone call" executable intent includes words and phrases such as "call", "make a phone call", "dial", "talk on the phone with", "call this number", "call", etc. The lexicon index 744 optionally includes words and phrases in different languages.

[0273] The natural language processing module 732 receives a candidate text representation (e.g., one or more text strings or one or more sequences of symbols) from the STT processing module 730, and for each candidate representation, determines which nodes the words in the candidate text representation pertain to. In some examples, if a word or phrase in the candidate text representation is associated (via the lexicon index 744) with one or more nodes in the ontology 760, then the word or phrase "triggers" or "activates" those nodes. Based on the number and / or relative importance of the activated nodes, the natural language processing module 732 selects one of the executable intents as the task for the digital assistant to perform as the user intent. In some examples, the domain with the most "triggered" nodes is selected. In some examples, the domain with the highest confidence (e.g., based on the relative importance of its respective triggered nodes) is selected. In some examples, the domain is selected based on a combination of the number and importance of the triggered nodes. In some examples, additional factors are also considered during the process of selecting nodes, such as whether the digital assistant has previously correctly interpreted similar requests from the user.

[0274] User data 748 includes user-specific information, such as user-specific vocabulary, user preferences, user address, the user's default language and second language, the user's contact list, and other short-term or long-term information for each user. In some examples, the natural language processing module 732 uses user-specific information to supplement the information contained in the user input to further qualify the user intent. For example, for the user request "invite my friends to my birthday party", the natural language processing module 732 can access the user data 748 to determine who the "friends" are and when and where the "birthday party" will be held, without the user having to explicitly provide such information in their request.

[0275] It should be recognized that, in some examples, one or more machine learning mechanisms (e.g., neural networks) are utilized to implement the natural language processing module 732. Specifically, one or more machine learning mechanisms are configured to receive a candidate text representation and context information associated with the candidate text representation. Based on the candidate text representation and the associated context information, one or more machine learning mechanisms are configured to determine an intent confidence score based on a set of candidate executable intents. The natural language processing module 732 may select one or more candidate executable intents from the set of candidate executable intents based on the determined intent confidence score. In some examples, a knowledge ontology (e.g., knowledge ontology 760) is also utilized to select one or more candidate executable intents from the set of candidate executable intents.

[0276] Other details of searching the knowledge ontology based on symbol strings are described in U.S. Utility Patent Application Serial No. 12 / 341,743, entitled "Method and Apparatus for Searching Using An Active Ontology", filed on December 22, 2008, the entire disclosure of which is incorporated herein by reference.

[0277] In some examples, once the natural language processing module 732 identifies an executable intent (or domain) based on a user request, the natural language processing module 732 generates a structured query to represent the identified executable intent. In some examples, the structured query includes parameters for one or more nodes within the domain of the executable intent, and at least some of the parameters are populated with specific information and requirements specified in the user request. For example, the user says "Help me reserve a seat at a sushi restaurant at 7 pm." In this case, the natural language processing module 732 can correctly identify the executable intent as "restaurant reservation" based on the user input. According to the knowledge ontology, the structured query for the "restaurant reservation" domain includes parameters such as {cuisine}, {time}, {date}, {number of companions}, etc. In some examples, based on the voice input and the text derived from the voice input using the STT processing module 730, the natural language processing module 732 generates a partially structured query for the restaurant reservation domain, where the partially structured query includes the parameters {cuisine = "sushi"} and {time = "7 pm"}. However, in this example, the user's utterance contains insufficient information to complete the structured query associated with the domain. Therefore, based on the currently available information, other necessary parameters such as {number of companions} and {date} are not specified in the structured query. In some examples, the natural language processing module 732 populates some of the parameters of the structured query with the received context information. For example, in some examples, if the request is for a "nearby" sushi restaurant, the natural language processing module 732 populates the {location} parameter in the structured query with the GPS coordinates from the user's device.

[0278] In some examples, the natural language processing module 732 identifies multiple candidate executable intents for each candidate text representation received from the STT processing module 730. Additionally, in some examples, a corresponding structured query (partially or fully) is generated for each identified candidate executable intent. The natural language processing module 732 determines an intent confidence score for each candidate executable intent and ranks the candidate executable intents based on the intent confidence scores. In some examples, the natural language processing module 732 transmits one or more generated structured queries (including any completed parameters) to the task flow processing module 736 ("task flow processor"). In some examples, one or more structured queries for the m best (e.g., m highest-ranked) candidate executable intents are provided to the task flow processing module 736, where m is a pre-determined integer greater than zero. In some examples, one or more structured queries for the m best candidate executable intents are provided to the task flow processing module 736 along with the corresponding candidate text representations.

[0279] Additional details regarding inferring user intent based on multiple candidate executable intents determined from multiple candidate text representations of a speech input are described in U.S. Utility Patent Application No. 14 / 298,725, filed on June 6, 2014, entitled "System and Method for Inferring User Intent From Speech Inputs", the entire disclosure of which is incorporated herein by reference.

[0280] The task flow processing module 736 is configured to receive one or more structured queries from the natural language processing module 732, complete the structured queries (if necessary), and perform the actions required to "fulfill" the user's final request. In some examples, the various processes necessary to complete these tasks are provided in the task flow model 754. In some examples, the task flow model 754 includes processes for obtaining additional information from the user and a task flow for performing actions associated with the executable intents.

[0281] As described above, to complete a structured query, the task flow processing module 736 needs to initiate an additional conversation with the user to obtain additional information and / or clarify potentially ambiguous utterances. When such an interaction is necessary, the task flow processing module 736 invokes the dialog flow processing module 734 to participate in the conversation with the user. In some examples, the dialog flow processor module 734 determines how (and / or when) to request additional information from the user and receives and processes the user response. The question is provided to the user and the answer is received from the user through the I / O processing module 728. In some examples, the dialog processing module 734 presents the dialog output to the user via audio and / or video output and receives input from the user via a verbal or physical (e.g., click) response. Continuing the above example, when the task flow processing module 736 invokes the dialog flow processing module 734 to determine the "number of people in the party" and "date" information for a structured query associated with the domain "restaurant reservation", the dialog flow processing module 734 generates questions such as "How many people in a row?" and "Which day to book?" and passes them to the user. Once the answer from the user is received, the dialog flow processing module 734 fills the structured query with the missing information or passes the information to the task flow processing module 736 to complete the missing information according to the structured query.

[0282] Once the task flow processing module 736 has completed the structured query for an executable intent, the task flow processing module 736 begins to execute the final task associated with the executable intent. Thus, the task flow processing module 736 executes the steps and instructions in the task flow model according to the specific parameters included in the structured query. For example, the task flow model for the executable intent "restaurant reservation" includes steps and instructions for contacting the restaurant and actually requesting a reservation for a specific number of people at a specific time. For example, using a structured query such as: restaurant reservation, {restaurant = ABC Cafe, date = 3 / 12 / 2012, time = 7 pm, number of people in the party = 5}, the task flow processing module 736 can execute the following steps: (1) log in to the server of ABC Cafe or a restaurant reservation system such as and (2) enter the date, time, and number of people in the party information in the form on the website, (3) submit the form, and (4) form a calendar entry for the reservation in the user's calendar.

[0283] In some examples, the task flow processing module 736, with the assistance of the service processing module 738 (the "service processing module"), completes the tasks requested in the user input or provides an informative response requested in the user input. For example, the service processing module 738 initiates a phone call, sets a calendar entry, invokes a map search, invokes other user applications installed on the user device or interacts with the other applications on behalf of the task flow processing module 736, and invokes or interacts with third-party services (such as a restaurant reservation portal, a social networking site, a bank portal, etc.). In some examples, the protocols and application programming interfaces (APIs) required for each service are specified by the corresponding service model in the service model 756. The service processing module 738 accesses the appropriate service model for the service and generates a request for the service according to the protocols and APIs required by the service based on the service model.

[0284] For example, if a restaurant has enabled an online reservation service, the restaurant submits a service model that specifies the necessary parameters for making a reservation and the API for transmitting the values of the necessary parameters to the online reservation service. When requested by the task flow processing module 736, the service processing module 738 can use the web address stored in the service model to establish a network connection with the online reservation service and send the necessary parameters for the reservation (such as time, date, number of people in the party) to the online reservation interface in a format according to the API of the online reservation service.

[0285] In some examples, the natural language processing module 732, the dialogue processing module 734, and the task flow processing module 736 are used jointly and repeatedly to infer and define the user's intention, obtain information to further clarify and refine the user's intention, and finally generate a response (i.e., output to the user or complete a task) to meet the user's intention. The generated response is a dialogue response to the voice input that at least partially meets the user's intention. Additionally, in some examples, the generated response is output as a voice output. In these examples, the generated response is sent to the speech synthesis processing module 740 (such as a speech synthesizer), where the generated response can be processed to synthesize the dialogue response in voice form. In other examples, the generated response is data content related to fulfilling the user's request in the voice input.

[0286] In an example where the task flow processing module 736 receives multiple structured queries from the natural language processing module 732, the task flow processing module 736 first processes the first structured query of the received structured queries to attempt to complete the first structured query and / or perform one or more tasks or actions represented by the first structured query. In some examples, the first structured query corresponds to the highest ranked executable intent. In other examples, the first structured query is selected from the received structured queries based on a combination of the corresponding speech recognition confidence score and the corresponding intent confidence score. In some examples, if the task flow processing module 736 encounters an error during the processing of the first structured query (e.g., due to inability to determine necessary parameters), the task flow processing module 736 may continue to select and process a second structured query of the received structured queries that corresponds to a lower ranked executable intent. The second structured query is selected, for example, based on the speech recognition confidence score corresponding to the corresponding candidate text representation, the intent confidence score of the corresponding candidate executable intent, the missing necessary parameters in the first structured query, or any combination thereof.

[0287] The speech synthesis processing module 740 is configured to synthesize a speech output for presentation to the user. The speech synthesis processing module 740 synthesizes the speech output based on the text provided by the digital assistant. For example, the generated conversation response is in the form of a text string. The speech synthesis processing module 740 converts the text string into an audible speech output. The speech synthesis processing module 740 uses any suitable speech synthesis technique to generate the speech output from the text, including but not limited to: concatenative synthesis, unit selection synthesis, diphone synthesis, domain specific synthesis, formant synthesis, articulatory synthesis, hidden Markov model (HMM)-based synthesis, and sine wave synthesis. In some examples, the speech synthesis processing module 740 is configured to synthesize individual words based on the phoneme strings corresponding to those words. For example, the phoneme strings are associated with the words in the generated conversation response. The phoneme strings are stored in the metadata associated with the words. The speech synthesis processing module 740 is configured to directly process the phoneme strings in the metadata to synthesize the words in speech form.

[0288] In some examples, instead of (or in addition to) using the speech synthesis processing module 740, speech synthesis is performed on a remote device (e.g., the server system 108), and the synthesized speech is sent to the user device for output to the user. For example, this may occur in some implementations where the output of the digital assistant is generated at the server system. Also, since the server system typically has more processing power or more resources than the user device, it is possible to obtain a higher quality speech output than would be achieved by client-side synthesis.

[0289] Additional details regarding digital assistants can be found in U.S. Utility Patent Application No. 12 / 987,982, entitled "Intelligent Automated Assistant," filed on January 10, 2011, and U.S. Utility Patent Application No. 13 / 251,088, entitled "Generating and Processing Task Items That Represent Tasks to Perform," filed on September 30, 2011, the entire disclosures of which are incorporated herein by reference.

[0290] 4. Accelerated task execution

[0291] Figures 8A - 8AF An exemplary user interface for providing suggestions on an electronic device (e.g., device 104, device 122, device 200, device 600, or device 700) is shown in accordance with some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including the processes Figures 9A - 9B described herein.

[0292] Figure 8A An electronic device 800 (e.g., device 104, device 122, device 200, device 600, or device 700) is shown. In Figures 8A - 8AF the non-limiting exemplary embodiment shown, the electronic device 800 is a smart phone. In other embodiments, the electronic device 800 can be a different type of electronic device, such as a wearable device (e.g., a smart watch). In some examples, the electronic device 800 has a display 801, one or more input devices (e.g., a touch screen of the display 801, buttons, a microphone), and wireless communication radio components. In some examples, the electronic device 800 includes multiple cameras. In some examples, the electronic device includes only one camera. In some examples, the electronic device includes one or more biometric sensors (e.g., biometric sensor 803), which optionally includes a camera, such as an infrared camera, a thermal imaging camera, or a combination thereof.

[0293] In Figure 8AIn this case, when the electronic device is in the locked state, the electronic device 800 displays a lock screen interface such as the lock screen interface 804 on the display 801. The lock screen interface 804 includes a suggestion enabling representation 806 and a notification 808. As shown in the figure, the suggestion enabling representation 806 is associated with an application named "Coffee", and the notification 808 is a message notification associated with an instant messaging application, indicating that the electronic device has received a new message from a contact ("John Appleseed") stored on the electronic device. In some examples, an application (e.g., a third-party application) can specify the way to display the suggestion enabling representation. For example, the application can specify the color of the suggestion enabling representation. In some examples, when in the locked state, the electronic device 800 operates in a secure manner. As an example, when operating in the locked state, the electronic device 800 does not display the content of the task suggestion associated with the suggestion enabling representation 806 or the message associated with the notification 808. In some embodiments, the locked state also corresponds to a constraint on accessing other data (including other applications) and / or a restriction on permitted inputs.

[0294] In some examples, the suggestion enabling representation 806 is displayed in a first manner, and the notification 808 is displayed in a second manner. As an example, the suggestion enabling representation 806 is displayed using a first color, and the notification 808 is displayed using a second color different from the first color. As another example, the suggestion enabling representation 806 can be displayed using a first shape, and the notification 808 can be displayed using a second shape different from the first shape.

[0295] In some examples, when operating in the locked state, the electronic device verifies the user of the electronic device. For example, a biometric sensor 803 (e.g., face recognition, fingerprint recognition) can be used to verify the user biometrically, or the user can be verified in response to entering a valid password (e.g., a password, a numeric passphrase). In some examples, in response to verifying the user, the electronic device 800 transitions to the unlocked state and displays the lock screen interface 810. In some examples, when displaying the lock screen interface 810, the electronic device 800 displays an animation indicating that the electronic device 800 is transitioning from the locked state to the unlocked state (e.g., the lock indicator 805 transitions from the locked state to the unlocked state).

[0296] In some examples, when operating in the unlocked state, the electronic device 800 operates in an insecure manner (e.g., the verified user can access secure data). As an example, as Figure 8BAs shown, the electronic device 800 displays content of a task suggestion associated with a suggestion enabling representation 806 and a message associated with a notification 808. As shown, the content of the task suggestion associated with the suggestion enabling representation includes an indicator 812 that indicates the task associated with the task suggestion and one or more parameters associated with the task.

[0297] When displaying a lock screen interface 810, the electronic device 800 detects a selection (e.g., activation) of the suggestion enabling representation 806. For example, as Figure 8C shown, the selection is a swipe gesture 816 on the suggestion enabling representation 806. As will be described in more detail below, in response to detecting the swipe gesture 816, the electronic device 800 selectively performs a task associated with the suggestion enabling representation 806. If the task is a first type of task (e.g., a background task), the electronic device 800 performs the task without further user input. The electronic device 800 may also stop displaying the suggestion enabling representation 806, as Figure 8H shown.

[0298] Referring Figure 8D , if the task is a second type of task different from the first type, the electronic device 800 displays a confirmation interface such as confirmation interface 820. The confirmation interface 820 includes task content 822, a confirmation enabling representation 824, and a cancellation enabling representation 826. In some examples, the task content 822 includes one or more of an application indicator 828, an application icon 830, and a task indicator 832. In some examples, the application indicator 828 indicates the application associated with the task. The application indicator 828 includes the name of the application (e.g., "Coffee") and / or an icon associated with the application. The application icon 830 includes an icon (or other image) associated with the task and / or the application associated with the task. The task indicator 832 indicates the task corresponding to the suggestion enabling representation ("Order") and / or one or more parameters associated with the task (small cup, latte, Homestead Rd. Cupertino CA).

[0299] In some examples, in response to a selection of the cancellation enabling representation 826, the electronic device 800 stops displaying the confirmation interface 820. In some examples, the application indicator 828 is implemented as an application enabling representation, and in response to a selection of the application indicator 828, the electronic device 800 opens the application associated with the task (e.g., "Coffee"). In some examples, the application icon 830 is implemented as an application enabling representation, and in response to a selection of the application icon 830, the electronic device 800 opens the application associated with the task.

[0300] In some examples, opening an application includes preloading the application with one or more parameters. For example, it is recommended that the suggestion affordance 806 be associated with the task of placing an order using the coffee application, and the parameters associated with this task include the size of the coffee (e.g., small cup), the type of coffee (e.g., latte), and the location where the order is picked up (Homestead Rd. in Cupertino, CA). Thus, in some examples, opening the application in this way includes inserting, on behalf of the user, one or more parameters of the task. As an example, opening the coffee application by selecting the application indicator 828 or the application icon 830 can cause the electronic device to open the coffee application and present an interface (e.g., a shopping cart interface) through which the user can confirm the order (small latte, Homestead Rd. location). In some examples, preloading the parameters causes the electronic device to perform the task in response to an input that confirms the intention to perform the task. In this way, the number of inputs required by the user to perform a specific task using the application can be reduced.

[0301] In some examples, an application (e.g., a third-party application) can specify the manner in which the confirmation affordance is displayed. For example, the application can specify the color of the confirmation affordance. In some examples, when the confirmation interface 820 is displayed, the electronic device 800 detects the selection of the confirmation affordance 824. For example, as Figure 8E shown, the selection is a swipe gesture 836 on the confirmation affordance 824. In response to detecting the swipe gesture 836, the electronic device 800 performs the task. In some examples, when performing the task, the electronic device 800 optionally displays a progress indicator 840 indicating that the task is being performed. In some examples, the display of the progress indicator 840 replaces the display of the confirmation affordance 824.

[0302] Once the task is performed, the electronic device 800 provides an output indicating whether the task was successfully performed. In Figure 8G the example, the task was successfully performed, so the electronic device 800 displays a success indicator 842 indicating that the task was successfully performed. In some examples, the display of the success indicator 842 replaces the display of the progress indicator 840. In some examples, after a predetermined amount of time after the task is completed, the electronic device replaces the display of the confirmation interface 820 with the lock screen interface 810. As shown, since the task associated with the suggestion affordance 806 was performed, the suggestion affordance 806 is not included in the Figure 8H lock screen interface 810.

[0303] In Figure 8IIn the example, the task was not successfully executed, so the electronic device 800 displays a failure interface 844. The failure interface 844 includes a retry enable indication 846, a cancel enable indication 848, and an application enable indication 850. The failure interface also includes content 852. In some examples, in response to the selection of the retry enable indication 846, the electronic device 800 executes the task again. In some examples, in response to the selection of the cancel enable indication, the electronic device 800 stops displaying the failure interface 844. In some examples, in response to the selection of the application enable indication 850, the electronic device 800 opens the application associated with the task. The content 852 may include information about the task, such as one or more parameters for executing the task. In some examples, the content 852 also specifies whether the task was successfully executed. For example, the content 852 may indicate that the electronic device failed to successfully execute the task (e.g., "There was a problem. Please try again.").

[0304] In Figure 8J , it is recommended that the selection of the enable indication 806 is a swipe gesture 854 on the enable indication 806. In response to detecting the swipe gesture 854, the electronic device 800 shifts (e.g., slides) the enable indication 806 in the left direction to display (e.g., show) a view enable indication 856 and a clear enable indication 858, as Figure 8K shown. In some examples, in response to the selection of the view enable indication 856, the electronic device 800 displays a confirmation interface such as the confirmation interface 820 ( Figure 8D ). In some examples, in response to the selection of the clear enable indication 858, the electronic device 800 stops displaying the enable indication 806.

[0305] In Figure 8L , it is recommended that the selection of the enable indication 806 is a swipe gesture 860 on the enable indication 806. In response to detecting the swipe gesture 860, the electronic device 800 shifts (e.g., slides) the enable indication 806 in the right direction to display (e.g., show) an open enable indication 862, as Figure 8M shown. In some examples, in response to the selection of the open enable indication 862, the electronic device opens the application associated with the task of the enable indication (e.g., "Coffee").

[0306] In Figure 8N , the electronic device 800 displays a search screen interface such as the search screen interface 866 on the display 801. The search screen interface 866 includes suggested applications 868 and enable indications 870, 872, 874. As shown in the figure, the enable indication 870 is associated with an instant messaging application, the enable indication 872 is associated with a phone application, and the enable indication 874 is associated with a media playback application.

[0307] In some examples, it is recommended that the suggestion affordance may optionally include symbol signs (e.g., symbol signs 876, 878, 880) indicating the category of the task associated with the suggestion affordance. In some examples, the categories specified in this way may include "currency", "message", "phone", "video", and "media". For example, if the task corresponding to the suggestion affordance does not correspond to a task category, the suggestion affordance may be displayed without a symbol sign (e.g., Figure 8A the suggestion affordance 806). In Figure 8N In the example shown in, the suggestion affordance 870 is associated with the task of sending a text message, and thus includes an instant message symbol sign 876 indicating that the task is associated with text instant messaging. As another example, the suggestion affordance 872 is associated with the task of initiating a phone call, and thus includes a phone symbol sign 878 indicating that the task is associated with the phone function. As yet another example, the suggestion affordance 874 is associated with the task of playing back a video, and thus includes a playback symbol sign 880 indicating that the task is associated with media playback. It should be understood that any number of types of symbol signs can be used, corresponding to any number of corresponding task categories.

[0308] Figures 8N - 8P Shows various ways in which the suggested application and the suggestion affordance can be displayed in the search screen interface. As Figure 8N shown, in at least one example, the suggested application and the suggestion affordance can be displayed in the corresponding parts of the search screen interface (e.g., parts 882, 884). As Figure 8O shown, in at least one example, the suggested application and the suggestion affordance can be displayed in the same part (e.g., part 886) of the search screen interface. As Figure 8P shown, in at least one example, the suggested application and the suggestion affordance can be displayed in the corresponding parts of the search screen interface (e.g., parts 888, 890, 892, 894).

[0309] In some examples, when the search screen interface 866 is displayed, the electronic device 800 detects the selection of a suggestion affordance, such as the suggestion affordance 870. In some examples, since the task associated with the suggestion affordance 870 is a pre-determined type of task (e.g., a background task), in response to the selection of the suggestion affordance 870 by a first type of input (e.g., a tap gesture), the electronic device executes (e.g., automatically executes) the task associated with the suggestion affordance 870. In addition or alternatively, in response to the selection of the affordance 870 by a second type of input different from the first type, the electronic device 800 displays a confirmation interface requesting task confirmation from the user.

[0310] For example, as Figure 8Q shown, the selection is to confirm a touch gesture 896 on the confirmation affordance 824. In some examples, the touch gesture 896 is a touch input that meets a threshold intensity and / or a threshold duration such that the touch gesture 896 can be distinguished from a tap gesture. As Figure 8R shown, in response to detecting the touch gesture 896, the electronic device 800 displays a confirmation interface 898. The confirmation interface 898 includes a confirmation affordance 802A, a cancel affordance 804A, an application indicator 806A, and content 808A. In some examples, the selection of the cancel affordance 804A causes the electronic device 800 to stop displaying the confirmation interface 898 and / or to abandon the execution of the task associated with the suggestion affordance 870. In some examples, the application indicator 806A indicates the application associated with the task. The application indicator 806A may include the name of the application (e.g., "Messages") and / or an icon associated with the application. The content 808A may include information for the task, such as one or more parameters for performing the task. For example, the content 808A may specify that the recipient of a text message is the contact "Mom" and the text of the text message is "Good morning". In some examples, the content 808A may be implemented as an affordance.

[0311] When displaying the confirmation interface 898, the electronic device 800 detects the selection of the confirmation affordance 802A. For example, as Figure 8S shown, the selection is a tap gesture 810A on the confirmation affordance 824. In response to detecting the tap gesture 810A, the electronic device 800 performs the task associated with the suggestion affordance 870. As Figure 8T shown, in some examples, when performing the task, the electronic device 800 optionally displays a progress indicator 812A indicating that the task is being performed. In some examples, the display of the progress indicator 812A replaces the display of the confirmation affordance 802A and the cancel affordance 804A.

[0312] Once the task is performed, the electronic device 800 provides an output indicating whether the task was successfully performed. In Figure 8U the example, the task was successfully performed, so the electronic device 800 displays a success indicator 814A indicating that the task was successfully performed (e.g., successfully sent a message to "Mom"). In some examples, the display of the success indicator 814A replaces the display of the progress indicator 812A. In some examples, after a predetermined amount of time after the task is completed, the electronic device replaces the display of the confirmation interface 898 with a search screen interface 866. As Figure 8V shown, since the task associated with the suggestion affordance 870 was performed, the suggestion affordance 870 is not included in the search screen interface 866.

[0313] In some examples, when displaying the confirmation interface 898, the electronic device 800 detects the selection of the content 808A. For example, in Figure 8W , the selection is a tap gesture 816A on the content 808A. In response to detecting the tap gesture 816A, the electronic device 800 opens the application associated with the suggested affordance 870, as Figure 8X shown.

[0314] In some examples, opening the application in this way includes preloading the application with one or more parameters associated with the task. In this way, the user can perform tasks within the application with a reduced number of inputs. As an example, in response to the selection of the content 808A associated with the suggested affordance 870, the electronic device 800 opens the instant messaging application and preloads the instant messaging application with the parameters specified by the suggested affordance 870. Specifically, the instant messaging application can be directed to the instant messaging interface 817A for providing a message to the recipient "Mom", and the input string "Good morning" can be inserted into the message composition field 818A of the instant messaging application.

[0315] When displaying the instant messaging interface 817A, the electronic device 800 detects the selection of the send affordance 820A. For example, as Figure 8Y shown, the selection is a tap gesture 822A on the send affordance 820A. In Figure 8Z , in response to detecting the tap gesture 822A, the electronic device 800 sends the preloaded message (e.g., "Good morning") to the recipient "Mom". By preloading the parameters in this way, the user can use the suggested affordance to open the application and perform tasks with fewer inputs than might otherwise be the case. For example, in the above example, the user sends a text message without having to select the recipient or enter a message for that recipient.

[0316] In some examples, the application can be opened without preloading parameters. In Figure 8AAIn this case, the electronic device 800 displays a search screen interface such as the search screen interface 826A on the display 801. The search screen interface 826A includes the suggested application 282A and the suggested enablement representations 830A, 832A. As shown, the suggested enablement representation 830A is associated with the notepad application, and the suggested enablement representation 832A is associated with the video call application. In addition, the suggested enablement representation 830A is associated with the task of opening the notepad application. In some examples, the task corresponding to the notepad application may not correspond to a task category, and thus the suggested enablement representation 830A does not include a symbol. The suggested enablement representation 832A is associated with the task of initiating a video call (e.g., a Skype call), and thus includes a video symbol 836A indicating that the task is associated with the video call function.

[0317] When displaying the search screen interface 826A, the electronic device 800 detects the selection of the suggested enablement representation 834A. For example, as Figure 8AA shown, the selection is a tap gesture 834A on the suggested enablement representation 834A. In Figure 8AB response to detecting the tap gesture 834A, the electronic device 800 opens the notepad application associated with the suggested enablement representation 830A.

[0318] In some examples, as described herein, the way the electronic device displays the interface depends on the type of the electronic device. In some examples, for example, the electronic device 800 may be implemented as a device with a relatively small display, such that interfaces such as the lock screen interface 804 or the search screen interface 866 may not be suitable for display. Thus, in some examples, the electronic device 800 may display an alternative interface to those previously described.

[0319] Referring to Figure 8AC , for example, the electronic device 800 displays a home screen interface 850A on the display 801. The home screen interface 850A includes the suggested enablement representation 852A and the notification 854A. As shown, the suggested enablement representation 852A is associated with an application named "Coffee", and the notification 854A is a calendar notification associated with the calendar application, which indicates that the user has an upcoming event ("Meeting").

[0320] It should be understood that although the home screen interface 850A is shown as including the suggested enablement representation 852A, in some examples, the home screen interface 850A includes multiple suggested enablement representations 852A. For example, in response to a user input such as a swipe gesture (e.g., an upward swipe gesture, a downward swipe gesture), the electronic device may display (e.g., show) one or more additional suggested enablement representations.

[0321] When displaying the home screen interface 850A, the electronic device 800 detects a selection of the suggestion enabling indication 852A. For example, as Figure 8AD shown, the selection is a tap gesture 852A on the suggestion enabling indication 858A. As will be described in more detail below, in response to detecting the tap gesture 858A, the electronic device 800 displays a confirmation interface such as the confirmation interface 820. The confirmation interface 820 includes an application indicator 861A, a task indicator 862A, a confirmation enabling indication 864A, and a cancellation enabling indication 866A. In some examples, the application indicator 861A indicates the application associated with the task. The application indicator 861A may include the name of the application (e.g., "Coffee") and / or an icon associated with the application. The task indicator 862A indicates the task associated with the application and one or more parameters associated with the task (small cup, latte, oat milk).

[0322] In some examples, in response to a selection of the cancellation enabling indication 866A, the electronic device 800 stops displaying the confirmation interface 860A. In some examples, when displaying the confirmation interface 860A, the electronic device 800 detects a selection of the confirmation enabling indication 864A. For example, as Figure 8AE shown, the selection is a tap gesture 868A on the confirmation enabling indication 864A. In response to detecting the swipe gesture 868A, the electronic device 800 selectively performs the task. If the task is a first type of task, the electronic device 800 performs the task without further user input and, optionally, replaces the display of the confirmation interface 860A with the home screen interface 850A, as Figure 8AF shown. Since the task associated with the suggestion enabling indication 852A has been performed, the suggestion interface 852A is not displayed in the home screen interface 850A. If the task is a second type of task, the electronic device 800 may request user confirmation of the task before performing the task, as described.

[0323] Figures 9A - 9Bis a flowchart showing a method 900 for providing suggestions according to some embodiments. The method 900 is performed at a device (e.g., device 104, device 122, device 200, device 600, device 700, device 800) having a display, one or more input devices (e.g., touchscreen, microphone, camera), and a wireless communication radio (e.g., Bluetooth connection, WiFi connection, mobile broadband connection such as 4G LTE connection). In some embodiments, the display is a touch-sensitive display. In some embodiments, the display is not a touch-sensitive display. In some embodiments, the electronic device includes multiple cameras. In some embodiments, the electronic device includes only one camera. In some embodiments, the device includes one or more biometric sensors, optionally including a camera, such as an infrared camera, a thermal imaging camera, or a combination thereof. Some operations in the method 900 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0324] As described below, displaying a user interface including a suggestion enablement representation and selectively requesting confirmation to perform a task in response to a selection of the suggestion enablement representation provides the user with an easy-to-recognize and intuitive way to perform tasks on an electronic device, thereby reducing the amount of user input otherwise required to perform such tasks. Thus, displaying the user interface in this manner enhances the operability of the device and makes the user device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn further reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0325] In some examples, the electronic device determines a first set of candidate tasks (902) and identifies a task from the first set of candidate tasks (904). In some examples, the task is identified based on determining that the context of the electronic device meets the task suggestion criteria.

[0326] In some examples, before displaying a user interface on a display, an electronic device determines whether a context criterion (e.g., a task recommendation criterion) has been met based on a first context of the electronic device (e.g., context data describing previous use of the electronic device). In some examples, the electronic device determines whether a task recommendation meets a confidence threshold for displaying the recommendation. In some examples, based on determining that the context criterion has been met (e.g., one or more task recommendations meet the confidence threshold), the electronic device determines a first set of candidate tasks. In some examples, the electronic device determines whether a heuristic criterion has been met based on a second context of the electronic device. In some examples, the second context of the electronic device indicates previous use of the device and / or context data associated with the user (e.g., contacts, calendar, location). In some examples, determining whether the heuristic criterion has been met includes determining whether a set of conditions for heuristic task recommendation has been met such that a heuristic task recommendation is provided in place of the recommended task. In some examples, the electronic device determines whether the context criterion has been met and then determines whether the heuristic criterion has been met. In some examples, the electronic device determines whether the heuristic criterion has been met and then determines whether the context criterion has been met. In some examples, the electronic device determines whether the context criterion and the heuristic criterion have been met simultaneously. In some examples, based on determining that the heuristic criterion has been met, the electronic device determines a second set of candidate tasks different from the first set of candidate tasks and identifies a task from the second set of candidate tasks. In some examples, based on determining that the heuristic criterion has not been met and the context criterion has been met, the electronic device identifies a task from the first set of candidate tasks. In some examples, based on determining that the heuristic criterion has not been met and the context criterion has not been met, the electronic device abandons determining the first set of candidate tasks and abandons determining the second set of candidate tasks.

[0327] Providing heuristic task recommendations in this manner allows the electronic device to provide task recommendations based on user-specific context data in addition to the context data of the electronic device (e.g., according to respective sets of conditions as described below). This allows the electronic device to provide prominent task recommendations for performing tasks on the electronic device, thereby reducing the amount of user input required to perform these tasks otherwise. Thus, providing heuristic task recommendations in this manner enhances the operability of the device and makes the use of the electronic device more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0328] In some examples, the electronic device determines whether the task is a task of a first type (906). In some examples, determining whether the task is a task of a first type includes determining whether the task is a background task (e.g., a task that can be performed without user confirmation and / or additional user input).

[0329] In some examples, determining whether the task is a task of a first type includes determining whether one or more parameters associated with the task are valid. In some examples, a parameter of a task is valid when each parameter required by the task is assigned a value within an allowed range or within a set of values for that parameter, and optionally, each optional parameter of the task has a value within an allowed range or set of parameters, or is unassigned.

[0330] In some examples, the electronic device (104, 200, 600, 700, 800) displays a user interface (804, 810, 866, 826A, 850A) including a suggested affordance representation (806, 870, 872, 874, 834A, 854A) associated with the task on a display of the electronic device (908). In some examples, the user interface is a lock screen interface (804, 810). In some examples, the user interface is a search screen interface (866, 826A). In some examples, the user interface is a digital assistant interface for a conversation between the user and the digital assistant. In some examples, the suggested affordance representation is an affordance representation corresponding to a task suggestion provided by the electronic device (and in some cases, corresponding to a digital assistant of the electronic device). In some examples, the suggestion is task-specific and / or parameter-specific. In some examples, task suggestions are provided based on the context of the electronic device (e.g., location, WiFi connection, WiFi network identifier (e.g., SSID), usage history, time / date, headphone connection, etc.). In some examples, the task suggestions are visually distinguishable from other notifications displayed by the electronic device in the user interface.

[0331] Providing task suggestions based on the context of the electronic device allows the electronic device to provide prominent task suggestions based on the user's previous use of the electronic device and / or the current state of the electronic device. As a result, the amount of input and time required to perform tasks on the electronic device is reduced, thus accelerating the interaction between the user and the electronic device. This in turn reduces power usage and extends the battery life of the device.

[0332] In some examples, displaying a user interface on a display that includes a suggested affordance representation associated with a task includes: displaying a suggested affordance representation with a symbol based on the task being a first type of task; and displaying a suggested affordance representation without a symbol based on the task being a second type of task. In some examples, the symbol indicates the type of task (e.g., background vs. non-background, whether the task requires user confirmation). In some examples, the symbol is an arrow indicating that the task requires user confirmation. In some examples, the symbol is a dollar sign (or other currency symbol) indicating that the task is a transaction. In some examples, the symbol is bounded by a circle.

[0333] In some examples, displaying a user interface on a display that includes a suggested affordance representation associated with a task includes: displaying a suggested affordance representation (910) with a first type of symbol based on determining that the task corresponds to a first set of tasks. In some examples, the set of tasks is a category of tasks. In some examples, the task categories include messaging tasks, phone tasks, video phone tasks, and media tasks. In some examples, each set of tasks corresponds to one or more respective first-party applications. In some examples, each set of tasks also includes tasks corresponding to one or more third-party applications. In some examples, if the task is a task corresponding to a particular category of tasks, the suggested affordance representation corresponding to the task includes a symbol (876, 878, 880, 836A) identifying the category. In some examples, displaying a user interface on a display that includes a suggested affordance representation associated with a task further includes: displaying a suggested affordance representation (912) with a second type of symbol different from the first type based on determining that the task corresponds to a second set of tasks different from the first set of tasks. In some examples, displaying a user interface on a display that includes a suggested affordance representation associated with a task further includes: displaying a suggested affordance representation (830A) without a symbol based on determining that the task does not correspond to the first set of tasks and does not correspond to the second set of tasks. In some examples, if the task does not correspond to one or more predetermined categories of tasks, the suggested affordance representation corresponding to the task is displayed without a symbol (914).

[0334] In some examples, the user interface (804, 810) includes a notification (808). The notification can be a notification of an event such as receiving a text message. In some examples, the suggested affordance representation (806) is displayed in a first manner, and the notification (808) is displayed in a second manner different from the first manner. For example, in some examples, the suggested affordance representation and the notification correspond to different colors, color schemes, and / or patterns. In some examples, the suggested affordance representation and the notification have different shapes and / or sizes.

[0335] As described herein, displaying notifications and suggestion affordances in different ways enables a user to easily distinguish between notifications and suggestion affordances on a display of an electronic device, thereby reducing the amount of time required to perform a task. Reducing time in this way enhances the operability of the device and makes the user device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn further reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0336] In some examples, displaying a user interface includes, in response to determining that the electronic device is in a locked state, displaying a suggestion affordance in a first visual state. In some examples, if the device is locked, the electronic device displays a reduced amount of information corresponding to the suggestion affordance. In some examples, displaying the user interface further includes, in response to determining that the electronic device is not in a locked state, displaying the suggestion affordance in a second visual state different from the first visual state. In some examples, if the device is unlocked, the electronic device displays content corresponding to the suggestion affordance. Content displayed in this way includes, but is not limited to, the name and / or icon of an application associated with the suggestion affordance, one or more parameters associated with the task of the suggestion affordance, and optionally, an indicia indicating the task category of the suggestion affordance.

[0337] In some examples, the electronic device detects a first user input (816, 834A, 858A)(916) corresponding to a selection of the suggestion affordance via one or more input devices. In some examples, the suggestion affordance is selected using a touch input, a gesture, or a combination thereof. In some examples, the suggestion affordance is selected using a voice input. In some examples, the touch input is a first type of input, such as a press having a relatively short duration or a relatively low intensity.

[0338] In some examples, in response to detecting the first user input (918), in response to determining that the task is a first type of task, the electronic device performs the task (920). In some examples, performing the task includes causing the task to be performed by another electronic device. In some examples, the electronic device is a first type of device (e.g., a smartwatch) and causes the task to be performed on a second type of device (e.g., a mobile phone).

[0339] In some examples, further in response to detecting a first user input (918), based on determining that the task is a second type of task different from the first type, the electronic device displays a confirmation interface (820, 898, 860A) (922) including a confirmation affordance representation. In some examples, the second type of task is a task that requires user confirmation and / or additional information from the user before performing the task, such as a task corresponding to a transaction. In some examples, the second type of task is a task performed by a particular type of device such as a smartwatch. In some examples, the confirmation interface is displayed simultaneously with the user interface. For example, the confirmation interface may be overlaid on the user interface. In some examples, the confirmation interface is displayed on a first portion of the user interface, and a second portion of the user interface is visually occluded (e.g., dimmed, blurred).

[0340] Responding to the selection of the suggestion affordance representation to selectively request confirmation to perform a task allows the electronic device to quickly perform tasks of the first type and confirm the user's intention before performing tasks of the second type. Thus, an intuitive and reliable method is provided for the user to quickly and reliably perform tasks on the electronic device, thereby reducing the amount of user input required to perform these tasks and accelerating task execution. Such advantages in turn reduce the amount of time required to perform the tasks and make the use of the electronic device more efficient, which in turn reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and effectively.

[0341] In some examples, when displaying the confirmation interface (926), the electronic device detects a second user input (836, 810A, 868A) (928) corresponding to the selection of the confirmation affordance representation (824, 802A, 864A). In some examples, the confirmation affordance representation is selected using a touch input, a gesture, or a combination thereof. In some examples, the confirmation affordance representation is selected using a voice input. In some examples, in response to detecting the second user input, the electronic device performs the task (930). In some examples, when performing the task, the electronic device displays a first progress indicator (840, 812A) to indicate that the electronic device is performing the task. In some examples, if the electronic device successfully performs the task, the electronic device displays a second progress indicator (842, 814A) to indicate that the task has been successfully performed. In some examples, after performing the task and / or displaying the second progress indicator, the electronic device displays an interface including one or more visual objects specified by the application (e.g., a message or an image that states "Thank you for your order"). In some examples, if the electronic device does not successfully perform the task, the electronic device provides a natural language output to the user indicating that the task was not successfully performed (e.g., "There was a problem. Try again."), and optionally, displays an affordance representation (846) through which the user can initiate an additional attempt to perform the task.

[0342] In some examples, the confirmation interface includes application-ability representations (828, 850). In some examples, an application-ability representation is a representation that indicates (e.g., identifies) an application and / or task associated with a suggested ability representation. In some examples, an application-ability representation is any part of the confirmation interface other than the confirmation ability representation and / or the cancellation ability representation. In some examples, the electronic device detects a third user input (932) corresponding to a selection of an application-ability representation. In some examples, an application-ability representation is selected using a touch input, a gesture, or a combination thereof. In some examples, an application-ability representation is selected using a voice input. In some examples, in response to detecting the third user input, the electronic device executes (e.g., launches, initiates) an application associated with the task (934). In some examples, the user selects an icon and / or name displayed for an application to launch the application corresponding to the selected ability representation for the task. In some examples, the application optionally pre-loads one or more parameters (e.g., the subject and / or body of an email).

[0343] In some examples, the suggested ability representation includes a visual indication of a parameter that affects task execution. In some examples, the suggested ability representation corresponds to a task to be executed using one or more specified parameters (e.g., ordering a specific coffee size and type using the Starbucks application, sending a text with a specific message). In some examples, executing an application associated with a task includes pre-loading the application with parameters. In some examples, executing the application causes parameters for the task to be input on behalf of the user (e.g., an order is already in the shopping cart and the user only needs to indicate the intent to order; a message is inserted into the message composition field and the user only needs to indicate the intent to send).

[0344] In some examples, the confirmation interface includes a cancellation ability representation. In some examples, the electronic device detects a fourth user input corresponding to a selection of the cancellation ability representation. In some examples, the cancellation ability representation is selected using a touch input, a gesture, or a combination thereof. In some examples, the cancellation ability representation is selected using a voice input. In some examples, in response to detecting the fourth user input, the electronic device abandons the execution of the task. In some examples, the electronic device also stops displaying the confirmation interface in response to detecting the fourth user input.

[0345] In some examples, the first user input is an input of a first type. In some examples, when displaying a user interface, the electronic device detects a second type of user input corresponding to a selection of a suggested affordance representation. In some examples, the suggested affordance representation is selected using a touch input, a gesture, or a combination thereof. In some examples, the suggested affordance representation is selected using a voice input. In some examples, the touch input is a second type of input, such as a press with a relatively long duration or a relatively high intensity. In some examples, the second type of input is different from the first type of input. In some examples, in response to detecting the second type of user input, the electronic device displays a confirmation interface.

[0346] Note that the details of the processes described above with respect to method 900 (e.g., Figures 9A - 9B ) also apply in a similar manner to the methods described below. For example, method 900 optionally includes one or more features of the various methods described with reference to methods 1100, 1300, 1500, 1700, and 1900.

[0347] The operations in the above methods are optionally implemented by running one or more functional modules in an information processing device, such as a general-purpose processor (e.g., as described with respect to Figure 2A 、 Figure 4 、 Figure 6A ) or an application-specific chip. Additionally, the operations described above with reference to Figures 8A - 8AF are optionally implemented by the components depicted in Figures 2A - 2B . For example, display operation 908, detection operation 916, execution operation 920, and display operation 922 are optionally implemented by event classifier 270, event recognizer 280, and event handler 290. Event monitor 271 in event classifier 270 detects a contact on touch-sensitive surface 604 ( Figure 6A ), and event distributor module 274 delivers the event information to application 236-1 ( Figure 2B ). The corresponding event recognizer 280 of application 236-1 compares the event information with the corresponding event definition 286 and determines whether the first contact at the first position on the touch-sensitive surface corresponds to a predefined event or sub-event, such as a selection of an object on the user interface. When the corresponding predefined event or sub-event is detected, event recognizer 280 activates event handler 290 associated with the detection of the event or sub-event. Event handler 290 optionally utilizes or invokes data updater 276 or object updater 277 to update the internal state 292 of the application. In some embodiments, event handler 290 accesses the corresponding GUI updater 278 to update the content displayed by the application. Similarly, those skilled in the art will clearly know how to based on Figures 2A - 2BThe components depicted are used to implement other processes.

[0348] Figures 10A - 10AJ FIG. shows an exemplary user interface for providing voice shortcuts on an electronic device (e.g., device 104, device 122, device 200, device 600, or device 700) according to some embodiments. The user interfaces in these figures are used to illustrate the processes described below, including the processes in Figures 11A - 11B the following description of the processes in.

[0349] Generally, the user interfaces described with reference to Figures 10A - 10AJ can be adopted such that a user can associate tasks with corresponding user-specific phrases. These phrases can then be used to cause the electronic device to perform the associated tasks.

[0350] Figure 10A FIG. shows an electronic device 1000 (e.g., device 104, device 122, device 200, device 600, or device 700). In the Figures 10A - 10AJ non-limiting exemplary embodiment shown in, the electronic device 1000 is a smart phone. In other embodiments, the electronic device 1000 can be a different type of electronic device, such as a wearable device (e.g., a smart watch). In some examples, the electronic device 1000 has a display 1001, one or more input devices (e.g., a touch screen of the display 1001, buttons, a microphone), and wireless communication radio components. In some examples, the electronic device 1000 includes multiple cameras. In some examples, the electronic device includes only one camera. In some examples, the electronic device includes one or more biometric sensors (e.g., biometric sensor 1003), which optionally includes a camera, such as an infrared camera, a thermal imaging camera, or a combination thereof.

[0351] In Figure 10A FIG., the electronic device 1000 displays a settings interface 1004 on the display 1001. The settings interface 1004 includes a candidate task section 1006 and an additional task enabling representation 1014. The candidate task section 1006 includes candidate task enabling representations 1008, 1010, and 1012. In some examples, in response to the selection of the additional task enabling representation 1014, the electronic device 1000 displays a global task interface such as global task interface 1018A, as described with respect to Figure 10S FIG..

[0352] In some examples, in response to a selection of a candidate task affordance representation, the electronic device 1000 displays a task-specific interface. In some examples, the task-specific interface is associated with the task of the candidate task affordance representation. As an example, when displaying the settings interface 1004, the electronic device 1000 detects a selection of the candidate task affordance representation 1008. In some examples, the selection is a tap gesture 1016 on the candidate task affordance representation 1008. As Figure 10B shown, in response to detecting the tap gesture 1016, the electronic device 1000 displays the task-specific interface 1018. The task-specific interface 1018 may be associated with the task of the candidate task affordance representation 1008 (e.g., viewing a side house camera). In some examples, selecting a candidate task affordance representation initiates a voice shortcut generation process corresponding to the task of the candidate task affordance representation. Thus, the selection of the candidate task affordance representation 1008 may initiate a voice shortcut generation process for the task of the candidate task affordance representation 1008.

[0353] The task-specific interface 1018 includes a task icon 1020, a task indicator 1022, a task descriptor 1024, an application indicator 1026, a candidate phrase 1028, and a recording affordance representation 1030. In some examples, the task icon 1020 includes an icon or image corresponding to the task. In some examples, the task indicator 1022 indicates the name and / or type of the task. In some examples, the task descriptor 1024 includes a description of the task and / or indicates that the user can record a command or phrase to link to the task. In some examples, the application indicator 1026 identifies the application corresponding to the task. For example, the application indicator 1026 may include the name of the application and / or an icon associated with the application. The candidate phrase 1028 includes suggested phrases that the user can select to be associated with the task.

[0354] When displaying the task-specific interface 1018, the electronic device 1000 detects a selection of the recording affordance representation 1030. As Figure 10C shown, the selection of the recording affordance representation 1030 is a tap gesture 1032. In response to the selection of the recording affordance representation 1030, the electronic device displays a recording interface 1034 on the display 1001 (e.g., replacing the display of the task-specific interface 1018).

[0355] Refer to Figure 10D, the recording interface 1034 includes a cancel enable indication 1036, a return enable indication 1038, a preview 1042, and a stop enable indication 1044. In some examples, in response to the selection of the cancel enable indication 1036, the electronic device stops displaying the recording interface 1034, and optionally, terminates the voice shortcut generation process. In some examples, in response to the selection of the return enable indication, the electronic device displays the task-specific interface 1018 (e.g., replacing the display of the recording interface 1038).

[0356] In some examples, when displaying the recording interface 1034, the electronic device 1000 uses the audio input device (e.g., microphone) of the electronic device 1000 to receive natural language voice input from the user. In some examples, when receiving natural language voice input, the electronic device 1000 provides a real-time preview of the natural language voice input such as the real-time preview 1042. As Figure 10D shown, in some examples, the real-time preview 1042 is a visual waveform indicating one or more auditory features of the natural language voice input.

[0357] Refer to Figure 10E , if after initially displaying the recording interface 1034, the user does not provide natural language voice input within a predetermined amount of time, the electronic device 1000 may optionally display a prompt to the user, including candidate phrases such as the candidate phrase 1046 (e.g., "View the side of the live house stream"). In some examples, the candidate phrase 1046 is the same as the candidate phrase 1028 ( Figure 10D ). In other examples, the candidate phrase 1046 is different from the candidate phrase 1028.

[0358] In some examples, when receiving natural language voice input, the electronic device 1000 performs speech-to-text conversion (e.g., natural language speech processing) on the natural language voice input to provide the candidate phrase 1048. As Figures 10F - 10G shown, since speech-to-text conversion is performed when receiving natural language voice input, the candidate phrase 1048 can be iteratively and / or continuously updated while receiving natural language voice input.

[0359] In some examples, the electronic device 1000 ensures that the candidate phrase is different from one or more predetermined phrases (e.g., "Call 911"). As an example, the electronic device 1000 determines whether the similarity between the candidate phrase and each of the one or more predetermined phrases exceeds a similarity threshold. If the similarity threshold is not met, the electronic device 1000 will notify the user that the provided candidate phrase is insufficient and / or not allowed. The electronic device 1000 may also request the user to provide another natural language voice input.

[0360] When displaying the recording interface 1034, the electronic device 1000 detects the selection of the stop enabling representation 1044. As Figure 10G shown, the selection of the recording enabling representation 1044 is a tap gesture 1050. In response to the selection of the stop enabling representation 1050, the electronic device 1000 displays a completion interface 1052 on the display 1001 (e.g., replacing the display of the recording interface 1034), as Figure 10H shown.

[0361] The completion interface 1052 includes a completion enabling representation 1054, a cancellation enabling representation 1056, a task icon 1058, a task indicator 1060, an application indicator 1062, a candidate phrase 1064, and an edit enabling representation 1066. In some examples, in response to the selection of the cancellation enabling representation 1056, the electronic device 1000 stops displaying the completion interface 1052 and, optionally, terminates the voice shortcut generation process. In some examples, the task icon 1058 includes an icon or image corresponding to the task. In some examples, the task indicator 1060 indicates the name and / or type of the task. In some examples, the application indicator 1062 identifies the application corresponding to the task. For example, the application indicator may include the name of the application and / or an icon associated with the application. The candidate phrase 1028 is a suggested phrase that the user can select to be associated with the task.

[0362] In some examples, when displaying the completion interface 1052, the electronic device 1000 detects the selection of the edit enabling representation 1066. As Figure 10I shown, the selection of the edit enabling representation 1066 is a tap gesture 1068. In response to the selection of the edit enabling representation 1068, the electronic device 1000 displays an edit interface 1070 on the display 1001 (e.g., replacing the display of the completion interface 1052), as Figure 10J shown.

[0363] The edit interface 1070 includes a completion enabling representation 1072, a cancellation enabling representation 1074, a task icon 1076, a task indicator 1078, a candidate phrase sorting 1080, and a re-record enabling representation 1088. In some examples, in response to the selection of the cancellation enabling representation 1056, the electronic device 1000 stops displaying the edit interface 1070 and, optionally, terminates the voice shortcut generation process. In some examples, the task icon 1076 includes an icon or image corresponding to the task. In some examples, the task indicator 1078 indicates the name and / or type of the task.

[0364] As described, in some examples, the electronic device 1000 provides candidate phrases (e.g., candidate phrase 1048) based on natural language speech input provided by a user. In some examples, providing candidate phrases in this manner includes generating a plurality of candidate phrases (e.g., candidate phrases 1082, 1084, 1086), and selecting the candidate phrase associated with the highest score (e.g., a text representation confidence score). In some examples, candidate phrase ranking 1080 includes a plurality of candidate phrases generated by the electronic device 1000 prior to selecting the candidate phrase associated with the highest score. In some examples, candidate phrase ranking 1080 includes a set (e.g., 3) of top-ranked candidate phrases, optionally listed according to the respective scores of each candidate phrase. For example, as Figure 10J shown, candidate phrase ranking 1080 includes candidate phrases 1082, 1084, and 1086. In some examples, candidate phrase 1082 may correspond to candidate phrase 1064.

[0365] In some examples, the user may select a new (or the same) candidate phrase from candidate phrases 1082, 1084, 1086 of candidate phrase ranking 1080. For example, as Figure 10K shown, the electronic device 1000 detects the selection of candidate phrase 1080 when displaying the editing interface 1070. The selection of candidate phrase 1080 is a tap gesture 1090. In some examples, in response to the selection of candidate phrase 1080, the electronic device 1000 selects candidate phrase 1080 as the new (or the same) candidate phrase. In some examples, in response to the selection of both candidate phrase 1080 and the completion affordance representation 1072, the electronic device selects candidate phrase 1080 as the new (or the same) candidate phrase.

[0366] In some examples, the candidate phrase ranking may not include the user's intended or preferred phrase. Thus, in some examples, in response to the selection of the re-record affordance representation 1088, the electronic device 1000 displays a recording interface such as recording interface 1034 (e.g., replacing the display of the editing interface 1070) to allow the user to provide new natural language speech input as described above.

[0367] In some examples, when displaying the completion interface 1052, the electronic device 1000 detects the selection of the completion affordance representation 1054. As Figure 10LAs shown, the selection that completes the enabling representation 1054 is the tap gesture 1092. In response to the selection that completes the enabling representation 1054, the electronic device 1000 associates the candidate phrase with the task of the candidate task enabling representation 1008. By associating the candidate phrase with the task in this way, the user can provide (e.g., speak out) the candidate phrase to the digital assistant of the electronic device to cause the device to perform the task associated with the candidate phrase. The candidate phrase associated with the corresponding task may be referred to herein as a voice shortcut. In some examples, further in response to the selection that completes the enabling representation 1054, the electronic device 1000 displays a settings interface 1004 on the display 1001 (e.g., replacing the display of the completion interface 1052), as Figure 10M shown.

[0368] In some examples, since the task of the candidate task enabling representation 1008 is assigned to a task, the candidate task enabling representation 1008 is not included in the candidate task portion 1006. In some examples, the candidate task portion 1006 alternatively includes a candidate task suggestion 1094 such that the candidate task portion 1006 includes at least a threshold number of candidate task enabling representations.

[0369] In some examples, if one or more candidate task enabling representations are associated with the corresponding task, the settings interface 1004 includes a user shortcut enabling representation 1096. In some examples, when displaying the settings interface 1004, the electronic device 1000 detects the selection of the user shortcut enabling representation 1096. As Figure 10N shown, the selection of the user shortcut enabling representation 1096 is the tap gesture 1098. In some examples, in response to the selection of the user shortcut enabling representation 1096, the electronic device 1000 displays a user shortcut interface 1002A on the display 1001 (e.g., replacing the display of the settings interface 1004), as Figure 10O shown.

[0370] The user shortcut interface 1002A includes an edit enabling representation 1004A, a return enabling representation 1006A, a shortcut enabling representation 1008A, and an additional task enabling representation 1010A. In some examples, in response to the selection of the return enabling representation 1006A, the electronic device 1000 displays the settings interface 1004( Figure 10N ). In some examples, in response to the selection of the edit enabling representation 1004A, the electronic device 1000 displays an interface through which the voice shortcut associated with the shortcut enabling representation 1008A can be deleted. In some examples, in response to the selection of the additional task enabling representation 1010A, the electronic device 1000 displays an interface including a global task interface 1018A including one or more candidate task enabling representations such as Figure 10S as will be described in further detail below.

[0371] In some examples, when displaying the user shortcut interface 1002A, the electronic device 1000 detects a selection of the shortcut enabling indication 1008A. As Figure 10P shown, the selection of the shortcut enabling indication 1008A is a tap gesture 1004A. In some examples, in response to the selection of the shortcut enabling indication 1008A, the electronic device 1000 displays, on the display 1001, a completion interface 1054 for the shortcut enabling indication 1008A (e.g., replacing the display of the user shortcut interface 1002A), as Figure 10Q shown. The completion interface 1054 may include a delete enabling indication 1014A. In response to the selection of the delete enabling indication 1014A, the electronic device 1000 deletes the voice shortcut associated with the task. Deleting the shortcut in this manner may include disassociating the voice shortcut from the task such that providing the voice shortcut to the digital assistant of the electronic device 1000 does not cause the electronic device 1000 to execute the task.

[0372] In some examples, when displaying the settings interface 1004, the electronic device 1000 detects a selection of the additional task enabling indication 1014. As Figure 10R shown, the selection of the additional task enabling indication 1014 is a tap gesture 1016A. In some examples, in response to the selection of the shortcut enabling indication 1014, the electronic device 1000 displays, on the display 1001, a global task interface 1018A (e.g., replacing the display of the settings interface 1004), as Figure 10S shown.

[0373] For each of the multiple applications, the global task interface 1018A includes a corresponding set of candidate task enabling indications. As an example, the global task interface 1018A includes a set of candidate task enabling indications 1020A associated with the active application, a set of candidate task enabling indications 1026A associated with the calendar application, and a set of candidate task enabling indications 1036A associated with the music application. The set of candidate task enabling indications 1020A may optionally include a "start fitness" candidate task enabling indication 1022A and a "view daily progress" candidate task enabling indication 1024A. The set of candidate task enabling indications 1026A may optionally include a "send lunch invitation" candidate task enabling indication 1028A, an "arrange meeting" candidate task enabling indication 1030A, and a "clear day's activities" candidate task enabling indication 1032A. The set of candidate task enabling indications 1036A may optionally include a "play fitness playlist" candidate task enabling indication 1038A and a "start R&B radio" candidate task enabling indication 1040A.

[0374] In some examples, the candidate task affordance representations of the global task interface 1018A are searchable. As an example, when the global task interface 1018A is being displayed, the electronic device detects a swipe gesture on the display 1001, such as Figure 10T the swipe gesture 1046A. In response to the swipe gesture 1046A, the electronic device slides the global task interface 1018A in a downward direction to display (e.g., present) a search field 1048A that can be used to search for the candidate task affordance representations of the global task interface 1018A, as Figure 10U shown.

[0375] In some examples, each set of candidate task affordance representations displayed by the electronic device 1000 can be a subset of all available candidate task affordance representations of the corresponding application. Thus, the user can select an application task column affordance representation such as the application-specific task column affordance representations 1034A, 1042A, to display one or more additional candidate task affordance representations of the application corresponding to the application task column affordance representation. For example, when the global task interface 1018A is being displayed, the electronic device 1000 detects a selection of the application task column affordance representation 1042A. As Figure 10V shown, the selection of the application task column affordance representation 1042A is a tap gesture 1050A. In some examples, in response to the selection of the application task column affordance representation 1042A, the electronic device 1000 displays an application task interface 1052A on the display 1001 (e.g., replacing the display of the global task interface 1018A), as Figure 10W shown. As shown, the application task interface 1052A includes a back affordance representation and candidate task affordance representations 1054A to 1070A, and the back affordance representation, when selected, causes the electronic device 1000 to display the global task interface 1018A.

[0376] In some examples, when the settings interface 1004 is being displayed, the electronic device 1000 detects a swipe gesture on the display 1001 such as Figure 10X the swipe gesture 1074A. As Figure 10Y shown, in response to the swipe gesture 1074A, the electronic device 1000 slides the settings interface 1004 in an upward direction to display (e.g., present) various settings. When enabled, these settings adjust the way the electronic device 1000 displays candidate task affordance representations and suggestion affordance representations, such as with reference to Figures 8A - 8AFThose described. For example, the settings interface 1004 includes a quick enable setting 1076A that, when enabled, allows the electronic device 1000 to display candidate task enabling representations as described herein. As another example, the settings interface 1004 includes a search suggestion enable setting 1078A that, when enabled, allows the electronic device 1000 to display suggestion enabling representations on a search screen interface. As another example, the settings interface 1004 includes a find suggestion enable setting 1080A that, when enabled, allows the electronic device 1000 to display suggestion enabling representations on a find results screen interface. As yet another example, the settings interface 1004 includes a lock screen suggestion enable setting 1082A that, when enabled, allows the electronic device 1000 to display suggestion enabling representations on a lock screen interface.

[0377] The settings interface 1004 also includes application enabling representations 1084A and 1086A, each associated with a respective application. In some examples, when the settings interface 1004 is displayed, the electronic device 1000 detects a selection of the application enabling representation 1086A. As Figure 10Y shown, the selection of the application enabling representation 1086A is a tap gesture 1088A. In some examples, in response to the selection of the application enabling representation 1086A, the electronic device 1000 displays an application settings interface 1090A on the display 1001 (e.g., replacing the display of the settings interface 1004), as Figure 10Z shown.

[0378] In some examples, the application settings interface 1090A includes an application user shortcut enabling representation 1092A that may correspond to a voice shortcut previously generated by the user, e.g., using the voice shortcut generation process described herein. The application settings interface 1090A may also include candidate task enabling representations 1094A through 1098A, where each may correspond to a respective task associated with the application of the application settings interface 1090A. In some examples, the application settings interface 1090A also includes an application task list enabling representation 1002B. In some examples, in response to a selection of the application task list enabling representation 1002B, the electronic device 1000 displays an application task interface such as the application task interface 1052A( Figure 10W(e.g., replacing the display of the application settings interface 1090A). In some examples, the application settings interface 1090A further includes an application suggestion enable setting 1004B, which when enabled allows the electronic device 1000 to display suggestion enabling representations associated with applications on the search screen interface, the search result screen interface, and / or the keyboard interface. In some examples, the application settings interface 1090A further includes an application suggestion enable setting 1006B, which when enabled allows the electronic device 1000 to display suggestion enabling representations associated with applications on the lock screen interface.

[0379] In some examples, the application settings interface 1090A further includes an edit enabling representation 1010B. In some examples, when the application settings interface 1090A is displayed, the electronic device 1000 detects a selection of the edit enabling representation 1010B. As Figure 10AA shown, the selection of the edit enabling representation 1010B is a tap gesture 1012B. In some examples, in response to the selection of the edit enabling representation 1010B, the electronic device 1000 displays an application-specific edit interface 1014B (e.g., replacing the display of the application settings interface 1090A) on the display 1001, as Figure 10AB shown. In some examples, the application-specific edit interface 1014B is displayed on the first part of the application settings interface 1090A and visually obscures the second part of the application settings interface 1090A (e.g., dims, blurs).

[0380] In some examples, displaying the application settings interface 1090A includes displaying a selection enabling representation (e.g., selection enabling representation 1020B) for each application user shortcut enabling representation (e.g., application user shortcut enabling representation 1092A) of the application settings interface 1090A. In an example operation, the electronic device 1000 detects a selection of the selection enabling representation 1020B, indicating a selection of the corresponding application user shortcut enabling representation 1092A and the user's intention to delete the application user shortcut enabling representation 1092A. In some examples, the user can confirm the deletion by selecting a confirmation enabling representation 1018B or abandon the deletion by selecting a cancellation enabling representation 1016B.

[0381] In Figure 10AC it, the electronic device 1000 displays an application permission interface 1022B for a specific application. The application permission interface 1022 includes a suggested permission enabling representation 1026B. When the application permission interface 1022 is displayed, the electronic device 1000 detects a selection of the suggested permission enabling representation 1026B. As Figure 10ACAs shown, it is recommended that the selection for enabling the permission indication 1026B is the tap gesture 1032B. In some examples, in response to the selection of the recommended permission enabling indication 1026B, the electronic device 1000 displays an application settings interface such as Figure 10AD the application settings interface 1034B (e.g., replacing the display of the application-specific permission interface 1022B).

[0382] In some examples, an application-specific settings interface can be used to initiate the voice shortcut generation process. Thus, in some examples, the application settings interface 1034B can include candidate task enabling indications 1038B. In some examples, when the application settings interface 1034B is displayed, the electronic device 1000 detects the selection of the candidate task enabling indication 1038B. As Figure 10AE shown, the selection of the candidate task enabling indication 1038B is the tap gesture 1050B. In response to the selection of the candidate recommendation enabling indication 1038B, the electronic device 1000 displays a task-specific interface 1054B associated with the task of the candidate task enabling indication 1038B, as Figure 10AF shown.

[0383] Referring to FIGS. AF - AH, as described above, the user can thereafter generate a voice shortcut by providing natural language voice input (e.g., "a cup of coffee") to the electronic device, and the electronic device then provides candidate phrases for association with the task of the candidate task enabling indication 1038B (e.g., ordering a large latte with oat milk from a nearby coffee shop). Once the voice shortcut for the task is generated, the electronic device 1000 displays the application settings interface 1034B. As Figure 10AI shown, the application settings interface 1034B includes an application user shortcut enabling indication 1036B. In some examples, in response to the selection of the application-specific user shortcut enabling indication 1036B, the electronic device displays the voice shortcut associated with the task of the application.

[0384] In some examples, an application interface such as a third-party application interface can be used to initiate the voice shortcut generation process. For example, as Figure 10AJ shown, the user can use the application to complete a task (e.g., order coffee). In response, the application can display a task completion interface 1060B including candidate task recommendation enabling indications 1062B. In some examples, in response to the selection of the candidate task recommendation enabling indication 1062B, the electronic device initiates the voice shortcut generation process.

[0385] Figures 11A - 11Bis a flowchart showing a method for providing voice shortcuts according to some embodiments. Method 1100 is performed at a device (e.g., device 104, device 122, device 200, device 600, device 700, device 1000) having a display, one or more input devices (e.g., touchscreen, microphone, camera), and a wireless communication radio (e.g., Bluetooth connection, WiFi connection, mobile broadband connection such as 4G LTE connection). In some embodiments, the display is a touch-sensitive display. In some embodiments, the display is not a touch-sensitive display. In some embodiments, the electronic device includes multiple cameras. In some embodiments, the electronic device includes only one camera. In some embodiments, the device includes one or more biometric sensors, optionally including a camera, such as an infrared camera, a thermal imaging camera, or a combination thereof. Some operations in method 1100 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0386] As described below, providing candidate phrases based on natural language voice input and associating the candidate phrases with corresponding tasks allows a user to accurately and efficiently generate user-specific voice shortcuts that can be used to perform tasks on an electronic device. For example, allowing the user to associate voice shortcuts with tasks in this way allows the user to visually confirm that the desired voice shortcut has been selected and assigned to the correct task, thereby reducing the likelihood of incorrect or unwanted associations. Thus, providing candidate phrases in the described manner makes the use of the electronic device more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn further reduces power usage and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0387] In some examples, the electronic device (104, 200, 600, 700, 1000) displays a plurality of candidate task enabling representations (1008, 1010, 1012, 1038B), including a candidate task enabling representation (1008, 1038B) associated with a task (1102). In some examples, the plurality of tasks are displayed in a settings interface or an application interface. In some examples, each of the plurality of tasks is a task recommended, for example, based on the context of the electronic device.

[0388] In some examples, an electronic device displays a plurality of candidate task affordance representations including candidate task affordance representations associated with a task, including displaying an application interface (1034B) including the plurality of candidate task affordance representations. In some examples, the application interface is an interface of a third-party application (1062B) that includes one or more affordance representations (corresponding to the respective one or more candidate tasks), and a user can select the affordance representations to cause a first user interface to be displayed and thereby create a voice shortcut for the selected task. In some examples, the first user interface (1018) is overlaid on the application interface. In some examples, the first user interface is overlaid on a portion of the application interface. In some examples, the first user interface is overlaid on the entire application interface. In some examples, the one or more affordance representations correspond to one or more respective tasks and are displayed in response to completion of the one or more tasks.

[0389] In some examples, an electronic device detects a set of inputs (1016, 1032) including a first user input (1016) that corresponds to a selection of a candidate task affordance representation associated with a task (1104).

[0390] In some examples, in response to the first user input, the electronic device displays a fourth interface (1018) associated with the task. In some examples, the fourth user interface is an initial task-specific interface for generating a voice shortcut. In some examples, the interface specifies the task associated with the voice shortcut (1022) and also includes an indication of the associated application (e.g., the name of the application) (1026) and / or an icon associated with the application (1020, 1026) (e.g., a donation). In some examples, the interface includes a cancel affordance representation. In response to a selection of the cancel affordance representation, the electronic device returns to the immediately previous interface. In some examples, the interface includes a record affordance representation (1030). In response to a selection of the record affordance representation, the electronic device records a voice input when displaying a voice recording interface (1034). In some examples, the fourth interface includes a first suggested voice shortcut phrase (1028). In some examples, the electronic device displays the suggested voice shortcut, and the user can adopt the shortcut as the voice shortcut for the relevant task.

[0391] In some examples, in response to detecting the set of user inputs, the electronic device displays a first interface (1034) for generating a voice shortcut associated with a task (1106). In some examples, the first interface is a voice recording interface. In some examples, the voice recording interface includes a prompt (1046) for the user to record a phrase that will be used as a voice shortcut for initiating the task. In some examples, the voice recording interface includes a real-time preview (1042) of the natural language voice input provided by the user. In some examples, the real-time preview is an audio waveform.

[0392] In some examples, when displaying the first interface (1108), the electronic device receives (e.g., samples, obtains, captures) natural language voice input (1110) via an audio input device. In some examples, the natural language voice input is a phrase spoken by a user of the device. In some examples, the electronic device receives the voice input for a pre-determined time and / or until the user selects a stop enabling indication (1044).

[0393] In some examples, when receiving the natural language voice input, the electronic device provides (e.g., displays) a real-time (e.g., immediate) preview of the natural language voice input. In some examples, the real-time preview is a waveform of the natural language input and / or an immediate display of the speech-to-text conversion of the natural language input. In some examples, the real-time preview of the natural language voice input is a visual waveform of the natural language voice input.

[0394] In some examples, in response to receiving the natural language voice input, the electronic device determines candidate phrases based on the natural language voice input. In some examples, determining candidate phrases based on the natural language voice input includes using natural language processing to determine the candidate phrases. In some examples, when the user provides the natural language voice input, the electronic device uses natural language processing to convert the natural language voice input into text.

[0395] In some examples, determining candidate phrases based on the natural language voice input includes providing the natural language voice input to another electronic device and receiving a representation (e.g., a text representation, a vector representation) of the candidate phrases from the other electronic device. In some examples, the electronic device provides the natural language voice input and / or a representation of the NL voice input to a server, and the server returns one or more candidate phrases.

[0396] In some examples, determining a candidate phrase based on a natural language voice input includes determining whether the natural language voice input meets a phrase similarity criterion. Based on determining that the natural language voice input meets the phrase similarity criterion, the electronic device indicates (e.g., displays an indication) that the natural language voice input meets the similarity criterion. In some examples, the electronic device ensures that the natural language voice input is different from one or more predetermined phrases (e.g., "call 911"). In some examples, the electronic device ensures that the natural language voice input is sufficiently different from one or more predetermined phrases. As an example, the electronic device may ensure that the similarity of the representation of the natural language voice input to each of the one or more predetermined phrases does not exceed a similarity threshold. In some examples, if the natural language voice input does not meet the similarity criterion, the electronic device will notify the user that the provided voice input is insufficient and / or not allowed. The electronic device may also request the user to provide a new natural language voice input. In some examples, the electronic device also compares the natural language input with phrases associated with one or more stored voice shortcuts. In some examples, the electronic device indicates to the user to provide a new natural language voice input and / or requests confirmation that the user intends for the natural language voice input to replace one or more other similarly worded voice shortcuts. In some examples, the replaced voice shortcuts are deleted.

[0397] In some examples, when the electronic device displays the first interface (1034), it determines whether a natural language voice input is received within a threshold amount of time. In some examples, based on determining that no natural language voice input is received within the threshold amount of time, the electronic device displays a second suggested voice shortcut phrase (1046) ("You can say something like 'Show me the side camera'") on the first interface. In some examples, the suggested voice shortcuts provided on the first interface (1034) are the same as those provided on the second interface (1028). In some examples, the suggested voice shortcuts for the first interface and the second interface are different.

[0398] In some examples, further when displaying the first interface, the electronic device displays a candidate phrase (1046) in the first interface (1034), where the candidate phrase is based on a natural language voice input (1120). In some examples, the candidate phrase is a speech-to-text conversion of the natural language voice input and is displayed when the electronic device receives the natural language input. In some examples, the electronic device displays the candidate phrase in response to the natural language input and / or converts the natural language voice input into text in real time. Thus, in some examples, the electronic device continuously updates the display of the candidate phrase while the electronic device receives the natural language.

[0399] In some examples, after presenting the candidate phrase, the electronic device detects a second user input (1092)(1122) via a touch-sensitive surface. In some examples, the second user input is a selection that completes an enabling representation (1072). In some examples, the second user input is a selection of an enabling representation such as an acknowledgement enabling representation for a completion interface (1054).

[0400] In some examples, in response to detecting the second user input, the electronic device associates the candidate phrase with a task (1124). In some examples, the electronic device generates a voice shortcut such that the user's narration of the candidate phrase to the digital assistant causes the execution of the task.

[0401] In some examples, associating the candidate phrase with a task includes associating the candidate phrase with a first action and associating the candidate phrase with a second action. In some examples, the voice shortcut corresponds to multiple tasks. In some examples, the digital assistant suggests a combination of tasks, and the user can assign a phrase to initiate the execution of the task combination. As an example, the voice shortcut "secure the house" may correspond to a task of turning off the lights and locking the door.

[0402] In some examples, the electronic device receives user voice input (e.g., natural language voice input) via an audio input device. In some examples, the electronic device determines whether the user voice input includes a candidate phrase. In some examples, based on determining that the user voice input includes a candidate phrase, the electronic device executes a task. In some examples, based on determining that the user voice input does not include a candidate phrase, the electronic device abandons the execution of the task.

[0403] In some examples, after associating the candidate phrase with a task, the electronic device presents a second interface that includes an edit enabling representation. In some examples, the electronic device presents the second interface at the end of the voice shortcut generation process. In some examples, the user can navigate to a voice shortcut in an interface (1004A) listing one or more stored voice shortcuts (1008A) and select the voice shortcut to present the second interface. In some examples, the second user interface (1054) includes a text representation (1060) of the task, an indication (1058) of the task, the candidate phrase (1064), an edit enabling representation (1066) that, when selected, causes the presentation of the candidate shortcut, a cancellation enabling representation (1056) that, when selected, causes the device to cancel the voice shortcut generation process, and a completion enabling representation (1054) that, when selected, causes the electronic device to associate the candidate phrase with the task or maintain the association of the candidate phrase with the task (if already associated).

[0404] In some examples, the electronic device detects a third user input indicating a selection of an edit enablement representation (1068). In some examples, in response to detecting the third user input, the electronic device displays a plurality of candidate phrase enablement representations (1082, 1084, 1086), the enablement representations including a first candidate phrase enablement representation corresponding to a candidate phrase and a second candidate phrase enablement representation corresponding to another candidate phrase. In some examples, the another candidate phrase is based on a natural language voice input. In some examples, in response to the selection of the edit enablement representation, the electronic device displays an edit interface (1070) including a plurality of candidate phrases. In some examples, the user can select one of the candidate phrases to associate with a task. In some examples, each of the plurality of candidate phrases is a speech-to-text candidate generated using one or more NLP methods based on a natural language voice input. In some examples, the NL voice input is provided to another device (e.g., a backend server), which returns one or more candidate phrases to the electronic device. In some examples, candidate phrases are generated on both the electronic device and the backend server, and the electronic device selects one or more "best" candidate phrases (e.g., candidate phrases associated with the highest corresponding confidence scores).

[0405] In some examples, the electronic device detects another set of inputs (1090, 1092), the set of inputs including a fourth user input indicating a selection of the second candidate phrase enablement representation. In some examples, in response to detecting the another set of inputs, the electronic device associates the another candidate phrase with the task. In some examples, the user selects a new candidate phrase to associate with the task such that providing the new candidate phrase causes the task to be executed. In some examples, associating the new candidate task causes the electronic device to disassociate the previously associated candidate phrase from the task.

[0406] In some examples, after associating a candidate phrase with the task, the electronic device detects a fifth input (1092). In some examples, in response to detecting the fifth user input, the electronic device displays another plurality of candidate task enablement representations, where the another plurality of candidate task enablement representations do not include a candidate task enablement representation associated with the task. In some examples, after creating a voice shortcut, the electronic device displays a settings interface (1004) that lists candidate tasks suggested by the digital assistant and / or one or more applications. In some examples, if a task was previously suggested and the user created a shortcut for the task, the task is removed from the list of suggested tasks.

[0407] In some examples, the electronic device provides candidate phrases to another electronic device. In some examples, the generated voice shortcuts are provided to a backend server for subsequent voice input matching. In some examples, the generated voice shortcuts and associated tasks are provided to the backend server. In some examples, input is provided from the electronic device to the backend server, and the backend server determines whether the input corresponds to a voice shortcut. If the backend server determines that the input corresponds to a voice shortcut, the electronic device performs the task. In some examples, providing each of the shortcuts and tasks to the server in this way also allows the same shortcut to be used to perform the task on other electronic devices subsequently.

[0408] Note that the details of the process described above with respect to method 1100 (e.g., Figures 11A - 11B ) also apply to the methods described below in a similar manner. For example, method 1100 optionally includes one or more features of the various methods described in reference methods 900, 1300, 1500, 1700, and 1900. For example, as described in method 1200, providing voice shortcuts can be applied to generating voice shortcuts for use as described in method 1300. For the sake of brevity, these details are not repeated below.

[0409] The operations in the above methods are optionally implemented by running one or more functional modules in an information processing device such as a general-purpose processor (e.g., as described with respect to Figure 2A , Figure 4 , Figure 6A ) or an application-specific chip. Additionally, the operations described above with reference to Figures 10A - 10AJ are optionally implemented by the components depicted in Figures 2A - 2B . For example, display operation 1102, detection operation 1104, display operation 1106, reception operation 1110, display operation 1120, detection operation 1122, and association operation 1124 are optionally implemented by event classifier 270, event recognizer 280, and event handler 290. The event monitor 271 in event classifier 270 detects a contact on the touch-sensitive surface 604 ( Figure 6A ), and event distributor module 274 delivers the event information to application 236-1 ( Figure 2B)。The corresponding event recognizer 280 of application 236-1 compares the event information with the corresponding event definition 286 and determines whether the first contact at the first position on the touch-sensitive surface corresponds to a predefined event or sub-event, such as the selection of an object on the user interface. When the corresponding predefined event or sub-event is detected, the event recognizer 280 activates an event handler 290 associated with the detection of the event or sub-event. The event handler 290 optionally utilizes or invokes a data updater 276 or an object updater 277 to update the internal state 292 of the application. In some embodiments, the event handler 290 accesses the corresponding GUI updater 278 to update the content displayed by the application. Similarly, those skilled in the art will clearly know how to implement other processes based on Figures 2A - 2B the components depicted therein.

[0410] Figure 12 is a block diagram of a task recommendation system 1200 according to various examples. Figure 12 It is also used to illustrate one or more processes described below, including Figure 13 method 1300.

[0411] Figure 12 illustrates a task recommendation system 1200, which can be implemented, for example, on the electronic devices described herein, including but not limited to devices 104, 200, 400, and 600 ( Figure 1 , Figure 2A , Figure 4 and 6A - Figure 6B ). It should be understood that although the task recommendation system 1200 is described herein as being implemented on a mobile device, the task recommendation system 1200 can be implemented on any type of device such as a phone, laptop computer, desktop computer, tablet, wearable device (e.g., smartwatch), set-top box, television, home automation device (e.g., thermostat), or any combination or sub-combination thereof.

[0412] Generally, the task recommendation system 1200 provides task recommendations based on the context data of the electronic device. The task recommendations provided by the task recommendation system 1200 can in turn be used to provide recommended affordance representations, such as regarding Figures 8A - 8AFThose described. In some examples, the task recommendation system 1200 is implemented as a probabilistic model. The model can include one or more stages, including but not limited to a first stage 1210, a second stage 1220, and a third stage 1230, each of which is described in further detail below. In some examples, one or more of the first stage 1210, the second stage 1220, and the third stage 1230 can be combined into a single stage and / or divided into multiple stages. As an example, in an alternative embodiment, the task recommendation system 1200 can include a first stage and a second stage, where the first stage implements the functions of both the first stage 1210 and the second stage 1220.

[0413] In some examples, the context data of the electronic device indicates the state or context of the electronic device. As an example, the context data can indicate various characteristics of the electronic device during the time of performing a particular task. For example, for each task performed by the electronic device, the context data can indicate the location of the electronic device, whether the electronic device is connected to a network (e.g., a WiFi network), whether the electronic device is connected to one or more other devices (e.g., headphones), and / or the time, date, and / or weekday when the task is performed. If the electronic device is connected to a network or a device during the execution of the task, the context data can further indicate the name and / or type of the network or the device, respectively. The context data can also indicate whether a suggestion enabling representation associated with the task has been previously presented to the user and the way the user responds to the suggestion enabling representation. For example, the context data can indicate whether the user has ignored the task recommendation or used the task recommendation to perform the corresponding task.

[0414] In some examples, the electronic device receives context data from an application during the operation of the electronic device. In some examples, the application provides context data for a task when the task is being performed or after the task has been performed. The context data provided in this way can indicate the task and, optionally, one or more parameters associated with the task. In some examples, the context data provided by the application also includes data indicating the state of the electronic device when the task is being performed, as described above (e.g., the location of the electronic device, whether the electronic device is connected to a network, etc.). In other examples, the context data provided by the application does not include data indicating the state of the electronic device when the task is being performed, and in response to receiving the context data from the application, the electronic device obtains data indicating the state of the electronic device.

[0415] In some examples, the context data is provided by an application using any of a plurality of context data donation mechanisms. In at least one example, context data is received from an application in response to an API call that causes the application to return information about the execution of the application. As an example, a word processing application (e.g., Notepad, Word) can return an identifier (e.g., a numeric ID, a name) of the word processing document accessed by the user. As another example, a media playback application (e.g., Music, Spotify) can return identifiers of the current media item (e.g., a song), album, playlist, collection, and one or more media items recommended by the media playback application to the user.

[0416] In some examples, context data is received from an application in response to an API call that causes the application to provide a data structure that indicates (e.g., identifies) a task performed by the application. For example, the data structure can include values of one or more parameters associated with the task. In some examples, the data structure is implemented as an intent object data structure. Additional exemplary descriptions of the operation of the intent object data structure can be found in U.S. Patent Application 15 / 269,728, "APPLICATION INTEGRATION WITH A DIGITAL ASSISTANT," filed on September 18, 2016, which is incorporated herein by reference in its entirety.

[0417] In some examples, context data is received from an application in response to an API call that causes the application to provide an application-specific (e.g., third-party application) data structure that indicates a task performed by the application. For example, the application-specific data structure can include values of one or more parameters associated with the task. The application-specific data structure can optionally further indicate which parameter values and combinations of parameter values of the task can be specified when providing task suggestions. Thus, while a task with M parameters can have up to 2 M -1 allowed parameter combinations, the application-specific data structure can explicitly reduce the number of allowed combinations. In some examples, the application-specific data structure indicates whether the task is a background task or a task that requires confirmation, as previously described with respect to Figures 8A - 8AF as described.

[0418] In some examples, context data is provided by an application whenever a task is executed on an electronic device. For example, the context data can be provided when the task is executed or after the task is executed. For example, in some examples, the electronic device detects that the application has been closed, and in response, requests context data from the application using one or more context data donation mechanisms.

[0419] In some examples, context data is selectively provided based on the type of task. As an example, context data can be provided in response to a user sending a text message or making a call, rather than in response to a user navigating to a web page using a browser or unlocking an electronic device. In some examples, context data is selectively provided based on the context of the electronic device. For example, if the charge level of the battery of the electronic device is below a threshold charge level, context data may not be provided.

[0420] In some examples, each type of context data (e.g., the location of the electronic device, whether the electronic device is connected to a network, etc.) is associated with a corresponding context weight. For example, the context weight can be used to influence the way in which context data is used to determine probabilities. As an example, a context type that is determined to be a stronger predictor of user behavior can be associated with a relatively large weight, and / or a context type that is determined to be a weaker predictor of user behavior can be associated with a relatively small weight. As an example, a task performed by a user when the electronic device is in a particular location can be a better indicator of user behavior than a task performed by the user on a particular workday. Thus, location context data can be associated with a relatively large weight, and workday context data can be associated with a relatively small weight. As another example, whether a user selected a suggested affordance for a task can be associated with a relatively large weight because such context data can strongly indicate whether the user is likely to select a subsequent suggested affordance for the task.

[0421] In some examples, the context weight is determined by the electronic device. In other examples, the weight is determined by another device using aggregated data. For example, in some examples, anonymous context data can be provided by the electronic device and / or one or more other devices to a data aggregation server. Subsequently, the data aggregation server can determine which type of context data is more discriminative and / or is a stronger predictor of user behavior. In some examples, the data aggregation server employs one or more machine learning techniques to determine which type of context data is relatively discriminative and / or is a stronger predictor of user behavior. Based on this determination, the data aggregation server can determine the weight for each type of context data and provide the weight to the electronic device. In some examples, a user can choose to opt out (e.g., opt out) of providing anonymous data to another device to determine context weights.

[0422] The privacy of a user of an electronic device can be protected by ensuring that the context data used by a data aggregation server is anonymous. The anonymous context data can include removing the user's name, a specific location recorded by the electronic device, the name of a WiFi network, and any other user-specific information. For example, the anonymous context data can specify that a user performs a certain task 20% of the time when in the same location, but does not specify which task or location.

[0423] In an example operation, a first stage 1210 performs task-specific modeling based on context data of an electronic device. In some examples, performing task-specific modeling includes determining one or more probabilities (e.g., task probabilities) for each of a plurality of tasks (e.g., one or more tasks previously performed by the electronic device and / or one or more tasks that can be performed by the electronic device). Each task probability determined in this way indicates the likelihood that the user will perform the task given the context of the electronic device.

[0424] In some examples, determining task probabilities in this way includes generating and / or updating one or more histograms based on the context data. For example, the first stage 1210 can generate a corresponding set of histograms for each task, indicating the probability that the user performs the task given various contexts of the electronic dev...

Claims

1. A method for displaying a suggested actionable representation, comprising: at an electronic device having a display and a touch-sensitive surface: receiving a plurality of media items from an application; receiving first context data associated with the electronic device; determining a first task based on the plurality of media items and the first context data; determining whether the first task meets a suggestion criterion; if it is determined that the first task meets the suggestion criterion, displaying a first suggested actionable representation corresponding to the first task on the display; if it is determined that the first task does not meet the suggestion criterion, forgoing displaying the first suggested actionable representation; determining a second task based on the first context data, wherein the second task is a task other than a task for playing back a media item; determining whether the second task meets the suggestion criterion; if it is determined that the second task meets the suggestion criterion, displaying a second suggested actionable representation corresponding to the second task on the display; and if it is determined that the second task does not meet the suggestion criterion, forgoing displaying the second suggested actionable representation.

2. The method according to claim 1, wherein the first task is a task for playing back a media item among the plurality of media items.

3. The method according to claim 2, wherein the first task specifies a specific playback time.

4. The method according to any one of claims 1-3, wherein receiving the plurality of media items from the application comprises: receiving a vector including the plurality of media items from the application.

5. The method according to any one of claims 1-3, wherein receiving the plurality of media items from the application comprises: receiving the plurality of media items when displaying another application different from the application.

6. The method according to any one of claims 1-3, further comprising: requesting the plurality of media items from the application before receiving the plurality of media items from the application.

7. The method according to any one of claims 1-3, wherein: the application is a first application; the first task is associated with the first application; and the first suggested actionable representation is associated with a second application different from the first application.

8. The method according to any one of claims 1-3, further comprising: at a second electronic device: using the application, determining one or more media items previously played on the electronic device; and generating the plurality of media items based on the determined one or more media items.

9. The method according to claim 8, wherein generating the plurality of media items comprises: identifying a category of at least one media item among the one or more media items previously played on the electronic device; and identifying media items associated with the identified category.

10. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device having a display, cause the electronic device to perform the following operations: Receive a plurality of media items from an application; Receive first context data associated with the electronic device; Determine a first task based on the plurality of media items and the first context data; Determine whether the first task meets a recommendation criterion; If it is determined that the first task meets the recommendation criterion, display a first recommended affordance representation corresponding to the first task on the display; If it is determined that the first task does not meet the recommendation criterion, forgo displaying the first recommended affordance representation; Determine a second task based on the first context data, wherein the second task is a task other than a task for playing back a media item; Determine whether the second task meets the recommendation criterion; If it is determined that the second task meets the recommendation criterion, display a second recommended affordance representation corresponding to the second task on the display; and If it is determined that the second task does not meet the recommendation criterion, forgo displaying the second recommended affordance representation.

11. The computer-readable storage medium according to claim 10, wherein the first task is a task for playing back a media item among the plurality of media items.

12. The computer-readable storage medium according to claim 11, wherein the first task specifies a specific playback time.

13. The computer-readable storage medium according to any one of claims 10-12, wherein receiving the plurality of media items from the application comprises: Receiving a vector including the plurality of media items from the application.

14. The computer-readable storage medium according to any one of claims 10-12, wherein receiving the plurality of media items from the application comprises: Receiving the plurality of media items when displaying another application different from the application.

15. The computer-readable storage medium according to any one of claims 10-12, wherein the instructions further cause the electronic device to perform the following operations: Request the plurality of media items from the application before receiving the plurality of media items from the application.

16. The computer-readable storage medium according to any one of claims 10-12, wherein: The application is a first application; The first task is associated with the first application; and The first recommended affordance representation is associated with a second application different from the first application.

17. The computer-readable storage medium according to any one of claims 10-12, wherein the instructions further cause a second electronic device to perform the following operations: Using the application, determine one or more media items previously played on the electronic device; and Generate the plurality of media items based on the determined one or more media items.

18. The computer-readable storage medium according to claim 17, wherein generating the plurality of media items comprises: Identifying a category of at least one media item among the one or more media items previously played on the electronic device; and Identifying media items associated with the identified category.

19. An electronic device, comprising: A display; One or more processors; Memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: Receiving a plurality of media items from an application; Receiving first context data associated with the electronic device; Determining a first task based on the plurality of media items and the first context data; Determining whether the first task meets a recommendation criterion; If it is determined that the first task meets the recommendation criterion, displaying a first recommended enabling representation corresponding to the first task on the display; If it is determined that the first task does not meet the recommendation criterion, forgoing displaying the first recommended enabling representation; Determining a second task based on the first context data, wherein the second task is a task other than a task for playing back a media item; Determining whether the second task meets a recommendation criterion; If it is determined that the second task meets the recommendation criterion, displaying a second recommended enabling representation corresponding to the second task on the display; and If it is determined that the second task does not meet the recommendation criterion, forgoing displaying the second recommended enabling representation.

20. The electronic device according to claim 19, wherein the task is a task for playing back a media item among the plurality of media items.

21. The electronic device according to claim 20, wherein the task specifies a specific playback time.

22. The electronic device according to any one of claims 19 - 21, wherein receiving the plurality of media items from the application comprises: Receiving a vector including the plurality of media items from the application.

23. The electronic device according to any one of claims 19 - 21, wherein receiving the plurality of media items from the application comprises: Receiving the plurality of media items when displaying another application different from the application.

24. The electronic device according to any one of claims 19 - 21, the one or more programs including instructions for performing the following operation: Requesting the plurality of media items from the application before receiving the plurality of media items from the application.

25. The electronic device according to any one of claims 19 - 21, wherein: The application is a first application; The task is associated with the first application; and The first recommended enabling representation is associated with a second application different from the first application.

26. The electronic device according to any one of claims 19 - 21, the one or more programs including instructions for performing the following operations: At a second electronic device: Using the application, determining one or more media items previously played on the electronic device; and Generating the plurality of media items based on the determined one or more media items.

27. The electronic device according to claim 26, wherein generating the plurality of media items comprises: Identifying a category of at least one media item among the one or more media items previously played on the electronic device; and Identify media items associated with the identified categories.

28. An electronic device comprising:[[]]END]] one or more processors; a memory; and means for performing the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • System and method for inferring user intent from speech inputs

    US10176167B2

  • Method and apparatus for integrating manual input

    US20020015024A1

  • Acceleration-based theft detection system for portable electronic devices

    US20050190059A1

  • Methods and apparatuses for operating a portable device based on an accelerometer

    US20060017692A1

  • Gestures for touch sensitive input devices

    US20060026521A1