Task modification after task initiation

By receiving the second user's speech input to modify or cancel the intelligent automated assistant's task, the problem of users consuming time to modify tasks after the task is initiated is solved, interaction efficiency is improved, and power is saved.

CN119137659BActive Publication Date: 2025-10-17APPLE INC
View PDF 24 Cites 0 Cited by

Patent Information

Application Number
CN202380037367.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-09-21
Filing Date
2023-05-05
Publication Date
2025-10-17
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

In the prior art, it is time-consuming and inefficient for a user to modify or cancel a task after it is initiated by an intelligent automated assistant, and it is difficult to efficiently determine and execute the need to modify or cancel a task.

Method used

By receiving the second user's speech input, determining the modification of the first task, and executing the second task to modify the execution part of the first task, efficient task modification or cancellation is achieved.

Benefits of technology

Improves user efficiency in interacting with intelligent automated assistants, reduces the number of inputs required, conserves power usage, and extends device battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119137659B_ABST
    Figure CN119137659B_ABST
Patent Text Reader

Abstract

Systems and processes are provided for operating an intelligent automated assistant. An example process includes, at an electronic device with one or more processors and memory: performing a first task specified in a first user verbal input; receiving a second user verbal input; and in accordance with a determination that the second user verbal input includes a modification to the first task, performing a second task, wherein performance of the second task modifies at least a portion of the performance of the first task.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 340,378, filed May 10, 2022, entitled “TASK MODIFICATION AFTER TASKINITIATION,” and U.S. Patent Application No. 17 / 949,941, filed September 21, 2022, entitled “TASK MODIFICATION AFTER TASK INITIATION.” The contents of each of these patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates generally to intelligent automated assistants, and more particularly, to modifying tasks after initiation by an intelligent automated assistant. Background Art

[0004] Intelligent automated assistants (or digital assistants) can provide a convenient interface between human users and electronic devices. Such assistants can allow users to interact with devices or systems using natural language in the form of speech and / or text. For example, a user can provide verbal input containing a user request to a digital assistant running on an electronic device. The digital assistant can interpret the user's intent from the verbal input and operationalize the user's intent into tasks. These tasks can then be performed by executing one or more services of the electronic device, and relevant output responsive to the user's request can be returned to the user.

[0005] In some instances, a user may want to modify or cancel a task after it has been initiated by a digital assistant. However, providing the digital assistant with several different commands to cancel or modify a previously given command can be time-consuming and inefficient. Therefore, an efficient way to process user input to determine when and how to modify a task after it has been performed is desired. Summary of the Invention

[0006] An example method is disclosed herein. An example method includes, at an electronic device having one or more processors and a memory, performing a first task specified in a first user speech input; receiving a second user speech input; and, based on determining that the second user speech input includes a modification to the first task, performing the second task, wherein the performing of the second task modifies at least a portion of the performing of the first task.

[0007] Example non-transitory computer-readable media are disclosed herein. An example non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device, the one or more programs including instructions for: performing a first task specified in a first user verbal input; receiving a second user verbal input; and in accordance with a determination that the second user verbal input includes a modification to the first task, performing a second task, wherein performance of the second task modifies at least a portion of the performance of the first task.

[0008] Example electronic devices are disclosed herein. An example electronic device includes one or more processors; memory; and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: performing a first task specified in a first user verbal input; receiving a second user verbal input; and in accordance with a determination that the second user verbal input includes a modification to the first task, performing a second task, wherein performance of the second task modifies at least a portion of the performance of the first task.

[0009] An example electronic device includes: means for performing a first task specified in a first user verbal input; means for receiving a second user verbal input; and in accordance with a determination that the second user verbal input includes a modification to the first task, means for performing a second task, wherein performance of the second task modifies at least a portion of the performance of the first task.

[0010] Performing a second task, wherein performance of the second task modifies at least a portion of the performance of the first task, allows for efficient determination and performance of a task that modifies or undoes at least a portion of a previously performed task. For example, a digital assistant can be able to respond to a user input requesting that a task be reversed without requiring several inputs from the user specifying how the task should be reversed. Thus, the digital assistant can provide a more efficient interaction with the user, thereby increasing user engagement. Thus, the user can interact with the digital assistant with fewer inputs, which additionally reduces power usage and improves battery life of the device. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 A block diagram of a system and environment for implementing a digital assistant in accordance with various examples is shown.

[0012] Figure 2A A block diagram of a portable multifunctional device implementing a client-side portion of a digital assistant in accordance with various examples is shown.

[0013] Figure 2B A block diagram of example components for event processing in accordance with various examples is shown.

[0014] Figure 3 A portable multifunction device implementing a client-side portion of a digital assistant is shown in accordance with various examples.

[0015] Figure 4 A block diagram of an example multifunction device having a display and a touch-sensitive surface in accordance with various examples.

[0016] Figure 5A An example user interface showing a menu of applications on a portable multifunction device in accordance with various examples.

[0017] Figure 5B An example user interface of a multifunction device having a touch-sensitive surface separate from the display in accordance with various examples.

[0018] Figure 6A A personal electronic device is shown in accordance with various examples.

[0019] Figure 6B A block diagram of a personal electronic device is shown in accordance with various examples.

[0020] Figure 7A A block diagram of a digital assistant system or server portion thereof is shown in accordance with various examples.

[0021] Figure 7B Functions of a digital assistant shown in Figure 7A are shown in accordance with various examples.

[0022] Figure 7C A portion of an ontology is shown in accordance with various examples.

[0023] Figure 8 A block diagram of a digital assistant for modifying a task after execution is illustrated in accordance with various examples.

[0024] Figures 9A-9C An example modification to a task after execution by a digital assistant is illustrated in accordance with various examples.

[0025] Figures 10A-10C An example modification to a task after execution by a digital assistant is illustrated in accordance with various examples.

[0026] Figures 11A-11D An example modification to a task after execution by a digital assistant is illustrated in accordance with various examples.

[0027] Figure 12 A process for operating a digital assistant to modify a task after execution is illustrated in accordance with various examples. DETAILED DESCRIPTION

[0028] In the following description of examples, reference is made to the accompanying drawings, which form a part hereof, and in which are shown, by way of illustration, specific examples that can be practiced. It is to be understood that other examples can be used and structural changes can be made without departing from the scope of the various examples.

[0029] Although the following description uses the terms "first," "second," etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, a first input could be termed a second input, and, similarly, a second input could be termed a first input, without departing from the scope of the various described examples. The first input and the second input are both inputs, and in some cases, independent and different inputs.

[0030] The terminology used in the description herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various described examples and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0031] The term "if can be interpreted as meaning "when," or "while," or "in response to determining," or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined," or "if [a stated condition or event] is detected," can be interpreted as meaning "upon determining," or "in response to determining," or "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]," depending on the context.

[0032] 1. Systems and Environments

[0033] Figure 1A block diagram of a system 100 is shown in accordance with various examples. In some examples, the system 100 implements a digital assistant. The terms "digital assistant," "virtual assistant," "intelligent automated assistant," or "automated digital assistant" refer to any information processing system that interprets natural language input in spoken and / or textual form to infer user intent and perform actions based on the inferred user intent. For example, to act on the inferred user intent, the system performs one or more of the following steps: identify a task flow having steps and parameters designed to implement the inferred user intent, input specific requirements into the task flow according to the inferred user intent; execute the task flow by invoking programs, methods, services, APIs, etc.; and generate an output response to the user in audible (e.g., spoken) and / or visual form.

[0034] In particular, a digital assistant can accept user requests that are at least partially in the form of natural language commands, requests, statements, narratives, and / or queries. Generally, a user request seeks an informational answer or performance of a task from the digital assistant. A satisfactory response to a user request includes providing the requested informational answer, performing the requested task, or a combination of the two. For example, a user asks a question of the digital assistant, such as "Where am I now?" Based on the user's current location, the digital assistant answers "You are near the Central Park West entrance." The user also requests performance of a task, such as "Please invite my friends to my girlfriend's birthday party next week." In response, the digital assistant can confirm the request by saying "Okay, right away," and then transmit appropriate calendar invitations to each of the user's friends listed in the user's electronic address book on behalf of the user. During performance of the requested task, the digital assistant sometimes interacts with the user in a continuing conversation involving multiple exchanges of information over a long period of time. There are many other ways in which a user interacts with a digital assistant to request information or perform various tasks. In addition to providing spoken responses and taking programmed actions, a digital assistant also provides responses in other video or audio forms, such as text, reminders, music, videos, animations, etc.

[0035] As Figure 1 shown, in some examples, a digital assistant is implemented according to a client-server model. The digital assistant includes a client-side portion 102 (hereinafter "DA client 102") that executes on a user device 104, and a server-side portion 106 (hereinafter "DA server 106") that executes on a server system 108. The DA client 102 communicates with the DA server 106 over one or more networks 110. The DA client 102 provides client-side functionality, such as user-facing input and output processing, and communication with the DA server 106. The DA server 106 provides server-side functionality for any number of DA clients 102 that are each located on a respective user device 104.

[0036] In some examples, the DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, and an external service I / O interface 118. The client-facing I / O interface 112 facilitates client-facing input and output processing of the DA server 106. The one or more processing modules 114 utilize the data and models 116 to process speech input and determine a user intent based on the natural language input. Further, the one or more processing modules 114 perform task execution based on the inferred user intent. In some examples, the DA server 106 communicates with external services 120 over the one or more networks 110 to complete tasks or gather information. The external service I / O interface 118 facilitates such communication.

[0037] The user device 104 can be any suitable electronic device. In some examples, the user device 104 is a portable multifunction device (e.g., a device 200 as described below with reference to Figure 2A The user device 104 can be any suitable electronic device. In some examples, the user device 104 is a portable multifunction device (e.g., a device 200 as described below with reference to Figure 4 The user device 104 can be any suitable electronic device. In some examples, the user device 104 is a portable multifunction device (e.g., a device 200 as described below with reference to Figures 6A-6B The user device 104 can be any suitable electronic device. In some examples, the user device 104 is a portable multifunction device (e.g., a device 200 as described below with reference to iPod and devices. Other examples of portable multifunction devices include, but are not limited to, earbuds / headphones, speakers, and laptop or tablet computers. In addition, in some examples, the user device 104 is a non-portable multifunction device. Specifically, the user device 104 is a desktop computer, a game console, a speaker, a television, or a television set-top box. In some examples, the user device 104 includes a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). In addition, the user device 104 optionally includes one or more other physical user-interface devices, such as a physical keyboard, a mouse, and / or a joystick. Various examples of electronic devices, such as multifunction devices, are described in more detail below.

[0038] Examples of the communication network 110 include a local area network (LAN) and a wide area network (WAN), such as the Internet. The communication network 110 is implemented using any known network protocol, including various wired or wireless protocols such as Ethernet, Universal Serial Bus (USB), FireWire, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi, Voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.

[0039] The server system 108 is implemented on one or more stand-alone data processing devices or a distributed computer network. In some examples, the server system 108 also uses various virtual devices and / or services of third-party service providers (e.g., third-party cloud service providers) to provide the potential computing resources and / or infrastructure resources of the server system 108.

[0040] In some examples, the user device 104 communicates with the DA server 106 via a second user device 122. The second user device 122 is similar or identical to the user device 104. For example, the second user device 122 is similar to the user device 104 described below with reference to FIG. Figure 2A 、 Figure 4 and Figures 6A-6B The device 200, 400 or 600 described. The user device 104 is configured to be communicatively coupled to the second user device 122 via a direct communication connection (such as Bluetooth, NFC, BTLE, etc.) or via a wired or wireless network (such as a local Wi-Fi network). In some examples, the second user device 122 is configured to act as a proxy between the user device 104 and the DA server 106. For example, the DA client 102 of the user device 104 is configured to send information (e.g., a user request received at the user device 104) to the DA server 106 via the second user device 122. The DA server 106 processes the information and returns relevant data (e.g., data content in response to the user request) to the user device 104 via the second user device 122.

[0041] In some examples, the user device 104 is configured to send a truncated request for data to the second user device 122 to reduce the amount of information sent from the user device 104. The second user device 122 is configured to determine supplemental information to add to the truncated request to generate a complete request to send to the DA server 106. This system architecture can advantageously allow a user device 104 with limited communication capabilities and / or limited battery power (e.g., a watch or similar compact electronic device) to access services provided by the DA server 106 by using a second user device 122 (e.g., a mobile phone, a laptop, a tablet, etc.) with stronger communication capabilities and / or battery power as a proxy to the DA server 106. Although Figure 1 Only two user devices 104 and 122 are shown in FIG. 1, but it should be understood that in some examples, the system 100 can include any number and type of user devices configured to communicate with the DA server system 106 in this proxy configuration.

[0042] Although Figure 1 The digital assistant shown in FIG. 1 includes both a client-side portion (e.g., the DA client 102) and a server-side portion (e.g., the DA server 106), but in some examples, the functionality of the digital assistant is implemented as a standalone application installed on a user device. Moreover, the division of functionality between the client portion and the server portion of the digital assistant can vary in different implementations. For example, in some examples, the DA client is a thin client that provides only user-facing input and output processing functionality and delegates all other functionality of the digital assistant to a backend server.

[0043] 2. Electronic device

[0044] Attention is now directed to embodiments of an electronic device for implementing the client-side portion of a digital assistant. Figure 2AA block diagram illustrating portable multifunction device 200 with touch- sensitive display 212 in accordance with some embodiments is shown in Figure 2. Touch- sensitive display 212 is sometimes called a "touch screen" for convenience and is sometimes known as or called a "touch- sensitive display system." Device 200 includes memory 202 (which optionally includes one or more computer-readable storage mediums), memory controller 222, one or more processing units (CPUs) 220, peripherals interface 218, RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, input / output (I / O) subsystem 206, other input control devices 216, and external port 224. Device 200 optionally includes one or more optical sensors 264. Device 200 optionally includes one or more contact intensity sensors 265 for detecting intensity of contacts on device 200 (e.g., a touch- sensitive surface such as touch- sensitive display system 212 of device 200). Device 200 optionally includes one or more tactile output generators 267 for generating tactile outputs on device 200 (e.g., generating tactile outputs on a touch- sensitive surface such as touch- sensitive display system 212 of device 200 or touchpad 455 of device 400). These components optionally communicate over one or more communication buses or signal lines 203.

[0045] As used in the specification and claims, the term "intensity" of a contact on a touch- sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch- sensitive surface, or to a proxy for the force or pressure of the contact on the touch- sensitive surface. The intensity of a contact has a range of values that includes at least four distinct values and more typically includes hundreds of distinct values (e.g., at least 256). Intensity of a contact is, optionally, determined (or measured) using various approaches and various sensors or combinations of sensors. For example, one or more force sensors underneath (or adjacent to) the touch- sensitive surface optionally are used to measure force at various points on the touch- sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., a weighted average) to determine an estimated force at the point of contact. Similarly, a pressure- sensitive tip of a stylus is, optionally, used to determine a pressure of the stylus on the touch- sensitive surface. Alternatively, the size of the contact area detected on the touch- sensitive surface and / or changes thereto, the capacitance of the touch- sensitive surface proximate to the contact and / or changes thereto, and / or the resistance of the touch- sensitive surface proximate to the contact and / or changes thereto are, optionally, used as a substitute for the force or pressure of the contact on the touch- sensitive surface. In some implementations, the substitute measurements for contact force or pressure are used directly to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the substitute measurements). In some implementations, the substitute measurements for contact force or pressure are converted to an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of a contact as an attribute of a user input allows for user access to additional device functions with additional device features that would otherwise not be accessible. For example, a

[0046] As used in the specification and claims, the term "tactile output" refers to physical displacement of a device relative to a previous positioning of the device by a user with the user's sense of touch, physical displacement of a component (e.g., a touch-sensitive surface) of a device relative to another component (e.g., a housing) of the device, or displacement of the component relative to a center of mass of the device that will be detected by a user with the user's sense of touch. For example, in the case of contact between a device or component of a device and a surface of a user that is sensitive to touch (e.g., a finger, palm, or other part of a user's hand), the tactile output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in physical characteristics of the device or component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) will optionally be interpreted by a user as a "down click" or "up click" of a physical button even if the touch-sensitive surface does not move. In some cases, a user will feel a tactile sensation such as an "down click" or "up click" even if there is no movement of a physical button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's movement. As another example, movement of the touch-sensitive surface will optionally be interpreted or sensed by a user as an alteration in "haptic texture" of the touch-sensitive surface even if there is no change in smoothness of the touch-sensitive surface. While such interpretations of touch by a user will be subject to the individualized sensory perceptions of the user, there are many tactile sensations that are commonly perceived by most users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., "down click," "up click," "haptic texture"), the generated tactile output corresponds to physical displacement of the device or a component thereof that will generate that particular sensory perception to a typical (or average) user unless otherwise stated.

[0047] It should be appreciated that device 200 is only one example of a portable multifunction device, and that device 200 optionally has more or fewer components than shown, optionally combines two or more components, or optionally has a different configuration or arrangement of the components. The various components illustrated for device 200 can be implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application specific integrated circuits. Figure 2A The various components illustrated for device 200 can be implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0048] Memory 202 includes one or more computer-readable storage mediums. The computer-readable storage mediums are for example, tangible and non-transitory. Memory 202 includes high-speed random access memory and can also include non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory controller 222 controls access to memory 202 by other components of device 200.

[0049] In some examples, the non-transitory computer-readable storage medium of memory 202 is for storing instructions (e.g., for performing aspects of the processes described below) for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. In other examples, the instructions (e.g., for performing aspects of the processes described below) are stored on a non-transitory computer- readable storage medium (not shown) of server system 108, or divided between the non- transitory computer-readable storage medium of memory 202 and the non-transitory computer-readable storage medium of server system 108.

[0050] Peripheral interface 218 is used to couple input and output peripherals of the device to CPU 220 and memory 202. One or more processors 220 run or execute various software programs and / or sets of instructions stored in memory 202 to perform various functions of device 200 and process data. In some embodiments, peripheral interface 218, CPU 220, and memory controller 222 are implemented on a single chip, such as chip 204. In some other embodiments, they are implemented on separate chips.

[0051] RF (radio frequency) circuitry 208 receives and sends RF signals, also called electromagnetic signals. RF circuitry 208 converts electrical signals to / from electromagnetic signals and communicates with communications networks and other communications devices, via the electromagnetic signals. RF circuitry 208 optionally includes well-known circuitry for detecting / handling the RF signals, including, for example, antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, and so forth. RF circuitry 208 optionally communicates with networks, such as the Internet, also referred to as the World Wide Web (WWW), an intranet and / or a wireless network, such as a cellular telephone network, a wireless local area network (LAN) and / or a metropolitan area network (MAN), and other devices by wireless communication using RF signals. RF circuitry 208 optionally includes well-known circuitry for detecting / handling the RF signals, such as a Bluetooth®, Bluetooth® Low Energy (LE), Zigbee®, near-field communication (NFC) circuitry, and / or a wireless local area network (WLAN) such as Wi-Fi circuitry. The wireless communication optionally uses any of a plurality of communications standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Wideband-CDMA (W-CDMA), Long Term Evolution (LTE), Near Field Communication (NFC), Bluetooth, Bluetooth Low Energy (BLE), a Wireless Fidelity (Wi-Fi) such as IEEE 802.11a, IEEE 802.10b, IEEE 802.10g, IEEE 802.10η, and / or IEEE 802.1 lac, a Wireless LAN (WLAN), a Wi-MAX, a protocol for e-mail (e.g., Internet Message Access Protocol (IMAP) and / or

[0052] Audio circuitry 210, speaker 211, and microphone 213 provide an audio interface between a user and device 200. Audio circuitry 210 receives audio data from peripherals interface 218, converts the audio data to an electrical signal, and transmits the electrical signal to speaker 211. Speaker 211 converts the electrical signal to human-audible sound waves. Audio circuitry 210 also receives electrical signals converted by microphone 213 from sound waves. Audio circuitry 210 converts the electrical signal to audio data and transmits the audio data to peripherals interface 218 for processing. Audio data are retrieved from and / or transmitted to memory 202 and / or RF circuitry 208 by peripherals interface 218 in some embodiments. In some embodiments, audio circuitry 210 also includes a headset jack (e.g., 312 in FIG. 3). The headset jack provides an interface between audio circuitry 210 and removable audio input / output peripherals, such as output-only headphones or a headset with both output (e.g., stereo Figure 3

[0053] I / O subsystem 206 couples input / output peripherals on device 200, such as touch screen 212 and other input control devices 216, with peripherals interface 218. I / O subsystem 206 optionally includes display controller 256, optical sensor controller 258, intensity sensor controller 259, haptic feedback controller 261, and one or more input controllers 260 for other input or control devices. The one or more input controllers 260 receive / send electrical signals from / to other input control devices 216. The other input control devices 216 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, and so forth. In some alternate embodiments, input controller(s) 260 are, optionally, coupled with any (or none) of the following: a keyboard, infrared port, USB port, and a pointer device such as a mouse. The one or more buttons (e.g., 308 in FIG. 3) optionally include an up / down button for Figure 3 Figure 3

[0054] ​​​A quick press of the down button disengages a lock of the touch screen 212 or begins a process that uses gestures on the touch screen to unlock the device, as described in U.S. Patent Application 11 / 322,549, "Unlocking a Device by Performing Gestures on an Unlock Image," filed December 23, 2005; U.S. Patent No. 7,657,849, which are hereby incorporated by reference in their entirety. A longer press of the down button (e.g., 306) turns power to the device 200 on or off. The user is able to customi ze a functionality of one or more of the buttons. The touch screen 212 is used to implement virtual or soft buttons and one or more soft keyboards.

[0055] The touch-sensitive display 212 provides an input interface and an output interface between the device and a user. The display controller 256 receives and / or sends electrical signals from / to the touch screen 212. The touch screen 212 displays visual output to the user. The visual output includes graphics, text, icons, video, and any combination thereof (collectively termed "graphics"). In some embodiments, some or all of the visual output or

[0056] The touch screen 212 has a touch-sensitive surface, sensor, or set of sensors that accepts input from the user based on haptic and / or tactile contact. The touch screen 212 and the display controller 256 (along with any associated modules and / or sets of instructions in memory 202) detect contact (and any movement or breaking of the contact) on the touch screen 212 and convert the

[0057] The touch screen 212 uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light-emitting diode) technology, although other display technologies can be used in other embodiments. The touch screen 212 and the display controller 256 can detect contact and any movement or breaking of the contact using any of a plurality of touch sensing technologies now known or later developed, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 212. In an example embodiment, projected mutual capacitance sensing technology is used, such as that found in the iPhone® and iPod touch® from Apple Inc. of Cupertino, California. and iPod touch® from Apple Inc. of Cupertino, California.

[0058] In some embodiments, the touch-sensitive display of touch screen 212 is analogous to the multi-touch- sensitive touchpads described in the following U.S. Patents: 6,323,846 (Westerman et al), 6,570,557 (Westerman et al), and / or 6,677,932 (Westerman) and / or U.S. Patent Publication 2002 / 0015024 Al, which are hereby incorporated by reference in their entirety. However, touch screen 212 displays visual output from device 200, whereas a touch-sensitive touchpad does not provide visual output.

[0059] The touch-sensitive display in some embodiments of touch screen 212 is described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, "Multipoint Touch Surface Controller," filed May 2, 2006; (2) U.S. Patent Application No. 10 / 840,862, "Multipoint Touchscreen," filed May 6, 2004; (3) U.S. Patent Application No. 10 / 903,964, "Gestures For Touch Sensitive Input Devices," filed July 30, 2004; (4) U.S. Patent Application No. 11 / 048,264, "Gestures For Touch Sensitive Input Devices," filed January 31, 2005; (5) U.S. Patent Application No. 11 / 038,590, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices," filed January 18, 2005; (6) U.S. Patent Application No. 11 / 228,758, "Virtual Input Device Placement On A Touch Screen User Interface," filed September 16, 2005; (7) U.S. Patent Application No. 11 / 228,700, "Operation Of A Computer With A Touch Screen Interface," filed September 16, 2005; (8) U.S. Patent Application No. 11 / 228,737, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard," filed September 16, 2005; and (9) U.S. Patent Application No. 11 / 367,749, "Multi-Functional Hand-Held Device," filed March 3, 2006. All of these applications are hereby incorporated by reference in their entirety.

[0060] Touch screen 212 has, for example, a video resolution in the range of 100 dpi. In some embodiments, the touch screen has a video resolution in the range of about 160 dpi. The user can make contact with touch screen 212 using any suitable object or appendage, such as a stylus, finger, etc. In some embodiments, the user interface is designed to work primarily with finger-based contacts and gestures, which can be less precise than stylus-based input due to the larger area of contact of a finger on the touch screen. In some embodiments, the device translates the rough finger- based input into a precise pointer / cursor position or command for performing the action desired by the user.

[0061] In some embodiments, in addition to the touch screen, device 200 includes a touchpad (not shown) for activating or deactivating particular functions. In some embodiments, the touchpad is a touch-sensitive area that is a part of the device that does not display visual output, unlike the touch screen. The touchpad is a touch-sensitive surface separate from the touch screen 212 or an extension of the touch-sensitive surface formed by the touch screen.

[0062] Device 200 also includes power system 262 for powering the various components of device 200. Power system 262 includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., a light-emitting diode (LED)) and any other components associated with the generation, management and distribution of power in portable devices.

[0063] Device 200 also includes one or more optical sensors 264. Figure 2A Optical sensor(s) 264 are coupled to optical sensor controller 258 in I / O subsystem 206. Optical sensor(s) 264 include charge-coupled device (CCD) or complementary metal-oxide semiconductor (CMOS) photoreceptor(s). Optical sensor(s) 264 receive light from the environment, projected through one or more lenses, and converts the light to data representing an image. In conjunction with imaging module 243 (also called a camera module), optical sensor(s) 264 capture still images or video. In some embodiments, an optical sensor is located on the back of device 200, opposite the touch screen display 212 on the front of the device, so that the touch screen display is used as the viewfinder for still and / or video image acquisition. In some embodiments, an optical sensor is located on the front of the device so that the user’s image is obtained for video conferencing while the user views the other video conference participants on the touch screen display. In some embodiments, the position of optical sensor(s) 264 can be changed by the user (e.g., by rotating the lens and the sensor in the device housing so that a single optical sensor 264 is used alongside the touch screen display for both video conferencing and still and / or video image acquisition.

[0064] Device 200 optionally also includes one or more contact intensity sensors 265. Figure 2A Contact intensity sensor(s) 265 are coupled to intensity sensor controller 259 in I / O subsystem 206. Contact intensity sensor(s) 265 optionally include one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of a contact on a touch-sensitive surface). Contact intensity sensor(s) 265 receive contact intensity information (e.g., pressure information or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is collocated with, or proximate to, a touch-sensitive surface (e.g., touch- sensitive display system 212). In some embodiments, at least one contact intensity sensor is located on the back of device 200, opposite touch screen display 212 which is located on the front of device 200.

[0065] Device 200 also includes one or more proximity sensors 266. Figure 2A Proximity sensor(s) 266 are coupled to peripherals interface 218. Alternatively, proximity sensor 266 is coupled to input controller 260 in I / O subsystem 206. Proximity sensor 266 detects the presence of a nearby object, such as a finger or other object, and generates a corresponding "intensity" data that is used to control and / or enable / disable operations such as touch screen 212. For example, if an object is detected as being close to the touch screen, ignoring touch screen input and enabling the proximity sensor 266. In some embodiments, the proximity sensor turns off and disables the touch screen when the device is placed in an ear of the user, during a call. In some embodiments, during a call if the user removes the device from their ear, the proximity sensor turns off and enables the touch screen. In some embodiments, the proximity sensor only disables the touch screen if certain features of the touch screen 212 are active; e.g., if the touch screen display 212 is locked.

[0066] Device 200 optionally also includes one or more tactile output generators 267. Figure 2AA tactile output generator coupled to tactile feedback controller 261 in I / O subsystem 206 is shown. Tactile output generator 267 optionally includes one or more electroacoustic devices such as speakers or other audio components and / or electromechanical devices such as a motor, a solenoid, an electroactive polymer, a piezoelectric actuator, an electrostatic actuator or other tactile output generating components (e.g., components used to convert electrical signals into tactile outputs on the device). Contact intensity sensor 265 receives tactile feedback generation instructions from haptic feedback module 233 and generates tactile outputs on device 200 that a user is able to feel, such as tactile vibrations. In some embodiments, at least one tactile output generator is collocated with, or proximate to, a touch-sensitive surface (e.g., touch- sensitive display system 212) and, optionally, generates a tactile output by moving the touch-sensitive surface vertically (e.g., in / out of a surface of device 200) or laterally (e.g., back and forth in the same plane as the touch-sensitive surface of device 200). In some embodiments, at least one tactile output generator sensor is located on the back of device 200, opposite touch screen display 212 which is located on the front of device 200.

[0067] Device 200 also includes one or more accelerometers 268. Figure 2A Accelerometer 268 coupled to peripherals interface 218. Alternatively, accelerometer 268 is coupled to input controller 260 in I / O subsystem 206. Accelerometer 268 performs as described in U.S. Patent Publication No. 20050190059, "Acceleration-based Theft Detection System for Portable Electronic Devices," and U.S. Patent Publication No. 20060017692, "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer," both of which are incorporated by reference herein in their entirety. In some embodiments, information is displayed on the touch screen display in a portrait view or a landscape view, based on a determination of an orientation of device 200 obtained from data received from one or more of the accelerometers. Device 200 optionally includes a magnetometer (not shown) and GPS (or GLONASS or other global navigation system) receivers (not shown), which are used

[0068] In some embodiments, the software components stored in memory 202 include operating system 226, communication module (or set of instructions) 228, contact / motion module (or set of instructions) 230, graphics module (or set of instructions) 232, text input module (or set of instructions) 234, Global Positioning System (GPS) module (or set of instructions) 235, digital assistant client module 229, and applications (or sets of instructions) 236. Further, memory 202 stores data and models, such as user data and models 231. Moreover, in some embodiments, memory 202 stores device / global internal state 257, as shown in FIGS. Figure 2A ) or 470( Figure 4 ) as shown in FIGS. Figure 2A and Figure 4 Device / global internal state 257 includes one or more of: active application state, indicating which applications, if any, are currently active; display state, indicating what applications, views or other information occupy various regions of touch screen display 212; sensor state, including information obtained from the device's various sensors and input control devices 216; and location information indicating the location and / or attitude of the device.

[0069] Operating system 226 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates wireless communications between various hardware and software components.

[0070] Communication module 228 facilitates communication with other devices over one or more external ports 224 and also includes various software components for handling data received by RF circuitry 208 and / or external port 224. External port 224 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) is adapted for coupling directly to other devices or indirectly over a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as, or similar to, and / or compatible with, the 30-pin connector used on iPod® (trademark of Apple Inc.), iPhone® and iPad® devices. (Apple Inc. of Cupertino, California) devices.

[0071] Contact / motion module 230 optionally detects contact with touch screen 212 (in conjunction with display controller 256) and other touch-sensitive devices (e.g., a touchpad or physical click wheel). Contact / motion module 230 includes various software components for performing various operations related to detection of contact, such as determining if contact has occurred (e.g., detecting a finger-down event), determining an intensity of the contact (e.g., the force or pressure of the contact or a substitute for the force or pressure of the contact), determining if movement of the contact has occurred (e.g., detecting a finger-dragging event), and determining if the contact has ceased (e.g., detecting a finger-up event or a break in contact). Contact / motion module 230 receives contact data from the touch-sensitive surface. Determining movement of the point of contact, which is represented by a series of contact data, optionally includes determining speed (magnitude), velocity (magnitude and direction), and / or acceleration (a change in magnitude and / or direction) of the point of contact. These operations are, optionally, applied to single contact events (e.g., one finger events) or to multiple simultaneous contact events (e.g., "multitouch" events).

[0072] In some embodiments, contact / motion module 230 uses a set of one or more intensity thresholds to determine whether operations have been performed (e.g., whether a user has "clicked" an icon). In some embodiments, at least a subset of the intensity thresholds are determined based on software parameters (e.g., the intensity thresholds are not determined by activation thresholds of specific physical actuators and can be adjusted without changing the physical hardware of device 200). For example, without changing touchpad or touch screen display hardware, a mouse "click" threshold of a touchpad or touch screen can be set to any one threshold in a large range of predefined thresholds. Additionally, in some implementations, a user of the device is provided with software settings to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once with a system-wide click on an "intensity" parameter).

[0073] Contact / motion module 230 optionally detects gestures on touch-sensitive surface. Different gestures on the touch-sensitive surface have different motion pattern (e.g., different trajectories of detected contact, timing, and / or intensity of the detected contact). Accordingly, the gesture is optionally detected by detecting a particular motion pattern. For example, detecting a finger tap gesture includes detecting a finger-down event, followed by a finger-up (lift off) event at the same location (or substantially the same location) as the finger-down event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on the touch- sensitive surface includes detecting a finger-down event, followed by one or more finger-dragging events, and followed by a finger-up (lift off) event.

[0074] Graphics module 232 includes various known software components for rendering and displaying graphics on touch screen 212 or other display, including components for changing the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual property) of graphics that are displayed. As used herein, the term "graphics" includes any object that can be displayed to a user, including, without limitation, text, web pages, icons (such as user-interface objects including soft keys), digital images, videos, animations, and the like.

[0075] In some embodiments, graphics module 232 stores data representing graphics to be used. Each graphic is, optionally, assigned a corresponding code. Graphics module 232 receives, from applications etc., one or more codes specifying graphics to be displayed, and then generates screen image data for a

[0076] Haptic feedback module 233 includes various software components for generating instructions used by haptic feedback mechanism(s) 267 to produce tactile outputs at one or more locations on device 200 in response to user interactions with device 200.

[0077] Text input module 234, which in some embodiments is a component of graphics module 232, provides soft keyboards for entering text in various applications (e.g., contacts 237, e-mail 240, IM 241, browser 247, and any other application that needs text input).

[0078] GPS module 235 determines the location of the device and provides this information for use in various applications (e.g., to telephone 238 for use in

[0079] The digital assistant client module 229 includes various client-side digital assistant instructions to provide client-side functionality of the digital assistant. For example, the digital assistant client module 229 can accept voice input (e.g., speech input), textual input, touch input, and / or gesture input through various user interfaces of the portable multifunction device 200 (e.g., microphone 213, one or more accelerometers 268, touch-sensitive display system 212, one or more optical sensors 264, other input controls 216, etc.). The digital assistant client module 229 can also provide output in the form of audio (e.g., speech output), visual, and / or tactile through various output interfaces of the portable multifunction device 200 (e.g., speaker 211, touch-sensitive display system 212, one or more tactile output generators 267, etc.). For example, output is provided as speech, sound, reminders, text messages, menus, graphics, videos, animations, vibrations, and / or combinations of two or more of the above. During operation, the digital assistant client module 229 communicates with the DA server 106 using the RF circuitry 208.

[0080] The user data and models 231 include various data associated with the user (e.g., user-specific vocabulary data, user preference data, user-specified name pronunciations, data from the user's electronic address book, to-do items, shopping lists, etc.) to provide client-side functionality of the digital assistant. In addition, the user data and models 231 include various models for processing user input and determining user intent (e.g., speech recognition models, statistical language models, natural language processing models, ontologies, task flow models, service models, etc.).

[0081] In some examples, the digital assistant client module 229 utilizes various sensors, subsystems, and peripherals of the portable multifunction device 200 to gather additional information from the surrounding environment of the portable multifunction device 200 to establish a context associated with the user, the current user interaction, and / or the current user input. In some examples, the digital assistant client module 229 provides the context information, or a subset thereof, to the DA server 106 along with the user input to help infer the user intent. In some examples, the digital assistant also uses the context information to determine how to prepare and deliver output to the user. The context information is referred to as context data.

[0082] In some examples, the contextual information that accompanies user input includes sensor information, such as lighting, ambient noise, ambient temperature, images or video of the surrounding environment, etc. In some examples, the contextual information can also include physical state of the device, such as device orientation, device location, device temperature, power level, velocity, acceleration, motion patterns, cellular signal strength, etc. In some examples, information related to the software state of the DA server 106, such as the running processes of the portable multifunctional device 200, installed programs, past and current network activity, background services, error logs, resource usage, etc., are provided to the DA server 106 as contextual information associated with user input.

[0083] In some examples, the digital assistant client module 229 selectively provides information stored on the portable multifunctional device 200 (e.g., user data 231) in response to requests from the DA server 106. In some examples, the digital assistant client module 229 also elicits additional input from the user via natural language dialog or other user interfaces upon request by the DA server 106. The digital assistant client module 229 transmits this additional input to the DA server 106 to assist the DA server 106 in intent inference and / or in fulfilling the user intent expressed in the user request.

[0084] Reference is made below to Figures 7A-7C The digital assistant is described in more detail. It should be appreciated that the digital assistant client module 229 can include any number of sub-modules of the digital assistant module 726 described below.

[0085] The applications 236 include the following modules (or sets of instructions) or a subset or superset thereof:

[0086] • a contacts module 237 (sometimes referred to as an address book or contact list);

[0087] • a telephony module 238;

[0088] • a video conferencing module 239;

[0089] • an email client module 240;

[0090] • an instant messaging (IM) module 241;

[0091] • a fitness support module 242;

[0092] • a camera module 243 for still and / or video images;

[0093] • an image management module 244;

[0094] • a video player module;

[0095] • a music player module;

[0096] • a browser module 247;

[0097] • a calendar module 248;

[0098] • a widget module 249, which in some examples includes one or more of the following:

[0099] a weather widget 249-1, a stocks widget 249-2, a calculator widget 249-3, a

[0100] • a widget creator module 250 for creating the user-created widgets 249-6;

[0101] • a search module 251;

[0102] • a video and music player module 252, which merges video player module and music player module;

[0103] • a notes module 253;

[0104] • a maps module 254; and / or

[0105] • an online video module 255.

[0106] Examples of other applications 236 stored on storage 202 include other word processing applications, other image editing applications, a drawing application, presentation

[0107] In conjunction with touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the contacts module 237 are used for managing a plurality of contacts stored in the phone's memory 202 or the device's memory 470, which are in an application internal state 292 in memory 202, and includes: adding a new contact to the phone's memory 202 and the device's memory 470; accessing a contact from the phone's memory 202 and the device's memory 470; modifying, or deleting a contact in the phone's memory 202 and the device's memory 470; and / or initiating communications, such as placing voice calls, sending text messages, or sending electronic mail messages to a contact, with the contact via the telephone 238, electronic mail 240, or instant messenger 241.

[0108] In conjunction with RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, telephone module 238 are used to place, conduct, and terminate a telephone call or to access other call-related services.

[0109] In conjunction with RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touch screen 212, display controller 256, optical sensor 264, optical sensor controller 258, contact / motion module 230, graphics module 232, text input module 234, contact module 237, and telephone module 238, video conference module 239 includes executable instructions to initiate, conduct, and terminate a video conference between a user and one or more other participants in accordance with user instructions.

[0110] In conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, e-mail client module 240 includes executable instructions to create, send, receive, and manage e-mail in response to user instructions. In conjunction with image management module 244, e-mail client module 240 makes it very easy to create and send e-mails with still or video images taken with camera module 243.

[0111] In conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the instant messaging module 241 includes executable instructions to enter a sequence of characters corresponding to a instant message, modify previously entered characters, transmit a respective instant message (e.g., using a Short Message Service (SMS) or Multimedia Message Service (MMS) protocol for a phone-based instant message or using XMPP, SIMPLE, or IMPS for an Internet-based instant message), receive instant messages, and view received instant messages. In some embodiments, an instant message that is sent and / or received includes graphics, photos, audio files, video files and / or other attachments as are supported in a MMS and / or an Enhanced Messaging Service (EMS). As used herein, "instant message" refers to both a phone-based message (e.g., a message sent using SMS or MMS) and an Internet-based message (e.g., a message sent using XMPP, SIMPLE, or IMPS).

[0112] In conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, GPS module 235, map module 254, and music player module, workout support module 242 includes instructions to create workouts (e.g., with time, distance, and / or calorie burning goals); communicate with workout sensors (sports devices); receive workout sensor data; calibrate sensors used for a workout; select and play music for a workout; and display, store, and send workout data.

[0113] In conjunction with touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and image management module 244, camera module 243 includes instructions to capture still images or video (including a video stream) and store them into memory 202, modify characteristics of a still image or video, or delete a still image or video from memory 202.

[0114] In conjunction with touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and camera module 243, image management module 244 includes instructions to arrange, modify (e.g., edit), or otherwise manipulate, label, delete, present (e.g., in a digital slide show or album), and store still and / or video images.

[0115] In conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, e-mail client module 240 includes instructions to create, send, receive, and manage e-mail

[0116] In conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, e-mail client module 240, and browser module 247, calendar module 248 includes instructions to create and display calendars and data associated with calendars (e.g., calendar entries, to-do lists, etc.), receive data

[0117] In conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and browser module 247, the widget module 249 is used by a user to create widgets (e.g., widgets 249-1 and 249-2, which respectively include weather and stock information) or to retrieve widgets for

[0118] In conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and browser module 247, the widget creator module 250 is used by a user to create widgets (e.g., to select the user- specified portion of a web page and to select the widget type, such as weather, stocks, or photos, in which the web page portion is provided).

[0119] In conjunction with touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the search module 251 includes executable instructions to search for text, music, sound, image, video, and / or other files in memory 202 that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with user instructions.

[0120] In conjunction with touch screen 212, display controller 256, contact / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, and browser module 247, the video and music player module 252 includes executable instructions that allow the user to download and play back recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, and executable instructions to display, present or otherwise play back videos (e.g., on the touch screen 212, or on an external, connected display via external port 224). In some embodiments, device 200 optionally includes the functionality of an MP3 player, such as an iPod (trademark of Apple Inc.).

[0121] In conjunction with touch screen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, the notes module 253 includes executable instructions to create and manage notes, to-do lists, and the like in accordance with user instructions.

[0122] In conjunction with RF circuitry 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, GPS module 235, and browser module 247, map module 254 can be used to receive, display, modify, and store maps and data associated with maps (e.g., driving directions, data on stores and other points of interest at or near a particular location, and other location-based data) in accordance with user instructions.

[0123] In conjunction with touch screen 212, display controller 256, contact / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, text input module 234, e-mail client module 240, and browser module 247, online video module 255 includes instructions that allow the user to access, browse, receive (e.g., by streaming and / or download), play back (e.g., on the touch screen or on an external, connected display via external port 224), send an e-mail with a link to a particular online video, and otherwise manage online videos in one or more file formats, such as H.264. In some embodiments, instant messaging module 241, rather than e-mail client module 240, is used to send a link to a particular online video. Additional descriptions of online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” filed June 20, 2007, and U.S. Patent Application No. 11 / 968,067, “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” filed December 31, 2007, the contents of which are hereby incorporated by reference in their entirety.

[0124] Each of the above-identified modules and applications corresponds to a set of executable instructions for performing one or more functions described above and the methods described in this application (e.g., the computer-implemented methods described herein and other information processing methods). These modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures or modules, and thus various subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, video Figure 2AThe memory 202 includes computer- readable instructions that, when executed by the processor 204, enable the device 200 to perform various functions (e.g., functions related to video and music player module 252, etc.). In some embodiments, the memory 202 stores a subset of the instructions described above. Alternatively, the memory 202 stores additional or different instructions.

[0125] In some embodiments, the device 200 is a device for which a predefined set of functions of the device are performed exclusively through touchscreens and / or touchpads. By using a touchscreen and / or touchpad as the primary input control device for operation of the device 200, the number of physical input control devices (such as push buttons, dials, etc.) on the device 200 is reduced.

[0126] The predefined set of functions performed exclusively through touchscreens and / or touchpads optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates the device 200 from any user interface displayed on the device 200 to a main menu, home menu, or root menu. In such embodiments, the touchpad is used to implement a "menu button." In some other embodiments, the menu button is a physical push button or other physical input control device, rather than a touchpad.

[0127] Figure 2B is a block diagram illustrating example components for event handling in accordance with some embodiments. In some embodiments, the memory 202 (e.g., memory 202 including event sorter 270) or the memory 470 (e.g., memory 470 including event sorter 270) includes event sorter 270 (e.g., in operating system 226) and respective application 236-1 (e.g., any of the aforementioned applications 237-251, 255, 480-490). Figure 2A ) or the memory 470 ( Figure 4 ) includes event sorter 270 (e.g., in operating system 226) and respective application 236-1 (e.g., any of the aforementioned applications 237-251, 255, 480-490).

[0128] The event sorter 270 receives event information and determines, with the assistance of the application internal state 292, which application 236-1, and of the application views 291, are to be delivered the event information. In some embodiments, the device / global internal state 257, is used by the event sorter 270 to determine which components (such as the applications 236-1, or portions thereof) should be notified of the event. In some embodiments, the event sorter 270 includes a notification component for discovering multiple component portions objects which handling a communication, and a messaging component for delivering notifications.

[0129] In some embodiments, the application internal state 292 includes additional information such as one or more of: resume information to be used when the application 236-1 resumes execution, user interface state information indicating that information is being displayed or is ready for display by the application 236-1, a state queue for enabling the user to return to a previous state or view of the application 236-1, and a repeat / undo queue of previous actions taken by the user.

[0130] The event monitor 271 receives event information from the peripherals interface 218. The event information includes information about sub-events (e.g., user touches on touch- sensitive display 212, motion sensor events such as free fall, wireless signal events such as Wi-Fi and Bluetooth beacons, and the like). The peripherals interface 218 sends events to the event monitor 271 in response to information it receives from I / O subsystem 206 or the various sensors 266, 268, and / or microphone 213 (through audio circuitry 210). The information from the peripherals interface 218 that the event monitor 271 receives includes information from the touch-sensitive display 212 or touchpad.

[0131] In some embodiments, the event monitor 271 transmits requests to the peripherals interface 218 at predetermined intervals. In response, the peripherals interface 218 sends event information. In other embodiments, the peripherals interface 218 sends event information only when there is a significant event (e.g., input that is above a predetermined noise threshold and / or input that exceeds a predetermined duration).

[0132] In some embodiments, the event classifier 270 also includes a hit view determination module 272 and / or an active event recognizer determination module 273.

[0133] When the touch-sensitive display 212 displays more than one view, the hit view determination module 272 provides a software process for determining where within one or more views a sub-event has occurred. A view is made up of controls and other elements that the user can see on the display.

[0134] Another aspect of a user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the respective application) in which a touch is detected corresponds to a programmed level within the application's programmed hierarchy of views or view hierarchy. For example, the lowest level view in which a touch is detected is referred to as the hit view, and the set of events that are correctly entered are determined based at least in part on the hit view of the initial touch that begins the touch-based gesture.

[0135] Hit view determination module 272 receives information related to touch-based gesture sub-events. When an application has multiple views organized in a hierarchy, hit view determination module 272 identifies the hit view as the lowest view in the hierarchy that should handle the sub-event. In most cases, the hit view is the lowest level view in which the sub-event (e.g., the first sub-event in a sequence of sub-events forming an event or potential event) originated. Once the hit view is identified by hit view determination module 272, the hit view typically receives all sub-events related to the same touch or input source for which it was identified as the hit view.

[0136] Active event recognizer determination module 273 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, active event recognizer determination module 273 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, active event recognizer determination module 273 determines that all views that include the physical location of the sub-events are actively engaged views, and thus determines that all actively engaged views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with one particular view, higher views in the hierarchy will remain actively engaged views.

[0137] Event dispatcher module 274 dispatches event information to event recognizers (e.g., event recognizers 280). In embodiments that include active event recognizer determination module 273, event dispatcher module 274 delivers event information to the event recognizers determined by active event recognizer determination module 273. In some embodiments, event dispatcher module 274 stores event information in an event queue that is retrieved by respective event receivers 282.

[0138] In some embodiments, operating system 226 includes event classifier 270. Alternatively, application 236-1 includes event classifier 270. In yet another embodiment, event classifier 270 is a standalone module, or part of another module stored in memory 202, such as contact / motion module 230.

[0139] In some embodiments, application 236-1 includes a plurality of event handlers 290 and one or more application views 291, each of which includes instructions for handling touch events that occur within a respective view of the user interface of the application. Each application view 291 of application 236-1 includes one or more event recognizers 280. Typically, a respective application view 291 includes multiple event recognizers 280. In other embodiments, one or more of the event recognizers 280 are part of a separate module, such as a user interface toolkit (not shown) or a higher level object from which the application 236-1 inherits methods and other attributes. In some embodiments, a respective event handler 290 includes one or more of the following: a data updater 276, an object updater 277, a GUI updater 278, and / or event data 279 received from event sorter 270. Event handler 290 utilizes or invokes the data updater 276, object updater 277, or GUI updater 278, as appropriate, to update the application internal state 292. Alternatively, one or more of the application views 291 includes one or more respective event handlers 290. In addition, in some embodiments, one or more of the data updater 276, object updater 277, and GUI updater 278 are included in respective application views 291.

[0140] A respective event recognizer 280 receives event information (e.g., event data 279) from event sorter 270 and identifies an event from the event information. Event recognizer 280 includes an event receiver 282 and an event comparator 284. In some embodiments, event recognizer 280 also includes at least a subset of metadata 283 and event delivery instructions 288 (which includes a sub-event delivery instructions).

[0141] Event receiver 282 receives event information from event sorter 270. The event information includes information about a sub-event, such as a touch or touch movement. Additionally, the event information also includes information about other touches such as one or more touches that occur while the sub-event is detected and other touches that occur after the sub-event but before the event is resolved. In some embodiments, the event information also includes information from event 271 about a device state immediately after event sorter 270 has sorted events. In some embodiments, a device state is indicated as a respective state of one or more device sensors, input control, and / or device context. For example, a device state indicates whether a device is in a ring, vibrate, or silent mode. A device state also indicates whether an input control is enabled or disabled (e.g., a button is pressed or unpressed, a touch is detected or not detected, etc.). A device state also indicates whether a device is in a prolonged duration input mode. In some embodiments, event information further includes information generated by one or more event recognizers 280.

[0142] Event comparator 284 compares event information to predefined event or sub-event definitions, and based on the comparison, determines an event or sub-event, or determines or updates a state of an event or sub-event. In some embodiments, event comparator 284 includes event definitions 286. Event definitions 286 contain definitions of events (e.g., predefined sequences of sub-events), such as event 1 (287-1), event 2 (287-2), and other events. In some embodiments, sub-events in an event (287) include, for example, touch begin, touch end, touch move, touch cancel, and multi-touch. In one example, the definition of event 1 (287-1) is a double tap on a displayed object. For example, a double tap includes a first touch (touch begin) on a displayed object for a predetermined duration, a first lift off (touch end) for a predetermined duration, a second touch (touch begin) on a displayed object for a predetermined duration, and a second lift off (touch end) for a predetermined duration. In another example, the definition of event 2 (287-2) is a drag on a displayed object. For example, a drag includes a touch (or contact) on a displayed object for a predetermined duration, movement of the touch on the touch-sensitive display 212, and a lift off of the touch (touch end). In some embodiments, an event also includes information for one or more associated event handlers 290.

[0143] In some embodiments, event definitions 287 include definitions of events for respective user interface objects. In some embodiments, event comparator 284 performs a hit test to determine which user interface object is associated with a sub-event. For example, in an application view that displays three user interface objects on touch-sensitive display 212, when a touch is detected on touch-sensitive display 212, event comparator 284 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a respective event handler 290, event comparator uses the results of the hit test to determine which event handler 290 should be activated. For example, event comparator 284 selects the event handler that is associated with the sub-event and the object that triggered the hit test.

[0144] In some embodiments, the definition of a respective event (287) also includes a deferred action that delays delivery of event information until it has been determined that a sequence of sub-events does or does not correspond to an event type of the event recognizer.

[0145] When a respective event recognizer 280 determines that a sub-event sequence does not match any event in the event definitions 286, the respective event recognizer 280 enters an event impossible, event failed, or event ended state, after which subsequent sub-events of the touch-based gesture are ignored. In this case, other event recognizers (if any) that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.

[0146] In some embodiments, the respective event recognizers 280 include metadata 283 having configurable properties, flags, and / or lists that indicate how the event delivery system should perform sub-event delivery to the active event recognizers. In some embodiments, the metadata 283 includes configurable properties, flags, and / or lists that indicate how the event recognizers interact or are able to interact with each other. In some embodiments, the metadata 283 includes configurable properties, flags, and / or lists that indicate whether sub-events are delivered to different levels in a view or programmatic hierarchy.

[0147] In some embodiments, when one or more particular sub-events of an event are recognized, the respective event recognizer 280 activates an event handler 290 associated with the event. In some embodiments, the respective event recognizer 280 delivers event information associated with the event to the event handler 290. Activating an event handler 290 is distinct from passing (and deferring passing) sub-events to a respective hit view. In some embodiments, the event recognizer 280 throws a token associated with the recognized event, and an event handler 290 associated with the token picks up the token and performs a pre-defined process.

[0148] In some embodiments, the event delivery instructions 288 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver the event information to an event handler associated with the sub-event sequence or to the active view. The event handler associated with the sub-event sequence or with the active view receives the event information and performs a pre-determined process.

[0149] In some embodiments, data updater 276 creates and updates data used in application 236-1. For example, data updater 276 updates phone numbers used in contact module 237 or stores video files used in video player module. In some embodiments, object updater 277 creates and updates objects used in application 236-1. For example, object updater 277 creates new user interface objects or updates the positioning of user interface objects. GUI updater 278 updates the GUI. For example, GUI updater 278 prepares display information and transmits the display information to graphics module 232 for display on a touch-sensitive display.

[0150] In some embodiments, event handler 290 includes or has access to data updater 276, object updater 277, and GUI updater 278. In some embodiments, data updater 276, object updater 277, and GUI updater 278 are included in a single module of the corresponding application 236-1 or application view 291. In other embodiments, they are included in two or more software modules.

[0151] It should be understood that the above discussion of event handling for user touches on a touch-sensitive display also applies to other forms of user input utilizing input devices to operate the multifunction device 200, and not all user input is initiated on a touch screen. For example, mouse movement and mouse button presses, optionally in conjunction with single or multiple keyboard presses or holddowns; contact movement on a touchpad, such as taps, drags, scrolls, etc.; stylus input; movement of the device; spoken commands; detected eye movement; biometric input; and / or any combination thereof, are optionally used as input corresponding to sub-events defining the event to be recognized.

[0152] Figure 3The portable multifunction device 200 with touch screen 212 is, optionally, a handheld device. Touch screen 212 optionally uses LCD (liquid crystal display) technology or OLED (organic light emitting diode) technology, although other technologies can be used. The touch screen 212 and the touch screen controller 256 optionally each include a plurality of pixels. The touch screen 212 optionally displays one or more graphics from the graphics memory 262 in accordance with various embodiments. Some embodiments include a touch screen controller 256 coupled with the touch screen 212. In some embodiments, the touch screen 212 displays one or more graphical items of a graphical user interface (GUI), various embodiments of which are described herein.

[0153] Device 200 also includes one or more physical buttons, such as "home" or menu button 304. As described previously, the menu button 304 is used to navigate to any application 236 in a set of applications on device 200. Alternatively, the menu button is implemented as a soft key in a GUI displayed by

[0154] In some embodiments, device 200 includes touch screen 212, menu button 304, push button 306 for powering the device on / off and placing device 200 in a sleep state, one or more volume adjustment buttons 308, a user identity module (SIM) card slot 310, a headphone jack 312, and a dock / charging external port 224. Push button 306 is, optionally, used to

[0155] Figure 4A block diagram of an exemplary multifunction device with a display and a touch-sensitive surface in accordance with some embodiments. Device 400 need not be portable. In some embodiments, device 400 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an education device (such as a child's learning toy), a gaming device, or a control device (e.g., a home or industrial controller). Device 400 typically includes one or more processing units (CPU's) 410, one or more network or other communications interfaces 460, memory 470, and one or more communication buses 420 for interconnecting these components. Communication buses 420 optionally include circuitry (sometimes called a chipset) that interconnects and controls communications between system components. Device 400 includes input / output (I / O) interface 430 comprising display 440, which is typically a touch screen display. I / O interface 430 also optionally includes a keyboard and / or mouse (or other pointing device) 450 and touchpad 455, tactile output generator 457 (e.g., one or more electroacoustic devices such as speakers for generating audio output, one or more electromechanical devices such as a vibration motor for generating tactile outputs, or one or more other tactile output mechanisms for generating tactile outputs), sensor 459 (e.g., one or more optical, touch, and / or contact intensity sensors such as contact intensity sensor 265 described above with reference to Figure 2A Figure 1C), and / or contact intensity sensor 265 described above with reference to Figure 2A Figure 1C). Memory 470 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid state memory devices; and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 470 optionally includes one or more storage devices remotely located from CPU(s) 410. In some embodiments, memory 470 stores programs, modules, and data structures similar to the programs, modules, and data structures stored in memory 202 of portable multifunction device 200 Figure 2A ), or a subset thereof. Furthermore, memory 470 optionally stores additional programs, modules, and data structures not present in the memory 202 of portable multifunction device 200. For example, the memory 470 of device 400 optionally stores an operating system 452, a communication module (or set of instructions) 454, a contact / motion module 456, a graphics module 462, a text input module 464, a Graphics Module 480, a Presentation Module 482, a Word Processing Module 484, a Website Creation Module 486, a Disk Editing Module 488, and / or a Spreadsheet Module 490, while memory 202 of portable multifunction device 200 Figure 2A ) optionally does not store these modules.

[0156] Figure 4Each of the above identified elements in the memory 470 can store data used by various applications run by the device 400. Furthermore, each of the above identified elements can also store computer-readable code for implementing one or more of the modules described above, such as the module for managing the display of the device 400. The computer-readable code can also be stored in the memory 470 and implemented as software instructions. Furthermore, the identified modules can also include the computer-readable code stored in the memory 470, which, when executed by the processor 460, can cause the processor 460 to perform one or more of the methods described herein. The modules described herein can control the processor 460 such that one or more of the methods described herein are performed by the device 400.

[0157] Attention is now directed towards embodiments of user interfaces ("UI") that can be implemented on portable multifunction devices 100.

[0158] Figure 5A An exemplary user interface that includes an application menu on a portable multifunction device 200 is shown. Similar user interfaces can be implemented on device 400. In some embodiments, user interface 500 includes the following elements, or a subset or superset thereof:

[0159] a signal strength indicator 502 for wireless communications such as cellular and Wi-Fi signals;

[0160] • a time 504;

[0161] • a Bluetooth indicator 505;

[0162] • a battery status indicator 506;

[0163] • a tray 508 with icons for frequently used applications, such as:

[0164] • an icon 516 for telephone module 238 labeled "Phone," which optionally includes an indicator 514 of the number of missed calls or voice mails;

[0165] • an icon 518 for email client module 240 labeled "Mail," which optionally includes an indicator 510 of the number of unread emails;

[0166] • an icon 520 for browser module 247 labeled "Browser"; and

[0167] • an icon 522 for video and music player module 252 (also referred to as iPod (Apple Inc.'s trademark) module 252) labeled "iPod"; and

[0168] • icons for other applications, such as:

[0169] • a menu button 512 for the above applications, which is used to navigate to a graphical listing of the applications;

[0170] • an icon 524 for the IM module 241, labeled "Messages";

[0171] • an icon 526 for the Calendar module 248, labeled "Calendar";

[0172] • an icon 528 for the Image Management module 244, labeled "Photos";

[0173] • an icon 530 for the Camera module 243, labeled "Camera";

[0174] • an icon 532 for the Online Video module 255, labeled "Online Videos";

[0175] • an icon 534 for the Stocks widget 249-2, labeled "Stocks";

[0176] • an icon 536 for the Maps module 254, labeled "Maps";

[0177] • an icon 538 for the Weather widget 249-1, labeled "Weather";

[0178] • an icon 540 for the Alarms widget 249-4, labeled "Clock";

[0179] • an icon 542 for the Fitness Support module 242, labeled "Fitness Support";

[0180] • an icon 544 for the Notes module 253, labeled "Notes"; and

[0181] • an icon 546 for setting applications or modules, labeled "Settings", which provides access to settings for the device 200 and its various applications 236.

[0182] It should be noted that Figure 5A The example icon labels are merely examples. For example, the icon 522 for the Video and Music Player module 252 is optionally labeled "Music" or "Music Player." Other labels for various application icons are used in some embodiments. In some embodiments, the label for a particular application icon includes the name of the application that corresponds to that particular application icon. In some embodiments, the label for a particular application icon is different than the name of the application that corresponds to that particular application icon.

[0183] Figure 5B A device (e.g., device 200) is shown having a touch-sensitive surface 551 (e.g., a touch screen display 212) separate from the display 550 (e.g., Figure 4 of FIG. 1A) of the device 100. Figure 4FIG. 1C shows an example user interface that is displayed on device 400). Device 400 also optionally includes one or more contact intensity sensors (e.g., one or more sensors of sensors 459) for detecting intensity of contacts on touch-sensitive surface 551 and / or one or more tactile output generators 457 for generating tactile outputs for the user of device 400.

[0184] Although some of the examples that follow will be given with reference to inputs on touch screen display 212 (in which a touch- sensitive surface and a display are combined), in some embodiments, the device detects inputs on a touch- sensitive surface that is separate from the display, as shown in FIG. 1A. In some embodiments, a touch- sensitive surface (e.g., 551 in FIG. 1C) has a major axis (e.g., 552 in FIG. 1C) that corresponds to a major axis (e.g., 553 in FIG. 1C) of the display (e.g., 550 in FIG. 1C). According to these embodiments, the device detects contacts with the touch- sensitive surface (e.g., 560 and 562 in FIG. 1C) at locations that correspond to respective locations on the display (e.g., 560 corresponds to 568 and 562 corresponds to 570 in FIG. 1C). In this way, when the touch- sensitive surface (e.g., 551 in FIG. 1C) is separate from the display (e.g., 550 in FIG. 1C) of a multifunction device, user inputs (e.g., contacts 560 and 562 and movements thereof) detected by the device on the touch- sensitive surface are used by the device to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein. Figure 5B Figure 5B Figure 5B Figure 5B Figure 5B Figure 5B Figure 5B Figure 5B

[0185]

[0186] Figure 6A ​​​​​​​​​An example personal electronic device 600 is shown. Device 600 includes a body 602. In some embodiments, device 600 includes some or all of the features described with respect to devices 200 and 400 (e.g., Figures 2A-4 In some embodiments, device 600 has a touch-sensitive display 604, referred to hereafter as touch screen 604. Alternatively, device 600 has a display and a touch-sensitive surface. As with device 200 and 400, in some embodiments, the touch screen 604 (or touch- sensitive surface) has one or more intensity sensors to detect intensity of contacts (e.g., touches) on the touch screen 604. The one or more intensity sensors of the touch screen 604 (or touch- sensitive surface) provide output data that represents the intensity of the touches. The user interface of device 600 responds to touches of different intensities by different user interface operations, which means that different intensity touches can invoke different user interface operations on device 600.

[0187] Techniques for detecting and processing touch intensity may, for example, be found in related applications: International Patent Application Serial No. PCT / US2013 / 040061, titled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," filed May 8, 2013, and International Patent Application Serial No. PCT / US2013 / 069483, titled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," filed November 11, 2013, each of which is hereby incorporated by reference in its entirety.

[0188] In some embodiments, device 600 has one or more input mechanisms 606 and 608. Input mechanisms 606 and 608, if included, are physical. Examples of physical input mechanisms include a depressible button and a rotatable mechanism. In some embodiments, device 600 has one or more attachment mechanisms. Such attachment mechanisms, if included, can allow device 600 to be attached to, for example, a hat, glasses, earrings, a necklace, a shirt, a jacket, a bracelet, a watchband, a necklace, pants, a belt, shoes, a purse, a backpack, and the like. These attachment mechanisms allow a user to wear device 600.

[0189] Figure 6BAn example personal electronic device 600 is shown. In some embodiments, device 600 includes some or all of the components described with respect to Figure 2A 、 Figure 2B and Figure 4 . Device 600 has bus 612 which operatively couples I / O section 614 with one or more computer processors 616 and memory 618. I / O section 614 is connected to display 604, which can have touch-sensitive component 622, and optionally also touch intensity sensor component 624. Furthermore, I / O section 614 is connected with communication unit 630 for receiving application and operating system data, such as data to create a

[0190] In some examples, input mechanism 608 is a microphone. Personal electronic device 600 includes various sensors, such as GPS sensor 632, accelerometer 634, directional sensor 640 (e.g., compass), gyroscope 636, motion sensor 638, and / or a combination thereof, all of which can be operatively connected to I / O section 614.

[0191] Memory 618 of personal electronic device 600 is a non-transitory computer-readable storage medium for storing computer-executable instructions, for example, that, when executed by one or more computer processors 616, cause the computer processors to perform the techniques and processes described above. The computer-executable instructions are also, for example, stored and / or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch instructions from the instruction execution system, apparatus, or device and execute instructions. The personal electronic device 600 is not limited to Figure 6B the components and configurations of FIGS. 1, 2, 4, 6, 9, 10, and / or 11, but can include other or additional components in multiple configurations.

[0192] As used herein, the term "affordance" refers to a graphical user interface object that is capable of being Figure 2A 、 Figure 4 、 Figures 6A-6B 、 Figures 9A-9C 、 Figures 10A-10C and Figures 11A-11D displayed on a display screen of a device 200, 400, 600, 900, 1000, and / or 1100, and that is interactive with a user. For example, an image (e.g., an icon), a button, and text (e.g., a hyperlink) each constitute an affordance.

[0193] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface with which a user is interacting. In some implementations that include a cursor or other position marker, the cursor acts as a "focus selector" such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), a focus selector is displayed on a touch-sensitive surface (e.g., Figure 4 Touchpad 455 or Figure 5B In the event that an input (e.g., a press input) is detected on the touch-sensitive surface 551 in FIG, the particular user interface element is adjusted according to the detected input. In the case of a touch screen display (e.g., a touch screen display) that enables direct interaction with user interface elements on the touch screen display Figure 2A touch-sensitive display system 212 or Figure 5A In some implementations of the touch screen 212 in FIG, 20 , a contact detected on the touch screen acts as a “focus selector” such that when input (e.g., a press input by the contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touch screen display, the particular user interface element is adjusted according to the detected input. In some implementations, the focus moves from one area of ​​the user interface to another area of ​​the user interface without corresponding movement of a cursor or movement of a contact on the touch screen display (e.g., by using a tab key or arrow keys to move the focus from one button to another); in these implementations, the focus selector moves according to the movement of the focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or contact on the touch screen display) that is controlled by the user to deliver the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface with which the user desires to interact). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or touch screen), the position of a focus selector (e.g., a cursor, contact, or selection box) over a corresponding button will indicate that the user intends to activate the corresponding button (rather than other user interface elements shown on the device display).

[0194] As used in the specification and claims, the term "feature strength" of a contact refers to a characteristic of the contact that is based on one or more strengths of the contact. In some embodiments, the feature strength is based on a plurality of strength samples. The feature strength is optionally based on a predefined number of strength samples or a set of strength samples that are acquired during a predetermined period of time (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) with respect to a predefined event (e.g., after detecting the contact, before detecting liftoff of the contact, before or after detecting the contact starting to move, before detecting the end of the contact, before detecting an increase in strength of the contact, before detecting a decrease in strength of the contact, and / or the like). The feature strength of the contact is optionally based on one or more of: a maximum value of the contact strength, a mean value of the contact strength, a median value of the contact strength, a value at the 10th percentile of the contact strength, a half-maximum value of the contact strength, a 90th maximum value of the contact strength, and / or the like. In some embodiments, the duration of the contact is used in determining the feature strength (e.g., when the feature strength is an average of the strength of the contact over time). In some embodiments, the feature strength is compared to a set of one or more strength thresholds to determine whether the user has performed an operation. For example, the set of one or more strength thresholds includes a first strength threshold and a second strength threshold. In this example, a contact with a feature strength that does not exceed the first threshold results in a first operation, a contact with a feature strength that exceeds the first strength threshold but does not exceed the second strength threshold results in a second operation, and a contact with a feature strength that exceeds the second threshold results in a third operation. In some embodiments, the comparison between the feature strength and the one or more thresholds is used to determine whether one or more operations are to be performed (e.g., whether a respective operation is performed or forgoen), rather than to determine whether a first operation or a second operation is performed.

[0195] In some embodiments, a portion of a gesture is identified for use in determining a feature strength. For example, a touch-sensitive surface receives a continuous swipe contact that transitions from a starting location and reaches an ending location at which the strength of the contact increases. In this example, the feature strength of the contact at the ending location is based only on a portion of the continuous swipe contact, rather than the entire swipe contact (e.g., the portion of the swipe contact that is at the ending location only). In some embodiments, a smoothing algorithm is applied to the strength of the swipe contact prior to determining the feature strength of the contact. For example, the smoothing algorithm optionally includes one or more of: an unweighted sliding average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow spikes or dips in the strength of the swipe contact for purposes of determining the feature strength.

[0196] The intensity of contact on the touch-sensitive surface is characterized relative to one or more intensity thresholds, such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, a light press intensity threshold corresponds to an intensity at which the device will perform an operation typically associated with clicking a button of a physical mouse or trackpad. In some embodiments, a deep press intensity threshold corresponds to an intensity at which the device will perform an operation different from that typically associated with clicking a button of a physical mouse or trackpad. In some embodiments, when a contact of feature intensity below the light press intensity threshold (e.g., and above the nominal contact detection intensity threshold, below which contacts are no longer detected) is detected, the device will move a focus selector in accordance with movement of the contact on the touch-sensitive surface without performing an operation associated with the light press intensity threshold or the deep press intensity threshold. Generally, unless otherwise stated, these intensity thresholds are consistent between different sets of user interface figures.

[0197] An increase in contact feature intensity from an intensity below the light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in contact feature intensity from an intensity below the deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in contact feature intensity from an intensity below the contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on the touch surface. A decrease in contact feature intensity from an intensity above the contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting liftoff of a contact from the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.

[0198] In some embodiments described herein, one or more operations are performed in response to detecting a gesture that includes a respective press input or in response to detecting a respective press input performed with a respective contact (or multiple contacts), where the respective press input is detected based at least in part on detecting an increase in intensity of the contact (or multiple contacts) above a press input intensity threshold. In some embodiments, a respective operation is performed in response to detecting an increase in intensity of the respective contact above the press input intensity threshold (e.g., a "downstroke" of the respective press input). In some embodiments, a press input includes an increase in intensity of a respective contact above a press input intensity threshold and a subsequent decrease in intensity of the contact below the press input intensity threshold, and a respective operation is performed in response to detecting the subsequent decrease in intensity of the respective contact below the press input threshold (e.g., an "upstroke" of the respective press input).

[0199] In some embodiments, the device employs intensity hysteresis to avoid an unintended input sometimes referred to as "bouncing," in which the device defines or selects a hysteresis intensity threshold that has a predefined relationship to the press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable proportion of the press input intensity threshold). Thus, in some embodiments, a press input includes an increase in intensity of a respective contact above the press input intensity threshold and a subsequent decrease in intensity of the contact below a hysteresis intensity threshold corresponding to the press input intensity threshold, and a respective operation is performed in response to detecting the subsequent decrease in intensity of the respective contact below the hysteresis intensity threshold (e.g., an "upstroke" of the respective press input). Similarly, in some embodiments, a press input is detected only when the device detects an increase in contact intensity from an intensity at or below the hysteresis intensity threshold to an intensity at or above the press input intensity threshold, and optionally a subsequent decrease in contact intensity to an intensity at or below the hysteresis intensity, and a respective operation is performed in response to detecting the press input (e.g., an increase in contact intensity or a decrease in contact intensity depending on the circumstances).

[0200] For ease of explanation, optionally, descriptions of operations performed in response to a press input associated with a press input intensity threshold or in response to a gesture that includes a press input are triggered in response to detecting any of a variety of conditions: an increase in contact intensity above the press input intensity threshold, an increase in contact intensity from an intensity below the hysteresis intensity threshold to an intensity above the press input intensity threshold, a decrease in contact intensity below the press input intensity threshold, and / or a decrease in contact intensity below a hysteresis intensity threshold corresponding to the press input intensity threshold. Additionally, in examples in which an operation is described as being performed in response to detecting a decrease in intensity of a contact below a press input intensity threshold, optionally the operation is performed in response to detecting a decrease in intensity of the contact below a hysteresis intensity threshold corresponding to and less than the press input intensity threshold.

[0201] 3. Digital assistant system

[0202] Figure 7A A block diagram of a digital assistant system 700 is shown in accordance with various examples. In some examples, the digital assistant system 700 is implemented on a standalone computer system. In some examples, the digital assistant system 700 is distributed across multiple computers. In some examples, some of the modules and functionality of the digital assistant are divided into server portions and client portions, where the client portions reside on one or more user devices (e.g., devices 104, 122, 200, 400, 600, 900, 1000, and / or 1100) and communicate with the server portions (e.g., server system 108) over one or more networks, e.g., as described in more detail below. Figure 1are shown. In some examples, the digital assistant system 700 is a Figure 1 implementation of the server system 108 (and / or DA server 106) shown in FIG. 1. It should be noted that the digital assistant system 700 is merely one example of a digital assistant system, and that the digital assistant system 700 has more or fewer components than shown, combines two or more components, or has a different configuration or arrangement of the components than shown. Figure 7A The various components shown in FIG. 7 are implemented in hardware, software (including one or more signal processing integrated circuits and / or application specific integrated circuits), firmware, or a combination thereof.

[0203] The digital assistant system 700 includes a memory 702, an input / output (I / O) interface 706, a network communication interface 708, and one or more processors 704. These components can communicate with one another over one or more communication buses or signal lines 710.

[0204] In some examples, the memory 702 includes non-transitory computer-readable media such as high-speed random access memory and / or non-volatile computer-readable storage media (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices).

[0205] In some examples, the I / O interface 706 couples input / output devices 716 of the digital assistant system 700, such as a display, a keyboard, a touchscreen, and a microphone, to the user interface module 722. The I / O interface 706, together with the user interface module 722, receives and processes user inputs (e.g., voice inputs, keyboard inputs, touch inputs, etc.) accordingly. In some examples, for instance, when the digital assistant is implemented on a standalone user device, the digital assistant system 700 includes any of the components and I / O communication interfaces described with respect to the devices 200, 400, 600, 900, 1000, and / or 1100 in FIGS. 2, 4, 6, 9, 10, and / or 11, respectively. Figure 2A 、 Figure 4 、 Figures 6A-6B 、 Figures 9A-9C 、 Figures 10A-10C and Figures 11A-11D In some examples, the digital assistant system 700 represents a server portion of a digital assistant implementation, and can interact with a user through a client-side portion that resides on a user device (e.g., the devices 104, 200, 400, 600, 900, 1000, and / or 1100).

[0206] In some examples, the network communication interface 708 includes one or more wired communication ports 712 and / or wireless transmit and receive circuitry 714. The one or more wired communication ports receive and transmit communication signals via one or more wired interfaces, such as Ethernet, Universal Serial Bus (USB), FIREWIRE, etc. The wireless circuitry 714 receives and transmits RF signals and / or optical signals from and to communication networks and other communication devices. The wireless communication uses any of a plurality of communications standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. The network communication interface 708 enables communication between the digital assistant system 700 and other devices over a network, such as the Internet, an intranet, and / or a wireless network, such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN).

[0207] In some examples, the memory 702 or the computer-readable storage medium of the memory 702 stores programs, modules, instructions, and data structures, including all or a subset of the following: an operating system 718, a communication module 720, a user interface module 722, one or more applications 724, and a digital assistant module 726. In particular, the memory 702 or the computer-readable storage medium of the memory 702 stores instructions for performing the processes described above. The one or more processors 704 execute these programs, modules, and instructions, and read from and write to the data structures.

[0208] The operating system 718 (e.g., Darwin, RTXC, LINUX, UNIX, iOS, OS X, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates wireless communications between various hardware and software components.

[0209] The communication module 720 facilitates communication between the digital assistant system 700 and other devices over the network communication interface 708. For example, the communication module 720 communicates with the RF circuitry 208 of the electronic devices (such as the devices 200, 400, or 600 shown in FIGS. 2, 4, and 6, respectively). Figure 2A , Figure 4 , Figures 6A-6B The communication module 720 also includes various components for processing data received by the wireless circuitry 714 and / or the wired communication ports 712.

[0210] The user interface module 722 receives commands and / or inputs from a user (e.g., from a keyboard, touch screen, pointing device, controller, and / or microphone) via the I / O interface 706 and generates user interface objects on a display. The user interface module 722 also prepares and transmits output (e.g., speech, sound, animation, text, icons, vibrations, haptic feedback, illumination, etc.) to the user via the I / O interface 706 (e.g., through a display, audio channel, speaker, touchpad, etc.).

[0211] The applications 724 include programs and / or modules configured for execution by the one or more processors 704. For example, if the digital assistant system is implemented on a standalone user device, the applications 724 include user applications such as games, calendar applications, navigation applications, or email applications. If the digital assistant system 700 is implemented on a server, the applications 724 include, for example, resource management applications, diagnostic applications, or scheduling applications.

[0212] The memory 702 also stores a digital assistant module 726 (or server portion of a digital assistant). In some examples, the digital assistant module 726 includes the following sub-modules or a subset or superset thereof: an input / output processing module 728, a speech-to-text (STT) processing module 730, a natural language processing module 732, a dialog flow processing module 734, a task flow processing module 736, a service processing module 738, and a speech synthesis processing module 740. Each of these modules has access to one or more of the following systems or data and models of the digital assistant module 726 or a subset or superset thereof: a knowledge ontology 760, a vocabulary index 744, user data 748, task flow models 754, service models 756, and an ASR system 758.

[0213] In some examples, using the processing modules, data, and models implemented in the digital assistant module 726, the digital assistant can perform at least some of the following: convert speech input into text; identify a user intent expressed in natural language input received from a user; proactively elicit and obtain information needed to fully infer the user intent (e.g., disambiguate words, games, intents, etc.); determine a task flow for satisfying the inferred intent; and execute the task flow to satisfy the inferred intent.

[0214] In some examples, as shown in Figure 7B the I / O processing module 728 can interact with a user through the I / O devices 716 in Figure 7A or through the I / O devices 716 in Figure 7AThe network communication interface 708 in the user device (e.g., device 104, device 200, device 400, or device 600) interacts with the user device to obtain user input (e.g., speech input) and provide responses to the user input (e.g., as speech output). The I / O processing module 728 optionally obtains contextual information associated with the user input from the user device along with or shortly after receiving the user input. The contextual information includes user-specific data, vocabulary, and / or preferences related to the user input. In some examples, the contextual information also includes the software state and hardware state of the user device at the time the user request is received, and / or information related to the user's surrounding environment at the time the user request is received. In some examples, the I / O processing module 728 also communicates follow-up questions to the user related to the user request and receives answers from the user. When a user request is received by the I / O processing module 728 and the user request includes speech input, the I / O processing module 728 forwards the speech input to the STT processing module 730 (or speech recognizer) for speech-to-text conversion.

[0215] The STT processing module 730 includes one or more ASR systems 758. The one or more ASR systems 758 can process speech input received through the I / O processing module 728 to produce recognition results. Each ASR system 758 includes a front-end speech preprocessor. The front-end speech preprocessor extracts representative features from the speech input. For example, the front-end speech preprocessor performs a Fourier transform on the speech input to extract spectral features characterizing the speech input as a sequence of representative multi-dimensional vectors. In addition, each ASR system 758 includes one or more speech recognition models (e.g., acoustic models and / or language models) and implements one or more speech recognition engines. Examples of speech recognition models include hidden Markov models, Gaussian mixture models, deep neural network models, n-gram language models, and other statistical models. Examples of speech recognition engines include engines based on dynamic time warping and engines based on weighted finite-state transducers (WFSTs). The extracted representative features of the front-end speech preprocessor are processed using the one or more speech recognition models and the one or more speech recognition engines to produce intermediate recognition results (e.g., phonemes, phoneme strings, and sub-words) and ultimately text recognition results (e.g., words, word strings, or symbol sequences). In some examples, the speech input is processed at least in part by a third-party service or on a user's device (e.g., device 104, device 200, device 400, or device 600) to produce the recognition results. Once the STT processing module 730 produces a recognition result containing a text string (e.g., a word, or a sequence of words, or a sequence of symbols), the recognition result is passed to the natural language processing module 732 for intent inference. In some examples, the STT processing module 730 produces multiple candidate text representations of the speech input. Each candidate text representation is a sequence of words or symbols corresponding to the speech input. In some examples, each candidate text representation is associated with a speech recognition confidence score. Based on the speech recognition confidence scores, the STT processing module 730 ranks the candidate text representations and provides the n-best (e.g., the n-highest ranked) candidate text representations to the natural language processing module 732 for intent inference, where n is a predetermined integer greater than zero. For example, in one example, only the highest-ranked (n = 1) candidate text representation is delivered to the natural language processing module 732 for intent inference. As another example, the five highest-ranked (n = 5) candidate text representations are passed to the natural language processing module 732 for intent inference.

[0216] Further details regarding speech-to-text processing are described in U.S. Utility Patent Application Serial No. 13 / 236,942, entitled "Consolidating Speech Recognition Results," filed September 20, 2011, the entire disclosure of which is incorporated by reference herein.

[0217] In some examples, the STT processing module 730 includes a vocabulary of recognizable words and / or accesses the vocabulary via a phonetic transcription module 731. Each vocabulary word is associated with one or more candidate pronunciations of the word represented in a speech recognition phonetic alphabet. In particular, the vocabulary of recognizable words includes words associated with multiple candidate pronunciations. For example, the vocabulary includes the word "tomato" associated with the candidate pronunciations and In addition, vocabulary words are associated with custom candidate pronunciations based on previous speech input from the user. Such custom candidate pronunciations are stored in the STT processing module 730 and are associated with a particular user via a user profile on the device. In some examples, candidate pronunciations of a word are determined based on the spelling of the word and one or more linguistic and / or phonetic rules. In some examples, candidate pronunciations are manually generated, e.g., based on known standard pronunciations.

[0218] In some examples, candidate pronunciations are ranked based on the prevalence of the candidate pronunciations. For example, the candidate pronunciation is ranked higher than the candidate pronunciation because the former is a more commonly used pronunciation (e.g., among all users, for users in a particular geographic region, or for any other suitable subset of users). In some examples, candidate pronunciations are ranked based on whether the candidate pronunciations are custom candidate pronunciations associated with the user. For example, custom candidate pronunciations are ranked higher than standard candidate pronunciations. This can be used to identify proper nouns with unique pronunciations that deviate from standard pronunciations. In some examples, candidate pronunciations are associated with one or more speech characteristics such as geographic origin, country, or ethnicity. For example, the candidate pronunciation is associated with the United States, while the candidate pronunciation is associated with the United Kingdom. Further, the ranking of candidate pronunciations is based on one or more characteristics (e.g., geographic origin, country, ethnicity, etc.) of the user stored in a user profile on the device. For example, it can be determined from the user profile that the user is associated with the United States. Based on the user being associated with the United States, the candidate pronunciation (associated with the United States) can be ranked higher than the candidate pronunciation (associated with the United Kingdom). In some examples, one of the ranked candidate pronunciations can be selected as the predicted pronunciation (e.g., the most likely pronunciation).

[0219] Upon receiving a speech input, the STT processing module 730 is used to determine the phonemes corresponding to the speech input (e.g., using an acoustic model), and then attempts to determine a word that matches the phonemes (e.g., using a language model). For example, if the STT processing module 730 first identifies a sequence of phonemes corresponding to a portion of the speech input It can then subsequently determine, based on the vocabulary index 744, that the sequence corresponds to the word "tomato".

[0220] In some examples, the STT processing module 730 uses fuzzy matching techniques to determine the words in the utterance. Thus, for example, the STT processing module 730 determines the phoneme sequence corresponds to the word "tomato", even though that particular phoneme sequence is not a candidate phoneme sequence for that word.

[0221] The natural language processing module 732 of the digital assistant ("natural language processor") takes the n-best candidate text representations ("word sequences" or "symbol sequences") generated by the STT processing module 730 and attempts to associate each candidate text representation with one or more "actionable intents" recognized by the digital assistant. An "actionable intent" (or "user intent") represents a task that can be performed by the digital assistant and that can have an associated task flow implemented in the task flow model 754. The associated task flow is a series of programmed actions and steps that the digital assistant takes in order to perform the task. The range of capabilities of the digital assistant depends on the number and variety of task flows that have been implemented and stored in the task flow model 754, or in other words, on the number and variety of "actionable intents" recognized by the digital assistant. However, the effectiveness of the digital assistant also depends on the ability of the assistant to infer the correct "one or more actionable intents" from a user request expressed in natural language.

[0222] In some examples, in addition to the sequence of words or symbols obtained from the STT processing module 730, the natural language processing module 732 also receives, e.g., from the I / O processing module 728, contextual information associated with the user request. The natural language processing module 732 optionally uses the contextual information to disambiguate, supplement, and / or further qualify the information contained in the candidate text representations received from the STT processing module 730. The contextual information includes, e.g., user preferences, hardware and / or software states of the user device, sensor information collected prior to, during, or shortly after the user request, prior interactions (e.g., conversations) between the digital assistant and the user, etc. As described herein, in some examples, the contextual information is dynamic and varies with the time, location, content, and other factors of the conversation.

[0223] In some examples, natural language processing is based on, for example, a knowledge ontology 760. The knowledge ontology 760 is a hierarchical structure containing a number of nodes, each of which represents a "actionable intent" or an "attribute" related to one or more of an "actionable intent" or another "attribute." As described above, an "actionable intent" represents a task that a digital assistant is capable of performing, i.e., the task is "actionable" or can be performed. An "attribute" represents a parameter associated with a sub- aspect of an actionable intent or another attribute. Connections between actionable intent nodes and attribute nodes in the knowledge ontology 760 define how the parameters represented by the attribute nodes pertain to the tasks represented by the actionable intent nodes.

[0224] In some examples, the knowledge ontology 760 is composed of actionable intent nodes and attribute nodes. Within the knowledge ontology 760, each actionable intent node is connected directly to or through one or more intermediate attribute nodes to one or more attribute nodes. Similarly, each attribute node is connected directly to or through one or more intermediate attribute nodes to one or more actionable intent nodes. For example, as shown in Figure 7C the knowledge ontology 760 includes a "restaurant reservation" node (i.e., an actionable intent node). The attribute nodes "cuisine," "price range," "phone number," and "location" are all sub-nodes of the attribute node "restaurant" and are each linked to the "restaurant reservation" node (i.e., the actionable intent node) through the intermediate attribute node "restaurant."

[0225] Further, the attribute nodes "cuisine," "price range," "phone number," and "location" are sub-nodes of the attribute node "restaurant" and are each linked to the "restaurant reservation" node (i.e., the actionable intent node) through the intermediate attribute node "restaurant." As another example, as shown in Figure 7C the knowledge ontology 760 also includes a "set reminder" node (i.e., another actionable intent node). The attribute nodes "date / time" (for setting a reminder) and "topic" (for the reminder) are each connected to the "set reminder" node. Because the attribute "date / time" is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node "date / time" is connected to both the "restaurant reservation" node and the "set reminder" node in the knowledge ontology 760.

[0226] An actionable intent node, along with its linked attribute nodes, is described as a "domain." In this discussion, each domain is associated with a respective actionable intent and refers to a set of nodes (and relationships between these nodes) associated with a particular actionable intent. For example, Figure 7CThe knowledge ontology 760 shown in includes an example of a restaurant reservation domain 762 and an example of a reminder domain 764 within the knowledge ontology 760. The restaurant reservation domain includes an executable intent node "restaurant reservation", attribute nodes "restaurant", "date / time" and "party size", and child attribute nodes "cuisine", "price range", "phone number" and "location". The reminder domain 764 includes an executable intent node "set reminder" and attribute nodes "subject" and "date / time". In some examples, the knowledge ontology 760 is composed of multiple domains. Each domain shares one or more attribute nodes with one or more other domains. For example, in addition to the restaurant reservation domain 762 and the reminder domain 764, the "date / time" attribute node is also associated with many different domains (e.g., a scheduling domain, a travel booking domain, a movie ticket domain, etc.).

[0227] although Figure 7C Two example domains within ontology 760 are shown, but other domains include, for example, "find a movie," "make a phone call," "find directions," "schedule a meeting," "send a message," and "provide answers to questions," "read a list," "provide navigation instructions," "provide instructions for a task," and the like. The "send message" domain is associated with the "send message" executable intent node and further includes attribute nodes such as "one or more recipients," "message type," and "message body." The attribute node "recipients" is further defined by, for example, child attribute nodes such as "recipient name" and "message address."

[0228] In some examples, ontology 760 includes all domains (and thus executable intents) that the digital assistant can understand and act upon. In some examples, ontology 760 is modified, such as by adding or removing entire domains or nodes, or by modifying the relationships between nodes within ontology 760.

[0229] In some examples, nodes associated with multiple related executable intents are clustered under a "superdomain" in the knowledge ontology 760. For example, the "travel" superdomain includes a cluster of attribute nodes and executable intent nodes related to travel. The executable intent nodes related to travel include "flight booking", "hotel booking", "car rental", "route planning", "find points of interest", etc. The executable intent nodes under the same superdomain (e.g., the "travel" superdomain) have multiple common attribute nodes. For example, the executable intent nodes for "flight booking", "hotel booking", "car rental", "get directions", and "find points of interest" share one or more of the attribute nodes "starting location", "destination", "departure date / time", "arrival date / time", and "party size".

[0230] In some examples, each node in the knowledge ontology 760 is associated with a set of words and / or phrases that are related to the attribute or actionable intent represented by the node. The respective set of words and / or phrases associated with each node is a so-called "vocabulary" associated with the node. The respective set of words and / or phrases associated with each node is stored in the vocabulary index 744 associated with the attribute or actionable intent represented by the node. For example, returning to the example above Figure 7B The vocabulary associated with the node for the "restaurant" attribute includes words such as "food," "drinks," "cuisine," "hungry," "eat," "pizza," "fast food," "meal," and the like. As another example, the vocabulary associated with the node for the "initiate a phone call" actionable intent includes words and phrases such as "call," "make a call," "dial," "talk to," "call that number," "call," and the like. The vocabulary index 744 optionally includes words and phrases in different languages.

[0231] The natural language processing module 732 receives the candidate text representations (e.g., one or more text strings or one or more symbol sequences) from the STT processing module 730 and, for each candidate representation, determines which nodes in the knowledge ontology 760 the words in the candidate text representation refer to. In some examples, if a word or phrase in the candidate text representation is found to be associated (via the vocabulary index 744) with one or more nodes in the knowledge ontology 760, the word or phrase "triggers" or "activates" those nodes. Based on the number and / or relative importance of the activated nodes, the natural language processing module 732 selects one of the actionable intents as the task that the user intends for the digital assistant to perform. In some examples, the domain with the most "triggered" nodes is selected. In some examples, the domain with the highest confidence (e.g., based on the relative importance of its individual triggered nodes) is selected. In some examples, the domain is selected based on a combination of the number and importance of the triggered nodes. In some examples, additional factors are also considered in the process of selecting the node, such as whether the digital assistant has previously correctly interpreted similar requests from the user.

[0232] The user data 748 includes information specific to the user, such as user-specific vocabulary, user preferences, user address, the user's default second language, the user's contact list, and other short-term or long-term information for each user. In some examples, the natural language processing module 732 uses the user-specific information to supplement the information contained in the user input to further qualify the user's intent. For example, for the user request "invite my friends to my birthday party," the natural language processing module 732 can access the user data 748 to determine which "friends" and when and where the "birthday party" will take place, without requiring the user to explicitly provide such information in their request.

[0233] It is recognized that, in some examples, the natural language processing module 732 is implemented with one or more machine learning mechanisms (e.g., neural networks). In particular, the one or more machine learning mechanisms are configured to receive a candidate text representation and context information associated with the candidate text representation. Based on the candidate text representation and the associated context information, the one or more machine learning mechanisms are configured to determine an intent confidence score based on a set of candidate actionable intents. The natural language processing module 732 can select one or more candidate actionable intents from the set of candidate actionable intents based on the determined intent confidence score. In some examples, the one or more candidate actionable intents are also selected from the set of candidate actionable intents with a knowledge ontology (e.g., the knowledge ontology 760).

[0234] Further details of searching a knowledge ontology based on a symbol string are described in U.S. Utility Patent Application Serial No. 12 / 341,743, entitled "Method and Apparatus for Searching Using An Active Ontology," filed December 22, 2008, the entire disclosure of which is incorporated by reference herein.

[0235] In some examples, once the natural language processing module 732 identifies an actionable intent (or domain) based on a user request, the natural language processing module 732 generates a structured query to represent the identified actionable intent. In some examples, the structured query includes parameters for one or more nodes within the domain of the actionable intent, and at least some of the parameters are populated with specific information and requirements specified in the user request. For example, a user says "Help me make a reservation for a sushi restaurant at 7 p.m. tonight." In this case, the natural language processing module 732 is able to correctly identify the actionable intent as "restaurant reservation" based on the user input. From the knowledge ontology, the structured query for the "restaurant reservation" domain includes parameters such as {cuisine}, {time}, {date}, {party size}, etc. In some examples, based on the spoken input and the text derived from the spoken input using the STT processing module 730, the natural language processing module 732 generates a partially structured query for the restaurant reservation domain, where the partially structured query includes the parameters {cuisine = "sushi"} and {time = "7 p.m."}. However, in this example, the user utterance contains insufficient information to complete the structured query associated with the domain. Thus, based on the currently available information, other necessary parameters such as {party size} and {date} are not specified in the structured query. In some examples, the natural language processing module 732 populates some of the parameters of the structured query with received context information. For example, in some examples, if the user requests a "nearby" sushi restaurant, the natural language processing module 732 populates the {location} parameter in the structured query with GPS coordinates from the user device.

[0236] In some examples, the natural language processing module 732 identifies a plurality of candidate actionable intents for each candidate text representation received from the STT processing module 730. Additionally, in some examples, a respective structured query is generated (partially or entirely) for each identified candidate actionable intent. The natural language processing module 732 determines an intent confidence score for each candidate actionable intent and ranks the candidate actionable intents based on the intent confidence scores. In some examples, the natural language processing module 732 communicates the generated structured query(s) (including any completed parameters) to the task flow processing module 736 ("task flow processor"). In some examples, one or more structured queries are provided to the task flow processing module 736 for the m best (e.g., m highest ranked) candidate actionable intents, where m is a predetermined integer greater than zero. In some examples, the one or more structured queries for the m best candidate actionable intents are provided to the task flow processing module 736 along with the corresponding candidate text representation(s).

[0237] Other details of inferring user intent based on a plurality of candidate actionable intents determined from a plurality of candidate text representations of speech inputs are described in U.S. Utility Patent Application Serial No. 14 / 298,725, filed June 6, 2014, entitled "System and Method for Inferring User Intent From Speech Inputs," the entire disclosure of which is incorporated by reference herein.

[0238] The task flow processing module 736 is configured to receive one or more structured queries from the natural language processing module 732, complete the structured query(s) (as necessary), and perform the actions necessary to "fulfill" the user's ultimate request. In some examples, various processes necessary to complete these tasks are provided in a task flow model 754. In some examples, the task flow model 754 includes processes for obtaining additional information from the user, as well as task flows for performing the actions associated with the actionable intent.

[0239] As described above, to complete the structured query, the task flow processing module 736 needs to initiate additional dialog with the user in order to obtain additional information and / or disambiguate potentially ambiguous utterances. When such interaction is necessary, the task flow processing module 736 invokes the dialog flow processing module 734 to engage in a dialog with the user. In some examples, the dialog flow processing module 734 determines how (and / or when) to request additional information from the user, and receives and processes the user's response. The questions are provided to the user and answers are received from the user through the I / O processing module 728. In some examples, the dialog flow processing module 734 presents dialog output to the user via audible output and / or visual output, and receives input from the user via spoken or physical (e.g., click) responses. Continuing the example described above, when the task flow processing module 736 invokes the dialog flow processing module 734 to determine "party size" and "date" information for the structured query associated with the domain "restaurant reservation," the dialog flow processing module 734 generates questions such as "How many in your party?" and "Which day for the reservation?" to pass to the user. Once answers are received from the user, the dialog flow processing module 734 fills in the structured query with the missing information, or passes the information to the task flow processing module 736 to complete the missing information from the structured query.

[0240] Once the task flow processing module 736 has completed the structured query for the executable intent, the task flow processing module 736 begins executing the final task associated with the executable intent. Thus, the task flow processing module 736 executes the steps and instructions in the task flow model according to the particular parameters contained in the structured query. For example, the task flow model for the executable intent "restaurant reservation" includes steps and instructions for contacting the restaurant and actually requesting a reservation for a particular party size at a particular time. For example, using a structured query such as: {restaurant reservation, restaurant = ABC Cafe, date = 3 / 12 / 2012, time = 7pm, party size = 5}, the task flow processing module 736 can perform the following steps: (1) log into the server for ABC Cafe or a restaurant reservation system such as (2) enter the date, time, and party size information in the form on the website, (3) submit the form, and (4) create a calendar entry for the reservation in the user's calendar.

[0241] In some examples, the task flow processing module 736, with the assistance of the service processing module 738 ("service processing module"), fulfills a task requested in the user input or provides an informational answer requested in the user input. For example, the service processing module 738 initiates a telephone call on behalf of the task flow processing module 736, sets a calendar entry, invokes a map search, invokes or interacts with other user applications installed on the user device, and invokes or interacts with third-party services (e.g., a restaurant reservation portal, a social networking site, a banking portal, etc.). In some examples, the protocols and application programming interfaces (APIs) required for each service are specified by a corresponding service model in the service models 756. The service processing module 738 accesses the appropriate service model for a service and generates a request for the service according to the protocols and APIs required for the service in accordance with the service model.

[0242] For example, if a restaurant has enabled an online reservation service, the restaurant submits a service model that specifies the necessary parameters for making a reservation and an API for transmitting values of the necessary parameters to the online reservation service. Upon request by the task flow processing module 736, the service processing module 738 can use the web address stored in the service model to establish a network connection with the online reservation service and transmit the necessary parameters for the reservation (e.g., time, date, number of party) to the online reservation interface in a format according to the API of the online reservation service.

[0243] In some examples, the natural language processing module 732, the dialog flow processing module 734, and the task flow processing module 736 are used collectively and iteratively to infer and qualify a user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response (i.e., output to the user, or fulfill a task) to satisfy the user's intent. The generated response is a dialog response to the speech input that satisfies the user's intent, at least in part. Additionally, in some examples, the generated response is output as speech output. In these examples, the generated response is transmitted to a speech synthesis processing module 740 (e.g., a speech synthesizer), where the generated response can be processed to synthesize the dialog response in speech form. In other examples, the generated response is data content related to satisfying the user's request in the speech input.

[0244] In examples where the task flow processing module 736 receives multiple structured queries from the natural language processing module 732, the task flow processing module 736 first processes a first structured query of the received structured queries to attempt to fulfill the first structured query and / or execute one or more tasks or actions represented by the first structured query. In some examples, the first structured query corresponds to the highest ranked executable intent. In other examples, the first structured query is selected from the received structured queries based on a combination of the corresponding speech recognition confidence scores and the corresponding intent confidence scores. In some examples, if the task flow processing module 736 encounters an error during processing of the first structured query (e.g., due to an inability to determine a necessary parameter), the task flow processing module 736 can proceed to select and process a second structured query of the received structured queries that corresponds to a lower ranked executable intent. The second structured query is selected, for example, based on a speech recognition confidence score of the corresponding candidate text representation, an intent confidence score of the corresponding candidate executable intent, a missing necessary parameter in the first structured query, or any combination thereof.

[0245] The speech synthesis processing module 740 is configured to synthesize speech output for presentation to the user. The speech synthesis processing module 740 synthesizes speech output based on text provided by the digital assistant. For example, the generated dialog response is in the form of a text string. The speech synthesis processing module 740 converts the text string into audible speech output. The speech synthesis processing module 740 uses any appropriate speech synthesis technique in order to generate speech output from text, including but not limited to: concatenative synthesis, unit selection synthesis, diphthong synthesis, domain-specific synthesis, formant synthesis, articulatory synthesis, Hidden Markov Model (HMM)-based synthesis, and sinusoidal synthesis. In some examples, the speech synthesis processing module 740 is configured to synthesize individual words based on a phoneme string corresponding to the words. For example, a phoneme string is associated with a word in the generated dialog response. The phoneme string is stored in metadata associated with the word. The speech synthesis processing module 740 is configured to directly process the phoneme string in the metadata to synthesize the word in speech form.

[0246] In some examples, instead of (or in addition to) using the speech synthesis processing module 740, speech synthesis is performed on a remote device (e.g., the server system 108) and the synthesized speech is transmitted to the user device for output to the user. For example, this can occur in some implementations where the output of the digital assistant is generated at the server system. Also, since the server system typically has more processing power or more resources than the user device, it is possible to obtain higher quality speech output than would be achieved with client-side synthesis.

[0247] Additional details regarding digital assistants can be found in U.S. Utility Patent Application No. 12 / 987,982, entitled "Intelligent Automated Assistant," filed on January 10, 2011, and U.S. Utility Patent Application No. 13 / 251,088, entitled "Generating and Processing Task Items That Represent Tasks to Perform," filed on September 30, 2011, the entire disclosures of which are incorporated by reference herein.

[0248] 4. System and process for modifying tasks after execution

[0249] Figure 8 A block diagram of a digital assistant 800 for modifying tasks after execution according to various examples is illustrated. In some examples, the digital assistant 800 is implemented on one or more electronic devices (e.g., devices 104, 122, 200, 400, 600, 900, 1000, or 1100), and the modules and functionality of the digital assistant 800 can be distributed among the devices in any fashion. In some examples, some of the modules and functionality of the digital assistant 800 are divided into a server portion and a client portion, where the client portion resides on one or more user devices (e.g., devices 104, 122, 200, 400, 600, 900, 1000, or 1100) and communicates with the server portion (e.g., server system 108) over one or more networks, for example, as shown. The digital assistant 800 is implemented using hardware, software, or a combination of hardware and software to perform the principles discussed herein. Figure 1 The digital assistant 800 is illustrated using a client-server architecture, where the client portion of the digital assistant 800 resides on one or more user devices (e.g., devices 104, 122, 200, 400, 600, 900, 1000, or 1100) and communicates with the server portion of the digital assistant 800 over one or more networks. In some examples, the digital assistant 800 is implemented using a peer-to-peer architecture, where the client portion of the digital assistant 800 resides on one or more user devices (e.g., devices 104, 122, 200, 400, 600, 900, 1000, or 1100) and communicates with other client portions of the digital assistant 800 over one or more networks. In some examples, the digital assistant 800 is implemented using a combination of a client-server architecture and a peer-to-peer architecture.

[0250] In some examples, a sub-module or a subset or superset of sub-modules described with respect to the digital assistant 800 can include one or more of the components discussed above, including components for automated speech recognition, natural language processing, or speech-to-text capabilities.

[0251] It should be noted that the digital assistant 800 is exemplary, and thus the digital assistant 800 can have more or less components than shown, can combine two or more components, or can have a different configuration or arrangement of components. Additionally, although the following discussion describes functionality performed at a single component of the digital assistant 800, it should be understood that such functionality can be performed at other components of the digital assistant 800, and that such functionality can be performed at more than one component of the digital assistant 800.

[0252] Figures 9A-9C Example modifications to tasks after execution by the digital assistant 800 according to various examples are illustrated.Figures 10A-10C Illustrated are exemplary modifications to a task after execution by the digital assistant 800 according to various examples. Figures 11A-11D Illustrated are exemplary modifications to a task after execution by the digital assistant 800 according to various examples. Figures 9A-9C 、 Figures 10A-10C and Figures 11A-11D Each of these will be discussed with the digital assistant 800.

[0253] like Figure 8 As illustrated, digital assistant 800 includes a modification module 810 and one or more domains 820. In some examples, modification module 810 is implemented using one of the other components of the digital assistant discussed above, including an input / output processing module, a speech-to-text processing module, a natural language processing module, a task flow processing module, etc. In some examples, domain 820 is included in an ontology (such as ontology 760), a natural language processing module, or another module of the digital assistant.

[0254] The digital assistant 800 receives user speech input 805 and provides the user speech input 805 to one or more components of the digital assistant 800, including a modification module 810, which determines a task 815 specified in the user speech input 805 and performs the task 815. In some examples, the modification module 810 determines that the task 815 is performing speech-to-text processing, natural language processing, and / or user intent determination. In some examples, another component of the digital assistant 800 discussed above determines the task 815 and provides the task 815 to the modification module 810. In some examples, the modification module 810 and another module of the digital assistant 800 perform the user intent determination concurrently. In some examples, as discussed further below, the modification module 810 performs a limited form of user intent determination to determine whether the user intends to modify or undo a previously performed task.

[0255] In some examples, the digital assistant 800 receives user speech input 805 from the user using a microphone of the electronic device. Figure 9A As shown, the user provides a user speech input 905 of "delete all my alarms", which is received by the microphone of the electronic device 900. As another example, Figure 10A As shown, the user provides a user speech input 1005 of “turn on the kitchen lights,” which is received by the microphone of the electronic device 900 .

[0256] In some examples, the digital assistant 800 receives the user verbal input 805 from another electronic device. For example, the digital assistant 800 can operate at least partially on a server, and thus can receive the user verbal input 905 of “delete all my alarms” from the electronic device 900 at the server. As another example, the digital assistant 800 can receive the user verbal input 1105 of “create an alarm for 2 PM” from the electronic device 1100 in a remote operation, or can receive the user verbal input 1105 at a microphone of the electronic device 1100 when operating on the electronic device 1100. Thus, in some examples, all processing of the verbal input and task determination can occur locally on the electronic device 1100.

[0257] After determining the task 815, the digital assistant 800 performs the task 815. For example, as shown, prior to receiving the user verbal input 905 of “delete all my alarms,” the alarm application displays two different alarms, one alarm at 7 AM and one alarm at 2 PM. After the digital assistant 800 receives the user verbal input 905 and determines the user intent to delete all currently existing alarms, the digital assistant 800 performs (or causes performance of) the task of deleting all currently existing alarms. Thus, the alarm application will display a user interface with no alarms, as shown. Figure 9A Figure 9B

[0258] As another example, after the digital assistant 800 receives the user verbal input 1005 and determines the user intent to turn on a smart light or bulb located in the user’s kitchen, the digital assistant 800 performs (or causes performance of) the task by activating the kitchen light in the home application, as shown. As yet another example, after the digital assistant 800 receives the user verbal input 1105 and determines the user intent to create an alarm, the digital assistant 800 performs the task by opening (e.g., displaying) the alarm application at 2 PM and creating an alarm, as shown. Figure 10B Figure 11B

[0259] In some examples, performing the task 815 includes initiating performance of the task 815. For example, when the digital assistant 800 receives the user verbal input of “text my sister,” the digital assistant 800 causes the electronic device to provide a user interface including a draft text message and indicating the recipient of the text message as the user’s sister. Thus, in performing the task 815, the digital assistant 800 has initiated performance but has not completed performance, as the text message has not yet been transmitted to the requested recipient.

[0260] In some examples, after performing the task 815, the digital assistant 800 provides an output indicating performance of the task 815. In some examples, the output is a visual output. For example, as shown, after the digital assistant 800 performs the task of deleting all currently existing alarms, the digital assistant 800 provides a visual output indicating that the task has been performed.​​​​Figure 9B As shown, the display 901 of the electronic device 900 is updated to present an output that the task of providing a deletion of the current alarm does not exist. As another example, as shown in FIG. 9B, the display 901 of the electronic device 900 is updated to present an output that the task of providing a deletion of the current alarm has been completed. Figure 10B As shown, changing the appearance of the affordance 1002 associated with the kitchen light indicates that the task of turning on the kitchen light has been completed.

[0261] In some examples, the visual output includes opening (e.g., displaying) an application associated with the task. For example, as shown in FIG. 9B, the digital assistant 800 causes the electronic device 900 to display the alarm application. Figures 11A-11B As shown, the home screen user interface is displayed when the user verbal input 1105 is received by the electronic device 1100. Once the digital assistant 800 processes the user verbal input 1105 and performs the task of creating an alarm at 2:00 PM, the digital assistant 800 causes the electronic device 1100 to display a user interface associated with the alarm application.

[0262] In some examples, the digital assistant 800 performs the task and provides the visual output without changing the user interface displayed on the electronic device. In some examples, the visual output indicating the performance of the task 815 is provided in an affordance or other user interface element associated with the digital assistant 800 and is displayed over (e.g., on top of) the currently displayed user interface. Thus, the digital assistant 800 can perform the requested task and provide an output notifying the user without interrupting the user’s current interaction with the electronic device.

[0263] In some examples, the output is an audio output. In some examples, the audio output includes a description of the task 815. For example, as shown in FIG. 9B, the digital assistant 800 causes the electronic device 900 to provide an audio output 925 of “I deleted your alarm” that indicates that the task has been completed and also describes the task. As another example, as shown in FIG. 9C, the digital assistant 800 can cause the electronic device 1000 to provide an audio output 1025 of “created an alarm at 2:00 PM” to indicate that the requested alarm has been created. Figure 9B As shown, the digital assistant 800 causes the electronic device 900 to provide an audio output 925 of “I deleted your alarm” that indicates that the task has been completed and also describes the task. As another example, as shown in FIG. 9C, the digital assistant 800 can cause the electronic device 1000 to provide an audio output 1025 of “created an alarm at 2:00 PM” to indicate that the requested alarm has been created. Figure 10B As shown, the digital assistant 800 can cause the electronic device 1000 to provide an audio output 1025 of “created an alarm at 2:00 PM” to indicate that the requested alarm has been created.

[0264] As the tasks are performed by the digital assistant 800, the domain 820 (or the domain associated with the particular task) will provide information related to the task to the modification module 810. In particular, the domain 820 will provide an indication to the modification module 810 of whether the task or a portion of the task can be modified and / or reversed after the performance of the task is completed. For example, as shown in FIG. 8, the domain 820 provides an indication to the modification module 810 that the task of deleting the alarm can be recreated if requested. Figure 9B As shown, the domain 820 can provide an indication to the modification module 810 that the alarm can be recreated if requested when the task of deleting the alarm is performed.

[0265] As another example, as shown in FIG. 8, the domain 820 provides an indication to the modification module 810 that the task of creating an alarm at 2:00 PM can be modified and / or reversed after the performance of the task is completed. Figure 10BAs shown, the task of opening the kitchen light, domain 820 provides an indication that the target device (e.g., the device that has been opened) can be changed. In particular, domain 820 can indicate that the state of the kitchen light is reversible (e.g., can be turned off), and that other lights or devices (e.g., lights in the bedroom, a smart lock, etc.) can be targets of similar commands.

[0266] Additionally, when domain 820 indicates that a task or a portion of a task can be modified and / or reversed after completion (and / or initiation) of the execution of the task, domain 820 provides the task that will be executed (or parameters used to execute the modification and / or reversal) to modification module 810. For example, when digital assistant 800 executes the task of deleting the user's alarm as shown in FIG. 7, domain 820 provides the task of recreating the alarm to modification module 810 for reference in the event that an additional user utterance is received. Figure 9B As shown, the task of opening the kitchen light, domain 820 provides an indication that the target device (e.g., the device that has been opened) can be changed. In particular, domain 820 can indicate that the state of the kitchen light is reversible (e.g., can be turned off), and that other lights or devices (e.g., lights in the bedroom, a smart lock, etc.) can be targets of similar commands.

[0267] As another example, when digital assistant 800 executes the task of opening the kitchen light as shown in FIG. 7, domain 820 provides the task of turning off the kitchen light to modification module 810. Additionally, because digital assistant 800 determines that there is another smart light in the user's system that has been indicated by the user, domain 820 can also provide the task of turning on the bedroom light to modification module 810. In this way, when an additional utterance is received by the user, as discussed further below, modification module 810 can reference the alternative tasks provided by domain 820 to quickly and efficiently determine the user's intent. Figure 10B As shown, the task of opening the kitchen light, domain 820 provides an indication that the target device (e.g., the device that has been opened) can be changed. In particular, domain 820 can indicate that the state of the kitchen light is reversible (e.g., can be turned off), and that other lights or devices (e.g., lights in the bedroom, a smart lock, etc.) can be targets of similar commands.

[0268] As yet another example, when digital assistant 800 executes the task of creating an alarm at 2 PM as shown in FIG. 7, domain 820 can provide the task of deleting the alarm to modification module 810. Additionally, domain 820 can identify tasks related to the task of creating the alarm, including tasks of creating a calendar appointment, starting a timer, etc. Accordingly, domain 820 provides the tasks of creating a calendar appointment at 2 PM and starting a timer at 2 PM to modification module 810. Figure 10B As another example, when digital assistant 800 initiates the execution of the task of sending a text message, domain 820 can provide data indicating that the recipient of the text message is a field that can be modified. Accordingly, if digital assistant 800 receives a later input related to the recipient of the text message, modification module 810 can quickly identify that the recipient is a field (or parameter) that can be modified.

[0269]

[0270] ​Domain 820 can provide modification module 810 with any number of related tasks that assist modification module 810 in processing future user speech input. In some examples, domain 820 provides a predetermined number of most recently accessed related tasks. In some examples, domain 820 provides a predetermined number of most closely related tasks (e.g., tasks that share a majority of nodes of the domain). In some examples, domain 820 provides a predetermined number of tasks that are most frequently accessed by other users. In some examples, domain 820 provides a predetermined number of tasks that are frequently accessed by a user associated with the electronic device (e.g., a user that owns or operates the device).

[0271] After performing task 815, digital assistant 800 receives another user speech input 805 and provides user speech input 805 to modification module 810 to determine a user intent and whether user speech input 805 includes a modification to task 815. For example, digital assistant 800 receives user speech input 915 of "no, no, no" and provides user speech input 915 to modification module 810, which can process user speech input 915 to determine a user intent and use the following processes to determine whether user speech input 915 includes a modification to task 815.

[0272] In some examples, digital assistant 800 receives another user speech input after initiating performance of a task. For example, while preparing a text message as requested by a user, digital assistant 800 can receive another user speech input indicating that the user wants to change certain aspects of the text message, such as a recipient, content, and / or attachments. Accordingly, before completing performance of the task, digital assistant 800 receives a user speech input that includes a modification and makes the appropriate changes.

[0273] In some examples, modification module 810 determines whether user speech input 805 includes an intent to undo an action (e.g., a task). In particular, modification module 810 can determine whether user speech input 805 includes a predetermined word, whether user speech input 805 is received within a predetermined time of performing task 815, and / or whether task 815 includes a reversible nature, as discussed further below.

[0274] In some examples, the intent to undo an action corresponds to a modification to task 815. For example, modification module 810 receives user speech input 1015 of "oh, no, I mean the bedroom" and determines that user speech input 1015 includes a modification to a previously performed task of turning on a light in the kitchen. In particular, modification module 810 determines that user speech input 1015 indicates that the user wants to turn on a light in the bedroom and does not want to keep the kitchen light on. Accordingly, user speech input 1015 indicates a user intent to undo a previous action (e.g., task) and also corresponds to a modification to a task.

[0275] As another example, the modification module 810 receives the user verbal input 1115 of "no, make it a calendar event," determines that the user verbal input 1115 includes a modification to a previously performed task of creating an alarm. In particular, the user verbal input 1115 indicates that the user does not want the previously created alarm, but wants a calendar reminder at the same time. Thus, the user verbal input 1115 indicates a user intent to undo the creation of the alarm and also corresponds to a modification because a calendar event or reminder at the same time should be created.

[0276] In some examples, determining whether the user verbal input 805 includes an intent to undo an action includes determining whether the user verbal input 805 includes a predetermined word or phrase. Example predetermined words or phrases include no, never mind, undo, just kidding, not like that, stop, another, etc. For example, the modification module 810 receives the user verbal input 915 of "no, never mind" and determines that the presence of "never mind" indicates an intent to undo the previous task. As another example, the modification module 810 receives the user verbal input 1015 of "oh no, I mean the bedroom" and determines that the presence of "oh no" indicates an intent to undo the previous task. As another example, the modification module 810 receives the user verbal input 1115 of "no, make it a calendar event" and determines that the presence of "no" indicates an intent to undo the previous task.

[0277] In some examples, determining whether the user verbal input 805 includes an intent to undo an action includes determining whether the user verbal input 805 was received within a predetermined time (e.g., five seconds, ten seconds, twenty seconds, thirty seconds, or forty-five seconds) of performing the task 815. For example, the modification module 810 determines that the user verbal input 915 was received 7 seconds after performing the task of deleting the alarm and thus determines that the user verbal input 915 is likely related to the previously performed task.

[0278] In some examples, determining whether the user verbal input 805 includes an intent to undo an action includes determining whether the task 815 includes a reversible nature. In particular, as discussed above, the domain 820 provides the task and a nature of the task to the modification module 810 when the task is performed. Thus, one of the natures provided by the domain 820 to the modification module 810 is whether a part of the task is reversible. For example, the domain 820 can indicate to the modification module 810 that an alarm can be recreated and / or deleted or that a light can be turned off. Each of these represents a reversible nature of a performed task that the modification module 810 can act upon.

[0279] In some examples, the modification module 810 and / or a portion of the modification module 810 that determines whether the user verbal input 805 includes an intent to undo an action is implemented with a machine learning model or neural network. In this way, the modification module 810 can be trained and / or refined over time to recognize subsequent utterances that a user is likely to provide and subsequent tasks that a user is likely to request. This allows for smoother determinations by the modification module 810 that can be adjusted over time, rather than relying solely on heuristics to make determinations

[0280] After determining that the user verbal input 805 includes a modification to the task 815, the digital assistant 800 performs (or causes performance of) a task 825 that modifies at least a portion of the task 815 or performance of the task 815.

[0281] In some examples, the task 825 undoes (e.g., reverses) performance of the task 815. For example, as shown in Figure 9C the digital assistant 800 undoes the deletion of the alarm that was performed with the previous task. In other words, the digital assistant 800 performs the task of creating an alarm, which reverses the previous deletion of the alarm. As another example, as shown in Figure 10C the digital assistant 800 undoes (e.g., reverses) turning on the kitchen light by turning off the kitchen light. As yet another example, as shown in Figure 11C the digital assistant 800 undoes the creation of the 2 PM alarm by deleting the 2 PM alarm.

[0282] In some examples, the task 825 includes information about the task 815 and / or the task 815 based on the nature of the domain 820. Thus, the task 825 is based on the task 815, and the domain 820 provides information that connects or links the two tasks. For example, when the task 825 is deleting an object that was previously created, the goal of the task 825 is the same as the goal of the task 815. Similarly, when the task 825 is turning off an electronic device, the nature of being disabled to turn off the device is the same as the nature of the task 815 being enabled to turn on the device.

[0283] In some examples, the task 825 includes deleting a created object. For example, as shown in Figure 11C the digital assistant 800 performs a task of deleting the previously created alarm at 2 PM in response to the user verbal input 1115.

[0284] In some examples, the task 825 includes passing data from a first application to a second application. For example, as shown in Figure 11D the digital assistant 800 not only deletes the previously created alarm, but also passes relevant information (e.g., the time of 2 PM) from the alarm application to the calendar application and creates the requested calendar event in response to the user verbal input 1115.

[0285] In some examples, the task 825 includes recreating a previously deleted object. For example, as shown, in response to the user verbal input 915, the digital assistant 800 recreates the previously deleted alarm to undo the previously performed task. Figure 9C

[0286] In some examples, the task 825 includes changing a state of the electronic device. For example, as shown, the digital assistant 800 performs (or causes to perform) a task that changes a kitchen light from an on state to an off state. Changing a state of the electronic device can include turning a device on or off, changing from a locked state to an unlocked state (or vice versa), changing a volume of the device, changing to or from a do not disturb mode, or any similar state change. Figure 10C

[0287] In some examples, the task 825 includes changing a parameter of the first task. For example, when receiving a user verbal input such as "Oh, no, I meant my mom," the digital assistant 800 can change the recipient of the draft text message from another person to the user's mom. As another example, when receiving a user verbal input such as "No problem, it's Paul McCartney," the digital assistant 800 can play a song by Paul McCartney instead of the previously provided artist.

[0288] In some examples, after performing the task 825, the digital assistant 800 causes the electronic device to provide an output indicating performance of the task 825. As discussed above with respect to the output indicating performance of the task 815, the output provided by the digital assistant can include a visual output and / or an audio output. For example, as shown, Figure 9C

[0289] As another example, as shown, Figure 10C Figure 11C As another example, as shown,

[0290] ​​​​In some examples, the indication of the performance of task 825 is an indication that task 815 was not performed. For example, as discussed above, the audio output 935 of "OK, I will not do that" indicates that the digital assistant 800 will not delete the alarm as requested by the user in the user utterance input 805. Thus, the audio output 935 indicates that the previously requested task of deleting will not occur and that the task of canceling (e.g., undoing) the deletion (or recreating the alarm) has been completed.

[0291] In some examples, in accordance with a determination that the user utterance input 805 includes a modification to task 815, the digital assistant 800 determines whether the user utterance input 805 also includes task 835 in addition to task 825. For example, when receiving the user utterance input 1015, the digital assistant 800 (and the modification module 810) determines that "Oh, no, I mean the bedroom" includes a modification to the previously performed task of turning on the kitchen light and thus includes two tasks. In particular, the digital assistant 800 determines that the two tasks are the task 825 of turning off the kitchen light and the task 825 of turning on the bedroom light.

[0292] As another example, when receiving the user utterance input 1115, the digital assistant 800 (and the modification module 81) determines that "No, make it a calendar event" includes a modification to the previously performed task of creating an alarm and thus includes two tasks. In particular, the digital assistant 800 determines that the two tasks are the task 825 of deleting the previously created alarm and the task 835 of creating a calendar event.

[0293] In some examples, the task 835 is related to the task 825. For example, the task of creating a calendar event at 2 PM includes the same parameter (e.g., 2 PM) as the deleted alarm. In some examples, the task 835 is related to the task 815. For example, the task of turning on the bedroom light is a task that is the same as turning on the kitchen light but has a different parameter (e.g., bedroom instead of kitchen).

[0294] In some examples, the task 835 continues or completes the performance of a previously initiated task. For example, when the digital assistant 800 receives an input requesting a change to a text message or similar communication (e.g., email, instant message, etc.), as a second task (e.g., task 825) in a sequence, the digital assistant 800 makes the desired change to the requested parameter. Thus, continuing or completing the performance of the task of transmitting the text message or other communication is a third task (e.g., task 835) in the sequence.

[0295] In some examples, after performing the task 835, the digital assistant 800 outputs (or causes the electronic device to output) an indication that the task 835 was performed. As discussed above with respect to tasks 815 and 825, the digital assistant 800 provides a visual output and / or an audio output. For example, as discussed above, the digital assistant 800 outputs an audio output 1035 of "OK, I turned off the kitchen light and turned on the bedroom light" and an audio output 1135 of "OK, I deleted the alarm and created a calendar event." Figure 10CAs shown, the digital assistant 800 causes the electronic device 1000 to display the affordance 1004 in a state that indicates the bedroom light is on. Additionally, the digital assistant 800 causes the electronic device 1000 to provide the audio output 1035 of "I turned off your kitchen light and turned on your bedroom light" to indicate to the user that the previous task has been undone and the task of turning on the bedroom light is complete.

[0296] As another example, as Figure 11D As shown, after deleting the previously created alarm to complete the task 825, the digital assistant 800 causes the electronic device 1100 to display the calendar application with the created calendar event to provide visual output indicating that the task has been completed. The digital assistant 800 can also provide the audio output 1125 of "created a calendar event at 2 PM" to indicate to the user that the task of creating the event has been completed.

[0297] In some examples, after performing the task 825 and / or the task 835, the digital assistant 800 receives the user verbal input 805 and determines that the user verbal input 805 includes a modification to the task 825 and / or 835. Accordingly, the digital assistant 800 can undo the previously performed task and / or perform another task to modify the previously performed task, as discussed above. For example, if the user verbal input 805 includes a request to modify the task 825, the digital assistant 800 can undo the task 825 and perform another task to modify the task 825, as shown. Figure 9C As shown, after recreating the alarm, the digital assistant 800 receives another user verbal input of "well, just delete them again," the digital assistant 800 (e.g., the modification module 810) determines that the user wants to undo the recent task of recreating the alarm and delete them again.

[0298] Accordingly, it will be appreciated that multiple user verbal inputs requesting reversal and / or modification of recently performed tasks can be received by and acted upon by the digital assistant 800. This allows the digital assistant 800 to provide efficient responses to user requests in a continuous manner to demonstrate to the user that they find the most helpful interactions. Accordingly, tasks can be reversed and / or modified as needed by the user without requiring the user to provide input requesting performance of a particular task.

[0299] Figure 12A process 1200 for operating a digital assistant to modify a task after execution in accordance with various examples is illustrated. For example, the process 1200 is performed using one or more electronic devices that implement a digital assistant. In some examples, the process 1200 is performed using a client-server system (e.g., the system 100), and the blocks of the process 1200 are divided between a server (e.g., the DA server 106) and a client device in any manner. In other examples, the blocks of the process 1200 are divided between a server and multiple client devices (e.g., a mobile phone and a smart watch). Thus, while portions of the process 1200 are described herein as being performed by particular devices of a client-server system, it should be understood that the process 1200 is not limited thereto. In other examples, the process 1200 is performed using only a client device (e.g., the user device 104) or only multiple client devices. In the process 1200, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some examples, additional steps can be performed in connection with the process 1200.

[0300] At block 1210, a first task (e.g., the task 815, 825, 835) specified in the first user utterance input (e.g., the user utterance input 805, 905, 915, 1005, 1015, 1105, 1115) is performed (and / or initiated). In some examples, the first user utterance input (e.g., "delete all my alarms," "text my mom," etc.) including the first task is received prior to performing the first task.

[0301] At block 1220, a second user utterance input (e.g., the user utterance input 805, 905, 915, 1005, 1015, 1105, 1115) (e.g., "oh, wait, never mind," "text my dad instead," etc.) is received.

[0302] At block 1230, in accordance with a determination that the second user utterance input (e.g., the user utterance input 805, 905, 915, 1005, 1015, 1105, 1115) includes a modification to the first task (e.g., the task 815, 825, 835), a second task (e.g., the task 815, 825, 835) is performed (and / or initiated), where performance of the second task modifies at least a portion of the performance of the first task (e.g., restoring the user's alarms, modifying the recipient of the text message, etc.).

[0303] In some examples, determining whether the second user speech input (e.g., user speech input 805, 905, 915, 1005, 1015, 1105, 1115) includes an intent to undo an action, where the intent to undo an action corresponds to a modification to the first task (e.g., task 815, 825, 835). In some examples, determining that the second user speech input includes the intent to undo an action further includes determining whether the second user speech input includes a predetermined word or phrase. In some examples, determining that the second user speech input includes the intent to undo an action further includes determining whether the second user speech input was received within a predetermined time of performing the first task. In some examples, determining that the second user speech input includes the intent to undo an action further includes determining whether the first task includes a reversible nature (e.g., turn on / turn off, cancel creation, delete, return to previous state). In some examples, the reversible nature of the first task is included in a domain (e.g., domain 820) associated with the first task. In some examples, a machine learning model (e.g., modification module 810) performs the determination of whether the second user input includes the intent to modify an action.

[0304] In some examples, in accordance with a determination that the second user speech input (e.g., user speech input 805, 905, 915, 1005, 1015, 1105, 1115) includes a modification to the first task (e.g., task 815, 825, 835) and in accordance with a determination that the second user speech input includes a third task (e.g., task 815, 825, 835), the third task is performed after performing the second task (e.g., task 815, 825, 835). In some examples, the third task is related to the second task. In some examples, the third task is related to the first task.

[0305] In some examples, the second task (e.g., task 815, 825, 835) includes undoing (e.g., reversing) performance of the first task (e.g., task 815, 825, 835). In some examples, the second task is based on a nature of a domain that includes the first task. In some examples, the second task includes deleting an object that was created. In some examples, the second task includes passing data from a first application to a second application. In some examples, the second task includes re-creating an object that was previously deleted. In some examples, the second task includes changing a state of the electronic device. In some examples, the second task includes changing a parameter of the first task.

[0306] In some examples, after completing a first task (e.g., tasks 815, 825, 835), a first audio output (e.g., audio outputs 925, 935, 1025, 1035, 1125) is provided that includes a description of the first task. In some examples, after completing a second task (e.g., tasks 815, 825, 835), a second audio output (e.g., audio outputs 925, 935, 1025, 1035, 1125) is provided. In some examples, the second audio output includes an indication that the first task was modified. In some examples, the second audio output includes an indication that a third task (e.g., tasks 815, 825, 835) was performed.

[0307] In some examples, a third user speech input is received (e.g., user speech input 805, 905, 915, 1005, 1015, 1105, 1115). In some examples, based on determining that the third user speech input includes a modification to the second task (e.g., tasks 815, 825, 835), a fourth task (e.g., tasks 815, 825, 835) is performed, where the performance of the fourth task modifies at least a portion of the second task.

[0308] References Figure 12 The described operations are optionally performed by Figures 1-4 、 Figures 6A-6B 、 Figures 7A-7C 、 Figure 8 、 Figures 9A-9C 、 Figures 10A-10C and Figures 11A-11D For example, the operations of process 1200 may be implemented by digital assistant 800 and electronic devices 900, 1000, and / or 1100. It will be clear to one of ordinary skill in the art how to implement the process 1200 based on the components depicted in FIG. Figures 1-4 、 Figures 6A-6B 、 Figures 7A-7C 、 Figure 8 、 Figures 9A-9C 、 Figures 10A-10C and Figures 11A-11D The components depicted in the implementation of other processes.

[0309] According to some specific implementations, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) is provided that stores one or more programs for execution by one or more processors of an electronic device, the one or more programs including instructions for performing any of the methods or processes described herein.

[0310] According to some implementations, an electronic device (eg, a portable electronic device) is provided that includes means for performing any of the methods or processes described herein.

[0311] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes a processing unit configured to perform any of the methods and processes described herein.

[0312] According to some implementations, an electronic device (e.g., a portable electronic device) is provided that includes one or more processors and memory storing one or more programs for execution by the one or more processors, the one or more programs including instructions for performing any of the methods or processes described herein.

[0313] The preceding description is presented to enable any person skilled in the art to practice the applications as described in the preceding specific embodiments. However, the preceding specific embodiments are not intended to be exhaustive or to be

[0314] Although the present disclosure and examples have been fully described with reference to the accompanying figures, it is to be noted that various changes and modifications will become apparent to those skilled in the art. It is therefore intended that the present disclosure and examples be taken as including all such changes and modifications as fall within the scope of the appended claims.

[0315] As described above, one aspect of the present technology is the collection and use of data available from a variety of sources to improve task performance and modification. The present disclosure contemplates that, in some instances, such collected data can include personal information data that uniquely identifies or can be used to contact or locate a specific person. Such personal information data can include demographic data, location-based data, telephone numbers, email addresses, twitter ID's, home addresses, data or records pertaining to a user's health or fitness level, date of birth, or any other identifying or personal information.

[0316] The present disclosure recognizes that the use of such personal information data, in the present technology, can be used to the benefit of users. For example, the personal information data can be used to determine which tasks a user typically modifies. Accordingly, use of such personal information data, enables users to make calculated control over interactions with the digital assistant. Furthermore, other uses for personal information data that benefit the user are also contemplated. For example, health and fitness data can be used to provide insights into a user's general wellness or can be used as positive feedback to individuals using the technology to pursue health goals.

[0317] The present disclosure contemplates that the entities responsible for the collection, analysis, disclosure, transfer, storage, or other use of such personal information data will comply with well-established privacy policies and / or privacy practices. In particular, such entities should implement and consistently use privacy policies and practices that are generally recognized as industry or governmental standards now or at the time of implementation. Such policies should include, for example, the practices set forth at the website of The National Institute of Standards and Technology, from and / or industry-recognized privacy policies and practices such as certified code of conduct from the Electronic Privacy Information Center (EPIC). In addition, such policies should be applicable to all personal information data uses by the entity. Specifically, such policies should be posted publicly and should be available via links at every point of interaction, and the policies applied should encompass all of the entity’s uses of personal information data. Additionally, the disclosure would be expected to apply at least to the extent of applicable law in the places where the entity operates, including any requirements and / or guidance from governments and their agencies available at the time of implementation. Such policies should be easily understood by the users of the entity’s services, consistent with the commercially reasonable practices of the entity. Additionally, policies should be followed by the entity’s employees and / or partners, and such employees and / or partners should be subject to disciplinary action if policies are not followed. Furthermore, entities implementing the present disclosure should commit to regularly reviewing and updating their privacy policies and / or practices to ensure compliance with the standards to which they commit. In addition, the policies should be applied to all of the entity’s services and / or products, including internet and mobile applications.

[0318] Regardless of the foregoing, the present disclosure also contemplates embodiments in which users selectively block the use of, or access to, personal information data. That is, the present disclosure contemplates that hardware and / or software elements can be provided to prevent or block access to such personal information data. For example, in the case of task execution, the present technology can be configured to allow users to select to "opt in" or "opt out" of permitting the collection of personal information data during registration for a service or at any other time. In addition to providing the "opt in" and "opt out" choice, the present disclosure contemplates providing notifications relating to the access or use of personal information. For instance, a user can be notified upon download of an application that their personal information data will be accessed. Additionally, a user can be notified that an application will access personal information data upon the application accessing the personal information data.

[0319] Moreover, it is the intent of the present disclosure that personal information data should be managed and processed in a manner that minimizes risks of unintentional or unauthorized access or use. Risk can be minimized by limiting the collection of personal information data, deleting such data when it is no longer needed, and by appropriately securing such data that is collected. In addition, and when applicable, data de-identification can be used to protect a user’s privacy. To the extent additional privacy protection is needed, hashed pointers can be generated to link together different pieces of information about a user (e.g., hashed pointers can be used to link together in-memory data about a user). De-identification can be facilitated in appropriate cases by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or

[0320] Thus, while the present disclosure broadly covers uses of personal information data to implement one or more various disclosed embodiments, the present disclosure also contemplates that the various embodiments can also be implemented without the need for accessing such personal information data. That is, the various embodiments of the present technology are not rendered inoperable due to the lack of access to such personal information data. For example, a task modification can be selected and delivered to a user based on non-personal information data or a minimum amount of personal information, such as content requested by a device associated with a user, other non-personal information available to a digital assistant, or information publicly available.

Claims

1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device, the one or more programs including instructions for: executing a first task specified in a first user speech input; receiving a second user speech input; and Based on determining that the second user speech input includes an intention to undo an action performed by the first task, performing a second task, wherein the execution of the second task returns the electronic device performing the first task to a state before the execution of the first task by recreating previously deleted objects.

2. The non-transitory computer-readable storage medium of claim 1 , wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the second user speech input includes a predetermined word or phrase.

3. The non-transitory computer-readable storage medium of claim 1 , wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the second user speech input is received within a predetermined time of performing the first task.

4. The non-transitory computer-readable storage medium of claim 1 , wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the first task includes a reversible property. 5 . The non-transitory computer-readable storage medium of claim 4 , wherein the reversible property of the first task is included in a domain associated with the first task.

6. The non-transitory computer-readable storage medium of claim 1 , wherein a machine learning model performs a determination of whether the second user speech input includes the intent to undo the action.

7. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions for: Based on determining that the second user speech input includes a modification to the first task: Based on determining that the second user speech input includes a third task, the third task is performed after performing the second task.

8. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions for: After completing the first task, a first audio output is provided that includes a description of the first task.

9. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions for: After completing the second task, a second audio output is provided.

10. The non-transitory computer-readable storage medium of claim 9, wherein the second audio output includes an indication that the first task is modified.

11. The non-transitory computer-readable storage medium of claim 9, wherein the second audio output includes an indication that a third task was performed.

12. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions for: receiving a third user speech input; and Based on determining that the third user speech input includes a modification to the second task, a fourth task is performed, wherein the performance of the fourth task modifies at least a portion of the performance of the second task.

13. The non-transitory computer-readable storage medium of claim 1, wherein the second task comprises transferring data from a first application to a second application.

14. The non-transitory computer-readable storage medium of claim 1, wherein the second task comprises changing a state of an electronic device.

15. An electronic device comprising: one or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: executing a first task specified in a first user speech input; receiving a second user speech input; and Based on determining that the second user speech input includes an intention to undo an action performed by the first task, performing a second task, wherein the execution of the second task returns the electronic device performing the first task to a state before the execution of the first task by recreating previously deleted objects.

16. A method comprising: At an electronic device having one or more processors and a memory: executing a first task specified in a first user speech input; receiving a second user speech input; as well as Based on determining that the second user speech input includes an intention to undo an action performed by the first task, performing a second task, wherein the execution of the second task returns the electronic device performing the first task to a state before the execution of the first task by recreating previously deleted objects.

17. The non-transitory computer-readable storage medium of claim 1, wherein determining that the second user speech input includes an intention to undo the action comprises: A determination is made as to whether the task of deleting the object includes a reversible property of recreating the deleted object.

18. The non-transitory computer-readable storage medium of claim 1, wherein the one or more programs further comprise instructions for: After performing the first task specified in the first user speech input, providing a possible task that reverses the first task specified in the first user speech input; and in, Recreating the previously deleted object includes performing the potential task that reverses the first task specified in the first user speech input.

19. The non-transitory computer-readable storage medium of claim 1 , wherein recreating the previously deleted object comprises: determining that the intent to undo the action performed by the first task corresponds to the second task to recreate the previously deleted object; as well as The second task is based on properties of a related domain, the related domain providing information linking the first task with the second task.

20. The non-transitory computer-readable storage medium of claim 1, wherein the one or more programs further comprise instructions for: The recreated object is displayed as an indication of completion of the second task.

21. The non-transitory computer-readable storage medium of claim 1, wherein the second task is provided by a domain associated with the first task.

22. The non-transitory computer-readable storage medium of claim 21, wherein the second task is linked to the first task in the domain.

23. The non-transitory computer-readable storage medium of claim 22, wherein the one or more programs further comprise instructions for: A goal of the second task is determined based on the link between the second task and the first task in the domain.

24. The electronic device of claim 15, wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the second user speech input includes a predetermined word or phrase.

25. The electronic device of claim 15, wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the second user speech input is received within a predetermined time of performing the first task.

26. The electronic device of claim 15, wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the first task includes a reversible property.

27. The electronic device of claim 26, wherein the reversible property of the first task is included in a domain associated with the first task.

28. The electronic device of claim 15, wherein a machine learning model performs a determination of whether the second user speech input includes the intent to undo the action.

29. The electronic device of claim 15, wherein the one or more programs further comprise instructions for: Based on determining that the second user speech input includes a modification to the first task: Based on determining that the second user speech input includes a third task, the third task is performed after performing the second task.

30. The electronic device of claim 15, wherein the one or more programs further comprise instructions for: After completing the first task, a first audio output is provided that includes a description of the first task.

31. The electronic device of claim 15, wherein the one or more programs further comprise instructions for: After completing the second task, a second audio output is provided.

32. The electronic device of claim 31, wherein the second audio output comprises an indication that the first task has been modified.

33. The electronic device of claim 31 , wherein the second audio output comprises an indication that a third task was performed.

34. The electronic device of claim 15, wherein the one or more programs further comprise instructions for: receiving a third user speech input; and Based on determining that the third user speech input includes a modification to the second task, a fourth task is performed, wherein the performance of the fourth task modifies at least a portion of the performance of the second task.

35. The electronic device of claim 15, wherein the second task comprises transferring data from a first application to a second application.

36. The electronic device of claim 15, wherein the second task comprises changing a state of the electronic device.

37. The electronic device of claim 15, wherein determining that the second user speech input includes an intention to undo the action comprises: A determination is made as to whether the task of deleting the object includes a reversible property of recreating the deleted object.

38. The electronic device of claim 15, wherein the one or more programs further comprise instructions for: After performing the first task specified in the first user speech input, providing a possible task that reverses the first task specified in the first user speech input; and in, Recreating the previously deleted object includes performing the potential task that reverses the first task specified in the first user speech input.

39. The electronic device of claim 15, wherein recreating the previously deleted object comprises: determining that the intent to undo the action performed by the first task corresponds to the second task to recreate the previously deleted object; as well as The second task is based on properties of a related domain, the related domain providing information linking the first task with the second task.

40. The electronic device of claim 15, wherein the one or more programs further comprise instructions for: The recreated object is displayed as an indication of completion of the second task.

41. The electronic device of claim 15, wherein the second task is provided by a domain associated with the first task.

42. The electronic device of claim 41, wherein the second task is linked to the first task in the domain.

43. The electronic device of claim 42, wherein the one or more programs further comprise instructions for: A goal of the second task is determined based on the link between the second task and the first task in the domain.

44. The method of claim 16, wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the second user speech input includes a predetermined word or phrase.

45. The method of claim 16, wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the second user speech input is received within a predetermined time of performing the first task.

46. ​​The method of claim 16, wherein determining that the second user speech input includes the intention to undo the action further comprises: A determination is made as to whether the first task includes a reversible property.

47. The method of claim 46, wherein the reversible property of the first task is included in a domain associated with the first task.

48. The method of claim 16, wherein a machine learning model performs a determination as to whether the second user speech input includes the intent to undo the action.

49. The method of claim 16, further comprising: Based on determining that the second user speech input includes a modification to the first task: Based on determining that the second user speech input includes a third task, the third task is performed after performing the second task.

50. The method of claim 16, further comprising: After completing the first task, a first audio output is provided that includes a description of the first task.

51. The method of claim 16, further comprising: After completing the second task, a second audio output is provided.

52. The method of claim 51, wherein the second audio output includes an indication that the first task has been modified.

53. The method of claim 51, wherein the second audio output includes an indication that a third task was performed.

54. The method of claim 16, further comprising: receiving a third user speech input; as well as Based on determining that the third user speech input includes a modification to the second task, a fourth task is performed, wherein the performance of the fourth task modifies at least a portion of the performance of the second task.

55. The method of claim 16, wherein the second task comprises transferring data from a first application to a second application.

56. The method of claim 16, wherein the second task comprises changing a state of an electronic device.

57. The method of claim 16, wherein determining that the second user speech input includes an intention to undo the action comprises: A determination is made as to whether the task of deleting the object includes a reversible property of recreating the deleted object.

58. The method of claim 16, further comprising: After performing the first task specified in the first user speech input, providing a possible task of reversing the first task specified in the first user speech input; and Wherein recreating the previously deleted object comprises performing the possible task that reverses the first task specified in the first user speech input.

59. The method of claim 16, wherein recreating the previously deleted object comprises: determining that the intent to undo the action performed by the first task corresponds to the second task to recreate the previously deleted object; as well as The second task is based on properties of a related domain, the related domain providing information linking the first task with the second task.

60. The method of claim 16, further comprising: The recreated object is displayed as an indication of completion of the second task.

61. The method of claim 16, wherein the second task is provided by a domain associated with the first task.

62. The method of claim 61, wherein the second task is linked to the first task in the domain.

63. The method of claim 62, further comprising: A goal of the second task is determined based on the link between the second task and the first task in the domain.

Citation Information

Patent Citations

  • System and method for inferring user intent from speech inputs

    US10176167B2

  • Method and apparatus for integrating manual input

    US20020015024A1

  • Acceleration-based theft detection system for portable electronic devices

    US20050190059A1

  • Methods and apparatuses for operating a portable device based on an accelerometer

    US20060017692A1

  • Gestures for touch sensitive input devices

    US20060026521A1