Digital assistant user interface and response patterns

By displaying digital assistant indicators and response capabilities in different parts of the display within the intelligent automation assistant user interface, while keeping the underlying interface visible, the problem of user interface interference is solved, and the efficiency of interaction and the effectiveness of device use are improved.

CN113703656BActive Publication Date: 2026-01-13APPLE INC
View PDF 26 Cites 0 Cited by

Patent Information

Application Number
CN202010977583.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-24
Filing Date
2020-09-16
Publication Date
2026-01-13
Estimated Expiration
2040-09-16

AI Technical Summary

Technical Problem

Existing intelligent automation assistant user interfaces may obscure other display elements that users are interested in and provide unexpected response formats, resulting in visual clutter and inefficient interaction.

Method used

When displaying a digital assistant user interface on a monitor, the user can interact without interfering with the underlying interface display by showing digital assistant indicators and responsive capabilities in different parts of the monitor while keeping a portion of the underlying user interface visible.

Benefits of technology

It improves the effectiveness and efficiency of digital assistants, reduces visual distractions, saves power consumption, and extends device battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113703656B_ABST
    Figure CN113703656B_ABST
Patent Text Reader

Abstract

This disclosure relates to digital assistant user interfaces and response modes. An example process includes: while displaying a user interface that is different from a digital assistant user interface, receiving user input; in accordance with a determination that the user input satisfies criteria for initiating a digital assistant: displaying the digital assistant user interface over the user interface, the digital assistant user interface including: a digital assistant indicator, the digital assistant indicator displayed at a first portion of a display; and a response affordance, the response affordance displayed at a second portion of the display; wherein: a portion of the user interface remains visible at a third portion of the display; and the third portion is located between the first portion and the second portion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates in general to intelligent automation assistants, and more specifically to user interfaces for intelligent automation assistants and the ways in which intelligent automation assistants can respond to user requests. Background Technology

[0002] Intelligent automated assistants (or digital assistants) provide a beneficial interface between human users and electronic devices. Such assistants allow users to interact with devices or systems using natural language in voice and / or text. For example, a user can provide voice input containing their request to a digital assistant running on an electronic device. The digital assistant can interpret the user's intent from this voice input and act it out as a task. These tasks can then be performed by executing one or more services of the electronic device, and relevant output in response to the user's request can be returned to the user.

[0003] The displayed user interface of a digital assistant may sometimes obscure other displayed elements that the user might be interested in. Furthermore, the digital assistant may sometimes provide responses in a format not expected by the user's current context. For example, the digital assistant may provide displayed output when the user does not expect (or cannot) view the device's display. Summary of the Invention

[0004] This document discloses an exemplary method. An exemplary method includes, at an electronic device having a display and a touch-sensitive surface: receiving user input while displaying a user interface different from a digital assistant user interface; and, based on determining that the user input meets criteria for initiating a digital assistant: displaying the digital assistant user interface on top of the user interface, the digital assistant user interface including: a digital assistant indicator displayed on a first portion of the display; and a responsive power indicator displayed on a second portion of the display; wherein: a portion of the user interface remains visible on a third portion of the display; and the third portion is located between the first portion and the second portion.

[0005] This document discloses an exemplary non-transitory computer-readable medium. An example non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device having a display and a touch-sensitive surface, cause the electronic device to: receive user input while displaying a user interface different from a digital assistant user interface; and, based on determining that the user input meets criteria for initiating a digital assistant: display the digital assistant user interface on top of the user interface, the digital assistant user interface including: a digital assistant indicator displayed at a first portion of the display; and a response enable indicator displayed at a second portion of the display; wherein: a portion of the user interface remains visible at a third portion of the display; and the third portion is located between the first portion and the second portion.

[0006] This document discloses an exemplary electronic device. An exemplary electronic device includes a display; a touch-sensitive surface; one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: receiving user input while displaying a user interface different from a digital assistant user interface; and, based on determining that the user input meets criteria for initiating a digital assistant: displaying the digital assistant user interface on top of the user interface, the digital assistant user interface including: a digital assistant indicator displayed at a first portion of the display; and a responsive enable indicator displayed at a second portion of the display; wherein: a portion of the user interface remains visible at a third portion of the display; and the third portion is located between the first portion and the second portion.

[0007] An exemplary electronic device includes means for performing the following operations: receiving user input while displaying a user interface different from a digital assistant user interface; and, based on determining that the user input meets criteria for initiating a digital assistant: displaying the digital assistant user interface on top of the user interface, the digital assistant user interface including: a digital assistant indicator displayed on a first portion of the display; and a response enable indicator displayed on a second portion of the display; wherein: a portion of the user interface remains visible on a third portion of the display; and the third portion is located between the first portion and the second portion.

[0008] Displaying a digital assistant user interface (where a portion of the user interface remains visible on a portion of the display) on top of the user interface improves the effectiveness of the digital assistant and reduces its visual interference with user-device interaction. For example, including information within the underlying visible user interface allows the user to more clearly define their requests to the digital assistant. Similarly, displaying the user interface in this manner facilitates interaction between elements of the digital assistant user interface and the underlying user interface (e.g., digital assistant responses within messages in the underlying messaging user interface). Furthermore, having the digital assistant user interface and the underlying user interface coexist on the display allows for simultaneous user interaction with both user interfaces, thus better integrating the digital assistant into the user-device interaction. This makes the user-device interface more efficient (e.g., by enabling the digital assistant to perform user-requested tasks more accurately and effectively, by reducing the visual interference of the digital assistant with what the user is viewing, and by reducing the amount of user input required to operate the device as needed), which in turn reduces power consumption and extends the device's battery life by allowing the user to use the device more quickly and efficiently.

[0009] This document discloses an exemplary method. An exemplary method includes, at an electronic device having a display and a touch-sensitive surface: displaying a digital assistant user interface (DUI) over a user interface, the DUI including: a DUI indicator displayed on a first portion of the display; and a responsive power indicator displayed on a second portion of the display; receiving user input corresponding to a selection of a third portion of the display, the third portion displaying a portion of the DUI, while the DUI is displayed over the user interface; stopping the display of the DUI and the responsive power indicator based on determining that the user input corresponds to a first type of input; and updating the display of the DUI at the third portion based on the user input while the responsive power indicator is displayed at the second portion, based on determining that the user input corresponds to a second type of input different from the first type of input; and updating the display of the DUI at the third portion based on the user input while the responsive power indicator is displayed at the second portion.

[0010] This document discloses an exemplary non-transitory computer-readable medium. An example non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device having a display and a touch-sensitive surface, cause the electronic device to: display a digital assistant user interface over a user interface, the digital assistant user interface including: a digital assistant indicator displayed at a first portion of the display; and a response enable indicator displayed at a second portion of the display; while the digital assistant user interface is displayed over the user interface, receive user input corresponding to a selection of a third portion of the display, the third portion displaying a portion of the user interface; based on determining that the user input corresponds to a first type of input: stop displaying the digital assistant indicator and the response enable indicator; and based on determining that the user input corresponds to a second type of input different from the first type of input: while displaying the response enable indicator at the second portion, update the display of the user interface at the third portion according to the user input.

[0011] This document discloses an exemplary electronic device. An exemplary electronic device includes a display; a touch-sensitive surface; one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for: displaying a digital assistant user interface on top of a user interface, the digital assistant user interface including: a digital assistant indicator displayed at a first portion of the display; and a response enable indicator displayed at a second portion of the display; receiving user input corresponding to a selection of a third portion of the display, the third portion displaying a portion of the user interface, when the digital assistant user interface is displayed on top of the user interface; stopping the display of the digital assistant indicator and the response enable indicator based on determining that the user input corresponds to a first type of input; and updating the display of the user interface at the third portion based on the user input when the response enable indicator is displayed at the second portion, based on determining that the user input corresponds to a second type of input different from the first type of input; and updating the display of the user interface at the third portion based on the user input when the response enable indicator is displayed at the second portion.

[0012] An exemplary electronic device includes means for operating as follows: displaying a digital assistant user interface on top of a user interface, the digital assistant user interface including: a digital assistant indicator displayed on a first portion of the display; and a response capability indicator displayed on a second portion of the display; receiving user input corresponding to a selection of a third portion of the display, the third portion displaying a portion of the user interface, when the digital assistant user interface is displayed on top of the user interface; stopping the display of the digital assistant indicator and the response capability indicator based on determining that the user input corresponds to a first type of input; and updating the display of the user interface at the third portion based on the user input when the response capability indicator is displayed at the second portion, based on determining that the user input corresponds to a second type of input different from the first type of input; and updating the display of the user interface at the third portion based on the user input when the response capability indicator is displayed at the second portion, based on determining that the user input corresponds to a second type of input different from the first type of input.

[0013] Stopping the display of the digital assistant indicator and the response power indicator based on determining that the user input corresponds to the first type of input provides an intuitive and effective way to eliminate the digital assistant. For example, the user can simply provide input to select the underlying user interface to eliminate the digital assistant user interface, thereby reducing the digital assistant's interference with user-device interaction. Updating the display of the user interface in the third part based on the user input when the response power indicator is displayed in the second part provides an intuitive way for the digital assistant user interface and the underlying user interface to coexist. For example, the user can provide input to select the underlying user interface to update the underlying user interface as if the digital assistant user interface were not displayed. Furthermore, retaining the digital assistant user interface (which may include information of interest to the user) while allowing the user to interact with the underlying user interface reduces the digital assistant's interference with the underlying user interface. In this way, the user-device interface can be more efficient (e.g., by allowing the user to interact with the underlying user interface when the digital assistant user interface is displayed, by reducing the visual interference of the digital assistant with the content the user is viewing, and by reducing the amount of user input required to operate the device as needed), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.

[0014] This document discloses an exemplary method. An exemplary method includes, at an electronic device having one or more processors, memory, and a display: receiving natural language input; initiating the digital assistant; obtaining a response packet in response to the natural language input based on initiating the digital assistant; after receiving the natural language input, selecting a first response mode of the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device; and, in response to selecting the first response mode, having the digital assistant present the response packet according to the first response mode.

[0015] This document discloses an exemplary non-transitory computer-readable medium. An example non-transitory computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by one or more processors of an electronic device having a display, cause the electronic device to: receive natural language input; initiate a digital assistant; obtain a response packet in response to the natural language input based on initiating the digital assistant; after receiving the natural language input, select a first response mode of the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device; and, in response to selecting the first response mode, have the digital assistant present the response packet according to the first response mode.

[0016] This document discloses an exemplary electronic device. An exemplary electronic device includes a display; one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: receiving natural language input; initiating the digital assistant; obtaining a response packet in response to the natural language input based on initiating the digital assistant; after receiving the natural language input, selecting a first response mode of the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device; and, in response to selecting the first response mode, having the digital assistant present the response packet according to the first response mode.

[0017] An exemplary electronic device includes means for performing the following operations: receiving natural language input; initiating the digital assistant; obtaining a response packet in response to the natural language input based on initiating the digital assistant; after receiving the natural language input, selecting a first response mode of the digital assistant from a plurality of digital assistant response modes based on context information associated with the electronic device; and in response to selecting the first response mode, having the digital assistant present the response packet according to the first response mode.

[0018] Presenting the response packet by the digital assistant according to the first response mode allows the digital assistant's response to be presented in an informational manner appropriate to the user's current context. For example, when the user's current context indicates that visual user-device interaction is not desired (or impossible), the digital assistant may present the response in audio format. Similarly, when the user's current context indicates that auditory user-device interaction is not desired, the digital assistant may present the response in visual format. Furthermore, when the user's current context indicates that both auditory and visual user-device interaction are desired, the digital assistant may present a response with visual components and concise audio components, thereby reducing the length of the digital assistant's audio output. Moreover, selecting the first response mode after receiving the natural language input (and before presenting the response packet) allows for a more accurate determination of the user's current context (and therefore a more accurate determination of the appropriate response mode). In this way, the user-device interface can be more efficient and secure (e.g., by reducing visual distractions from the digital assistant, by presenting responses effectively in an informative manner, and by intelligently adjusting responses based on the user's current context), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently. Attached Figure Description

[0019] Figure 1 Block diagrams are shown for systems and environments used to implement digital assistants, based on various examples.

[0020] Figure 2A A block diagram is shown for a portable multi-functional device that implements the client-side portion of a digital assistant according to various examples.

[0021] Figure 2B A block diagram illustrating exemplary components for event handling, based on various examples.

[0022] Figure 3 Portable multi-functional devices are shown that implement the client-side portion of a digital assistant according to various examples.

[0023] Figure 4 A block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface, according to various examples.

[0024] Figure 5A An exemplary user interface for the menu of an application on a portable multi-functional device, based on various examples, is shown.

[0025] Figure 5B Exemplary user interfaces of multifunctional devices with touch-sensitive surfaces separate from the display are shown according to various examples.

[0026] Figure 6AThe images show personal electronic devices based on various examples.

[0027] Figure 6B A block diagram illustrating a personal electronic device based on various examples is provided.

[0028] Figure 7A A block diagram illustrating a digital assistant system or its server portion, based on various examples, is provided.

[0029] Figure 7B Examples are shown in Figure 7A The digital assistant functions shown.

[0030] Figure 7C A portion of the knowledge ontology is shown based on various examples.

[0031] Figures 8A to 8CT The user interface and digital assistant user interface are shown according to various examples.

[0032] Figures 9A to 9C Multiple devices are shown, based on various examples, for determining which device should respond to voice input.

[0033] Figures 10A to 10V The user interface and digital assistant user interface are shown according to various examples.

[0034] Figure 11 The diagram illustrates various examples of systems for selecting a digital assistant response mode and for presenting a response based on the selected digital assistant response mode.

[0035] Figure 12 The diagram illustrates devices that present responses to received natural language input according to different digital assistant response modes, based on various examples.

[0036] Figure 13 Exemplary processes for selecting a digital assistant response mode are shown, based on various examples.

[0037] Figure 14 The diagram illustrates devices that, based on various examples, present responses according to voice response patterns when it is determined that a user is in a vehicle (e.g., driving).

[0038] Figure 15 The diagram illustrates devices that present responses based on voice response patterns when the device is running a navigation application, according to various examples.

[0039] Figure 16 The response pattern changes throughout the multi-turn DA interaction process are shown based on various examples.

[0040] Figures 17A to 17FThe process for operating a digital assistant is illustrated with various examples.

[0041] Figures 18A to 18B The process for operating a digital assistant is illustrated with various examples.

[0042] Figures 19A to 19E The process for selecting a digital assistant response mode is illustrated with various examples. Detailed Implementation

[0043] The accompanying drawings will be referenced in the following description of the examples, which illustrate specific examples that can be implemented by way of example. It should be understood that other examples may be used and structural changes may be made without departing from the scope of the individual examples.

[0044] Although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various examples described, a first input may be referred to as a second input, and similarly, a second input may be referred to as a first input. Both the first and second inputs are inputs, and in some cases, they are independent and distinct inputs.

[0045] The terminology used in the description of the various examples herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “includes”, “including”, “comprises”, and / or “comprising”, when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0046] Depending on the context, the term "if" can be interpreted as meaning "when" or "upon" or "in response to determination" or "in response to detection." Similarly, depending on the context, the phrases "if determination..." or "if detection [the stated condition or event]" can be interpreted as meaning "when determination..." or "in response to determination..." or "when detection [the stated condition or event]" or "in response to detection [the stated condition or event]."

[0047] 1. System and Environment

[0048] Figure 1 A block diagram of system 100 according to various examples is shown. In some examples, system 100 implements a digital assistant. The terms "digital assistant," "virtual assistant," "intelligent automated assistant," or "automated digital assistant" refer to any information processing system that interprets natural language input in spoken and / or textual form to infer user intent and performs actions based on the inferred user intent. For example, to act on an inferred user intent, the system performs one or more of the following steps: identifying a task flow having steps and parameters designed to achieve the inferred user intent; inputting a specific request into the task flow based on the inferred user intent; executing the task flow by invoking programs, methods, services, APIs, etc.; and generating an output response to the user in an audible (e.g., voice) and / or visual form.

[0049] Specifically, a digital assistant can accept user requests, at least in part, in the form of natural language commands, requests, statements, narration, and / or inquiries. Typically, user requests seek an informational response or task from the digital assistant. A satisfactory response to a user request includes providing the requested informational response, performing the requested task, or a combination of both. For example, a user asks a digital assistant a question such as, “Where am I now?” Based on the user’s current location, the digital assistant answers, “You are near the west entrance of Central Park.” The user also requests a task, such as, “Please invite my friends to my girlfriend’s birthday party next week.” In response, the digital assistant confirms the request by saying “Okay, coming right away,” and then sends the appropriate calendar invitations to each of the user’s friends listed in the user’s electronic address book. During the performance of the requested task, the digital assistant sometimes interacts with the user in a sustained dialogue involving multiple exchanges of information over extended periods. Many other methods exist for interacting with a digital assistant to request information or perform various tasks. In addition to providing verbal responses and taking programmed actions, digital assistants also provide responses in other forms of video or audio, such as text, alerts, music, video, animation, etc.

[0050] like Figure 1As shown, in some examples, the digital assistant is implemented according to a client-server model. The digital assistant includes a client-side portion 102 (hereinafter referred to as "DA client 102") executing on user device 104 and a server-side portion 106 (hereinafter referred to as "DA server 106") executing on server system 108. DA client 102 communicates with DA server 106 via one or more networks 110. DA client 102 provides client-side functionality, such as user-oriented input and output processing, and communication with DA server 106. DA server 106 provides server-side functionality for any number of DA clients 102, each residing on a corresponding user device 104.

[0051] In some examples, DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, and an I / O interface 118 to external services. The client-facing I / O interface 112 facilitates client-facing input and output processing of DA server 106. One or more processing modules 114 utilize data and models 116 to process voice input and determine user intent based on natural language input. Furthermore, one or more processing modules 114 perform task execution based on the inferred user intent. In some examples, DA server 106 communicates with external services 120 via one or more networks 110 to complete tasks or collect information. The I / O interface 118 to external services facilitates such communication.

[0052] User equipment 104 can be any suitable electronic device. In some examples, user equipment 104 is a portable multi-functional device (e.g., see reference below). Figure 2A The aforementioned device 200), multifunctional device (for example, see below for reference) Figure 4 The device 400) or personal electronic device (e.g., referred to below) Figures 6A to 6B The device 600 is described above. A portable multifunction device is, for example, a mobile phone that also includes other functions such as a PDA and / or music player. Specific examples of portable multifunction devices include Apple Inc. (Cupertino, California). iPod and Devices. Other examples of portable multifunction devices include, but are not limited to, earbuds / headphones, speakers, and laptops or tablets. Additionally, in some examples, user device 104 is a non-portable multifunction device. Specifically, user device 104 is a desktop computer, game console, speaker, television, or set-top box. In some examples, user device 104 includes a touch-sensitive surface (e.g., a touchscreen display and / or touchpad). Furthermore, user device 104 optionally includes one or more other physical user interface devices, such as a physical keyboard, mouse, and / or joystick. Various examples of electronic devices such as multifunction devices are described in more detail below.

[0053] Examples of one or more communication networks 110 include local area networks (LANs) and wide area networks (WANs), such as the Internet. One or more communication networks 110 are implemented using any known network protocol, including various wired or wireless protocols such as Ethernet, Universal Serial Bus (USB), FireWire, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi, Internet Protocol Voice (VoIP), Wi-MAX, or any other suitable communication protocol.

[0054] Server system 108 is implemented on one or more stand-alone data processing devices or a distributed computer network. In some examples, server system 108 also utilizes various virtual devices and / or services from third-party service providers (e.g., third-party cloud service providers) to provide potential computing and / or infrastructure resources for server system 108.

[0055] In some examples, user equipment 104 communicates with DA server 106 via a second user equipment 122. The second user equipment 122 is similar to or identical to user equipment 104. For example, the second user equipment 122 is similar to the one described below. Figure 2A , Figure 4 and Figures 6A to 6BThe devices 200, 400, or 600 are described above. User equipment 104 is configured to be communicatively coupled to a second user equipment 122 via a direct communication connection (such as Bluetooth, NFC, BTLE, etc.) or via a wired or wireless network (such as a local Wi-Fi network). In some examples, the second user equipment 122 is configured to act as a proxy between user equipment 104 and DA server 106. For example, a DA client 102 of user equipment 104 is configured to transmit information (e.g., a user request received at user equipment 104) to DA server 106 via the second user equipment 122. DA server 106 processes this information and returns relevant data (e.g., data content in response to the user request) to user equipment 104 via the second user equipment 122.

[0056] In some examples, user equipment 104 is configured to send a shortened request for data to a second user equipment 122 to reduce the amount of information transmitted from user equipment 104. The second user equipment 122 is configured to determine supplementary information to be added to the shortened request to generate a complete request to be transmitted to DA server 106. This system architecture can advantageously allow user equipment 104 (e.g., a watch or similar compact electronic device) with limited communication capabilities and / or limited battery power (e.g., a second user equipment 122 with strong communication capabilities and / or battery power, such as a mobile phone, laptop computer, tablet computer, etc.) acting as a proxy to DA server 106 to access the services provided by DA server 106. Although Figure 1 Only two user devices, 104 and 122, are shown in this document, but it should be understood that in some examples, system 100 may include any number and type of user devices configured in this agent configuration to communicate with DA server system 106.

[0057] Although Figure 1 The digital assistant shown includes both a client-side component (e.g., DA client 102) and a server-side component (e.g., DA server 106), but in some examples, the digital assistant's functionality is implemented as a standalone application installed on the user's device. Furthermore, the functional division between the client and server components of the digital assistant can vary in different implementations. For example, in some examples, the DA client is a thin client that only provides user-facing input and output processing functions and delegates all other functions of the digital assistant to the backend server.

[0058] 2. Electronic equipment

[0059] Now let’s turn our attention to the implementation of electronic devices for the client-side portion of a digital assistant. Figure 2AThis is a block diagram illustrating a portable multi-functional device 200 with a touch-sensitive display system 212 according to some embodiments. The touch-sensitive display 212 is sometimes referred to as a “touchscreen” for convenience, and is sometimes referred to as or called a “touch-sensitive display system.” Device 200 includes a memory 202 (which optionally includes one or more computer-readable storage media), a memory controller 222, one or more processing units (CPUs) 220, a peripheral interface 218, RF circuitry 208, audio circuitry 210, a speaker 211, a microphone 213, an input / output (I / O) subsystem 206, other input control devices 216, and an external port 224. Device 200 optionally includes one or more optical sensors 264. Device 200 optionally includes one or more contact strength sensors 265 for detecting the intensity of contact on the device 200 (e.g., a touch-sensitive surface of the device 200 such as the touch-sensitive display system 212). Device 200 optionally includes one or more haptic output generators 267 for generating haptic outputs on device 200 (e.g., generating haptic outputs on a touch-sensitive surface such as the touch-sensitive display system 212 of device 200 or the touchpad 455 of device 400). These components optionally communicate via one or more communication buses or signal lines 203.

[0060] As used in this specification and claims, the term "intensity" of contact on a tactile surface refers to the force or pressure (force per unit area) of a contact (e.g., finger contact) on a tactile surface, or to a substitute (alternative) for the force or pressure of a contact on a tactile surface. The intensity of contact has a range of values ​​that includes at least four different values ​​and more typically hundreds of different values ​​(e.g., at least 256). The intensity of contact is optionally determined (or measured) using various methods and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the tactile surface are optionally used to measure the force at different points on the tactile surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted average) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the tactile surface. Alternatively, the size and / or variation of the contact area detected on the touch-sensitive surface, the capacitance and / or variation of the touch-sensitive surface near the contact, and / or the resistance and / or variation of the touch-sensitive surface near the contact may optionally be used as substitutes for the force or pressure of the contact on the touch-sensitive surface. In some embodiments, the substitute measurement of the contact force or pressure is used directly to determine whether an intensity threshold (e.g., the intensity threshold is described in units corresponding to the substitute measurement) has been exceeded. In some embodiments, the substitute measurement of the contact force or pressure is converted into an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold (e.g., the intensity threshold is a pressure threshold measured in units of pressure) has been exceeded. Using the intensity of the contact as an attribute of user input allows the user to access additional device functions that would otherwise be inaccessible to the user on smaller devices with limited physical space, such smaller devices being used (e.g., on a touch-sensitive display) to display power indications and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls, such as knobs or buttons).

[0061] As used in this specification and claims, the term "haptic output" refers to a physical displacement of the device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., the housing), or a displacement of a component relative to the center of mass of the device, which is detected by the user using the user's tactile sense. For example, when the device or a component of the device comes into contact with a touch-sensitive surface (e.g., a finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or touchpad) may optionally be interpreted by the user as a "press-click" or "release-click" on a physically actuated button. In some cases, the user will feel a tactile sensation, such as a "press-click" or "release-click," even when a physically actuated button associated with a touch-sensitive surface that has been physically pressed (e.g., displaced) by the user's movement does not move. For example, even when the smoothness of the tactile surface remains unchanged, the movement of the tactile surface can optionally be interpreted or sensed by the user as the "roughness" of the tactile surface. While such interpretations of touch by users will be limited by the individualized sensory perceptions of the user, many sensory perceptions of touch are common to most users. Therefore, when a tactile output is described as corresponding to a specific sensory perception of a user (e.g., "press click", "release click", "roughness"), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or its components that will generate the sensory perception of a typical (or ordinary) user.

[0062] It should be understood that device 200 is merely an example of a portable multifunctional device, and device 200 may optionally have more or fewer components than shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of these components. Figure 2A The various components shown are implemented in hardware, software, or a combination of both, including one or more signal processing and / or application-specific integrated circuits.

[0063] Memory 202 includes one or more computer-readable storage media. These computer-readable storage media are, for example, tangible and non-transitory. Memory 202 includes high-speed random access memory and also includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 222 controls other components of device 200 to access memory 202.

[0064] In some examples, the non-transitory computer-readable storage medium of memory 202 is used to store instructions (e.g., aspects of the processes described below) for use by or in conjunction with an instruction execution system, apparatus, or device, such as a computer-based system, a processor-integrated system, or other system from which instructions can be fetched and executed. In other examples, instructions (e.g., aspects of the processes described below) are stored on a non-transitory computer-readable storage medium (not shown) of server system 108, or partitioned between the non-transitory computer-readable storage medium of memory 202 and the non-transitory computer-readable storage medium of server system 108.

[0065] Peripheral interface 218 is used to couple the input and output peripherals of the device to CPU 220 and memory 202. One or more processors 220 run or execute various software programs and / or instruction sets stored in memory 202 to perform various functions of device 200 and process data. In some embodiments, peripheral interface 218, CPU 220, and memory controller 222 are implemented on a single chip, such as chip 204. In some other embodiments, they are implemented on separate chips.

[0066] RF (Radio Frequency) circuit 208 receives and transmits RF signals, also known as electromagnetic signals. RF circuit 208 converts electrical signals into electromagnetic signals and vice versa, and communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 208 optionally includes well-known circuitry for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 208 optionally communicates wirelessly with networks and other devices, such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs)). RF circuit 208 optionally includes well-known circuitry for detecting near-field communication (NFC) fields, such as via near-field communication radio components. Wireless communication may optionally employ any of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed ​​Downlink Packet Access (HSDPA), High-Speed ​​Uplink Packet Access (HSUPA), Evolution, Pure Data (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), and Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE...). 802.11n and / or IEEE 802.11ac), Internet Protocol Voice (VoIP), Wi-MAX, email protocols (e.g., Internet Messaging Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Utilizing Extended Protocol (SIMPLE), Instant Messaging and Presence Service (IMPS)) and / or Short Message Service (SMS), or any other suitable communication protocol, including communication protocols that have not yet been developed as of the date of this document submission.

[0067] Audio circuitry 210, speaker 211, and microphone 213 provide an audio interface between the user and device 200. Audio circuitry 210 receives audio data from peripheral interface 218, converts the audio data into electrical signals, and transmits the electrical signals to speaker 211. Speaker 211 converts the electrical signals into sound waves that are audible to humans. Audio circuitry 210 also receives electrical signals converted from sound waves by microphone 213. Audio circuitry 210 converts the electrical signals into audio data and transmits the audio data to peripheral interface 218 for processing. Audio data is retrieved from and / or transmitted to memory 202 and / or RF circuitry 208 via peripheral interface 218. In some embodiments, audio circuitry 210 also includes a headset jack (e.g., ...). Figure 3 (312 in the text). The headset jack provides an interface between the audio circuitry 210 and a removable audio input / output peripheral device, such as an output-only headphone or a headset with both output (e.g., a mono or binaural headphone) and input (e.g., a microphone).

[0068] I / O subsystem 206 couples input / output peripherals on device 200, such as touchscreen 212 and other input control devices 216, to peripheral interface 218. I / O subsystem 206 optionally includes display controller 256, optical sensor controller 258, intensity sensor controller 259, haptic feedback controller 261, and one or more input controllers 260 for other input or control devices. One or more input controllers 260 receive electrical signals from / send electrical signals to other input control devices 216. Other input control devices 216 optionally include physical buttons (e.g., push-buttons, rocker buttons, etc.), dial pads, slide switches, joysticks, click wheels, etc. In some alternative embodiments, input controllers 260 are optionally coupled to (or not coupled to) any of the following: keyboard, infrared port, USB port, and pointing devices such as a mouse. One or more buttons (e.g., Figure 3 Optionally, 308) includes volume up / down buttons for volume control of speaker 211 and / or microphone 213. One or more buttons optionally include push-button buttons (e.g., Figure 3 (306 in the middle).

[0069] A rapid press of the down button disengages the touchscreen 212 from its lock or initiates a process of unlocking the device using gestures on the touchscreen, as described in U.S. Patent Application 11 / 322,549, filed December 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," the entire contents of which are incorporated herein by reference. A longer press of the down button (e.g., 306) powers the device 200 on or off. The user can customize the function of one or more buttons. The touchscreen 212 is used to implement virtual buttons or soft buttons and one or more soft keyboards.

[0070] The touch-sensitive display 212 provides input and output interfaces between the device and the user. The display controller 256 receives electrical signals from and / or sends electrical signals to the touchscreen 212. The touchscreen 212 displays visual output to the user. Visual output includes graphics, text, icons, video, and any combination thereof (collectively, "graphics"). In some embodiments, some or all of the visual output corresponds to user interface objects.

[0071] Touchscreen 212 has a touch-sensitive surface, sensor, or sensor array that accepts input from a user based on tactile and / or haptic contact. Touchscreen 212 and display controller 256 (along with any associated modules and / or instruction set in memory 202) detect contact on touchscreen 212 (and any movement or interruption of that contact) and translate the detected contact into interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on touchscreen 212. In an exemplary embodiment, the contact point between touchscreen 212 and the user corresponds to the user's finger.

[0072] Touchscreen 212 uses LCD (Liquid Crystal Display) technology, LPD (Light Emitting Polymer Display) technology, or LED (Light Emitting Diode) technology, but other display technologies may be used in other embodiments. Touchscreen 212 and display controller 256 use any of a variety of touch sensing technologies currently known or to be developed thereafter, as well as other proximity sensor arrays or other elements for determining one or more points of contact with touchscreen 212 to detect contact and any movement or interruption thereto. These various touch sensing technologies include, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that from Apple Inc. (Cupertino, California). and iPod The technology used.

[0073] In some embodiments, the touchscreen 212's touch-sensitive display is similar to the multi-touch pad described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman) and / or U.S. Patent Publication 2002 / 0015024A1, all of which are incorporated herein by reference in their entirety. However, the touchscreen 212 displays visual output from the device 200, while the touch-sensitive pad does not provide visual output.

[0074] In some embodiments, the touchscreen 212 has a touch-sensitive display as described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, filed May 2, 2006, entitled “Multipoint Touch Surface Controller”; (2) U.S. Patent Application No. 10 / 840,862, filed May 6, 2004, entitled “Multipoint Touchscreen”; (3) U.S. Patent Application No. 10 / 903,964, filed July 30, 2004, entitled “Gestures For Touch Sensitive Input Devices”; (4) U.S. Patent Application No. 11 / 048,264, filed January 31, 2005, entitled “Gestures For Touch Sensitive Input Devices”; and (5) U.S. Patent Application No. 18, 2005, entitled “Mode-Based Graphical User Interfaces For Touch Sensitive Input”. U.S. Patent Application No. 11 / 038,590, entitled “Virtual Input Device Placement On A Touch Screen User Interface”, filed September 16, 2005; U.S. Patent Application No. 11 / 228,758, entitled “Virtual Input Device Placement On A Touch Screen User Interface”, filed September 16, 2005; U.S. Patent Application No. 11 / 228,700, entitled “Operation Of A Computer With A Touch Screen Interface”, filed September 16, 2005; U.S. Patent Application No. 11 / 228,737, entitled “Activating Virtual Keys Of A Touch-Screen Virtual Keyboard”, filed September 16, 2005; and U.S. Patent Application No. 11 / 367,749, entitled “Multi-Functional Hand-Held Device”, filed March 3, 2006. The full text of all these applications is incorporated herein by reference.

[0075] Touchscreen 212 has a video resolution of over 100 dpi, for example. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. The user interacts with touchscreen 212 using any suitable object or accessory such as a stylus, finger, etc. In some embodiments, the user interface is designed to operate primarily through finger-based touch and gestures, which may be less precise than stylus-based input due to the larger contact area of ​​a finger on the touchscreen. In some embodiments, the device translates coarse finger-based input into precise pointer / cursor positions or commands to perform the user-desired actions.

[0076] In some embodiments, in addition to the touchscreen, device 200 also includes a touchpad (not shown) for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of ​​the device that, unlike the touchscreen, does not display visual output. The touchpad is a touch-sensitive surface separate from the touchscreen 212, or an extension of the touch-sensitive surface formed by the touchscreen.

[0077] The device 200 also includes a power system 262 for supplying power to various components. The power system 262 includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in the portable device.

[0078] The device 200 also includes one or more optical sensors 264. Figure 2A An optical sensor 264 is shown coupled to an optical sensor controller 258 in the I / O subsystem 206. The optical sensor 264 includes a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS) phototransistor. The optical sensor 264 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In conjunction with an imaging module 243 (also called a camera module), the optical sensor 264 captures still images or video. In some embodiments, the optical sensor is located at the rear of the device 200, opposite to the touchscreen display 212 at the front of the device, such that the touchscreen display is used as a viewfinder for still image and / or video image acquisition. In some embodiments, the optical sensor is located at the front of the device, such that an image of the user is acquired for use in video conferencing while the user views other video conferencing participants on the touchscreen display. In some embodiments, the position of the optical sensor 264 can be changed by the user (e.g., by rotating the lenses and sensors within the device housing), such that a single optical sensor 264 is used in conjunction with the touchscreen display for both video conferencing and still image and / or video image acquisition.

[0079] The device 200 may optionally also include one or more contact strength sensors 265. Figure 2A A contact strength sensor 265 is shown coupled to a strength sensor controller 259 in I / O subsystem 206. The contact strength sensor 265 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric sensors, optical force sensors, capacitive touch-sensitive surfaces, or other strength sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). The contact strength sensor 265 receives contact strength information (e.g., pressure information or a substitute for pressure information) from the environment. In some embodiments, at least one contact strength sensor is arranged juxtaposed with or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 212). In some embodiments, at least one contact strength sensor is located on the rear of device 200, opposite to the touchscreen display 212 located on the front of device 200.

[0080] The device 200 also includes one or more proximity sensors 266. Figure 2A A proximity sensor 266 coupled to a peripheral device interface 218 is shown. Alternatively, the proximity sensor 266 is coupled to an input controller 260 in an I / O subsystem 206. The proximity sensor 266 performs as described in the following U.S. patent applications: 11 / 241,839, entitled "Proximity Detector In Handheld Device"; 11 / 240,788, entitled "Proximity Detector In Handheld Device"; 11 / 620,702, entitled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; 11 / 586,862, entitled "Automated Response To And Sensing Of User Activity In Portable Devices"; and 11 / 638,251, entitled "Methods And Systems For Automatic Configuration Of Peripherals," the entire contents of which are incorporated herein by reference. In some implementations, the proximity sensor is turned off and the touchscreen 212 is disabled when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call).

[0081] The device 200 optionally also includes one or more tactile output generators 267. Figure 2AA haptic output generator coupled to a haptic feedback controller 261 in I / O subsystem 206 is shown. The haptic output generator 267 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion such as motors, solenoids, electroactive polymerizers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components for converting electrical signals into haptic outputs on the device). A contact intensity sensor 265 receives haptic feedback generation instructions from a haptic feedback module 233 and generates a haptic output on device 200 that can be felt by a user of device 200. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a haptic surface (e.g., haptic display system 212) and optionally generates the haptic output by moving the haptic surface vertically (e.g., in / outward from the surface of device 200) or laterally (e.g., backward and forward in the same plane as the surface of device 200). In some implementations, at least one haptic output generator sensor is located on the rear of the device 200, opposite to the touch screen display 212 located on the front of the device 200.

[0082] The device 200 also includes one or more accelerometers 268. Figure 2A An accelerometer 268 coupled to a peripheral device interface 218 is shown. Alternatively, the accelerometer 268 is coupled to an input controller 260 in an I / O subsystem 206. The accelerometer 268 performs as described in the following U.S. patent publications: U.S. Patent Publication 20050190059, “Acceleration-based Theft Detection System for Portable Electronic Devices” and U.S. Patent Publication 20060017692, “Methods and Apparatuses For Operating A Portable Device Based On An Accelerometer,” the entire contents of which are incorporated herein by reference. In some embodiments, information is displayed on a touchscreen display in portrait or landscape view based on analysis of data received from one or more accelerometers. Device 200 optionally includes, in addition to one or more accelerometers 268, a magnetometer (not shown) and a GPS (or GLONASS or other global navigation system) receiver (not shown) for acquiring information about the position and orientation (e.g., portrait or landscape) of device 200.

[0083] In some embodiments, software components stored in memory 202 include an operating system 226, a communication module (or instruction set) 228, a contact / motion module (or instruction set) 230, a graphics module (or instruction set) 232, a text input module (or instruction set) 234, a Global Positioning System (GPS) module (or instruction set) 235, a digital assistant client module 229, and an application program (or instruction set) 236. Furthermore, memory 202 stores data and models, such as user data and models 231. Additionally, in some embodiments, memory 202 ( Figure 2A ) or 470 ( Figure 4 Storage device / global internal state 257, such as Figure 2A and Figure 4 As shown in the diagram. Device / global internal state 257 includes one or more of the following: active application state, which indicates which applications (if any) are currently active; display state, which indicates what applications, views or other information occupy various areas of the touchscreen display 212; sensor state, including information obtained from the device's various sensors and input control devices 216; and position information about the device's position and / or orientation.

[0084] Operating system 226 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or embedded operating systems such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware and software components.

[0085] The communication module 228 facilitates communication with other devices via one or more external ports 224 and includes various software components for processing data received by the RF circuitry 208 and / or the external ports 224. The external ports 224 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices or indirectly coupled via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is connected to… (Trademark of Apple Inc.) The same or similar and / or compatible multi-pin (e.g., 30-pin) connectors used in Apple Inc. devices.

[0086] The contact / motion module 230 optionally detects contact with the touchscreen 212 (in conjunction with the display controller 256) and other touch-sensitive devices (e.g., touchpads or physical click-based rotary dials). The contact / motion module 230 includes various software components for performing various operations related to contact detection, such as determining whether contact has occurred (e.g., detecting a finger press event), determining the intensity of contact (e.g., the force or pressure of the contact, or an alternative to force or pressure), determining whether there is movement of the contact and tracking movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or contact disconnection). The contact / motion module 230 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations are optionally applied to single-point contact (e.g., single-finger contact) or multi-point simultaneous contact (e.g., "multi-touch" / multiple-finger contact). In some implementations, the contact / motion module 230 and the display controller 256 detect contact on the touchpad.

[0087] In some implementations, the contact / motion module 230 uses a set of one or more intensity thresholds to determine whether an operation has been performed by a user (e.g., determining whether the user has “clicked” an icon). In some implementations, at least a subset of the intensity thresholds is determined based on software parameters (e.g., the intensity thresholds are not determined by the activation thresholds of a specific physical actuator and can be adjusted without changing the physical hardware of the device 200). For example, the mouse “click” threshold for a touchpad or touchscreen can be set to any threshold in a wide range of predefined thresholds without changing the touchpad or touchscreen display hardware. Additionally, in some implementations, the user of the device is provided with software settings for adjusting one or more intensity thresholds in a set (e.g., by adjusting the individual intensity thresholds and / or by adjusting multiple intensity thresholds at once using a system-level click on the “intensity” parameter).

[0088] The touch / motion module 230 optionally detects the user's gesture input. Different gestures on a touch-sensitive surface have different contact patterns (e.g., different movements, timings, and / or intensities of the detected contact). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a finger tap gesture includes detecting a finger press event, and then detecting a finger lift-off (lift-away) event at the same (or substantially the same) location as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on a touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift-off (lift-away) event.

[0089] The graphics module 232 includes various known software components for rendering and displaying graphics on the touchscreen 212 or other displays, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual characteristics). As used herein, the term "graphics" includes any object that can be displayed to a user, and non-limitingly includes text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.

[0090] In some implementations, the graphics module 232 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 232 receives one or more codes specifying the graphics to be displayed from an application or the like, and, if necessary, also receives coordinate data and other graphic attribute data, and then generates screen image data for output to the display controller 256.

[0091] The haptic feedback module 233 includes various software components for generating instructions that are used by one or more haptic output generators 267 to produce haptic output at one or more locations on the device 200 in response to user interaction with the device 200.

[0092] In some examples, the text input module 234, which is a component of the graphics module 232, provides a soft keyboard for entering text in various applications (e.g., contacts 237, email 240, IM 241, browser 247, and any other application that requires text input).

[0093] GPS module 235 determines the location of the device and provides that information for use in various applications (e.g., to phone 238 for use in location-based dialing; to camera 243 as image / video metadata; and to applications that provide location-based services, such as weather desktop apps, local yellow pages desktop apps, and map / navigation desktop apps).

[0094] The digital assistant client module 229 includes various client-side digital assistant commands to provide client-side functionality for the digital assistant. For example, the digital assistant client module 229 can accept voice input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces of the portable multifunction device 200 (e.g., microphone 213, one or more accelerometers 268, touch-sensitive display system 212, one or more optical sensors 264, other input control devices 216, etc.). The digital assistant client module 229 can also provide audio output (e.g., voice output), visual output, and / or haptic output through various output interfaces of the portable multifunction device 200 (e.g., speaker 211, touch-sensitive display system 212, one or more haptic output generators 267, etc.). For example, output can be provided as voice, sound, alarms, text messages, menus, graphics, video, animation, vibration, and / or combinations of both or more of these. During operation, the digital assistant client module 229 communicates with the DA server 106 using RF circuitry 208.

[0095] User data and models 231 include various data associated with the user (e.g., user-specific vocabulary data, user preference data, user-specified name pronunciations, data from the user's electronic address book, to-do lists, shopping lists, etc.) to provide client-side functionality for the digital assistant. Furthermore, user data and models 231 include various models for processing user input and determining user intent (e.g., speech recognition models, statistical language models, natural language processing models, knowledge ontology, task flow models, service models, etc.).

[0096] In some examples, the digital assistant client module 229 utilizes various sensors, subsystems, and peripherals of the portable multifunction device 200 to collect additional information from the surrounding environment of the portable multifunction device 200 to establish a context associated with the user, the current user interaction, and / or the current user input. In some examples, the digital assistant client module 229 provides the contextual information, or a subset thereof, along with the user input to the DA server 106 to help infer the user's intent. In some examples, the digital assistant also uses the contextual information to determine how to prepare output and deliver it to the user. This contextual information is referred to as contextual data.

[0097] In some examples, the contextual information accompanying user input includes sensor information such as lighting, ambient noise, ambient temperature, and images or videos of the surrounding environment. In some examples, the contextual information may also include the physical state of the device, such as device orientation, device location, device temperature, power level, speed, acceleration, motion mode, and cellular signal strength. In some examples, information related to the software state of the DA server 106, such as the operation of the portable multifunction device 200, installed programs, past and current network activity, background services, error logs, and resource usage, is provided to the DA server 106 as contextual information associated with the user input.

[0098] In some examples, the digital assistant client module 229 selectively provides information (e.g., user data 231) stored on the portable multifunction device 200 in response to a request from the DA server 106. In some examples, the digital assistant client module 229 also elicits additional input from the user via natural language dialogue or other user interfaces when requested by the DA server 106. The digital assistant client module 229 transmits this additional input to the DA server 106 to assist the DA server 106 in intent inference and / or in realizing the user intent expressed in the user request.

[0099] The following is for reference. Figures 7A to 7C A more detailed description of the digital assistant follows. It should be understood that the digital assistant client module 229 may include any number of sub-modules of the digital assistant module 726 described below.

[0100] Application 236 includes the following modules (or instruction sets) or subsets or supersets thereof:

[0101] • Contacts module 237 (sometimes called address book or contact list);

[0102] • Telephone module 238;

[0103] Video conferencing module 239;

[0104] • Email client module 240;

[0105] • Instant Messaging (IM) module 241;

[0106] Fitness support module 242;

[0107] • Camera module 243 for still images and / or video images;

[0108] • Image management module 244;

[0109] • Video player module;

[0110] Music player module;

[0111] • Browser module 247;

[0112] • Calendar module 248;

[0113] • Desktop mini-program module 249, in some examples, includes one or more of the following: weather desktop mini-program 249-1, stock market desktop mini-program 249-2, calculator desktop mini-program 249-3, alarm clock desktop mini-program 249-4, dictionary desktop mini-program 249-5 and other desktop mini-programs obtained by the user and desktop mini-programs created by the user 249-6;

[0114] • Desktop app creator module 250 for creating user-created desktop apps 249-6;

[0115] • Search module 251;

[0116] • Video and music player module 252, which combines the video player module and the music player module;

[0117] Notepad module 253;

[0118] • Map module 254; and / or

[0119] • Online video module 255.

[0120] Examples of other applications 236 stored in memory 202 include other word processing applications, other image editing applications, drawing applications, rendering applications, Java-enabled applications, encryption, digital access control, speech recognition, and speech duplication.

[0121] In conjunction with touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, contact module 237 manages an address book or contact list (e.g., in application internal state 292 of contact module 237 stored in memory 202 or memory 470), including: adding one or more names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers or email addresses to initiate and / or facilitate communications via telephone 238, video conferencing module 239, email 240, or IM 241; and so on.

[0122] Combining RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, telephone module 238 is used to input character sequences corresponding to telephone numbers, access one or more telephone numbers in contact module 237, modify already entered telephone numbers, dial corresponding telephone numbers, initiate conversations, and disconnect or hang up when a conversation is completed. As described above, wireless communication uses any of a variety of communication standards, protocols, and technologies.

[0123] Combining RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touchscreen 212, display controller 256, optical sensor 264, optical sensor controller 258, contact / motion module 230, graphics module 232, text input module 234, contact module 237, and telephone module 238, video conferencing module 239 includes executable instructions to initiate, conduct, and terminate video conferences between the user and one or more other participants based on user instructions.

[0124] Incorporating RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, email client module 240 includes executable instructions for creating, sending, receiving, and managing emails in response to user commands. Combined with image management module 244, email client module 240 makes it very easy to create and send emails containing still images or video images captured by camera module 243.

[0125] In conjunction with RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, instant messaging module 241 includes executable instructions for: inputting a character sequence corresponding to an instant message, modifying previously input characters, transmitting a corresponding instant message (e.g., using Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocols for telephone-based instant messaging or using XMPP, SIMPLE, or IMPS for internet-based instant messaging), receiving an instant message, and viewing received instant messages. In some embodiments, the transmitted and / or received instant messages include graphics, photographs, audio files, video files, and / or other attachments such as those supported in MMS and / or Enhanced Messaging Services (EMS). As used herein, "instant message" refers to both telephone-based messages (e.g., messages sent using SMS or MMS) and internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).

[0126] Incorporating RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, GPS module 235, map module 254, and music player module, fitness support module 242 includes executable instructions for: creating fitness activities (e.g., with time, distance, and / or calorie burning goals); communicating with fitness sensors (exercise equipment); receiving fitness sensor data; calibrating sensors used to monitor fitness; selecting and playing music for fitness; and displaying, storing, and transmitting fitness data.

[0127] In conjunction with the touchscreen 212, display controller 256, one or more optical sensors 264, optical sensor controller 258, contact / motion module 230, graphics module 232, and image management module 244, camera module 243 includes executable instructions for: capturing still images or videos (including video streams) and storing them in memory 202, modifying the characteristics of still images or videos, or deleting still images or videos from memory 202.

[0128] Incorporating touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, and camera module 243, image management module 244 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, tagging, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or video images.

[0129] Combining RF circuitry 208, touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, browser module 247 includes executable instructions for browsing the Internet according to user instructions, including searching, linking to, receiving, and displaying web pages or portions thereof, as well as links to attachments and other files on web pages.

[0130] Combining RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, email client module 240, and browser module 247, calendar module 248 includes executable instructions to create, display, modify, and store calendars and associated data (e.g., calendar entries, to-dos, etc.) according to user instructions.

[0131] In conjunction with RF circuitry 208, touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and browser module 247, the desktop applet module 249 is a micro-application that can be downloaded and used by a user (e.g., weather desktop applet 249-1, stock market desktop applet 249-2, calculator desktop applet 249-3, alarm clock desktop applet 249-4, and dictionary desktop applet 249-5) or a user-created micro-application (e.g., user-created desktop applet 249-6). In some embodiments, the desktop applet includes HTML (Hypertext Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the desktop applet includes XML (Extensible Markup Language) files and JavaScript files (e.g., Yahoo! desktop applet).

[0132] Combining RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, and browser module 247, the desktop applet creator module 250 is used by the user to create desktop applets (e.g., to turn a user-specified portion of a webpage into a desktop applet).

[0133] In conjunction with the touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, the search module 251 includes executable instructions for searching the memory 202 for text, music, sound, images, videos, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) according to user instructions.

[0134] Incorporating touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, and browser module 247, the video and music player module 252 includes executable instructions allowing users to download and play back recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touchscreen 212 or on an external display connected via external port 224). In some embodiments, device 200 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0135] Combining the touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, the notepad module 253 includes executable instructions for creating and managing notes, to-do items, etc., according to user instructions.

[0136] Combining RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, GPS module 235, and browser module 247, map module 254 is used to receive, display, modify, and store maps and data associated with the maps (e.g., driving directions, data related to shops and other points of interest at or near a specific location, and other location-based data) according to user instructions.

[0137] Incorporating touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, text input module 234, email client module 240, and browser module 247, the online video module 255 includes instructions allowing users to access, browse, receive (e.g., via streaming and / or downloading), play back (e.g., on the touchscreen or on a connected external display via external port 224), send emails with links to specific online videos, and otherwise manage online videos in one or more file formats (such as H.264). In some embodiments, instant messaging module 241 is used instead of email client module 240 to send links to specific online videos. Further descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” the contents of which are incorporated herein by reference in their entirety.

[0138] Each of the modules and applications described above corresponds to an executable set of instructions for performing one or more of the functions described above and the methods described in this patent application (e.g., computer-implemented methods and other information processing methods as described herein). These modules (e.g., instruction sets) need not be implemented as standalone software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, a video player module can be combined with a music player module into a single module (e.g., Figure 2A(e.g., video and music player module 252). In some embodiments, memory 202 stores a subset of the aforementioned modules and data structures. Additionally, memory 202 stores additional modules and data structures not described above.

[0139] In some implementations, device 200 is a device on which the operation of a predefined set of functions is performed solely via a touchscreen and / or touchpad. By using a touchscreen and / or touchpad as the primary input control device for the operation of device 200, the number of physical input control devices (such as push-buttons, dials, etc.) on device 200 is reduced.

[0140] A predefined set of functions, uniquely performed via a touchscreen and / or touchpad, optionally includes navigation between user interfaces. In some implementations, the touchpad, when touched by a user, navigates device 200 from any user interface displayed on device 200 to a main menu, home menu, or root menu. In such implementations, a touchpad is used to implement a "menu button." In some other implementations, the menu button is a physical push-button or other physical input control device, rather than a touchpad.

[0141] Figure 2B A block diagram illustrating exemplary components for event processing according to some embodiments is provided. In some embodiments, memory 202 ( Figure 2A ) or memory 470 ( Figure 4 This includes an event classifier 270 (e.g., in operating system 226) and a corresponding application 236-1 (e.g., any one of the aforementioned applications 237 to 251, 255, 480 to 490).

[0142] Event classifier 270 receives event information and determines the application 236-1 to which the event information should be delivered and the application view 291 of application 236-1. Event classifier 270 includes event monitor 271 and event dispatcher module 274. In some embodiments, application 236-1 includes application internal state 292, which indicates one or more current application views displayed on touch-sensitive display 212 when the application is active or executing. In some embodiments, device / global internal state 257 is used by event classifier 270 to determine which application(s) is currently active, and application internal state 292 is used by event classifier 270 to determine the application view 291 to which the event information should be delivered.

[0143] In some implementations, the application internal state 292 includes additional information such as one or more of the following: recovery information to be used when the application 236-1 resumes execution, user interface state information indicating that information is being displayed or ready to be displayed by the application 236-1, a state queue for enabling the user to return to the previous state or view of the application 236-1, and a repeat / undo queue for the user's previous actions.

[0144] Event monitor 271 receives event information from peripheral device interface 218. The event information includes information about sub-events (e.g., user touches on touch-sensitive display 212 as part of a multi-touch gesture). Peripheral device interface 218 transmits information it receives from I / O subsystem 206 or sensors such as proximity sensor 266, one or more accelerometers 268, and / or microphone 213 (via audio circuitry 210). The information received by peripheral device interface 218 from I / O subsystem 206 includes information from touch-sensitive display 212 or touch-sensitive surfaces.

[0145] In some implementations, event monitor 271 sends requests to peripheral device interface 218 at predetermined intervals. In response, peripheral device interface 218 transmits event information. In other implementations, peripheral device interface 218 transmits event information only when a significant event occurs (e.g., receiving input above a predetermined noise threshold and / or receiving input for a predetermined duration).

[0146] In some implementations, the event classifier 270 also includes a hit view determination module 272 and / or an activity event recognizer determination module 273.

[0147] When the touch-sensitive display 212 displays more than one view, the hit view determination module 272 provides a software process for determining where a sub-event has occurred within one or more views. A view consists of controls and other elements that the user can see on the display.

[0148] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected corresponds to a procedural level within the application's procedural hierarchy or view hierarchy. For example, the lowest-level view in which a touch is detected is called the hit view, and the set of events considered as correct input is determined at least in part based on the hit view of the initial touch that initiates the touch-based gesture.

[0149] The hit view determination module 272 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchical structure, the hit view determination module 272 identifies the hit view as the lowest-level view in the hierarchical structure from which the sub-events should be processed. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events forming an event or potential event) occurs. Once the hit view is identified by the hit view determination module 272, the hit view typically receives all sub-events related to the same touch or input source to which it was identified as the hit view.

[0150] The activity event recognizer determination module 273 determines which views(s) within the view hierarchy should receive a specific sub-event sequence. In some embodiments, the activity event recognizer determination module 273 determines that only the hit view should receive the specific sub-event sequence. In other embodiments, the activity event recognizer determination module 273 determines that all views including the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive the specific sub-event sequence. In other embodiments, even if the touch sub-event is entirely confined to the area associated with a particular view, the higher-level views in the hierarchy will still remain actively participating views.

[0151] Event assigner module 274 assigns event information to event identifiers (e.g., event identifier 280). In embodiments that include active event identifier determination module 273, event assigner module 274 delivers event information to the event identifier determined by active event identifier determination module 273. In some embodiments, event assigner module 274 stores event information in an event queue, which is retrieved by the corresponding event receiver 282.

[0152] In some embodiments, operating system 226 includes event classifier 270. Alternatively, application 236-1 includes event classifier 270. In yet another embodiment, event classifier 270 is a standalone module or part of another module (such as contact / motion module 230) stored in memory 202.

[0153] In some embodiments, application 236-1 includes a plurality of event handlers 290 and one or more application views 291, each of which includes instructions for handling touch events occurring within a corresponding view of the application's user interface. Each application view 291 of application 236-1 includes one or more event recognizers 280. Typically, a corresponding application view 291 includes a plurality of event recognizers 280. In other embodiments, one or more event recognizers among the event recognizers 280 are part of a separate module, which is a higher-level object such as a user interface toolkit (not shown) from which application 236-1 inherits methods and other properties. In some embodiments, a corresponding event handler 290 includes one or more of the following: a data updater 276, an object updater 277, a GUI updater 278, and / or event data 279 received from an event classifier 270. The event handler 290 utilizes or invokes the data updater 276, the object updater 277, or the GUI updater 278 to update the application's internal state 292. Alternatively, one or more application views in application view 291 include one or more corresponding event handlers 290. Additionally, in some embodiments, one or more of data updater 276, object updater 277, and GUI updater 278 are included in the corresponding application view 291.

[0154] The corresponding event recognizer 280 receives event information (e.g., event data 279) from the event classifier 270 and identifies events from the event information. The event recognizer 280 includes an event receiver 282 and an event comparator 284. In some embodiments, the event recognizer 280 also includes at least one subset of metadata 283 and event delivery instructions 288 (which includes sub-event delivery instructions).

[0155] Event receiver 282 receives event information from event classifier 270. The event information includes information about sub-events such as touch or touch movement. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves touch movement, the event information also includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a lateral orientation, or vice versa), and the event information includes corresponding information about the device's current orientation (also referred to as device pose).

[0156] Event comparator 284 compares event information with predefined event or sub-event definitions and, based on the comparison, determines the event or sub-event, or determines or updates the state of the event or sub-event. In some embodiments, event comparator 284 includes event definition 286. Event definition 286 contains definitions of events (e.g., predefined sequences of sub-events), such as event 1 (287-1), event 2 (287-2), and other events. In some embodiments, sub-events in event (287) include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, event 1 (287-1) is defined as a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift-off of a predetermined duration (touch end), a second touch (touch start) of a predetermined duration on the displayed object, and a second lift-off of a predetermined duration (touch end). In another example, event 2 (287-2) is defined as a drag on a displayed object. For example, dragging includes a touch (or contact) on the displayed object for a predetermined duration, movement of the touch on the touch-sensitive display 212, and lifting off the touch (end of touch). In some embodiments, the event also includes information for one or more associated event handlers 290.

[0157] In some implementations, event definition 287 includes definitions of events for corresponding user interface objects. In some implementations, event comparator 284 performs a hit test to determine which user interface object is associated with the sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display 212, when a touch is detected on touch-sensitive display 212, event comparator 284 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 290, the event comparator uses the result of the hit test to determine which event handler 290 should be activated. For example, event comparator 284 selects the event handler associated with the sub-event and the object that triggered the hit test.

[0158] In some implementations, the definition of the corresponding event (287) also includes a delay action that delays the delivery of event information until it has been determined whether the sub-event sequence actually corresponds to or does not correspond to the event type of the event recognizer.

[0159] When the corresponding event recognizer 280 determines that the sub-event sequence does not match any event in event definition 286, the corresponding event recognizer 280 enters an event impossible, event failed, or event ended state, after which subsequent sub-events based on touch gestures are ignored. In this case, other event recognizers (if any) that remain active in the hit view continue to track and process the ongoing sub-events based on touch gestures.

[0160] In some embodiments, the corresponding event recognizer 280 includes metadata 283 having configurable attributes, flags, and / or lists instructing how the event delivery system should perform sub-event delivery to actively participating event recognizers. In some embodiments, the metadata 283 includes configurable attributes, flags, and / or lists instructing how or how likely event recognizers can interact with each other. In some embodiments, the metadata 283 includes configurable attributes, flags, and / or lists instructing whether sub-events are delivered to different levels in a view or programmatic hierarchy.

[0161] In some implementations, when one or more specific sub-events of an event are identified, the corresponding event recognizer 280 activates the event handler 290 associated with the event. In some implementations, the corresponding event recognizer 280 delivers event information associated with the event to the event handler 290. Activating the event handler 290 is different from sending (and delaying) the sub-event to the corresponding hit view. In some implementations, the event recognizer 280 throws a tag associated with the identified event, and the event handler 290 associated with that tag retrieves the tag and executes a predefined procedure.

[0162] In some implementations, event delivery instruction 288 includes a sub-event delivery instruction that delivers event information about a sub-event without activating an event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with the sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or the actively participating view receives the event information and executes a predetermined procedure.

[0163] In some implementations, data updater 276 creates and updates data used in application 236-1. For example, data updater 276 updates phone numbers used in contact module 237 or stores video files used in video player module. In some implementations, object updater 277 creates and updates objects used in application 236-1. For example, object updater 277 creates new user interface objects or updates the location of user interface objects. GUI updater 278 updates the GUI. For example, GUI updater 278 prepares display information and sends the display information to graphics module 232 for display on touch-sensitive display.

[0164] In some implementations, event handler 290 includes, or has access to, a data updater 276, an object updater 277, and a GUI updater 278. In some implementations, data updater 276, object updater 277, and GUI updater 278 are included in a single module of the corresponding application 236-1 or application view 291. In other implementations, they are included in two or more software modules.

[0165] It should be understood that the above discussion regarding event handling for user touch on a touch-sensitive display also applies to other forms of user input used to operate the multifunction device 200 using an input device, and not all user input is initiated on the touchscreen. For example, mouse movement and mouse button presses optionally in conjunction with single or multiple keyboard presses or holds; touch movements on the touchpad, such as taps, drags, scrolls, etc.; stylus input; device movement; verbal commands; detected eye movements; biometric input; and / or any combination thereof may optionally be used as input corresponding to sub-events that define the event to be identified.

[0166] Figure 3A portable multifunction device 200 with a touchscreen 212 is shown according to some embodiments. The touchscreen optionally displays one or more graphics within a user interface (UI) 300. In this embodiment and other embodiments described below, a user can select one or more graphics by gesturing over the graphics, for example, using one or more fingers 302 (not drawn to scale in the figure) or one or more styluses 303 (not drawn to scale in the figure). In some embodiments, selection of one or more graphics occurs when the user breaks contact with one or more graphics. In some embodiments, gestures optionally include one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or scrolling (from right to left, from left to right, up and / or down) of a finger already in contact with the device 200. In some specific embodiments or in some cases, unintentional contact with a graphic does not select the graphic. For example, a swipe gesture over an application icon optionally does not select the corresponding application when the gesture corresponding to selection is a tap.

[0167] Device 200 also includes one or more physical buttons, such as a "home" or menu button 304. As previously described, menu button 304 is used to navigate to any application 236 of a set of applications running on device 200. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on touchscreen 212.

[0168] In some embodiments, device 200 includes a touchscreen 212, a menu button 304, a push-button 306 for powering on / off the device and locking the device, one or more volume control buttons 308, a SIM card slot 310, a headset jack 312, and a docking / charging external port 224. The push-button 306 is optionally used to power on / off the device by pressing the button and holding it in the pressed state for a predefined time interval; to lock the device by pressing the button and releasing it before the predefined time interval has elapsed; and / or to unlock the device or initiate an unlocking process. In another embodiment, device 200 also accepts verbal input via microphone 213 for activating or deactivating certain functions. Device 200 also optionally includes one or more contact strength sensors 265 for detecting the intensity of contact on the touchscreen 212, and / or one or more haptic output generators 267 for generating haptic outputs for the user of device 200.

[0169] Figure 4This is a block diagram of an exemplary multifunctional device with a display and a touch-sensitive surface according to some embodiments. Device 400 need not be portable. In some embodiments, device 400 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a children's learning toy), gaming system, or control device (e.g., a home controller or industrial controller). Device 400 typically includes one or more processing units (CPUs) 410, one or more network or other communication interfaces 460, memory 470, and one or more communication buses 420 for interconnecting these components. Communication bus 420 optionally includes circuitry (sometimes referred to as a chipset) that interconnects system components and controls communication between system components. Device 400 includes an input / output (I / O) interface 430 with a display 440, which is typically a touchscreen display. I / O interface 430 also optionally includes a keyboard and / or mouse (or other pointing device) 450 and a touchpad 455, and a haptic output generator 457 for generating haptic output on device 400 (e.g., similar to the reference above). Figure 2A The one or more tactile output generators 267 and sensors 459 (e.g., optical sensors, accelerometers, proximity sensors, touch sensors, and / or contact intensity sensors similar to those mentioned above) Figure 2A The one or more contact strength sensors 265 mentioned above). Memory 470 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 470 optionally includes one or more storage devices located remotely from CPU 410. In some embodiments, memory 470 stores data with portable multifunction device 200 (…). Figure 2A The memory 470 stores programs, modules, and data structures similar to those in the memory 202 of the portable multifunction device 200, or subsets thereof. Additionally, the memory 470 optionally stores additional programs, modules, and data structures not present in the memory 202 of the portable multifunction device 200. For example, the memory 470 of the device 400 optionally stores a drawing module 480, a rendering module 482, a word processing module 484, a website creation module 486, a disk editing module 488, and / or a spreadsheet module 490, while the portable multifunction device 200 ( Figure 2A The memory 202 optionally does not store these modules.

[0170] Figure 4Each of the aforementioned elements is stored in one or more of the previously mentioned memory devices in some examples. Each of the aforementioned modules corresponds to an instruction set for performing the functions described above. The aforementioned modules or programs (e.g., instruction sets) need not be implemented as standalone software programs, processes, or modules; therefore, various subsets of these modules are combined or otherwise rearranged in various embodiments. In some embodiments, memory 470 stores a subset of the aforementioned modules and data structures. Furthermore, memory 470 stores additional modules and data structures not described above.

[0171] Now let’s turn our attention to implementations of user interfaces that can be implemented, for example, on a portable multi-functional device 200.

[0172] Figure 5A An exemplary user interface for an application menu on a portable multifunction device 200 according to some embodiments is shown. A similar user interface is implemented on device 400. In some embodiments, user interface 500 includes the following elements or a subset or superset thereof:

[0173] One or more signal strength indicators 502 for wireless communications such as cellular signals and Wi-Fi signals;

[0174] Time 504;

[0175] Bluetooth indicator 505;

[0176] • Battery status indicator 506;

[0177] • Tray icon 508 with icons of commonly used applications, such as:

[0178] The telephone module 238 has an icon 516 labeled "telephone", which optionally includes an indicator 514 indicating the number of missed calls or voicemail messages;

[0179] The email client module 240 has an icon 518 labeled "Mail", which optionally includes an indicator 510 for the number of unread emails;

[0180] The icon 520 labeled "Browser" in browser module 247; and

[0181] The video and music player module 252 (also known as the iPod (a trademark of Apple Inc.) module 252) has an icon 522 labeled "iPod"; and

[0182] • Icons of other applications, such as:

[0183] ο IM module 241's icon 524 marked as "message";

[0184] The icon 526 of the calendar module 248 is labeled "Calendar";

[0185] The icon 528 of the image management module 244 is labeled "photo";

[0186] The icon 530 of the camera module 243 is labeled "camera";

[0187] The icon 532 of the online video module 255 is labeled "Online Video";

[0188] The icon labeled "Stock Market" in the Stock Market Desktop Mini Program 249-2 is 534.

[0189] The icon 536 of the map module 254 is labeled "map";

[0190] The icon labeled "Weather" in the Weather Desktop Mini Program 249-1 is 538.

[0191] The icon labeled "Clock" in the alarm clock desktop mini-program 249-4 is 540;

[0192] The icon 542 of the fitness support module 242 is labeled "fitness support";

[0193] The icon 544 labeled "Notepad" in the Notepad module 253; and

[0194] An icon 546 labeled "Settings" is used to set the settings of an application or module, providing access to the settings of the device 200 and its various applications 236.

[0195] It should be pointed out that, Figure 5A The icon labels shown are merely exemplary. For example, the icon 522 of the video and music player module 252 is optionally labeled "Music" or "Music Player". Other labels are optionally used for various application icons. In some embodiments, the label of a particular application icon includes the name of the application corresponding to that particular application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to that particular application icon.

[0196] Figure 5B A touch-sensitive surface 551 (e.g., separate from the display 550 (e.g., touchscreen display 212)) is shown. Figure 4 Devices (e.g., tablets or touchpads 455) Figure 4An exemplary user interface on device 400. Device 400 also optionally includes one or more contact intensity sensors (e.g., one or more sensors in sensor 457) for detecting the intensity of contact on tactile surface 551 and / or one or more tactile output generators 459 for generating tactile output for the user of device 400.

[0197] While some examples of input on a reference touchscreen display 212 (which combines a touch-sensitive surface and a display) are given in the following examples, in some implementations, the device detects input on a touch-sensitive surface separate from the display, such as... Figure 5B As shown in the diagram. In some embodiments, the touch-sensitive surface (e.g., Figure 5B 551) has a spindle (e.g., on the display (e.g., 550) that is aligned with the main axis on the display (e.g., Figure 5B The spindle corresponding to 553 in the figure (e.g., Figure 5B (552 in the example). According to these embodiments, the device detects the position corresponding to the corresponding position on the display (e.g., in the example). Figure 5B In the middle, 560 corresponds to 568 and 562 corresponds to 570) the contact with the touch-sensitive surface 551 at the location (e.g., Figure 5B (560 and 562 in the example). Thus, on touch-sensitive surfaces (e.g., Figure 5B 551 in the middle) and the display of a multi-functional device (e.g., Figure 5B When 550 is separated from 560, user input detected by the device on the touch-sensitive surface (e.g., contact with 560 and 562 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods may be optionally used for other user interfaces described herein.

[0198] Additionally, while the examples below are primarily given with reference to finger input (e.g., finger touch, single-finger tap, finger swipe), it should be understood that in some implementations, one or more of these finger inputs may be replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be replaced by a mouse click (e.g., instead of a touch), followed by movement of the cursor along the swipe path (e.g., instead of movement of the touch). Similarly, a tap gesture may optionally be replaced by a mouse click while the cursor is over the location of the tap gesture (e.g., instead of detection of touch, followed by cessation of touch detection). Likewise, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice may optionally be used simultaneously, or mouse and finger touch may optionally be used simultaneously.

[0199] Figure 6AAn exemplary personal electronic device 600 is illustrated. Device 600 includes a body 602. In some embodiments, device 600 includes components relative to devices 200 and 400 (e.g., Figures 2A to 4 Some or all of the features described herein. In some embodiments, device 600 has a touch-sensitive display 604, hereinafter referred to as touchscreen 604. As an alternative to or complement to touchscreen 604, device 600 has a display and a touch-sensitive surface. Similar to devices 200 and 400, in some embodiments, touchscreen 604 (or touch-sensitive surface) has one or more intensity sensors for detecting the intensity of an applied contact (e.g., a touch). The one or more intensity sensors of touchscreen 604 (or touch-sensitive surface) provide output data representing the intensity of the touch. The user interface of device 600 responds to touches based on touch intensity, meaning that touches of different intensities may invoke different user interface operations on device 600.

[0200] Techniques for detecting and processing touch intensity may exist, for example, in the relevant applications: international patent application PCT / US2013 / 040061, filed May 8, 2013, entitled “Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application,” and international patent application PCT / US2013 / 069483, filed November 11, 2013, entitled “Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships,” each of which is incorporated herein by reference in its entirety.

[0201] In some embodiments, device 600 has one or more input mechanisms 606 and 608. Input mechanisms 606 and 608 (if included) are physical in form. Examples of physical input mechanisms include push-buttons and rotatable mechanisms. In some embodiments, device 600 has one or more attachment mechanisms. Such attachment mechanisms (if included) allow device 600 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch straps, bangles, trousers, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow a user to wear device 600.

[0202] Figure 6BAn exemplary personal electronic device 600 is illustrated. In some embodiments, device 600 includes, relative to... Figure 2A , Figure 2B and Figure 4 Some or all of the components described herein. Device 600 has a bus 612 that operatively couples I / O portion 614 to one or more computer processors 616 and memory 618. I / O portion 614 is connected to display 604, which may have touch-sensitive component 622 and optionally also has touch intensity-sensitive component 624. Furthermore, I / O portion 614 is connected to communication unit 630 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular and / or other wireless communication technologies. Device 600 includes input mechanisms 606 and / or 608. For example, input mechanism 606 is a rotatable input device or a pressable input device and a rotatable input device. In some examples, input mechanism 608 is a button.

[0203] In some examples, the input mechanism 608 is a microphone. The personal electronic device 600 includes, for example, various sensors such as a GPS sensor 632, an accelerometer 634, an orientation sensor 640 (e.g., a compass), a gyroscope 636, a motion sensor 638, and / or combinations thereof, all of which are operatively connected to the I / O section 614.

[0204] The memory 618 of the personal electronic device 600 is a non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors 616, cause the computer processors to perform the techniques and processes described above. The computer-executable instructions are also stored and / or transmitted, for example, in any non-transitory computer-readable storage medium, for use by or in conjunction with an instruction execution system, apparatus, or device, such as a computer-based system, a processor-containing system, or other system capable of retrieving and executing instructions from and from an instruction execution system, apparatus, or device. The personal electronic device 600 is not limited to... Figure 6B It can be the components and configurations, or it can include other components or additional components in a variety of configurations.

[0205] As used herein, the term "power indication" refers, for example, in devices 200, 400, 600, 800, 900, 902, or 904 (… Figure 2A , Figure 4 , Figures 6A to 6B , Figures 8A to 8CT , Figures 9A to 9C , Figures 10A to 10V , Figure 12 , Figure 14 , Figure 15 and Figure 16A graphical user interface object displayed on a screen. For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each constitute a display representation.

[0206] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface with which a user is interacting. In some specific implementations that include a cursor or other positional marker, the cursor acts as a "focus selector," such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the cursor is positioned on a touch-sensitive surface (e.g., a...). Figure 4 The touchpad 455 or Figure 5B When an input (e.g., a press input) is detected on the touch-sensitive surface 551 of the display, the specific user interface element is adjusted according to the detected input. This applies to touchscreen displays (e.g., those capable of direct interaction with user interface elements on a touchscreen display) that enable direct interaction with user interface elements on the touchscreen display. Figure 2A The touch-sensitive display system 212 or Figure 5A In some embodiments of the touchscreen 212, a touch detected on the touchscreen acts as a "focus selector," such that when input (e.g., a press input by touch) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, that particular user interface element is adjusted according to the detected input. In some embodiments, focus moves from one area of ​​the user interface to another without corresponding movement of the cursor or movement of a touch on the touchscreen display (e.g., moving focus from one button to another using tab keys or arrow keys); in these embodiments, the focus selector moves according to the movement of focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is typically a user-controlled user interface element (or a touch on the touchscreen display) that delivers the user-expected interaction with the user interface (e.g., by indicating to the device the element of the user interface that the user expects to interact with). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen), the position of the focus selector (e.g., a cursor, touch, or selection box) above the corresponding button will indicate to the user that they expect to activate the corresponding button (rather than other user interface elements shown on the device's display).

[0207] As used in the specification and claims, the term "characteristic intensity" of a contact refers to a characteristic of the contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on multiple intensity samples. The characteristic intensity is optionally based on a predefined number of intensity samples or a set of intensity samples collected over a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after contact is detected, before contact is detected to be lifted, before or after contact begins to move, before contact ends, before or after contact intensity is detected to increase and / or before or after contact intensity decreases). The characteristic intensity of the contact is optionally based on one or more of the following: the maximum value of the contact intensity, the mean value of the contact intensity, the average value of the contact intensity, the value at the top 10% of the contact intensity, the half maximum value of the contact intensity, the 90% maximum value of the contact intensity, etc. In some embodiments, the duration of the contact is used when determining the characteristic intensity (e.g., when the characteristic intensity is the average value of the contact intensity over time). In some implementations, the feature intensity is compared to a set of one or more intensity thresholds to determine whether a user has performed an action. For example, the set of one or more intensity thresholds may include a first intensity threshold and a second intensity threshold. In this example, contact with a feature intensity not exceeding the first threshold results in a first action, contact with a feature intensity exceeding the first intensity threshold but not exceeding the second intensity threshold results in a second action, and contact with a feature intensity exceeding the second threshold results in a third action. In some implementations, a comparison between the feature intensity and one or more thresholds is used to determine whether to perform one or more actions (e.g., whether to perform the corresponding action or abort performing the corresponding action), rather than to determine whether to perform the first or second action.

[0208] In some implementations, a portion of the gesture is identified to determine the characteristic intensity. For example, a touch-sensitive surface receives a series of swipes that transition from a starting position to an ending position, where the intensity of the contact increases. In this example, the characteristic intensity of the contact at the ending position is based only on a portion of the series of swipes, rather than the entire swipe (e.g., the swipe contact is only the portion at the ending position). In some implementations, a smoothing algorithm is applied to the intensity of the swipe contact before determining its characteristic intensity. For example, the smoothing algorithm optionally includes one or more of the following: unweighted moving average smoothing algorithm, triangular smoothing algorithm, median filter smoothing algorithm, and / or exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow spikes or dips in the intensity of the swipe contact to achieve the purpose of determining the characteristic intensity.

[0209] The intensity of a contact on a touch-sensitive surface is characterized relative to one or more intensity thresholds, such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device performs an operation typically associated with clicking a button on a physical mouse or touchpad. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device performs an operation different from the operation typically associated with clicking a button on a physical mouse or touchpad. In some embodiments, when a contact with a characteristic intensity lower than the light press intensity threshold (e.g., and higher than the nominal contact detection intensity threshold, where contacts lower than the nominal contact detection intensity threshold are no longer detected) is detected, the device will move the focus selector based on the movement of the contact on the touch-sensitive surface without performing the operation associated with the light press intensity threshold or the deep press intensity threshold. Generally, unless otherwise stated, these intensity thresholds are consistent across different groups of user interface figures.

[0210] An increase in contact intensity from below a light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in contact intensity from below a deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in contact intensity from below a contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on the touch surface. A decrease in contact intensity from above a contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting a contact being lifted off the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.

[0211] In some embodiments described herein, one or more operations are performed in response to detecting a gesture including a corresponding press input or in response to detecting a corresponding press input performed using a corresponding contact (or multiple contacts), wherein the corresponding press input is detected at least in part based on detecting that the intensity of the contact (or multiple contacts) increases to above a press input intensity threshold. In some embodiments, the corresponding operation is performed in response to detecting that the intensity of the corresponding contact increases to above a press input intensity threshold (e.g., a "downward stroke" of the corresponding press input). In some embodiments, the press input includes the intensity of the corresponding contact increasing to above a press input intensity threshold and the intensity of the contact subsequently decreasing to below the press input intensity threshold, and the corresponding operation is performed in response to detecting that the intensity of the corresponding contact subsequently decreases to below the press input threshold (e.g., an "upward stroke" of the corresponding press input).

[0212] In some implementations, the device employs intensity hysteresis to avoid unintended inputs sometimes referred to as "jitter," wherein the device defines or selects a hysteresis intensity threshold that has a predefined relationship with a press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the press input intensity threshold). Therefore, in some implementations, a press input includes an increase in the intensity of the corresponding contact above the press input intensity threshold and a subsequent decrease in the intensity of that contact below the hysteresis intensity threshold corresponding to the press input intensity threshold, and a corresponding operation is performed in response to detecting that the intensity of the corresponding contact subsequently decreases below the hysteresis intensity threshold (e.g., an "upstroke" of the corresponding press input). Similarly, in some implementations, a press input is detected only when the device detects that the contact intensity increases from an intensity equal to or below the hysteresis intensity threshold to an intensity equal to or above the press input intensity threshold and optionally the contact intensity subsequently decreases to an intensity equal to or below the hysteresis intensity threshold, and a corresponding operation is performed in response to detecting a press input (e.g., an increase or decrease in contact intensity depending on the environment).

[0213] For ease of explanation, optionally, the description of an operation triggered in response to a press input associated with a press input strength threshold or in response to a gesture including a press input is provided in response to detecting any of the following conditions: the contact strength increases to above the press input strength threshold, the contact strength increases from below a hysteresis strength threshold to above the press input strength threshold, the contact strength decreases to below the press input strength threshold, and / or the contact strength decreases to below the hysteresis strength threshold corresponding to the press input strength threshold. Additionally, in the example where the operation is described as being performed in response to detecting a decrease in contact strength below the press input strength threshold, the operation is optionally performed in response to detecting a decrease in contact strength below a hysteresis strength threshold corresponding to and less than the press input strength threshold.

[0214] 3. Digital Assistant System

[0215] Figure 7A A block diagram of a digital assistant system 700 according to various examples is shown. In some examples, the digital assistant system 700 is implemented on a standalone computer system. In some examples, the digital assistant system 700 is distributed across multiple computers. In some examples, some modules and functions of the digital assistant are divided into server and client parts, wherein the client part resides on one or more user devices (e.g., devices 104, 122, 200, 400, 600, 800, 900, 902, or 904) and communicates with the server part (e.g., server system 108) via one or more networks, for example, as... Figure 1As shown. In some examples, the digital assistant system 700 is Figure 1 The specific implementation of the server system 108 (and / or DA server 106) shown is illustrated. It should be noted that the digital assistant system 700 is merely an example of a digital assistant system, and the digital assistant system 700 may have more or fewer components than shown, combine two or more components, or have different configurations or layouts of components. Figure 7A The various components shown are implemented in hardware, software instructions for execution by one or more processors, firmware (including one or more signal processing integrated circuits and / or application-specific integrated circuits), or a combination thereof.

[0216] The digital assistant system 700 includes a memory 702, an input / output (I / O) interface 706, a network communication interface 708, and one or more processors 704. These components can communicate with each other via one or more communication buses or signal lines 710.

[0217] In some examples, memory 702 includes non-transitory computer-readable media, such as high-speed random access memory and / or non-volatile computer-readable storage media (e.g., one or more disk storage devices, flash memory devices or other non-volatile solid-state memory devices).

[0218] In some examples, I / O interface 706 couples input / output devices 716 of digital assistant system 700, such as a display, keyboard, touchscreen, and microphone, to user interface module 722. I / O interface 706, together with user interface module 722, receives user input (e.g., voice input, keyboard input, touch input, etc.) and processes this input accordingly. In some examples, such as when the digital assistant is implemented on a standalone user device, digital assistant system 700 includes a user interface module 722. Figure 2A , Figure 4 , Figures 6A to 6B , Figures 8A to 8CT , Figures 9A to 9C , Figures 10A to 10V , Figure 12 , Figure 14 , Figure 15 and Figure 16 Any of the components and I / O communication interfaces described in devices 200, 400, 600, 800, 900, 902, or 904. In some examples, digital assistant system 700 represents the server portion of a digital assistant implementation and can interact with the user through a client-side portion located on a user device (e.g., device 104, 200, 400, 600, 800, 900, 902, or 904).

[0219] In some examples, the network communication interface 708 includes one or more wired communication ports 712 and / or wireless transmission and reception circuitry 714. The wired communication ports receive and transmit communication signals via one or more wired interfaces such as Ethernet, Universal Serial Bus (USB), FireWire, etc. The wireless circuitry 714 receives RF signals and / or optical signals from the communication network and other communication devices, and transmits RF signals and / or optical signals to the communication network and other communication devices. Wireless communication uses any of a variety of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. The network communication interface 708 enables the digital assistant system 700 to communicate with other devices via networks such as the Internet, intranets, and / or wireless networks such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs).

[0220] In some examples, memory 702 or its computer-readable storage medium stores programs, modules, instructions, and data structures, including all or a subset of the following: operating system 718, communication module 720, user interface module 722, one or more application programs 724, and digital assistant module 726. Specifically, memory 702 or its computer-readable storage medium stores instructions for performing the above-described processes. One or more processors 704 execute these programs, modules, and instructions, and read data from or write data to data structures.

[0221] Operating systems 718 (e.g., Darwin, RTXC, LINUX, UNIX, iOS, OSX, WINDOWS, or embedded operating systems such as VxWorks) include various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware, firmware, and software components.

[0222] The communication module 720 facilitates communication between the digital assistant system 700 and other devices via the network communication interface 708. For example, the communication module 720 communicates with electronic devices (such as those in…) Figure 2A , Figure 4 , Figures 6A to 6B The device 200, 400, or 600 shown communicates with the RF circuit 208. The communication module 720 also includes various components for processing data received by the wireless circuit 714 and / or the wired communication port 712.

[0223] The user interface module 722 receives commands and / or input from the user (e.g., from a keyboard, touchscreen, pointing device, controller, and / or microphone) via the I / O interface 706 and generates user interface objects on the display. The user interface module 722 also prepares output (e.g., voice, sound, animation, text, icons, vibration, haptic feedback, lighting, etc.) and transmits it to the user via the I / O interface 706 (e.g., through a display, audio channel, speaker, touchpad, etc.).

[0224] Application 724 includes programs and / or modules configured to be executed by the one or more processors 704. For example, if the digital assistant system is implemented on a standalone user device, application 724 includes user applications such as games, calendar applications, navigation applications, or email applications. If the digital assistant system 700 is implemented on a server, application 724 includes, for example, resource management applications, diagnostic applications, or scheduling applications.

[0225] The memory 702 also stores the digital assistant module 726 (or the server portion of the digital assistant). In some examples, the digital assistant module 726 includes the following submodules or subsets or supersets thereof: input / output processing module 728, speech-to-text (STT) processing module 730, natural language processing module 732, dialogue flow processing module 734, task flow processing module 736, service processing module 738, and speech synthesis processing module 740. Each of these modules has access to one or more, or subsets or supersets of, the following systems or data and models of the digital assistant module 726: knowledge ontology 760, vocabulary index 744, user data 748, task flow model 754, service model 756, and ASR system 758.

[0226] In some examples, using the processing modules, data, and models implemented in the digital assistant module 726, the digital assistant can perform at least some of the following: converting voice input into text; recognizing user intent expressed in natural language input received from the user; proactively eliciting and obtaining the information needed to fully infer the user intent (e.g., by disambiguating words, games, intents, etc.); determining a task flow to satisfy the inferred intent; and executing the task flow to satisfy the inferred intent.

[0227] In some examples, such as Figure 7B As shown, the I / O processing module 728 can... Figure 7A The I / O device 716 in the middle interacts with the user or through Figure 7AThe network communication interface 708 interacts with user equipment (e.g., device 104, 200, 400, 600, or 800) to receive user input (e.g., voice input) and provide a response to the user input (e.g., as voice output). The I / O processing module 728 optionally obtains contextual information associated with the user input from the user equipment along with or shortly after receiving the user input. Contextual information includes user-specific data, vocabulary, and / or preferences associated with the user input. In some examples, this contextual information also includes the software and hardware states of the user equipment at the time the user request is received, and / or information related to the user's surrounding environment at the time the user request is received. In some examples, the I / O processing module 728 also sends follow-up questions related to the user request to the user and receives answers from the user. When a user request is received by the I / O processing module 728 and the user request includes voice input, the I / O processing module 728 forwards the voice input to the STT processing module 730 (or a speech recognizer) for speech-to-text conversion.

[0228] STT processing module 730 includes one or more ASR systems 758. The one or more ASR systems 758 can process speech input received through I / O processing module 728 to produce recognition results. Each ASR system 758 includes a front-end speech preprocessor. The front-end speech preprocessor extracts representative features from the speech input. For example, the front-end speech preprocessor performs a Fourier transform on the speech input to extract spectral features characterizing the speech input as a sequence of representative multidimensional vectors. Additionally, each ASR system 758 includes one or more speech recognition models (e.g., acoustic models and / or language models) and implements one or more speech recognition engines. Examples of speech recognition models include Hidden Markov Models, Gaussian Mixture Models, Deep Neural Network Models, n-gram grammar language models, and other statistical models. Examples of speech recognition engines include engines based on Dynamic Time Warping (VTW) and engines based on Weighted Finite State Transformers (WFST). One or more speech recognition models and one or more speech recognition engines are used to process the representative features extracted by the front-end speech preprocessor to produce intermediate recognition results (e.g., phonemes, phoneme strings, and sub-words) and ultimately produce text recognition results (e.g., words, word strings, or symbol sequences). In some examples, the speech input is processed at least in part by a third-party service or on the user's device (e.g., device 104, 200, 400, 600, or 800) to produce the recognition results. Once the STT processing module 730 produces a recognition result containing a text string (e.g., words, or sequences of words, or sequences of symbols), the recognition result is passed to the natural language processing module 732 for intent inference. In some examples, the STT processing module 730 produces multiple candidate text representations of the speech input. Each candidate text representation is a sequence of words or symbols corresponding to the speech input. In some examples, each candidate text representation is associated with a speech recognition confidence score. Based on the speech recognition confidence score, the STT processing module 730 ranks the candidate text representations and provides the n best (e.g., the n highest-ranked) candidate text representations to the natural language processing module 732 for intent inference, where n is a predetermined integer greater than zero. For example, in one example, only the highest-ranked (n=1) candidate text representation is delivered to the natural language processing module 732 for intent inference. Alternatively, the five highest-ranked (n=5) candidate text representations are passed to the natural language processing module 732 for intent inference.

[0229] Further details regarding speech-to-text processing are described in U.S. Utility Model Patent Application Serial No. 13 / 236,942, entitled "Consolidating Speech Recognition Results," filed on September 20, 2011, the entire disclosure of which is incorporated herein by reference.

[0230] In some examples, the STT processing module 730 includes a vocabulary of recognizable words and / or accesses that vocabulary via the speech-to-letter conversion module 731. Each vocabulary word is associated with one or more candidate pronunciations of a word represented in the speech recognition alphabet. Specifically, the vocabulary of recognizable words includes words associated with multiple candidate pronunciations. For example, the vocabulary includes words associated with... and The candidate pronunciations are associated with the word "tomato". Additionally, vocabulary words are associated with custom candidate pronunciations based on previous speech input from the user. These custom candidate pronunciations are stored in the STT processing module 730 and associated with a specific user via a user profile on the device. In some examples, candidate pronunciations are determined based on the spelling of the word and one or more linguistic and / or phonetic rules. In some examples, candidate pronunciations are generated manually, for example, based on known standard pronunciations.

[0231] In some examples, candidate pronunciations are ranked based on their prevalence. For example, candidate pronunciations... The ranking is higher than This is because the former is a more commonly used pronunciation (e.g., among all users, for users in a specific geographic region, or for any other suitable subset of users). In some examples, candidate pronunciations are ranked based on whether they are custom candidate pronunciations associated with a user. For example, custom candidate pronunciations rank higher than standard candidate pronunciations. This can be used to identify proper nouns with unique pronunciations that deviate from the canonical pronunciation. In some examples, candidate pronunciations are associated with one or more phonological features such as geographic origin, country, or ethnicity. For example, candidate pronunciations... Associated with the United States, and candidate pronunciation The candidate pronunciations are associated with the United Kingdom. Furthermore, the ranking of candidate pronunciations is based on one or more characteristics of the user (e.g., geographic origin, country, ethnicity, etc.) stored in the user profile on the device. For example, it can be determined from the user profile that the user is associated with the United States. Based on the user's association with the United States, candidate pronunciations... (US-related) Comparable candidate pronunciations (Related to the UK) It ranks higher. In some examples, one of the ranked candidate pronunciations can be selected as the predicted pronunciation (e.g., the most likely pronunciation).

[0232] Upon receiving voice input, the STT processing module 730 is used (e.g., using a sound model) to determine the phonemes corresponding to the voice input, and then attempts (e.g., using a language model) to determine the words that match those phonemes. For example, if the STT processing module 730 first identifies a sequence of phonemes corresponding to a portion of the voice input... It can then determine, based on vocabulary index 744, that the sequence corresponds to the word "tomato".

[0233] In some examples, the STT processing module 730 uses fuzzy matching techniques to determine words in a utterance. Therefore, for example, the STT processing module 730 determines phoneme sequences. This corresponds to the word "tomato," even if the specific phoneme sequence is not a candidate phoneme sequence for that word.

[0234] The digital assistant's natural language processing module 732 ("natural language processor") acquires n best candidate text representations ("word sequences" or "symbol sequences") generated by the STT processing module 730 and attempts to associate each candidate text representation with one or more "executable intentions" recognized by the digital assistant. An "executable intention" (or "user intention") represents a task that can be performed by the digital assistant and may have an associated task flow implemented in the task flow model 754. An associated task flow is a series of programmed actions and steps taken by the digital assistant to perform the task. The capabilities of the digital assistant depend on the number and type of task flows implemented and stored in the task flow model 754, or in other words, on the number and type of "executable intentions" recognized by the digital assistant. However, the effectiveness of the digital assistant also depends on its ability to infer the correct "one or more executable intentions" from user requests expressed in natural language.

[0235] In some examples, in addition to the sequence of words or symbols obtained from the STT processing module 730, the natural language processing module 732 also receives, for example, contextual information associated with the user request from the I / O processing module 728. The natural language processing module 732 optionally uses the contextual information to clarify, supplement, and / or further define the information contained in the candidate text representation received from the STT processing module 730. Contextual information includes, for example, user preferences, the hardware and / or software state of the user's device, sensor information collected before, during, or shortly after the user request, previous interactions (e.g., conversations) between the digital assistant and the user, and so on. As described herein, in some examples, the contextual information is dynamic and varies with the time, location, content, and other factors of the conversation.

[0236] In some examples, natural language processing is based on, for example, a knowledge ontology 760. Knowledge ontology 760 is a hierarchical structure containing many nodes, each node representing an "executable intent" or an "attribute" associated with one or more of the "executable intent" or other "attributes." As mentioned above, an "executable intent" represents a task that a digital assistant can perform; that is, the task is "executable" or can be done. An "attribute" represents a parameter associated with a sub-aspect of an executable intent or another attribute. The connections between executable intent nodes and attribute nodes in knowledge ontology 760 define how the parameters represented by the attribute nodes are subordinate to the task represented by the executable intent nodes.

[0237] In some examples, the knowledge ontology 760 consists of executable intent nodes and attribute nodes. Within the knowledge ontology 760, each executable intent node is directly connected to or connected to one or more attribute nodes via one or more intermediate attribute nodes. Similarly, each attribute node is directly connected to or connected to one or more executable intent nodes via one or more intermediate attribute nodes. For example, as... Figure 7C As shown, knowledge ontology 760 includes a "Restaurant Reservation" node (i.e., an executable intent node). The attribute nodes "Restaurant", "Date / Time" (for reservations) and "Party Attendees" are all directly connected to the executable intent node (i.e., the "Restaurant Reservation" node).

[0238] Furthermore, the attribute nodes "Cuisine," "Price Range," "Phone Number," and "Location" are child nodes of the attribute node "Restaurant," and all are connected to the "Restaurant Reservation" node (i.e., the executable intent node) through the intermediate attribute node "Restaurant." For example, ... Figure 7C As shown, knowledge ontology 760 also includes a "Set Reminder" node (i.e., another executable intent node). The attribute nodes "Date / Time" (for setting reminders) and "Topic" (for reminders) are both connected to the "Set Reminder" node. Since the attribute "Date / Time" is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node "Date / Time" is connected to both the "Restaurant Reservation" node and the "Set Reminder" node in knowledge ontology 760.

[0239] An executable intent node, along with its linked attribute nodes, is described as a "domain." In this discussion, each domain is associated with a corresponding executable intent and refers to a set of nodes (and the relationships between these nodes) associated with a particular executable intent. For example, Figure 7CThe knowledge ontology 760 shown includes examples of a restaurant reservation domain 762 and a reminder domain 764 within the knowledge ontology 760. The restaurant reservation domain includes an actionable intent node “Restaurant Reservation”, attribute nodes “Restaurant”, “Date / Time”, and “Participant Size”, and sub-attribute nodes “Cuisine”, “Price Range”, “Phone Number”, and “Location”. The reminder domain 764 includes an actionable intent node “Set Reminder” and attribute nodes “Topic” and “Date / Time”. In some examples, the knowledge ontology 760 consists of multiple domains. Each domain shares one or more attribute nodes with one or more other domains. For example, in addition to the restaurant reservation domain 762 and the reminder domain 764, the “Date / Time” attribute node is associated with many different domains (e.g., itinerary domain, travel booking domain, movie ticket domain, etc.).

[0240] although Figure 7C Two exemplary fields within knowledge ontology 760 are shown, but other fields include, for example, “Find a movie,” “Initiate a phone call,” “Find directions,” “Schedule a meeting,” “Send a message,” and “Provide answers to questions,” “Reading lists,” “Provide navigation instructions,” “Provide instructions for a task,” etc. The “Send a message” field is associated with the “Send a message” executable intent node and further includes attribute nodes such as “one or more recipients,” “message type,” and “message body.” The attribute node “Recipients” is further defined, for example, by sub-attribute nodes such as “Recipient name” and “Message address.”

[0241] In some examples, knowledge ontology 760 includes all domains (and thus executable intents) that a digital assistant can understand and act upon. In some examples, knowledge ontology 760 is modified, such as by adding or removing entire domains or nodes, or by modifying the relationships between nodes within knowledge ontology 760.

[0242] In some examples, nodes associated with multiple related executable intents are clustered under a “superdomain” in knowledge ontology 760. For example, the “Travel” superdomain includes clusters of travel-related attribute nodes and executable intent nodes. Travel-related executable intent nodes include “flight booking,” “hotel booking,” “car rental,” “route planning,” “finding points of interest,” and so on. Executable intent nodes under the same superdomain (e.g., the “Travel” superdomain) have multiple shared attribute nodes. For example, executable intent nodes for “flight booking,” “hotel booking,” “car rental,” “get route,” and “finding points of interest” share one or more of the attribute nodes “starting location,” “destination,” “departure date / time,” “arrival date / time,” and “party size.”

[0243] In some examples, each node in the knowledge ontology 760 is associated with a set of words and / or phrases related to the attribute or executable intent represented by the node. The corresponding set of words and / or phrases associated with each node is called the "vocabulary" associated with the node. The corresponding set of words and / or phrases associated with each node is stored in the vocabulary index 744 associated with the attribute or executable intent represented by the node. For example, returning... Figure 7B The vocabulary associated with nodes of the "restaurant" attribute includes words such as "food," "drinks," "cuisine," "hunger," "eat," "pizza," "fast food," and "meals." Similarly, the vocabulary associated with nodes of the "initiate a phone call" action includes words and phrases such as "call," "make a phone call," "dial," "talk to," "call this number," and "make a phone call." The vocabulary index 744 optionally includes words and phrases from different languages.

[0244] Natural Language Processing (NLP) module 732 receives candidate text representations (e.g., one or more text strings or one or more sequences of symbols) from STT processing module 730 and, for each candidate representation, determines which nodes the words in the candidate text representation relate to. In some examples, if a word or phrase in a candidate text representation is found to be associated with one or more nodes in knowledge ontology 760 (via lexical index 744), the word or phrase "triggers" or "activates" those nodes. Based on the number and / or relative importance of the activated nodes, NLP module 732 selects one executable intent as the task the user intends the digital assistant to perform. In some examples, the domain with the most "triggered" nodes is selected. In some examples, the domain with the highest confidence (e.g., based on the relative importance of its individual triggered nodes) is selected. In some examples, the domain is selected based on a combination of the number and importance of the triggered nodes. In some examples, additional factors, such as whether the digital assistant has previously correctly interpreted similar requests from the user, are also considered in the node selection process.

[0245] User data 748 includes user-specific information such as user-specific vocabulary, user preferences, user address, user's default second language, user's contact list, and other short- or long-term information for each user. In some examples, the natural language processing module 732 uses user-specific information to supplement the information contained in the user input to further refine the user's intent. For example, in response to a user request "Invite my friends to my birthday party," the natural language processing module 732 can access user data 748 to determine who the "friends" are and when and where the "birthday party" will be held, without requiring the user to explicitly provide such information in their request.

[0246] It should be recognized that, in some examples, the natural language processing module 732 is implemented using one or more machine learning agencies (e.g., neural networks). Specifically, the one or more machine learning agencies are configured to receive candidate text representations and contextual information associated with the candidate text representations. Based on the candidate text representations and the associated contextual information, the one or more machine learning agencies are configured to determine an intent confidence score based on a set of candidate executable intents. The natural language processing module 732 can select one or more candidate executable intents from the set of candidate executable intents based on the determined intent confidence score. In some examples, a knowledge ontology (e.g., knowledge ontology 760) is also utilized to select one or more candidate executable intents from the set of candidate executable intents.

[0247] Further details regarding the symbol string-based search of knowledge ontology are described in U.S. Utility Model Patent Application Serial No. 12 / 341,743, entitled “Method and Apparatus for Searching Using An Active Ontology,” filed on December 22, 2008, the entire disclosure of which is incorporated herein by reference.

[0248] In some examples, once the natural language processing module 732 identifies an executable intent (or domain) based on a user request, it generates a structured query to represent the identified executable intent. In some examples, the structured query includes parameters for one or more nodes within the domain of the executable intent, and at least some of these parameters are populated with specific information and requirements specified in the user request. For example, a user says, “Reserve a table at a sushi restaurant for 7 pm.” In this case, the natural language processing module 732 is able to correctly identify the executable intent as “restaurant reservation” based on the user input. According to the knowledge ontology, the structured query for the “restaurant reservation” domain includes parameters such as {cuisine}, {time}, {date}, {number of people}, etc. In some examples, based on voice input and text derived from the voice input using the STT processing module 730, the natural language processing module 732 generates a partially structured query for the restaurant reservation domain, where the partially structured query includes the parameters {cuisine = “sushi”} and {time = “7 pm”}. However, in this example, the user's utterance contains insufficient information to complete a structured query associated with the domain. Therefore, based on the currently available information, no other necessary parameters such as {number of people at the party} and {date} are specified in the structured query. In some examples, the natural language processing module 732 uses the received context information to populate some parameters of the structured query. For example, in some examples, if a user requests a "nearby" sushi restaurant, the natural language processing module 732 uses GPS coordinates from the user's device to populate the {location} parameter in the structured query.

[0249] In some examples, the Natural Language Processing (NLP) module 732 identifies multiple candidate executable intents for each candidate text representation received from the STT processing module 730. Additionally, in some examples, a corresponding structured query (partially or entirely) is generated for each identified candidate executable intent. The NLP module 732 determines an intent confidence score for each candidate executable intent and ranks the candidate executable intents based on the intent confidence scores. In some examples, the NLP module 732 transmits one or more of the generated structured queries (including any completed parameters) to the task flow processing module 736 (“task flow processor”). In some examples, one or more structured queries for the m best (e.g., the m highest-ranked) candidate executable intents are provided to the task flow processing module 736, where m is a predetermined integer greater than zero. In some examples, one or more structured queries for the m best candidate executable intents, along with corresponding one or more candidate text representations, are provided to the task flow processing module 736.

[0250] Further details regarding the inference of user intent based on multiple candidate executable intents determined from multiple candidate text representations of speech input are described in U.S. Utility Model Patent Application 14 / 298,725, filed June 6, 2014, entitled “System and Method for Inferring UserIntent From Speech Inputs,” the entire disclosure of which is incorporated herein by reference.

[0251] Task flow processing module 736 is configured to receive one or more structured queries from natural language processing module 732, complete the structured queries (if necessary), and perform the actions required to "complete" the user's final request. In some examples, the various processes necessary to complete these tasks are provided in task flow model 754. In some examples, task flow model 754 includes processes for obtaining additional information from the user, and task flows for performing actions associated with executable intentions.

[0252] As described above, to complete a structured query, the task flow processing module 736 needs to initiate additional dialogue with the user to obtain additional information and / or clarify potentially ambiguous statements. When such interaction is necessary, the task flow processing module 736 invokes the dialogue flow processing module 734 to participate in the dialogue with the user. In some examples, the dialogue flow processing module 734 determines how (and / or when) to request additional information from the user and receives and processes the user's response. The I / O processing module 728 presents questions to the user and receives answers from the user. In some examples, the dialogue flow processing module 734 presents dialogue output to the user via audible and / or visual output and receives input from the user via verbal or physical (e.g., click) responses. Continuing with the above example, when the task flow processing module 736 invokes the dialogue flow processing module 734 to determine the "party size" and "date" information for a structured query associated with the domain "restaurant reservation," the dialogue flow processing module 734 generates questions such as "How many people in a row?" and "Which day to book?" and presents them to the user. Once a response is received from the user, the dialogue flow processing module 734 either fills the structured query with the missing information or passes the information to the task flow processing module 736 to complete the missing information based on the structured query.

[0253] Once the task flow processing module 736 has completed a structured query for the executable intent, it begins executing the final task associated with the executable intent. Therefore, the task flow processing module 736 executes the steps and instructions in the task flow model based on the specific parameters contained in the structured query. For example, the task flow model for the executable intent "restaurant reservation" includes steps and instructions for contacting the restaurant and actually requesting a reservation for a specific number of people at a specific time for a specific party. For example, using a structured query such as: {restaurant reservation, restaurant = ABC Cafe, date = 3 / 12 / 2012, time = 7 pm, number of people = 5}, the task flow processing module 736 can perform the following steps: (1) log in to ABC Cafe's server or such The restaurant reservation system, (2) inputs date, time and party number information on the website, (3) submits the form, and (4) creates a calendar entry for the reservation in the user's calendar.

[0254] In some examples, task flow processing module 736, with the assistance of service processing module 738 (“service processing module”), completes the task requested in the user input or provides the informational answer requested in the user input. For example, service processing module 738, on behalf of task flow processing module 736, initiates a phone call, sets a calendar entry, invokes a map search, invokes or interacts with other user applications installed on the user's device, and invokes or interacts with third-party services (e.g., restaurant reservation portals, social networking sites, bank portals, etc.). In some examples, the protocols and application programming interfaces (APIs) required for each service are specified through the corresponding service model in service model 756. Service processing module 738 accesses the appropriate service model for a service and, based on the service model, generates a request for that service according to the protocols and APIs required by that service.

[0255] For example, if a restaurant has enabled an online reservation service, it submits a service model that specifies the necessary parameters for making a reservation and the values ​​of those parameters to be sent to the online reservation service's API. When requested by the task flow processing module 736, the service processing module 738 can use the web address stored in the service model to establish a network connection with the online reservation service and send the necessary reservation parameters (e.g., time, date, number of party members) to the online reservation interface in a format appropriate to the online reservation service's API.

[0256] In some examples, the natural language processing module 732, the dialogue flow processing module 734, and the task flow processing module 736 are used together and repeatedly to infer and define the user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response (i.e., output to the user, or complete the task) to satisfy the user's intent. The generated response is a dialogue response to the voice input that at least partially satisfies the user's intent. Additionally, in some examples, the generated response is output as voice output. In these examples, the generated response is sent to the speech synthesis processing module 740 (e.g., a speech synthesizer), which processes the generated response to synthesize the dialogue response in speech form. In other examples, the generated response is data content related to satisfying the user's request in the voice input.

[0257] In an example where the task flow processing module 736 receives multiple structured queries from the natural language processing module 732, the task flow processing module 736 first processes a first structured query of the received structured queries to attempt to complete the first structured query and / or execute one or more tasks or actions represented by the first structured query. In some examples, the first structured query corresponds to the highest-ranking executable intent. In other examples, the first structured query is selected from structured queries received based on a combination of the corresponding speech recognition confidence score and the corresponding intent confidence score. In some examples, if the task flow processing module 736 encounters an error during the processing of the first structured query (e.g., due to the inability to determine necessary parameters), the task flow processing module 736 may continue to select and process a second structured query from the received structured queries that corresponds to a lower-ranking executable intent. For example, the second structured query may be selected based on the speech recognition confidence score of the corresponding candidate text representation, the intent confidence score of the corresponding candidate executable intent, missing necessary parameters in the first structured query, or any combination thereof.

[0258] The speech synthesis processing module 740 is configured to synthesize speech output for presentation to a user. The speech synthesis processing module 740 synthesizes speech output based on text provided by a digital assistant. For example, the generated dialogue response is in the form of a text string. The speech synthesis processing module 740 converts the text string into audible speech output. The speech synthesis processing module 740 uses any appropriate speech synthesis techniques to generate speech output from text, including but not limited to: concatenation synthesis, unit selection synthesis, diphone synthesis, domain-specific synthesis, formant synthesis, articulation synthesis, Hidden Markov Model (HMM) based synthesis, and sine wave synthesis. In some examples, the speech synthesis processing module 740 is configured to synthesize individual words based on phoneme strings corresponding to those words. For example, phoneme strings are associated with words in the generated dialogue response. The phoneme strings are stored in metadata associated with the words. The speech synthesis processing module 740 is configured to directly process the phoneme strings in the metadata to synthesize words in speech form.

[0259] In some examples, instead of using a speech synthesis processing module 740 (or other alternatives), speech synthesis is performed on a remote device (e.g., server system 108), and the synthesized speech is sent to a user device for output to the user. For example, this could occur in some implementations where the output of a digital assistant is generated at the server system. And since server systems typically have greater processing power or more resources than user devices, they are likely to achieve higher quality speech output than client-side synthesis would achieve.

[0260] Additional details regarding digital assistants can be found in U.S. Utility Model Patent Application 12 / 987,982, entitled “Intelligent Automated Assistant,” filed January 10, 2011, and U.S. Utility Model Patent Application 13 / 251,088, entitled “Generating and Processing Task Items That Represent Tasks to Perform,” filed September 30, 2011, the entire disclosure of which is incorporated herein by reference.

[0261] 4. Digital Assistant User Interface

[0262] Figures 8A to 8CT The user interface and digital assistant user interface are shown according to various examples. Figures 8A to 8CT Used to illustrate the processes described below, these processes include Figures 17A to 17F The process in.

[0263] Figure 8A An electronic device 800 is shown. Device 800 is implemented as device 104, device 122, device 200, or device 600. In some examples, device 800 at least partially implements a digital assistant system 700. Figure 8A In one example, device 800 is a smartphone with a display and a touch-sensitive surface. In other examples, device 800 is a different type of device, such as a wearable device (e.g., a smartwatch), a tablet, a laptop, or a desktop computer.

[0264] exist Figure 8A In this process, device 800 displays a user interface 802 on display 801 that is different from the digital assistant (DA) user interface 803, as described below. Figure 8A In the example, user interface 802 is the home screen user interface. In other examples, the user interface is another type of user interface, such as a lock screen user interface or an application-specific user interface, such as a map application user interface, a weather application user interface, a messaging application user interface, a music application user interface, a movie application user interface, etc.

[0265] In some examples, device 800 receives user input when displaying a user interface different from DA user interface 803. Device 800 determines whether the user input meets the criteria for initiating DA. Exemplary user inputs that meet the criteria for initiating DA include: pre-determined type of voice input (e.g., “Hey Siri”); input that selects a virtual or physical button on device 800 (or input that selects such a button for a pre-determined duration); input received at an external device coupled to device 800; user gestures performed on display 801 (e.g., a drag or swipe gesture from a corner of display 801 toward the center of display 801); and inputs indicating movement of device 800 (e.g., raising device 800 to a viewing position).

[0266] In some examples, based on the determination that the user input meets the criteria for initiating DA, device 800 displays DA user interface 803 on top of the user interface. In some examples, displaying DA user interface 803 (or another displayed element) on top of the user interface includes replacing at least a portion of the display of the user interface with the display of DA user interface 803 (or the display of another graphical element). In some examples, based on the determination that the user input does not meet the criteria for initiating DA, device 800 abandons the display of DA user interface 803 and instead performs an action (e.g., updating user interface 802) in response to the user input.

[0267] Figure 8B The diagram shows the DA user interface 803 displayed on top of the user interface 802. In some examples, such as... Figure 8B As shown, the DA user interface 803 includes a DA indicator 804. In some examples, the indicator 804 is displayed in different states to indicate the corresponding state of the DA. DA states include listening state (indicating that the DA is sampling voice input), processing state (indicating that the DA is processing a natural language request), speaking state (indicating that the DA is providing audio and / or text output), and idle state. In some examples, the indicator 804 includes different visualizations indicating different DA states. Figure 8B The indicator 804 shows that after DA is initiated based on the detection that the user input meets the criteria, DA is in a listening state because DA is ready to accept voice input.

[0268] In some examples, the size of the indicator 804 in listening mode varies based on the received natural language input. For example, the indicator 804 expands and contracts in real time according to the amplitude of the received speech input. Figure 8C Indicator 804 shows the status of being in listening mode. Figure 8CIn the process, device 800 receives natural language voice input "How is the weather today?", and indicator 804 expands and collapses in real time according to the voice input.

[0269] Figure 8D Indicator 804 is shown as being in a processing state, for example, indicating that DA is processing the request "How's the weather today?". Figure 8E An indicator 804 is shown in a speaking state, for example, indicating that DA is currently providing audio output "The weather is nice today" in response to a request. Figure 8F Indicator 804 is shown in an idle state. In some examples, user input that selects indicator 804 in an idle state causes DA (and indicator 804) to enter a listening state, for example, by activating one or more microphones to sample audio input.

[0270] In some examples, DA provides audio output in response to a user request, while device 800 provides other audio outputs. In some examples, while providing audio output in response to a user request and other audio outputs simultaneously, DA lowers the volume of the other audio outputs. For example, DA user interface 803 is displayed on top of a user interface that includes currently playing media (e.g., a movie or song). When DA provides audio output in response to a user request, DA lowers the volume of the audio output of the playing media.

[0271] In some examples, the DA user interface 803 includes a DA response capability representation. In some examples, the response capability representation corresponds to the DA's response to received natural language input. For example, Figure 8E A device 800 is shown that displays a response display 805, including weather information, in response to received voice input.

[0272] like Figures 8E to 8F As shown, device 800 displays an indicator 804 on a first portion of display 801 and a response power indicator 805 on a second portion of display 801. A portion of the DA user interface 803 displayed above user interface 802 remains visible on a third portion of display 801 (e.g., not visually obscured). For example, before receiving user input that initiated the digital assistant (e.g., ...). Figure 8A The portion of the user interface 802 that remains visible is displayed in the third portion of the display 801. In some examples, the first, second, and third portions of the display 801 are referred to as the “indicator portion,” the “response portion,” and the “user interface (UI) portion,” respectively.

[0273] In some examples, the UI section is located between the indicator section (display indicator 804) and the response section (display response indicator 805). For example, in Figure 8F In this context, the UI portion includes (or) a display area 8011 (e.g., a rectangular area) between the bottom of the responsive display 805 and the top of the indicator 804, wherein the side edges of the display area 8011 are defined by the side edges of the responsive display 805 (or the display 801). In some examples, the portion of the user interface 802 that remains visible at the UI portion of the display 801 includes one or more user-selectable graphical elements, such as links and / or displays, such as... Figure 8F The home screen in the application displays the functionality.

[0274] In some examples, device 800 displays a response power indicator 805 in a first state. In some examples, the first state includes a compact state, wherein the display size of the response power indicator 805 is smaller (e.g., compared to the expanded response power indicator state described below), and / or the response power indicator 805 displays information in a compact (e.g., overview) format (compared to the expanded response power indicator state). In some examples, device 800 receives user input corresponding to a selection of the response power indicator 805 in the first state, and in response, replaces the display of the response power indicator 805 in the first state with a display of the response power indicator 805 in a second state. In some examples, the second state is an expanded state, wherein the display size of the response power indicator 805 is larger (e.g., compared to the compact state), and / or the response power indicator 805 displays (e.g., compared to the compact state) a greater amount of information / more detailed information. In some examples, device 800 defaults to displaying the response power indicator 805 in the first state, for example, such that device 800 initially displays (…) in the first state. Figures 8E to 8G The response capability is indicated by 805.

[0275] Figures 8E to 8G The response energy display 805 in its first state is shown. As shown, the response energy display 805 provides weather information compactly, for example, by providing the current temperature and state and omitting more detailed weather information (e.g., hourly weather information). Figure 8G The diagram shows device 800 receiving user input 806 (e.g., a tap gesture) corresponding to a selection made by a response enable display 805 in a first state. Although Figures 8G to 8P Overall, the user input corresponding to the selection of the response capability representation is touch input. However, in other examples, the user input corresponding to the selection of the response capability representation is another type of input, such as voice input (e.g., “Show me more information”) or peripheral device input (e.g., input from a mouse or touchpad). Figure 8HThe diagram illustrates that in response to receiving user input 806, device 800 replaces the display of response display 805 in its first state with a display of response display 805 in its second state. As shown, the response display 805 in the second state now includes more detailed weather information.

[0276] In some examples, when the response capability indicator 805 is displayed in the second state, device 800 receives user input requesting that the response capability indicator 805 be displayed in the first state. In some examples, in response to receiving user input, device 800 replaces the display of the response capability indicator 805 in the second state with the display of the response capability indicator 805 in the first state. For example, in Figure 8H In the DA user interface 803, selectable elements (e.g., a back button) 807 are included. User input that selects the selectable element 807 causes the device 800 to return to the previous state. Figure 8F The display.

[0277] In some examples, when the response power indicator 805 is displayed in a second state, device 800 receives user input corresponding to a selection of the response power indicator 805. In response to receiving the user input, device 800 displays a user interface corresponding to the application on the response power indicator 805. For example, Figure 8I The device 800 is shown receiving user input 808 (e.g., a tap gesture) corresponding to a selection on a response enable display 805. Figure 8J The device 800 displays a user interface 809 for a weather application in response to receiving user input 808.

[0278] In some examples, when displaying the application's user interface, device 800 displays a selectable DA indicator. For example, Figure 8J An optional DA indicator 810 is shown. In some examples, when displaying the user interface of an application, the device 800 additionally or alternatively displays an indicator 804, such as an idle indicator, on a first portion of the display 801.

[0279] In some examples, when displaying the application's user interface, device 800 receives user input selecting a selectable DA indicator. In some examples, in response to receiving user input, device 800 replaces the display of the application's user interface with a DA user interface 803. In some examples, DA user interface 803 is the DA user interface displayed immediately preceding the display of the application's user interface. For example, Figure 8K The device 800 is shown receiving user input 811 (e.g., a tap gesture) to select the DA indicator 810. Figure 8LIt is shown that in response to receiving user input 811, device 800 replaces the display of weather application user interface 809 with the display of DA user interface 803.

[0280] Figure 8G User input 806 corresponds to a selection of a first portion of the response display 805. In some examples, when device 800 displays the response display 805 in a first state (e.g., a compact state), device 800 receives user input corresponding to a selection of a second portion of the response display 805. In some examples, the first portion of the response display 805 (e.g., the bottom portion) includes information intended to answer a user request. In some examples, the second portion of the response display 805 (e.g., the top portion) includes a marker indicating the category of the response display 805 and / or associated text. Exemplary categories of the response display include weather, stocks, knowledge, calculator, messages, music, maps, etc. These categories may correspond to categories of services provided by DA. In some examples, the first portion of the response display 805 occupies a larger display area than the second portion of the response display 805.

[0281] In some examples, in response to receiving user input corresponding to a selection of the second portion of the response capability representation 805, device 800 displays a user interface corresponding to the application of the response capability representation 805 (e.g., without displaying the response capability representation 805 in its second state). For example, Figure 8M The device 800 is shown receiving a second portion of a response enable representation 805 (e.g., a tap gesture) to select a first state for display. Figure 8N A user interface 809 for a weather application is shown in response to receiving user input 812 (e.g., without displaying the expanded weather display 805). This allows the user to provide input to select different portions of the weather display 805, to expand the weather display 805, or to display an application corresponding to the weather display 805, such as... Figures 8G to 8H and Figures 8M to 8N As shown.

[0282] Figure 8N It is also shown that when displaying user interface 809, device 800 displays a selectable DA indicator 810. User input selecting DA indicator 810 causes device 800 to return to [the previous state]. Figure 8M The display, for example, is similar to that of... Figures 8K to 8L Examples are shown. In some examples, when displaying user interface 809, device 800 displays DA indicator 804 (e.g., DA indicator in an idle state) on a first portion of display 801.

[0283] In some examples, for certain types of response capability representations, user input corresponding to a selection of any part of the response capability representation causes device 800 to display a user interface for the application corresponding to the response capability representation. In some examples, this is because the response capability representation cannot be displayed in a more granular manner (e.g., in a second state). For example, there may not be additional information that the DA can provide in response to natural language input. For example, consider the natural language input "What is 5 times 6?". Figure 8O A DA user interface 803 is shown that displays in response to natural language input. The DA user interface 803 includes a response capability representation 813 displayed in a first state. The response capability representation 813 includes the answer "5×6=30", but no additional information that DA can provide is present. Figure 8O It is also shown that device 800 receives user input 814 (e.g., tap gesture) in the first part of selection response indication 813. Figure 8P It is shown that in response to receiving user input 814, device 800 displays a user interface 815 corresponding to an application, such as a calculator application user interface, for an application that responds to a power indication 813.

[0284] In some examples, the response enablement includes selectable elements, such as selectable text indicating a link. Figure 8Q A DA user interface 803 is shown in response to the natural language input "Tell me more about Famous Band". The DA user interface 803 includes a response capability representation 816. The response capability representation 816 includes information about "Famous Band" and a selectable element 817 corresponding to member #1 of "Famous Band". In some examples, the device 800 receives user input corresponding to a selection of the selectable element, and in response, displays a capability representation (a second response capability representation) corresponding to the selectable element above the response capability representation. Figure 8R The device 800 is shown receiving user input 818 (e.g., a tap gesture) to select selectable element 817. Figure 8S In response to receiving user input 818, device 800 displays a second response display 819, including information about member #1, on top of response display display 816 to form a response display display overlay.

[0285] In some examples, when a second response power indicator is displayed above the response power indicator, device 800 visually masks the user interface at or a portion of a third portion of display 801 (e.g., a portion that does not display any response power indicator or indicator 804). In some examples, visually masking the user interface includes darkening or blurring the user interface. Figure 8SThe user interface 802 of the device 800 is visually obscured in the third part of the display 801 when the second response power indicator 819 is displayed above the response power indicator 816.

[0286] Figure 8S This illustration shows that when a second response power indicator 819 is displayed over a portion of a response power indicator 816, that portion remains visible. In other examples, the second response power indicator 819 replaces the display of the response power indicator 816, such that any portion of the response power indicator 816 is not visible.

[0287] Figure 8T The device 800 is shown receiving user input 820 (e.g., a tap gesture) to select a selectable element 821 (“Detroit”) in a second response enable representation 819. Figure 8U In response to receiving user input 820, device 800 displays a third response display 822 above the second response display 819. The third response display 822 includes information about Detroit (the birthplace of member #1). Figure 8U The user interface 802 is shown to remain visually obscured in the third part of the display 801.

[0288] Figure 8U It is also shown that although there are three response power representations (e.g., 816, 819, and 822) in the response power representation stack, device 800 only indicates two response power representations in the stack. For example, a portion of the third response power representation 822 and the second response power representation 819 is shown, but no portion of response power representation 816 is shown. Therefore, in some examples, when more than two response power representations are stacked, device 800 only visually indicates the presence of two response power representations in the stack. In other examples, when response power representations are stacked, device 800 only visually indicates a single response power representation in the stack (e.g., such that the display of the next response power representation completely replaces the display of the previous response power representation).

[0289] Figures 8V to 8Y This illustrates how a user provides input to return a previous response energy representation in the stack. Specifically, in Figure 8V In the process, device 800 receives a request on third response enablement 822 to return to user input 823 (e.g., a swipe gesture) on second response enablement 819. Figure 8W It is shown that in response to receiving user input 823, device 800 stops displaying the third response capability representation 822 and fully displays the second response capability representation 819. Device 800 also displays (e.g., exposes) a portion of the response capability representation 816. Figure 8XThe device 800 is shown receiving a request on a second response enable display 819 and returning to a response enable display 816 via user input 824 (e.g., a swipe gesture). Figure 8Y The diagram shows that in response to receiving user input 824, device 800 stops displaying the second response capability representation 819 and displays the full response capability representation 816. In some examples, device 800 receives input to display the next response capability representation in the overlay (e.g., a swipe gesture in the opposite direction) and, in response, displays the next response capability representation in the overlay in a manner similar to that described above. In other examples, navigation within the response capability representations in the overlay relies on other input methods (e.g., the user's selection of a displayed "back" or "next" button), in a manner similar to that described above.

[0290] Figure 8Y It is also shown that the user interface 802 is no longer visually obscured at the third portion of the display 801. Therefore, in some examples, such as... Figures 8Q to 8Y As shown, the user interface 802 is visually masked when the response power indicators are stacked, but not visually masked when the power indicators are not stacked. For example, when the initial response power indicator 816 is not displayed (or only partially displayed), the user interface 802 is visually masked, while when the initial response power indicator 816 is fully displayed, the user interface 802 is not visually masked.

[0291] In some examples, the user interface (e.g., the user interface displayed on top of the DA user interface 803) includes an input field that occupies a fourth portion (e.g., an "input field portion") of the display 801. The input field includes an area where the user can provide natural language input. In some examples, the input field corresponds to an application, such as a messaging application, email application, note-taking application, reminder application, calendar application, etc. Figure 8Z A user interface 825 for a messaging application is shown, which includes an input field 826 occupying the fourth portion of the display 801.

[0292] Figure 8AA A DA user interface 803 is shown displayed on top of a user interface 825. The device 800 displays the user interface 803 in response to a natural language input, “What is the name of this song?” The DA user interface 803 includes an indicator 804 on a first portion of a display 801 and a response indication 827 (indicating the song recognized by the DA) on a second portion of the display 801.

[0293] In some examples, device 800 receives user input corresponding to a displacement of the response indication from a first portion of display 801 to a fourth portion of display 801. In response to receiving the user input, device 800 replaces the display of the response indication at the first portion of display 801 with the display of the response indication in the input area field. For example, Figures 8AB to 8AD The device 800 is shown to receive user input 828 that moves a response enable display 827 from a first portion of display 801 to an input field 826. User input 828 corresponds to a drag gesture from the first portion of display 801 to a fourth portion of display 801 and ends with a lift-off event (e.g., a finger lift-off event) at the display in input field 826.

[0294] In some examples, such as Figures 8AB to 8AD As shown, upon receiving user input 828, device 800 continuously shifts the response indication 827 from a first portion of display 801 to a fourth portion of display 801. For example, while the response indication 827 is shifted, device 800 displays the response indication 827 at a position corresponding to the current display contact position of user input 828. In some examples, while the response indication 827 is shifted, the display size of the response indication 827 decreases, for example, causing the response indication 827 to retract under the user's finger (or other input device) while being shifted. Figures 8AB to 8AD It is also shown that when the continuous shift response indicator 827 is displayed, the indicator 804 stops displaying.

[0295] Figure 8AD A response enable indicator 827 is now displayed in the input field 826 of the messaging application. Figure 8AE The device 800 is shown receiving user input 830 (e.g., a tap gesture) corresponding to a selection of a send message enablement display 829. Figure 8AF The diagram illustrates that in response to receiving user input 830, device 800 sends a response enable representation 827 as a message. Thus, a user can send a response enable representation in communication (e.g., text message, email) by providing input (e.g., drag and drop) to shift the response enable representation into the appropriate input field. In other examples, a user can include response enable representations in notes, calendar entries, word processing documents, reminder entries, etc., in a similar manner.

[0296] In some examples, user input corresponding to the displacement of the response power representation from the first portion of display 801 to the fourth portion of display 801 (of the display input field) corresponds to a selection of the power representation. In some examples, the power representation is either a shared power representation (e.g., sharing the response power representation in communication) or a saved power representation (e.g., saving the power representation in a note or reminder entry). For example, when device 800 displays DA user interface 803 on top of a user interface including an input field, the response power representation includes either a shared power representation or a saved power representation, depending on the type of user interface. For example, when the user interface corresponds to a communication application (e.g., messaging or email), the response power representation includes a shared power representation, and when the user interface corresponds to another type of application with an input field (e.g., word processing, reminders, calendar, notes), the response power representation includes a saved power representation. User input selecting a shared or saved power representation causes device 800 to replace the display of the response power representation at the first portion of display 801 with the display of the response power representation in the input field in a manner similar to that described above. For example, when a response indication is displayed in the input area, device 800 stops displaying indicator 804.

[0297] In some examples, the user interface (e.g., the user interface displayed on top of the DA user interface 803) includes a desktop applet area that occupies a fifth portion (e.g., a "desktop applet portion") of the display 801. Figure 8AG In this example, device 800 is a tablet device. Device 800 displays a user interface 831 on display 801, which includes a desktop applet area 832 occupying a fifth portion of display 801. Device 800 also displays a DA user interface 803 on top of user interface 831. DA user interface 803 is displayed in response to natural language input “track flight 23”. DA user interface 803 includes an indicator 804 displayed on a first portion of display 801 and a response display 833 (including information about flight 23) displayed on a second portion of display 801.

[0298] In some examples, device 800 receives user input corresponding to a displacement of the response indication from a first portion of display 801 to a fifth portion of display 801. In some examples, in response to receiving user input, device 800 replaces the display of the response indication at the first portion of the display with the display of the response indication in a desktop applet area. For example, Figures 8AH to 8AJThe device 800 is shown receiving user input 834 that shifts a response power indicator 833 from a first portion of display 801 to a desktop applet area 832. User input 834 corresponds to a drag gesture from the first portion of display 801 to a fifth portion of display 801 and ends with a lift-off event at the display of desktop applet area 832. In some examples, shifting the response power indicator 833 from the first portion of display 801 to the fifth portion of display 801 is performed in a manner similar to the shift of the response power indicator 827 described above. For example, when the response power indicator 833 is shifted, the indicator 804 stops displaying.

[0299] Figure 8AJ The responsive power indicator 833 is now displayed in the desktop applet area 832 along with the displayed calendar and music desktop applets. Thus, the user can provide input to the desktop applet area 832 by shifting the responsive power indicator 833 (e.g., dragging and dropping) to add the responsive power indicator 833 as a desktop applet.

[0300] In some examples, user input corresponding to the displacement of the response power indicator from the first portion to the fifth portion of display 801 corresponds to a selection of the power indicator. In some examples, the power indicator is a "Show in Desktop Applets" power indicator. For example, when device 800 displays DA user interface 803 on top of a user interface that includes a desktop applet area, the response power indicator includes a "Show in Desktop Applets" power indicator. User input selecting the "Show in Desktop Applets" power indicator causes device 800 to replace the display of the response power indicator at the first portion of display 801 with the display of the response power indicator in the desktop applet area in a manner similar to that described above.

[0301] In some examples, the response indicator corresponds to an event, and device 800 determines the completion of the event. In some examples, in response to determining the completion of the event, device 800 stops displaying the response indicator in the applet area (e.g., for a predetermined duration after determining completion). For example, response indicator 833 corresponds to a flight, and in response to determining that the flight has completed (e.g., landed), device 800 stops displaying response indicator 833 in applet area 832. As another example, the response indicator corresponds to a sporting event, and in response to determining that the sporting event has ended, device 800 stops displaying the response indicator in the applet area.

[0302] Figure 8AK to Figure 8AN Various exemplary types of response energy representations are shown. Specifically, Figure 8AKA compact response display 835 is shown in response to a natural language request, “How old is celebrity X?”. The compact response display 835 includes a direct answer to the request (e.g., “30 years old”) without including further information (e.g., additional information about celebrity X). In some examples, all compact response displays have the same maximum size, such that the compact response display occupies only a (relatively small) area of ​​display 801. Figure 8AL A detailed response power display 836 is shown, displayed in response to a natural language request to “provide statistics about team #1”. The detailed response power display 836 includes detailed information about team #1 (e.g., various statistics) and has a larger display size than the compact response power display 835. Figure 8AM A list response display 837 is shown that displays in response to the natural language "Show me a list of nearby restaurants". The list response display 837 includes a list of options (e.g., restaurants) and has a larger display size than the compact response display 835. Figure 8AN A disambiguation response display 838 is shown in response to a natural language request “Call Neal”. The disambiguation response display includes selectable disambiguation options: (1) Neal Ellis, (2) Neal Smith, and (3) Neal Johnson. Device 800 also provides an audio output asking “Which Neal?”

[0303] like Figure 8AK to Figure 8AN As shown, the type of response enablement displayed (e.g., compact, detailed, list, disambiguation) depends on the content of the natural language input and / or the DA's interpretation of the natural language input. In some examples, enablement construction rules specify a particular type of response enablement to be displayed for a particular type of natural language input. In some examples, the construction rule specifies that a compact response enablement should be displayed by default, for example, causing device 800 to display a compact response enablement in response to natural language input that can be adequately answered by a compact response enablement. In some examples, when the response enablement can be displayed in different states (e.g., a first compact state and a second expanded (detailed) state), the construction rule specifies that the response enablement should initially be displayed as a compact enablement. As relative to... Figures 8G to 8H As stated above, in response to receiving appropriate user input, a more detailed version of the compact energy representation can be displayed. It should be understood that some natural language inputs (e.g., “Give me statistics about team #1” and “Show me a list of nearby restaurants”) cannot be adequately answered with a compact energy representation (or may not be expected to be answered with a compact energy representation). Therefore, construction rules can specify a particular type of energy representation (e.g., a detailed list) to be displayed for such inputs.

[0304] In some examples, the DA determines multiple results corresponding to the received natural language input. In some examples, device 800 displays a response capability representation that includes a single result from these multiple results. In some examples, when displaying the response capability representation, the other results from these multiple results are not displayed. For example, consider the natural language input "nearest coffee". The DA determines multiple results corresponding to the input (multiple nearby coffee shops). Figure 8AO A response power representation 839 (e.g., a compact power representation) is shown that displays in response to input. The response power representation 839 includes a single result among the multiple results (the coffee shop closest to the location of device 800). Device 800 also provides voice output "This is the nearest coffee shop". Thus, for a natural language request involving multiple results, DA can initially provide a single result, such as the most relevant result.

[0305] In some examples, after providing a single result (e.g., displaying a response energy representation 839), DA provides the next of those multiple results. For example, in Figure 8 AP In this context, device 800 replaces response indication 839 with response indication 840, which includes the second nearest coffee shop. Device 800 also provides voice output, "This is the second nearest coffee shop." In some examples, in response to receiving user input that rejects a single result (e.g., "I don't want to go to that coffee shop") or user input instructing for the next result, device 800... Figure 8AO Transform into Figure 8 AP In some examples, after displaying power indicator 839 and / or for a predetermined duration after providing the voice output "This is the nearest coffee shop," if no user input to select power indicator 839 is received, device 800 will... Figure 8AO Transform into Figure 8 AP In this way, device 800 can sequentially provide results for natural language inputs involving multiple outcomes.

[0306] In some examples, the response capability representation includes one or more task capability representations. User input (e.g., a tap gesture) selects a task capability representation, causing device 800 to perform the corresponding task. For example, in Figure 8ANIn this context, response capability representation 838 includes task capability representations 841, 842, and 843. A user's selection of task capability representation 841 causes device 800 to initiate a phone call to Neal Ellis, and a user's selection of task capability representation 842 causes device 800 to initiate a phone call to Neal Smith, and so on. Similarly, response capability representation 839 includes task capability representation 844, and response capability representation 840 includes task capability representation 845. A user's selection of task capability representation 844 causes device 800 to launch a map application displaying directions to the nearest coffee shop, while a user's selection of task capability representation 845 causes device 800 to launch a map application displaying directions to the second nearest coffee shop.

[0307] In some examples, device 800 displays multiple response capability representations simultaneously in response to natural language input. In some examples, each of the multiple response capability representations corresponds to a different possible domain of the natural language input. In some examples, device 800 displays the multiple response capability representations when the natural language input is determined to be ambiguous (e.g., corresponding to multiple domains).

[0308] For example, consider the natural language input "Beyoncé". Figure 8AQ Response capability representations 846, 847, and 848, displayed simultaneously in response to natural language input, are shown. Response capability representations 846, 847, and 848 correspond to the news domain (e.g., a user requests news about Beyoncé), the music domain (e.g., a user requests to play music by Beyoncé), and the knowledge domain (e.g., a user requests information about Beyoncé), respectively. In some examples, corresponding user input to the selection of response capability representations 846, 847, and 848 causes device 800 to perform the corresponding action. For example, selecting response capability representation 846 causes the display of a detailed response capability representation including news about Beyoncé, selecting response capability representation 847 causes device 800 to launch a music application including songs by Beyoncé, and selecting response capability representation 848 causes the display of a detailed response capability representation including information about Beyoncé.

[0309] In some examples, the response capability representation includes an editable text field, which comprises text determined based on natural language input. For example, Figure 8ARA response capability representation 849 is shown that displays in response to the natural language voice input “text mom I'm home”. The response capability representation 849 includes an editable text field 850 containing the text “I'm hole”, for example, because DA incorrectly recognized “I'm home” as “I'm hole”. The response capability representation also includes a task capability representation 851. User input selecting the task capability representation 851 causes device 800 to send a text message.

[0310] In some examples, device 800 receives user input corresponding to a selection of an editable text field, and in response, displays a keyboard while displaying a response enable display. For example, Figure 8AS The device 800 is shown receiving user input 852 (e.g., a tap gesture) to select an editable text field 850. Figure 8AT It is shown that in response to receiving user input 852, device 800 displays keyboard 853 while displaying a response enable display 849. As shown, device 800 displays keyboard 853 on top of user interface 802 (e.g., a user interface displayed on top of DA user interface 803). Although Figure 8AT to Figure 8AV It is shown that when the response power indicator and keyboard are displayed on top of the user interface 802, a portion of the user interface 802 is not visually obscured, but in other examples, at least a portion of the user interface 802 (e.g., the portion of the display 801 that does not display the keyboard or response power indicator 849) is visually obscured.

[0311] In some examples, device 800 receives one or more keyboard inputs and, in response, updates the text in the editable text field based on those keyboard inputs. For example, Figure 8AU This indicates that device 800 has received keyboard input correcting "hole" to "home". Device 800 displays the corrected text in the editable text field 850 of the response display 849.

[0312] In other examples, device 800 receives voice input requesting editing of text displayed in an editable text field. In response to receiving the voice input, device 800 updates the text in the editable text field based on the voice input. For example, in... Figure 8AR In this context, the user can provide voice input such as "No, I said I'm home" to cause the device 800 to update the text in the editable text field 850 accordingly.

[0313] In some examples, after updating the text in the editable text field, device 800 receives user input requesting the execution of a task associated with the display. In response to receiving the user input, device 800 performs the requested task based on the updated text. For example, Figure 8AV The diagram shows that after “hole” is edited to “home”, device 800 receives user input 854 (e.g., a tap gesture) corresponding to the selection of task enablement representation 851. Figure 8AW The device 800 is shown sending the message "I'm home" to the user's mother in response to receiving user input 854. The device 800 also displays a symbol 855 indicating the completion of the task. Figure 8AW Further illustrated is that in response to receiving user input 854, device 800 stops displaying keyboard 853 to display (e.g., expose) a portion of user interface 802, and device 800 displays indicator 804.

[0314] In this way, the user can edit the text included in the response capability representation (e.g., if the DA incorrectly recognizes the user's voice input) and cause the DA to perform an action using the correct text. Although Figures 8AR to 8AW An example of editing and sending text messages is shown, but in other examples, users can edit and save (or send) notes, calendar entries, reminder entries, email entries, etc., in a similar way.

[0315] In some examples, device 800 receives user input to cancel DA. In some examples, canceling DA includes stopping the display of the DA user interface 803. The following is relative to... Figures 10A to 10V DA cancellation is discussed in more detail. In some examples, after DA cancellation, device 800 receives user input to re-initiate DA (e.g., user input that meets the criteria used to initiate DA). In some examples, based on the received user input to re-initiate DA, device 800 displays a DA user interface that includes the same response capability representation, e.g., the response capability representation displayed before DA cancellation.

[0316] In some examples, device 800 displays the same response capability representation based on determining that the same response capability representation corresponds to a response to a received natural language input (e.g., input intended for re-initiated DA). For example, Figure 8AX A DA user interface 803 including a response power indicator 856 is shown. The device 800 displays the DA user interface 803 in response to the natural language input “How’s the weather?”. Figure 8AY The device 800 is shown receiving user input 857 to cancel DA, for example, a tap gesture corresponding to a selection of user interface 802. Figure 8AZIt is shown that in response to receiving user input 857, device 800 cancels DA, for example, stops displaying DA user interface 803. Figure 8BA The device 800 has received input to re-initiate DA and is currently receiving natural language input "Will it be windy?". Figure 8BB A device 800 is shown displaying a DA user interface 803, which includes the same response power indication 856 and provides the voice output “Yes, it will be windy.” For example, the DA has determined that the same response power indication 856 corresponds to the natural language inputs “How’s the weather?” and “Will it be windy?”. Thus, if a previous response power indication is relevant to the current natural language request, the previous response power indication can be included in the subsequently initiated DA user interface.

[0317] In some examples, device 800 displays the same response capability representation based on whether user input to re-initiate DA is received within a predetermined duration after DA elimination. For example, Figure 8BC The DA user interface 803 is shown in response to the natural language input "What is 3 multiplied by 5?". The DA user interface 803 includes a responsive power display 858. Figure 8BD This shows that DA was eliminated immediately. Figure 8BE It is shown that within a predetermined duration (e.g., 5 seconds) of the first time, device 800 has received user input that re-initiates DA. For example, device 800 has received any type of input that meets the criteria for initiating DA as described above, but has not yet received another natural language input that includes a different request for DA. Therefore, in Figure 8BE In this device 800, a DA user interface 803 is displayed, which includes the same response power indicator 858 and a listening indicator 804. Thus, if the user quickly re-initiates DA, for example, due to the user previously accidentally canceling DA, the previous response power indicator can be included in the subsequently initiated DA user interface.

[0318] Figure 8BF A device 800 is shown in landscape orientation. In some examples, because device 800 is landscape oriented, the user interface is displayed in landscape mode. For example, Figure 8BF A messaging application user interface 859 is shown displayed in landscape mode. In some examples, device 800 displays a DA user interface 803 in landscape mode on top of the user interface in landscape mode. For example, Figure 8BG The DA user interface 803 is shown in landscape mode on top of user interface 859. It should be understood that users can provide one or more inputs to interact with the DA user interface 803 in landscape mode in a manner consistent with the techniques discussed herein.

[0319] In some examples, certain user interfaces do not have a landscape mode. For instance, the user interface is displayed the same regardless of whether the device 800 is in landscape or portrait orientation. Exemplary user interfaces that do not have a landscape mode include the home screen user interface and the lock screen user interface. Figure 8BH The home screen user interface 860 displayed when the device 800 is in landscape orientation is shown (it does not have a landscape mode).

[0320] In some examples, when device 800 is in landscape orientation, device 800 displays DA user interface 803 on top of a user interface that does not have a landscape mode. In some examples, when displaying DA user interface 803 (in landscape mode) on top of a user interface that does not have a landscape mode, device 800 visually masks the user interface, for example, visually masking portions of the user interface that are not displayed above the user interface. Figure 8BI A landscape-oriented device 800 displays a DA user interface 803 in landscape mode above a home screen user interface 860. The home screen user interface 860 is displayed in portrait mode (even though the device 800 is landscape-oriented) because the home screen user interface 860 does not have a landscape mode. As shown, the device 800 visually masks the home screen user interface 860. In this way, the device 800 avoids simultaneously displaying the landscape-mode DA user interface 803 and the not-visually-masked portrait-mode user interface (e.g., the home screen user interface 860), which could provide a confusing user visual experience.

[0321] In some examples, when device 800 displays DA user interface 803 on top of a predefined type of user interface, device 800 visually masks the predefined type of user interface. Exemplary predefined type of user interface includes a lock screen user interface. Figure 8BJ A device 800 displaying an exemplary lock screen user interface 861 is shown. Figure 8BK A device 800 is shown displaying a DA user interface 803 on top of a lock screen user interface 861. As shown, the portion of the device 800 above the lock screen user interface 861 where the DA user interface 803 is not displayed visually obscures the lock screen user interface 861.

[0322] In some examples, the DA user interface 803 includes a dialogue enablement. In some examples, the dialogue enablement includes a dialogue generated by the DA in response to received natural language input. In some examples, the dialogue enablement is displayed in a sixth portion (e.g., the "conversation portion") of display 801, located between the first portion of display 801 (displaying the DA indicator 804) and the second portion of display 801 (displaying the response enablement). For example, Figure 8BL Dialogue enablement representation 862 is shown, which includes a dialogue generated by the DA in response to the natural language input "play Frozen", and will be discussed further below. Figure 8BM Dialogue enablement representation 863 is shown, which includes a dialogue generated by DA in response to the natural language input “Delete meeting #1”, and will be discussed further below. Figure 8BM The device 800 also shows a dialog enable representation 863 displayed in the sixth part of the display 801, the sixth part being between the display of indicator 804 and the display of response enable representation 864.

[0323] In some examples, DA determines multiple alternative disambiguation options for the received natural language input. In some examples, the dialogue enabled by the dialogue includes these multiple alternative disambiguation options. In some examples, the multiple disambiguation options are determined based on DA's determination that the natural language input is ambiguous. The ambiguous natural language input corresponds to multiple possible executable intentions, for example, each executable intention having a relatively high (and / or equal) confidence score. For example, consider... Figure 8BL The natural language input is "Play Frozen". DA determines two disambiguation options: option 865 "Play Movie" (e.g., the user wants to play the movie "Frozen") and option 866 "Play Music" (e.g., the user wants to play music from the movie "Frozen"). Dialogue enablement representation 862 includes options 865 and 866, where the user's selection of option 865 causes device 800 to play the movie "Frozen", and the user's selection of option 866 causes device 800 to play music from the movie "Frozen". For example, consider... Figure 8BM The natural language input is "Delete Meeting #1", where "Meeting #1" is a duplicate meeting. DA determines two selectable disambiguation options: option 867 "Delete Single" (e.g., the user wants to delete a single instance in Meeting #1), and option 868 "Delete All" (e.g., the user wants to delete all instances in Meeting #1). Dialogue enablement representation 863 includes options 867 and 868, and a cancel option 869.

[0324] In some examples, the DA determines the additional information required to perform the task based on the received natural language input. In some examples, the dialogue capability representation includes one or more optional options recommended by the DA for the required additional information. For example, the DA may have determined the domain of the received natural language input but cannot determine the parameters required to complete the task associated with that domain. For example, consider the natural language input "call". The DA determines that the domain of the natural language input is the telephone call domain (e.g., the domain associated with the executable intent to make a telephone call), but cannot determine the parameters (i.e., who to call). In some examples, the DA therefore determines one or more optional options as recommendations for the parameters. For example, device 800 displays optional options in the dialogue capability representation corresponding to the contact the user most frequently calls. The user's selection of any of the optional options causes device 800 to call the corresponding contact.

[0325] In some examples, DA determines the primary user intent based on the received natural language input, and determines alternative user intents based on the received natural language input. In some examples, the primary intent is the highest-ranking executable intent, and the alternative user intent is the second-ranking executable intent. In some examples, the displayed response enablement corresponds to the primary user intent, while the simultaneously displayed dialog enablement includes selectable options corresponding to the alternative user intents. For example, Figure 8BN A DA user interface 803 is shown in response to the natural language input "directions to Phil". The DA determines a primary user intent, i.e., the user wants directions to "Phil Coffee", and a secondary user intent, i.e., the user wants directions to the home of a contact named "Phil". The DA user interface 803 includes a response enablement representation 870 corresponding to the primary user intent and a dialogue enablement representation 871. Dialogue 872 of dialogue enablement representation 871 corresponds to the secondary user intent. User input selecting dialogue 872 causes device 800 to obtain directions to the home of a contact named "Phil", while user input selecting response enablement representation 870 causes device 800 to obtain directions to "Phil Coffee".

[0326] In some examples, the dialog capability representation is displayed in a first state. In other examples, the first state is an initial state, such as a description of the dialog capability representation that is initially displayed before user input is received to interact with the dialog capability representation. Figure 8BO A DA user interface 803 is shown, including a dialog power indicator 873 displayed in its initial state. Device 800 displays the DA user interface 803 in response to natural language input “What’s the weather like?”. The dialog power indicator 873 includes at least a portion of a dialogue generated by the DA in response to the input, such as “The current temperature is 70 degrees, and it’s windy…”. The following will be relative to… Figures 11 to 16 Let's discuss further details regarding whether to display the dialog generated by DA.

[0327] In some examples, device 800 receives user input corresponding to a selection of a dialog display shown in a first state. In response to receiving the user input, device 800 replaces the display of the dialog display in the first state with a display of the dialog display in a second state. In some examples, the second state is an expanded state, wherein the display size of the dialog display in the expanded state is larger than the display size of the dialog display in the initial state, and / or wherein the amount of content displayed by the dialog display in the expanded state is larger than that of the dialog display in the initial state. Figure 8BP The device 800 is shown receiving user input 874 (e.g., a drag gesture) corresponding to a selection of a dialog enable representation 873 displayed in its initial state. Figure 8BQ This illustrates that, in response to receiving user input 874 (or a portion thereof), device 800 replaces the display of the initial dialog display 873 with the display of the expanded dialog display 873. As shown in the figure, with... Figure 8BP Compared to the dialogue capabilities shown in the text, Figure 8BQ The dialog display in the 873 has a larger display size and includes a larger amount of text.

[0328] In some examples, the display size of the dialog enablement (in the second state) is proportional to the length of the user input that causes the dialog enablement to be displayed in the second state. For example, in Figure 8BP to Figure 8BQ In this context, the display size of the dialogue display 873 is increased proportionally to the length (e.g., physical distance) of the drag gesture 874. This allows the user to provide continuous drag gestures, expanding the responsive display 873 according to the drag length of the gesture. Furthermore, although... Figures 8BO to 8BQ The device 800 initially displayed as shown Figure 8BO The dialogue is shown as 873, then expanded. Figure 8BQ The dialog functionality is indicated in the text, but in other examples, device 800 initially displays as... Figure 8BQ The dialog capability representation 873 is shown. Therefore, in some examples, the device 800 initially displays the dialog capability representation such that the dialog capability representation displays the maximum amount of content, for example, without masking (overwriting) any concurrently displayed response capability representations.

[0329] In some examples, the display of the dialog capability representation masks the display of the simultaneously displayed response capability representation. Specifically, in some examples, the display of the dialog capability representation in a second (e.g., expanded) state occupies at least a portion of the second portion of the display 801 (where the response capability representation is displayed). In some examples, displaying the dialog capability representation in the second state also includes displaying the dialog capability representation over at least a portion of the response capability representation. For example, Figure 8BQ The drag gesture 874 is shown to continue. Figure 8BR It is shown that in response to receiving a continued drag gesture 874, device 800 displays a dialog display 873 on top of the display of response display 875.

[0330] In some examples, the response capability representation is displayed in its initial state before receiving user input that causes the dialogue capability representation to be displayed in a second state (e.g., expanded). In some examples, the initial state describes the state of the response capability representation before the dialogue capability representation (or a portion thereof) is displayed on top of the response capability representation. For example, Figures 8BO to 8BQ A response capability representation 875 is shown in its initial state. In some examples, displaying a dialog capability representation in a second (e.g., expanded) state over at least a portion of the response capability representation includes replacing the display of the response capability representation in its initial state with the display of the response capability representation in its overlaid state. Figure 8BR A response power display 875 is shown in an overlay state. In some examples, when displayed in an overlay state, the display size of the response power display (e.g., relative to the initial state) shrinks and / or dims (e.g., the displayed color is darker than the initial state). In some examples, the degree to which the response power display shrinks and / or dims is proportional to the amount of dialog power display displayed on top of the response power display.

[0331] In some examples, the dialog display has a maximum display size, and the second (e.g., expanded) state of the dialog display corresponds to the maximum display size. In some examples, the dialog display shown at the maximum display size cannot expand further in response to user input such as a drag gesture. In some examples, the dialog display shown at the maximum display size shows the entire content of the dialog display. In other examples, the dialog display shown at the maximum display size does not show the entire content of the dialog display. Therefore, in some examples, when device 800 displays the dialog display in the second state (which has the maximum display size), device 800 enables user input (e.g., a drag gesture / swipe gesture) to scroll through the content of the dialog display. Figure 8BS The dialog display 873 is shown at its maximum display size. Specifically, in Figure 8BRIn the middle, drag gesture 874 continues. In response to receiving the continue drag gesture 874, in Figure 8BS In this device 800, the dialog functionality representation 873 is displayed (e.g., expanded) to its maximum display size. The dialog functionality representation 873 includes a scroll indicator 876 that indicates to the user that input can be provided to scroll through the content of the dialog functionality representation 873.

[0332] In some examples, when the dialog power indicator is displayed in its second state (and at its maximum size), a portion of the response power indicator remains visible. Therefore, in some examples, device 800 constrains the maximum size of the dialog power indicator displayed above the response power indicator, such that the dialog power indicator does not completely cover the response power indicator. In some examples, the portion of the response power indicator that remains visible is relative to the preceding text. Figure 8M The second part of the response capability representation. For example, this part is the top portion of the response capability representation that includes a flag symbol indicating the category of the response capability representation and / or associated text. Figure 8BS It is shown that when the device 800 displays the dialog power indicator 873 at its maximum size above the response power indicator 875, the top portion of the response power indicator 875 remains visible.

[0333] In some examples, device 800 receives user input corresponding to a selection of a portion of the response display that remains visible (when the dialog display is displayed in a second state above the response display). In response to receiving user input, the device displays the response display at a first portion of display 801, for example, in its initial state. In response to receiving user input, device 800 further replaces the display of the dialog display in the second (e.g., expanded) state with a display of the dialog display in the third state. In some examples, the third state is a compressed state, wherein the dialog display in the third state (compared to the dialog display in the initial or expanded state) has a smaller display size and / or includes less content (compared to the dialog display in the initial or expanded state). In other examples, the third state is the first state (e.g., the initial state). Figure 8BT The device 800 is shown receiving user input 877 (e.g., a tap gesture) on the top portion of a selection response enable display 875. Figure 8BU This demonstrates that in response to receiving user input 877, device 800 replaces the expanded display of dialog display 873 with a compressed display. Figure 8BT The device 800 further displays a response indicator 875 in its initial state.

[0334] In some examples, device 800 receives user input corresponding to a selection of a dialog capability representation displayed in a third state. In response to receiving the user input, device 800 replaces the display of a response capability representation in the third state with the display of a dialog capability representation in a first state. For example, in Figure 8BU In this context, the user can provide input (e.g., a tap gesture) to select the dialogue capability representation 873 to be displayed in a compressed state. In response to receiving input, the device 800 displays the dialogue capability representation in its initial state, for example, reverting to... Figure 8BO The display.

[0335] In some examples, when the dialog power indicator is displayed in a first or second state (e.g., an initial state or an expanded state), device 800 receives user input corresponding to a selection of a simultaneously displayed response power indicator. In response to receiving the user input, device 800 replaces the display of the dialog power indicator in the first or second state with the display of the dialog power indicator in a third (e.g., a compressed) state. For example, Figure 8BV The DA user interface 803 is shown in response to the natural language input “Show me the list of team #1”. The DA user interface 803 includes a detailed response capability representation 878 and a dialog capability representation 879 displayed in its initial state. Figure 8BV It is also shown that device 800 receives user input 880 (e.g., drag gesture) from selection response enable display 878. Figure 8BW It is shown that in response to receiving user input 880, device 800 replaces the display of dialog display 879 in its initial state with the display of dialog display 879 in a compressed state.

[0336] In some examples, when the dialog capability representation is displayed in a first or second state (e.g., an initial state or an expanded state), device 800 receives user input corresponding to a selection of the dialog capability representation. In response to receiving the user input, device 800 replaces the display of the dialog capability representation in the first or second state with the display of the dialog capability representation in a third (e.g., a compressed) state. For example, Figure 8BX The diagram shows a DA user interface 803 displayed in response to the natural language input “What music can you provide for me?”. The DA user interface 803 includes a response capability representation 881 and a dialog capability representation 882 displayed in its initial state. Figure 8BX The device 800 is also shown receiving user input 883 (e.g., a drag-down or swipe gesture) from a selection dialog enablement 882. Figure 8 BY It is shown that in response to receiving user input 883, device 800 replaces the display of dialog display 882 in its initial state with the display of dialog display 882 in a compressed state. Although Figures 8BX to 8BYThe user input shown corresponds to a drag or swipe gesture for selecting a dialog power indicator; however, in other examples, the user input is a selection of a power indicator included in the dialog power indicator. For example, user input (e.g., a tap gesture) selecting a "compressed" power indicator in a dialog power indicator displayed in a first or second state causes device 800 to replace the display of the dialog power indicator in the first or second state with the display of the dialog power indicator in the third state.

[0337] In some examples, device 800 displays a transcription of the received natural language speech input in a dialogue capability representation. The transcription is obtained by performing automatic speech recognition (ASR) on the natural language speech input. Figure 8BZ The diagram illustrates a DA user interface 803 displayed in response to a natural language speech input, "How's the weather?". The DA user interface includes a response capability representation 884 and a dialogue capability representation 885. The dialogue capability representation 885 includes a transcription 886 of the speech input and a dialogue 887 generated by the DA in response to the speech input.

[0338] In some examples, device 800 does not display the transcription of received natural language speech input by default. In some examples, device 800 includes a setting that causes device 800 to always display the transcription of natural language speech input when activated. Various other instances in which device 800 may display the transcription of received natural language speech input will now be discussed.

[0339] In some examples, natural language speech input (with the displayed transcription) follows a second natural language speech input received prior to the natural language speech input. In some examples, the displayed transcription is performed based on the determination that the DA cannot determine the user intent of the natural language speech input and cannot determine the second user intent of the second natural language speech input. Therefore, in some examples, if the DA cannot determine the executable intent of the two consecutive natural language inputs, the device 800 displays the transcription of the natural language input.

[0340] For example, Figure 8CA The device 800 has received the voice input "how far to Dish n'Dash?", and DA cannot determine the user's intent regarding the natural language input. For example, device 800 provides the audio output "I'm not sure I understand, can you please say that again?". Therefore, the user repeats the voice input. Figure 8CB The device 800 is shown receiving the continuous voice input "how far to Dish n'Dash?". Figure 8CCThis illustrates that the DA (Dialogue Awareness) still cannot determine the user's intent for consecutive speech inputs. For example, device 800 provides the audio output "I'm not sure I understand." Therefore, device 800 also displays a dialogue awareness representation 888 including the transcription 889 "how far to Rish and Rash?" of the subsequent speech input. In this example, transcription 889 reveals that the DA incorrectly identifies "how far to Dish n'Dash?" twice as "how far to Rish and Rash?". Since "Rish and Rash" may not be the actual location, the DA cannot determine the user's intent for the two speech inputs.

[0341] In some examples, transcription of the received natural language speech input is performed based on the determination that the natural language speech input repeats the previous natural language speech input. For example, Figure 8CD The DA user interface 803 is shown in response to the voice input (previous voice input) "where is Starbucks?". The DA incorrectly recognizes the voice input as "where is Star Mall?", and therefore displays a response display 890 that includes "Star Mall". Because the DA incorrectly understands the voice input, the user repeats the voice input. For example, Figure 8CE The diagram illustrates that device 800 receives a repetition (e.g., a continuation repetition) of the previous voice input "where is Starbucks?". DA determines that the voice input is a repetition of the previous voice input. Figure 8CF Based on this determination, device 800 displays a dialogue energization representation 891 including transcription 892. Transcription 892 reveals that DA (e.g., twice) incorrectly identifies “where is Starbucks?” as “where is Star Mall?”.

[0342] In some examples, after receiving natural language speech input (e.g., for which a transcription will be displayed), the device receives a second natural language speech input following the natural language speech input. In some examples, the display transcription is performed based on the determination that the second natural language speech input indicates a speech recognition error. Therefore, in some examples, if subsequent speech input indicates that the DA incorrectly recognized the previous speech input, the device 800 displays a transcription of the previous speech input. For example, Figure 8CGA DA user interface 803 is shown in response to the voice input "set a timer for 15 minutes". The DA incorrectly recognizes "15 minutes" as "50 minutes". The DA user interface 803 therefore includes a response indication 893 indicating that the timer has been set to 50 minutes. Because the DA incorrectly recognizes the voice input, the user provides a second voice input indicating the voice recognition error (e.g., "that's not what I said", "you heard me wrong", "that's incorrect", etc.). Figure 8CH The device 800 receives a second voice input, “that’s not what Isaid.” The DA determines that the second voice input indicates a voice recognition error. Figure 8CI Based on this determination, device 800 displays a dialogue energization representation 894 including transcription 895. Transcription 895 reveals that DA incorrectly identifies "15 minutes" as "50 minutes".

[0343] In some examples, device 800 receives user input corresponding to a selection of the displayed transcription. In response to receiving user input, device 800 simultaneously displays a keyboard and an editable text field including the transcription, for example, displaying the keyboard and editable text field on top of the user interface displayed on top of DA user interface 803. In some examples, device 800 also visually masks at least a portion of the user interface (e.g., the portion of display 801 that does not display the keyboard or editable text field). Figure 8CI Example, Figure 8CJ The device 800 is shown receiving user input 896 (e.g., a tap gesture) for selecting transcription 895. Figure 8CK The device 800 is shown displaying a keyboard 897 and an editable text field 898 including transcription 895 in response to receiving user input 896. Figure 8CK The device 800 is also shown to visually conceal a portion of the user interface 802.

[0344] Figure 8CL It is shown that device 800 has received one or more keyboard inputs and has edited the transcription 895 based on the one or more keyboard inputs, for example, editing "set a timer for 50 minutes" to "set a timer for 15 minutes". Figure 8CL It is also shown that device 800 receives user input 899 (e.g., a tap gesture) corresponding to the selection of the completion key 8001 on keyboard 897. Figure 8CMThe diagram illustrates how, in response to receiving user input 899, the DA performs a task based on the current (e.g., edited) transcription 895. For example, device 800 displays a DA user interface 803, which includes a response indicator 8002 indicating that a timer has been set to 15 minutes. Device 800 also provides the voice output “Ok, I set the timer for 15 minutes.” This allows the user to manually correct incorrect transcriptions (e.g., using keyboard input) to ensure the correct task is performed.

[0345] In some examples, while displaying the keyboard and editable text field, device 800 receives user input corresponding to a selection of the visually obscured user interface. In some examples, in response to receiving user input, device 800 stops displaying the keyboard and editable text field. In some examples, device 800 additionally or alternatively stops displaying the DA user interface 803. For example, in Figure 8CK to Figure 8CL In this context, selecting user input (e.g., a tap gesture) that is visually hidden in the user interface 802 allows the device 800 to revert to its previous state. Figure 8CI The display or device 800 stops displaying the DA user interface 803 and fully displays the user interface 802, such as Figure 8A As shown.

[0346] In some examples, device 800 presents a digital assistant result (e.g., a response power indicator and / or audio output) immediately. In some examples, based on determining that the digital assistant result corresponds to a pre-defined type of digital assistant result, device 800 automatically stops displaying the DA user interface 803 for a predetermined duration after the initial presentation. Therefore, in some examples, device 800 may quickly (e.g., within 5 seconds) eliminate the DA user interface 803 after providing a pre-defined type of result. Exemplary pre-defined type results correspond to completed tasks that do not require further user input (or further user interaction). For example, such results include confirmation that a timer has been set, a message has been sent, and a household appliance (e.g., a light) has changed state. Examples of results that do not correspond to a pre-defined type include results where the DA requires further user input, and results where the DA provides information (e.g., news, Wikipedia articles, locations) in response to a user's information request.

[0347] For example, Figure 8CM The device 800 is shown to present the result immediately, for example, providing a voice output "OK, I set the timer for 15 minutes." Because the result corresponds to a predetermined type, Figure 8CNThe device 800 is shown to automatically eliminate DA within a predetermined duration (e.g., 5 seconds) after the first time (e.g., without further user input).

[0348] Figure 8CO to Figure 8CT An example of the DA user interface 803 and an exemplary user interface is shown when device 800 is a tablet device. It should be understood that when device 800 is another type of device, the same applies to any technology discussed herein with respect to device 800 as a tablet device (and vice versa).

[0349] Figure 8CO A device 800 displaying a user interface 8003 is shown. The user interface 8003 includes a taskbar area 8004. Figure 8CO In this configuration, device 800 displays a DA user interface 803 on top of a user interface 8003. The DA user interface 803 includes an indicator 804 displayed on a first portion of display 801 and a responsive power indicator 8005 displayed on a second portion of display 801. As shown, a portion of the user interface 8003 remains visible (e.g., not visually obscured) on a third portion of display 801. In some examples, the third portion is located between the first and second portions of display 801. In some examples, such as... Figure 8CO As shown, the display of the DA user interface 803 will not visually obscure the taskbar area 8004; for example, no part of the DA user interface 803 will be displayed on top of the taskbar area 8004.

[0350] Figure 8CP A device 800 is shown displaying a DA user interface 803 including a dialog power indicator 8006. As shown, the dialog power indicator 8006 is displayed in a portion of a display 801, between a first portion of the display 801 (displaying an indicator 804) and a second portion of the display (displaying a response power indicator 8005). Displaying the dialog power indicator 8006 also causes the response power indicator 8005 (from...) Figure 8CO It is shifted toward the top portion of the display 801.

[0351] Figure 8CQ A device 800 is shown displaying a user interface 8003 including a media panel 8007 that indicates that media is currently playing. Figure 8 CRA device 800 is shown displaying a DA user interface 803 on top of a user interface 8003. The DA user interface 803 includes a responsive power indicator 8008 and an indicator 804. As shown, the display of the DA user interface 803 does not visually obscure the media panel 8007. For example, as shown, the display elements of the DA user interface 803 (e.g., indicator 804, responsive power indicator 8008, dialog power indicator) cause the media panel 8007 to shift toward the top portion of the display 801.

[0352] Figure 8CS A device 800 is shown displaying a user interface 8009 including a keyboard 8010. Figure 8 CT The DA user interface 803 is shown above the user interface 8009. (See diagram.) Figure 8 CT As shown, in some examples, displaying the DA user interface 803 on top of the user interface 8009, which includes the keyboard 8010, causes the device 800 to visually obscure the keys of the keyboard 8010 (e.g., by graying out the keys).

[0353] Figures 9A to 9C Multiple devices are illustrated, based on various examples, for determining which device should respond to voice input. Specifically, Figure 9A Devices 900, 902, and 904 are shown. Devices 900, 902, and 904 are each implemented as device 104, device 122, device 200, or device 600. In some examples, devices 900, 902, and 904 each at least partially implement DA system 700.

[0354] exist Figure 9A In some examples, when a user provides voice input including a trigger phrase for initiating DA (e.g., “Hey Siri”), such as “Hey Siri, what’s the weather?”, the corresponding displays of devices 900, 902, and 904 do not show anything. In some examples, when a user provides voice input, the corresponding display of at least one of devices 900, 902, and 904 displays a user interface (e.g., a home screen user interface, an application-specific user interface). Figure 9B The diagram illustrates that in response to receiving voice input including a trigger phrase, devices 900, 902, and 904 each display an indicator 804. In some examples, each indicator 804 is displayed in a listening state, for example, indicating that the corresponding device is sampling the audio input.

[0355] exist Figure 9BIn this embodiment, devices 900, 902, and 904 coordinate with each other (or via a fourth device) to determine which device should respond to a user request. Exemplary techniques for device coordination to determine which device should respond to a user request are described in U.S. Patent No. 10,089,072, entitled “INTELLIGENT DEVICE ARBITRATION AND CONTROL,” filed October 2, 2018, and in U.S. Patent Application No. 63 / 022,942, filed May 11, 2020, entitled “DIGITAL ASSISTANT HARDWARE ABSTRACTION,” the contents of which are incorporated herein by reference in their entirety. Figure 9B As shown, each device displays indicator 804 only when it determines whether to respond to a user request. For example, the corresponding portions of the non-display indicator 804 on the displays of devices 900, 902, and 904 are not displayed. In some examples, when at least one of devices 900, 902, and 904 displays a user interface (the previous user interface) when the user provides voice input, that at least one device displays indicator 804 only on top of the previous user interface when it determines whether to respond to a user request.

[0356] Figure 9C The diagram illustrates device 902 being identified as the device responding to a user request. As shown, in response to determining that another device (e.g., device 902) should respond to the user request, the displays of devices 900 and 904 cease displaying (or the stop display indicator 804 is used to fully display the previous user interface). As further shown, in response to determining that device 902 should respond to the user request, device 902 displays user interface 906 (e.g., a lock screen user interface) and a DA user interface 803 on top of user interface 906. The DA user interface 803 includes the response to the user request. Thus, visual interference is minimized when it is determined which of multiple devices should respond to voice input. For example, in Figure 9B In this case, the display of a device determined not to respond to user requests only shows indicator 804, for example, the opposite of displaying the user interface with the entire display.

[0357] Determining which of a plurality of devices should respond to voice input in the manner described above provides the user with feedback that the voice input has been received and is being processed. Furthermore, providing feedback in this manner advantageously reduces unnecessary visual or auditory distractions in response to voice input. For example, it eliminates the need for the user to manually stop the display and / or audible output of unselected devices and minimizes visual distractions to the user interface of unselected devices (e.g., if the user was previously interacting with the user interface of an unselected device). Providing the user with improved visual feedback enhances device operability and makes the user-device interface more efficient (e.g., by reducing the amount of user input expected to perform the requested task), which in turn reduces power consumption and extends device battery life by enabling the user to use the device more quickly and effectively.

[0358] Figures 10A to 10V The user interface and digital assistant user interface are shown according to various examples. Figures 10A to 10V Used to illustrate the processes described below, these processes include Figures 18A to 18B The process in.

[0359] Figure 10A Device 800 is shown. Device 800 displays a DA user interface 803 on display 801, above the user interface. Figure 10A In this example, device 800 displays DA user interface 803 above home screen user interface 1001. In other examples, the user interface is another type of user interface, such as a lock screen user interface or an application-specific user interface.

[0360] In some examples, the DA user interface 803 includes an indicator 804 displayed in a first portion (e.g., the "indicator portion") of the display 801 and a responsive power indicator displayed in a second portion (e.g., the "response portion") of the display 801. A third portion (e.g., the "UI portion") of the display 801 displays a portion of the user interface (the user interface displayed on top of the DA user interface 803). For example, in Figure 10A In the display 801, the first part displays an indicator 804, the second part displays a response enable indicator 1002, and the third part displays a portion of the home screen user interface 1001.

[0361] In some examples, when the DA user interface 803 is displayed on top of the user interface, the device 800 receives user input corresponding to a selection of a third portion of the display 801. The device 800 determines whether the user input corresponds to a first type of input or a second type of input. In some examples, the first type of user input includes a tap gesture, and the second type of user input includes a drag or swipe gesture.

[0362] In some examples, device 800 stops displaying DA user interface 803 based on determining that the user input corresponds to a first type of input. Stopping the display of DA user interface 803 includes stopping the display of any part of DA user interface 803, such as indicator 804, response enable representation, and dialog enable representation (if included). In some examples, stopping the display of DA user interface 803 includes replacing the display of elements of DA user interface 803 at corresponding portions of display 801 with the display of the user interface at the corresponding portions. For example, device 800 replaces the display of indicator 804 at a first portion of display 801 with the display of a first portion of user interface, and replaces the display of response enable representation at a second portion of display 801 with the display of a second portion of user interface.

[0363] For example, Figure 10B The illustration shows that device 800 receives user input 1003 (e.g., a tap gesture) corresponding to a selection of a third portion of display 801. Device 800 determines that user input 1003 corresponds to a first type of input. Figure 10C It is shown that, based on this determination, device 800 stops displaying DA user interface 803 and displays user interface 1001 in full.

[0364] In this way, the user can eliminate the DA user interface 803 by providing input that selects the display 801 to not display any part of the DA user interface 803. For example, in the above... Figures 8S to 8X In the process, a tap gesture on a portion of the user interface 802 that visually obscures the display of the selected monitor 801 causes the device 800 to return to the home screen. Figure 8A The display.

[0365] In some examples, the user input corresponds to a selection of a selectable element displayed in the third portion of display 801. In some examples, based on determining that the user input corresponds to a first type of input, device 800 displays a user interface corresponding to the selectable element. For example, device 800 replaces the display of that portion of the user interface (displayed in the third portion of display 801), the display of the response enable indicator, and the display of indicator 804 with the display of the user interface corresponding to the selectable element.

[0366] In some examples, the user interface is a home screen user interface 1001, the selectable elements are application functional representations of the home screen user interface 1001, and the user interface corresponding to the selectable elements is the user interface corresponding to the application functional representation. For example, Figure 10D A DA user interface 803 is shown displayed on top of the home screen user interface 1001. The display 801 displays an indicator 804 in a first part, a response power indicator 1004 in a second part, and a portion of the user interface 1001 in a third part. Figure 10E The device 800 is shown receiving user input 1005 (e.g., a tap gesture) in the health application display 1006 shown in the third section, where the device 800 selects the health application display 1006. Figure 10F The diagram shows that, based on the determination by device 800 that user input 1005 corresponds to a first type of input, device 800 stops displaying this portion of the indicator 804, the response power indicator 1004, and the user interface 1001. Device 800 also displays a user interface 1007 corresponding to a health application.

[0367] In some examples, the selectable element is a link, and the user interface corresponding to the selectable element is the same as the user interface corresponding to the link. For example, Figure 10G A DA user interface 803 is shown displayed on top of a web browsing application user interface 1008. A display 801 shows an indicator 804 in a first portion, a response power indicator 1009 in a second portion, and a portion of the user interface 1008 in a third portion. Figure 10G It is also shown that device 800 receives user input 1010 (e.g., tap gesture) for selecting a link 1011 (e.g., a web link) displayed in the third part. Figure 10H The diagram shows that, based on the determination by device 800 that user input 1010 corresponds to a first type of input, device 800 stops displaying this portion of indicator 804, response power indicator 1009, and user interface 1008. Device 800 also displays user interface 1012 corresponding to web page link 1011.

[0368] In this way, user input that selects the third part of the display 801 eliminates the DA user interface 803 and also enables an action to be performed based on the user's selection (e.g., updating the display 801).

[0369] In some examples, based on determining that the user input corresponds to a second type of input (e.g., a drag or swipe gesture), device 800 updates the display of the user interface at a third portion of display 801 according to the user input. In some examples, while device 800 updates the display of the user interface at the third portion of display 801, device 800 continues to display at least some elements of the DA user interface 803 at corresponding display portions of the elements. For example, device 800 displays (e.g., continues displaying) a responsive enable indicator at a second portion of display 801. In some examples, device 800 also displays (e.g., continues displaying) an indicator 804 at a first portion of display 801. In some examples, updating the display of the user interface at the third portion includes scrolling the content of the user interface.

[0370] For example, Figure 10I A DA user interface 803 is shown displayed on top of a web browser application user interface 1013 that displays a webpage. The display 801 displays an indicator 804 in a first part, a response enable indicator 1014 in a second part, and a portion of the user interface 1013 in a third part. Figure 10I The device 800 is also shown receiving user input 1015 (e.g., a drag gesture) to select a third part. Figure 10J It is shown that the device 800 determines that the user input 1015 corresponds to a second type of input, and the device 800 updates (e.g., scrolls over) the content of the user interface 1013 based on the user input 1015, such as the content of a webpage. Figures 10I to 10J As shown, while updating the user interface 1013 (at the third part of display 801), device 800 continues to display indicator 804 at the first part of display 801 and displays response power indication 1014 at the second part of display 801.

[0371] For example, Figure 10K A DA user interface 803 is shown displayed on the home screen user interface 1001. The display 801 displays an indicator 804 in a first part, a response power indicator 1016 in a second part, and a portion of the user interface 1001 in a third part. Figure 10K It is also shown that device 800 receives user input 1017 (e.g., a swipe gesture) to select a third part. Figure 10LThe diagram illustrates that device 800 determines user input 1017 corresponds to a second type of input, and device 800 updates the content of user interface 1001 based on user input 1017. For example, as shown, device 800 updates user interface 1001 to display a secondary home screen user interface 1018, which includes one or more application display representations different from the application display representations of home screen user interface 1001. Figures 10K to 10L It is shown that when updating the user interface 1001, the device 800 continues to display an indicator 804 on a first portion of the display 801 and a response power indicator 1016 on a second portion of the display 801.

[0372] In this way, the user can provide input to update the user interface displayed on the DA user interface 803 without causing the DA user interface 803 to be eliminated.

[0373] In some examples, the display of the user interface at the third portion of display 801 is updated based on determining that DA is in a listening state. Therefore, device 800 may enable drag or swipe gestures to update the user interface (displayed on top of DA user interface 803) only when DA is in a listening state. In such examples, if DA is not in a listening state, in response to receiving user input corresponding to the second type (and corresponding to selection of the third portion of display 801), device 800 does not update display 801 or stops displaying DA user interface 803 in response to user input. In some examples, when updating the display of the user interface while DA is in a listening state, the display size of indicator 804 varies based on the amplitude of the received voice input, as described above.

[0374] In some examples, when device 800 displays DA user interface 803 on top of the user interface, device 800 receives a second user input. In some examples, based on determining that the second user input corresponds to a third type of input, device 800 stops displaying DA user interface 803. In some examples, the third type of input includes a swipe gesture originating from the bottom of display 801 toward the top of display 801. The third type of input is sometimes considered a "home swipe" because receiving such input while device 800 displays a user interface different from the home screen user interface (and does not display DA user interface 803) causes device 800 to return to the display of the home screen user interface.

[0375] Figure 10M A device 800 is shown displaying a DA user interface 803 on top of a home screen user interface 1001. The DA user interface 803 includes a responsive power indicator 1020 and an indicator 804. Figure 10MThe device 800 is also shown receiving user input 1019, namely a swipe gesture from the bottom of the display 801 toward the top of the display 801. Figure 10N It is shown that, based on the determination by device 800 that user input 1019 corresponds to the third type of input, device 800 stops displaying response power indicator 1020 and indicator 804.

[0376] In some examples, the user interface (displayed on top of the DA user interface 803) is an application-specific user interface. In some examples, when the device 800 displays the DA user interface 803 on top of the application-specific user interface, the device 800 receives a second user input. In some examples, based on determining that the second user input corresponds to a third type of input, the device stops displaying the DA user interface 803 and instead displays the home screen user interface. For example, Figure 10O A device 800 is shown displaying a DA user interface 803 on top of a health application user interface 1022. The DA user interface 803 includes a responsive power indicator 1021 and an indicator 804. Figure 10O The device 800 is also shown receiving user input 1023, namely a swipe gesture from the bottom of the display 801 toward the top of the display 801. Figure 10P It is shown that, based on the determination that user input 1023 corresponds to a third type of input according to device 800, device 800 displays home screen user interface 1001. For example, as shown, device 800 replaces the display of indicator 804, response enable indicator 1021, and messaging application user interface 1022 with the display of home screen user interface 1001.

[0377] In some examples, when device 800 displays DA user interface 803 on top of the user interface, device 800 receives third user input corresponding to a selection of the response enable representation. In response to receiving the third user input, device 800 stops displaying DA user interface 803. For example, Figure 10Q The DA user interface 803 displayed on top of the home screen user interface 1001 is shown. The DA user interface 803 includes a responsive enable indicator 1024, a dialog enable indicator 1025, and an indicator 804. Figure 10Q The device 800 is also shown receiving user input 1026 (e.g., swipe up or drag gesture) from the selection response enable display 1024. Figure 10R The diagram shows that in response to receiving user input 1026, device 800 stops displaying the DA user interface 803.

[0378] In some examples, when device 800 displays user interface 803 on top of user interface 803, device 800 receives a fourth user input corresponding to the displacement of indicator 804 from the first portion of display 801 to the edge of display 801. In response to receiving the fourth user input, device 800 stops displaying DA user interface 803. For example, Figure 10S The DA user interface 803 displayed above the home screen user interface 1001 is shown. Figure 10S In this process, device 800 receives user input 1027 (e.g., a drag or swipe gesture) that moves an indicator from a first portion of display 801 to the edge of display 801. Figures 10S to 10V It is shown that in response to receiving user input 1027 (e.g., in response to indicator 804 reaching the edge of display 801), device 800 stops displaying DA user interface 803.

[0379] 5. Digital Assistant Response Mode

[0380] Figure 11 A system 1100 for selecting a DA response mode and for presenting a response based on the selected DA response mode is illustrated according to various examples. In some examples, system 1100 is implemented on a stand-alone computer system (e.g., device 104, 122, 200, 400, 600, 800, 900, 902, or 904). System 1100 is implemented using hardware, software, or a combination of hardware and software to perform the principles discussed herein. In some examples, modules and functions of system 1100 are implemented within a DA system, as described above relative to... Figures 7A to 7C As mentioned above.

[0381] System 1100 is exemplary, and therefore system 1100 may have more or fewer components than those illustrated, may combine two or more components, or may have different component configurations or arrangements. Furthermore, although the following discussion describes functions performed at a single component of system 1100, it should be understood that these functions may be performed at other components of system 1100, and these functions may be performed at more than one component of system 1100.

[0382] Figure 12 A device 800 is shown that presents a response to received natural language input according to different DA response modes, based on various examples. Figure 12 In each instance of device 800, device 800 has initiated DA and presented a response to the voice input “How’s the weather?” according to the silence response mode, mixed response mode, or voice response mode described below. Device 800 implementing system 1100 selects a DA response mode and presents a response according to the selected response mode using the techniques described below.

[0383] System 1100 includes an acquisition module 1102. Acquisition module 1102 acquires a response packet in response to natural language input. The response packet includes content (e.g., conversational text) intended as a response to the natural language input. In some examples, the response packet includes a first text (content text) associated with a digital assistant response capability representation (e.g., response capability representation 1202) and a second text (title text) associated with the response capability representation. In some examples, the title text is less verbose than the content text (e.g., includes fewer words). The content text provides a complete response to the user's request, while the title text provides a brief (e.g., incomplete) response. For a complete response to the request, device 800 may render both the title text and the response capability representation simultaneously; for example, the rendering of the content text may not require the rendering of the response capability representation for a complete response.

[0384] For example, consider Figure 12 The natural language input is "How's the weather?". The content text is "Current temperature 70 degrees Celsius, sunny, zero chance of rain. Today's high will reach 75 degrees Celsius, and the low will reach 60 degrees Celsius." The title text is simply "The weather is great today." As shown in the figure, the title text is intended to be presented together with the responsive display representation 1202, which visually indicates the information in the content text. Therefore, presenting the content text alone can fully answer the request, while presenting both the title text and the responsive display representation is necessary to fully answer the request.

[0385] In some examples, the acquisition module 1102 is obtained, for example, via device 800 according to... Figures 7A to 7C The natural language input is processed as described above to retrieve the response packet from the local machine. In some examples, the retrieval module 1102 retrieves the response packet from an external device such as DA server 106. In such examples, DA server 106 processes the response packet according to... Figures 7A to 7C Natural language input is processed in the manner described above to determine the response packet. In some examples, the acquisition module 1102 acquires a portion of the response packet locally and another portion of the response packet from an external device.

[0386] System 1100 includes a mode selection module 1104. Selection module 1104 selects a DA response mode from multiple DA response modes based on context information associated with device 800. A DA response mode specifies how the DA presents a response to natural language input (e.g., a response packet) (e.g., a format).

[0387] In some examples, after device 800 receives natural language input, selection module 1104 selects a DA response mode, for example, based on current context information acquired after receiving the natural language input. In some examples, after acquisition module 1102 acquires the response packet, selection module 1104 selects a DA response mode, for example, based on current context information acquired after acquiring the response packet. The current context information describes the context information at the time selection module 1104 selects the DA response mode. In some examples, this time is after receiving the natural language input and before presenting the response to the natural language input. In some examples, the multiple DA response modes include a silence response mode, a mixed response mode, and a voice response mode, which will be discussed further below.

[0388] System 1100 includes a formatting module 1106. In response to selection module 1104 selecting a DA response mode, formatting module 1106 causes the DA to present a response packet according to the selected DA response mode (e.g., in a format consistent with the selected DA response mode). In some examples, the selected DA response mode is a silent response mode. In some examples, presenting a response packet according to a silent response mode includes displaying a response enable indicator and displaying title text, without providing audio output indicating (e.g., speaking) the title text (and without providing content text). In some examples, the selected DA response mode is a mixed response mode. In some examples, presenting a response packet according to a mixed response mode includes displaying a response enable indicator and speaking the title text, without displaying the title text (and without providing context text). In some examples, the selected DA response mode is a voice response mode. In some examples, presenting a response packet according to a voice response mode includes speaking the content text, for example, without presenting the title text and / or without displaying the response enable indicator.

[0389] For example, in Figure 12 In the silent response mode, the response package includes displaying response indicator 1202 and displaying the title text "The weather is nice today" in the dialogue indicator 1204 without speaking the title text. In the mixed response mode, the response package includes displaying response indicator 1202 and speaking the title text "The weather is nice today" without displaying the title text. In the voice response mode, the response package includes speaking the content text "Current temperature 70 degrees, sunny, zero probability of rain. Today's high will reach 75 degrees, and the low will reach 60 degrees." Although... Figure 12 The example shown illustrates that device 800 displays a response enable representation 1202 when presenting a response packet according to a voice response pattern, but in other examples, a response enable representation is not displayed when presenting a response packet according to a voice response pattern.

[0390] In some examples, when DA renders a response according to a silent response mode, device 800 displays a response capability representation but not a dialogue capability representation (e.g., including text). In some examples, device 800 omits providing text based on determining that the response capability representation includes a direct answer to the natural language request. For example, device 800 determines that the title text and the response capability representation each include corresponding matching text that answers the user's request (thus rendering the title text redundant). For example, for the natural language request "What is the temperature?", if the response capability representation includes the current temperature, then in silent mode, device 800 does not display any title text because the title text including the current temperature is redundant for the response capability representation. In contrast, consider the exemplary natural language request "Is it cold?" The response capability representation for the request might include the current temperature and weather conditions, but might not include a direct (e.g., explicit) answer to the request, such as "Yes" or "No". Therefore, for such natural language input, in silent mode, device 800 displays both the response capability representation and the title text that includes a direct answer to the request, such as "No, it's not cold."

[0391] Figure 12 In some examples, selecting the DA response mode involves determining whether to (1) display the title text without speaking it or (2) speak the title text without displaying it. In some examples, selecting the response mode involves determining whether to speak the content text.

[0392] Generally speaking, a mute response mode may be appropriate when the user expects to view the display but not audio output. A mixed response mode may be appropriate when the user expects to view the display and also expects audio output. A voice response mode may be appropriate when the user does not expect (or cannot) view the display. The various techniques and contextual information selection module 1104 used for selecting the DA response mode will now be discussed.

[0393] Figure 13 Exemplary process 1300, implemented by selection module 1104 to select a DA response mode according to various examples, is shown. In some examples, selection module 1104 implements process 1300 as computer-executable instructions, such as those stored in the memory of device 800.

[0394] At box 1302, module 1104 selects (e.g., determines) the current context information. At box 1304, module 1104 determines whether to select a voice mode based on the current context information. If module 1104 determines to select a voice mode, then at box 1306, module 1104 selects a voice mode. If module 1104 determines not to select a voice mode, process 1300 proceeds to box 1308. At box 1308, module 1104 selects between a mute mode and a mixed mode. If module 1104 determines to select a mute mode, then at box 1310, module 1104 selects a mute mode. If module 1104 determines to select a mixed mode, then at box 1312, module 1104 selects a mixed mode.

[0395] In some examples, boxes 1304 and 1308 are implemented using a rule-based system. For example, at box 1304, module 1104 determines whether the current context information satisfies specific conditions for selecting a voice mode. If the specific conditions are met, module 1104 selects a voice mode. If the specific conditions are not met (meaning the current context information satisfies the conditions for selecting a mixed mode or a voice mode), module 1104 proceeds to box 1308. Similarly, at box 1308, module 1104 determines whether the current context information satisfies specific conditions for selecting a silence mode or a mixed mode, and selects a silence mode or a mixed mode accordingly.

[0396] In some examples, boxes 1304 and 1308 are implemented using a probabilistic (e.g., machine learning) system. For example, at box 1304, module 1104 determines the probability of selecting a voice mode and the probability of not selecting a voice mode (e.g., the probability of selecting a silent mode or a mixed mode) based on the current context information, and selects the branch with the highest probability. At box 1308, module 1104 determines the probability of selecting a mixed mode and the probability of selecting a silent mode based on the current context information, and selects the mode with the highest probability. In some examples, the sum of the probabilities of voice mode, mixed mode, and silent mode is 1.

[0397] We will now discuss the various types of current context information used to determine boxes 1304 and / or 1308.

[0398] In some examples, the contextual information includes whether device 800 has a display. In a rule-based system, determining that device 800 does not have a display satisfies the condition used for selecting a voice mode. In a probabilistic system, determining that device 800 does not have a display increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode.

[0399] In some examples, contextual information includes whether device 800 detected a voice input initiating DA (e.g., "Hey Siri"). In a rule-based system, detecting a voice input initiating DA satisfies the conditions for selecting a voice mode. In a rule-based system, not detecting a voice input initiating DA does not satisfy the conditions for selecting a voice mode (and therefore satisfies the conditions for selecting a mixed mode or a silent mode). In a probabilistic system, in some examples, detecting a voice input initiating DA increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, not detecting a voice input initiating DA decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0400] In some examples, the context information includes whether device 800 detected physical contact with device 800, which is used to initiate DA. In a rule-based system, no physical contact is detected, which satisfies the conditions for selecting a voice mode. In a rule-based system, detected physical contact does not satisfy the conditions for selecting a voice mode. In a probabilistic system, in some examples, no physical contact increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, detected physical contact decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0401] In some examples, the context information includes whether device 800 is in a locked state. In a rule-based system, determining that device 800 is in a locked state satisfies the conditions for selecting a voice mode. In a rule-based system, determining that device 800 is not in a locked state does not satisfy the conditions for selecting a voice mode. In a probabilistic system, in some examples, determining that device 800 is in a locked state increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, determining that device 800 is not in a locked state decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0402] In some examples, contextual information includes whether the display of device 800 was displaying before initiating DA. In a rule-based system, it is determined that the display did not display conditions for selecting a voice mode before initiating DA. In a rule-based system, it is determined that the display was displaying conditions for selecting a voice mode not being met before initiating DA. In a probabilistic system, in some examples, it is determined that the display did not display an increased probability of voice mode and / or a decreased probability of mixed mode and a decreased probability of silent mode before initiating DA. In a probabilistic system, in some examples, it is determined that the display was displaying a decreased probability of voice mode and / or an increased probability of mixed mode and an increased probability of silent mode before initiating DA.

[0403] In some examples, the context information includes the display orientation of device 800. In a rule-based system, it is determined that a display facing down satisfies the condition for selecting a voice mode. In a rule-based system, it is determined that a display facing up does not satisfy the condition for selecting a voice mode. In a probabilistic system, in some examples, it is determined that a display facing down increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, it is determined that a display facing up decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0404] In some examples, contextual information includes whether the display of device 800 is obstructed. For example, device 800 uses one or more sensors (e.g., a light sensor, a microphone, a proximity sensor) to determine whether a user cannot view the display. For example, the display may be located in a space that is at least partially enclosed (e.g., a pocket, bag, or drawer) or may be covered by an object. In a rule-based system, determining that the display is obstructed satisfies the conditions for selecting a voice mode. In a rule-based system, determining that the display is not obstructed does not satisfy the conditions for selecting a voice mode. In a probabilistic system, in some examples, determining that the display is obstructed increases the probability of a voice mode and / or decreases the probability of a mixed mode and decreases the probability of a silent mode. In a probabilistic system, in some examples, determining that the display is not obstructed decreases the probability of a voice mode and / or increases the probability of a mixed mode and increases the probability of a silent mode.

[0405] In some examples, contextual information includes whether device 800 is coupled to an external audio output device (e.g., headphones, Bluetooth device, speaker). In a rule-based system, determining that device 800 is coupled to an external device satisfies the conditions for selecting a voice mode. In a rule-based system, determining that device 800 is not coupled to an external device does not satisfy the conditions for selecting a voice mode. In a probabilistic system, in some examples, determining that device 800 being coupled to an external device increases the probability of a voice mode and / or decreases the probability of a mixed mode and decreases the probability of a silent mode. In a probabilistic system, in some examples, determining that device 800 not being coupled to an external device decreases the probability of a voice mode and / or increases the probability of a mixed mode and increases the probability of a silent mode.

[0406] In some examples, the contextual information includes whether the user's gaze is directed towards device 800. In a rule-based system, determining that the user's gaze is not directed towards device 800 satisfies the conditions for selecting a voice mode. In a rule-based system, determining that the user's gaze is directed towards device 800 does not satisfy the conditions for selecting a voice mode. In a probabilistic system, in some examples, determining that the user's gaze is not directed towards device 800 increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, determining that the user's gaze is directed towards device 800 decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0407] In some examples, the contextual information includes whether a predetermined type of gesture of device 800 was detected within a predetermined duration prior to the selection of a response mode. Predetermined type gestures include, for example, lift and / or rotate gestures that cause device 800 to turn on a display. In a rule-based system, failure to detect a predetermined type of gesture within a predetermined duration satisfies the conditions for selecting a voice mode. In a rule-based system, detecting a predetermined type of gesture within a predetermined duration does not satisfy the conditions for selecting a voice mode. In a probabilistic system, in some examples, failure to detect a predetermined type of gesture within a predetermined duration increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, detecting a predetermined type of gesture within a predetermined duration decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0408] In some examples, contextual information includes the direction of the natural language input. In a rule-based system, determining that the direction of the natural language input is not towards the device 800 orientation satisfies the condition used for selecting a voice mode. In a rule-based system, determining that the direction of the natural language input is towards the device 800 orientation does not satisfy the condition used for selecting a voice mode. In a probabilistic system, in some examples, determining that the direction of the natural language input is not towards the device 800 orientation increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, determining that the direction of the natural language input is towards the device 800 orientation decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0409] In some examples, contextual information includes whether a touch performed on device 800 (e.g., user input selecting a response enablement) is detected within a predetermined duration prior to selecting a response mode. In a rule-based system, failure to detect a touch within the predetermined duration satisfies the conditions for selecting a voice mode. In a rule-based system, detecting a touch within the predetermined duration does not satisfy the conditions for selecting a voice mode. In a probabilistic system, in some examples, failure to detect a touch within the predetermined duration increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, detecting a touch within the predetermined duration decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0410] In some examples, contextual information includes whether the natural language input is typed input, for example, the opposite of spoken input. In rule-based systems, determining that the natural language input is not typed input satisfies the conditions used to select a voice mode. In rule-based systems, determining that the natural language input is typed input does not satisfy the conditions used to select a voice mode. In probabilistic systems, in some examples, determining that the natural language input is not typed input increases the probability of a voice mode and / or decreases the probability of a mixed mode and decreases the probability of a silent mode. In probabilistic systems, in some examples, determining that the natural language input is typed input decreases the probability of a voice mode and / or increases the probability of a mixed mode and increases the probability of a silent mode.

[0411] In some examples, contextual information includes whether device 800 received a notification (e.g., text message, email message, application notification, system notification) within a predetermined duration (e.g., 10, 15, 30 seconds) before selecting a response mode. In a rule-based system, not receiving a notification within the predetermined duration satisfies the condition for selecting a voice mode. In a rule-based system, receiving a notification within the predetermined duration does not satisfy the condition for selecting a voice mode. In a probabilistic system, in some examples, not receiving a notification within the predetermined duration increases the probability of voice mode and / or decreases the probability of mixed mode and decreases the probability of silent mode. In a probabilistic system, in some examples, receiving a notification within the predetermined duration decreases the probability of voice mode and / or increases the probability of mixed mode and increases the probability of silent mode.

[0412] In some examples, contextual information includes the ambient noise level detected by device 800. An ambient noise level above a threshold may indicate that the user cannot hear the audio output, for example, because the user is in a noisy environment. Therefore, detecting an ambient noise level above the threshold may recommend selecting a silent mode (because device 800 provides audio output in both voice and mixed modes). Thus, in a rule-based system, determining that an ambient noise level below the threshold satisfies the conditions for selecting voice mode, satisfies the conditions for selecting mixed mode (at box 1308), and does not satisfy the conditions for selecting silent mode (at box 1308). In a rule-based system, determining that an ambient noise level above the threshold does not satisfy the conditions for selecting voice mode, does not satisfy the conditions for selecting mixed mode (at box 1308), and satisfies the conditions for selecting silent mode (at box 1308). In a probabilistic system, in some examples, determining that an ambient noise level below the threshold increases the probability of voice mode, increases the probability of mixed mode, and decreases the probability of silent mode. In probabilistic systems, in some examples, determining that an ambient noise level above a threshold reduces the probability of a voice mode, reduces the probability of a mixed mode, and increases the probability of a silent mode.

[0413] In some examples, contextual information includes whether the natural language input corresponds to a whispered input. A user's whispered natural language input may indicate that the user does not expect audio output, for example, because the user is in a quiet environment such as a movie theater. Therefore, determining that the natural language input corresponds to a whispered input may recommend selecting a silent mode. Thus, in a rule-based system, determining that the natural language input does not correspond to a whispered input satisfies the conditions for selecting a voice mode, satisfies the conditions for selecting a mixed mode (in box 1308), and does not satisfy the conditions for selecting a silent mode (in box 1308). In a rule-based system, determining that the natural language input corresponds to a whispered input does not satisfy the conditions for selecting a voice mode, does not satisfy the conditions for selecting a mixed mode (in box 1308), and satisfies the conditions for selecting a silent mode (in box 1308). In a probabilistic system, in some examples, determining that the natural language input does not correspond to a whispered input increases the probability of a voice mode, increases the probability of a mixed mode, and decreases the probability of a silent mode. In probabilistic systems, in some examples, determining that a natural language input corresponds to a whispered input reduces the probability of a speech mode, reduces the probability of a mixed mode, and increases the probability of a silent mode.

[0414] In some examples, contextual information includes whether the user's schedule information indicates that the user is busy (e.g., in a meeting). Schedule information indicating that the user is busy may recommend selecting the mute mode. Therefore, in a rule-based system, determining that the schedule information indicates the user is not busy satisfies the conditions for selecting the voice mode, satisfies the conditions for selecting the mixed mode (at box 1308), and does not satisfy the conditions for selecting the mute mode (at box 1308). In a rule-based system, determining that the schedule information indicates the user is busy does not satisfy the conditions for selecting the voice mode, does not satisfy the conditions for selecting the mixed mode (at box 1308), and satisfies the conditions for selecting the mute mode (at box 1308). In a probabilistic system, in some examples, determining that the schedule information indicating the user is not busy increases the probability of the voice mode, increases the probability of the mixed mode, and decreases the probability of the mute mode. In a probabilistic system, in some examples, determining that the schedule information indicating the user is busy decreases the probability of the voice mode, decreases the probability of the mixed mode, and increases the probability of the mute mode.

[0415] In some examples, contextual information includes whether device 800 is in a vehicle. In some examples, device 800 determines whether it is in a vehicle by detecting pairing with the vehicle (e.g., via Bluetooth or via Apple Inc.'s CarPlay) or by determining the activation of a setting that indicates device 800 is in a vehicle (e.g., a Do Not Disturb setting while driving). In some examples, device 800 uses its location and / or speed to determine whether it is in a vehicle. For example, data indicating that device 800 is traveling at 65 miles per hour on a highway could indicate that device 800 is in a vehicle.

[0416] In a rule-based system, it is determined that device 800 is in the vehicle and meets the conditions for selecting a voice mode. In a rule-based system, it is determined that device 800 is not in the vehicle and does not meet the conditions for selecting a voice mode. In a probabilistic system, in some examples, it is determined that device 800 being in the vehicle increases the probability of the voice mode and / or decreases the probability of the mixed mode and decreases the probability of the silent mode. In a probabilistic system, in some examples, it is determined that device 800 not being in the vehicle decreases the probability of the voice mode and / or increases the probability of the mixed mode and increases the probability of the silent mode.

[0417] Figure 14 A device 800 is shown that presents a response based on a voice response pattern when it is determined that the user is in a vehicle (e.g., driving), according to various examples. As shown, device 800 displays a DA use...

Claims

1. A method for operating a digital assistant, the method comprising: In electronic devices having one or more processors, memory, and displays: Receive natural language input; Initiate the digital assistant; Based on the initiation of the digital assistant, a response packet in response to the natural language input is obtained; After receiving the natural language input, the first response mode of the digital assistant is selected from a plurality of digital assistant response modes based on one or more conditions for selecting the first response mode, based on context information associated with the electronic device, wherein the context information includes whether the user's gaze is directed toward the electronic device, and wherein the context information includes the physical state of the electronic device. as well as In response to selecting the first response mode, the digital assistant presents the response package according to the first response mode, wherein the response package includes: A first text associated with the digital assistant's response capability representation, wherein the first text is a complete response to the natural language input; as well as A second text associated with the digital assistant's response capability representation, the second text being a simplified response to the natural language input.

2. The method of claim 1, wherein selecting the first response mode includes determining: Is it to display the second text without providing audio output representing the second text; or... Provide the audio output representing the second text without displaying the second text.

3. The method of claim 1, wherein selecting the first response mode includes determining whether to provide audio output representing the first text.

4. The method according to claim 1, wherein: The first response mode is a silent response mode; and The digital assistant presents the response package according to the first response pattern, including: Display the digital assistant's response capability indication; and The second text is displayed without providing a second audio output representing the second text.

5. The method according to claim 4, wherein: The context information includes digital assistant voice feedback settings; and Based on the determination that the digital assistant's voice feedback setting indicates that no voice feedback is provided, the first response mode is selected.

6. The method according to claim 4, wherein: The context information includes the detection of physical contact with the electronic device, the physical contact being used to initiate the digital assistant; and Based on the detection of the physical contact, the first response mode is selected.

7. The method according to claim 4, wherein: The context information includes whether the electronic device is in a locked state; and Based on the determination that the electronic device is not in the locked state, the first response mode is selected.

8. The method according to claim 4, wherein: The context information includes whether the display of the electronic device was showing something before the digital assistant was initiated; and Based on the determination that the display was showing before the digital assistant was initiated, the first response mode is selected.

9. The method according to claim 4, wherein: The context information includes the detection of touches performed on the electronic device within a predetermined duration prior to the selection of the first response mode; and Based on the detected touch, the first response mode is selected.

10. The method according to claim 4, wherein: The context information includes the detection of predetermined gestures of the electronic device during a second predetermined duration prior to the selection of the first response mode; and Based on the detected pre-determined gesture, the first response mode is selected.

11. The method according to claim 1, wherein: The first response mode is a hybrid response mode; and The digital assistant presenting the response package according to the first response mode includes: displaying a digital assistant response enable representation and providing a second audio output representing the second text without displaying the second text.

12. The method according to claim 11, wherein: The context information includes digital assistant voice feedback settings; and Based on the determination that the digital assistant's voice feedback settings indicate that voice feedback should be provided, the first response mode is selected.

13. The method according to claim 11, wherein: The context information includes the detection of physical contact with the electronic device, the physical contact being used to initiate the digital assistant; and Based on the detection of the physical contact, the first response mode is selected.

14. The method of claim 11, wherein: The context information includes whether the electronic device is in a locked state; and Based on the determination that the electronic device is not in the locked state, the first response mode is selected.

15. The method according to claim 11, wherein: The context information includes whether the display of the electronic device was showing something before the digital assistant was initiated; and Based on the determination that the display was showing before the digital assistant was initiated, the first response mode is selected.

16. The method of claim 11, wherein: The context information includes the detection of touches performed on the electronic device within a predetermined duration prior to the selection of the first response mode; and Based on the detected touch, the first response mode is selected.

17. The method of claim 11, wherein: The context information includes the detection of predetermined gestures of the electronic device during a second predetermined duration prior to the selection of the first response mode; and Based on the detected pre-determined gesture, the first response mode is selected.

18. The method according to claim 1, wherein: The first response mode is the voice response mode; and The digital assistant presenting the response package according to the first response pattern includes providing audio output representing the first text.

19. The method of claim 18, wherein: The context information includes determining that the electronic device is in a vehicle; and Based on the determination that the electronic device is in the vehicle, the first response mode is selected.

20. The method of claim 18, wherein: The context information includes determining that the electronic device is coupled to an external audio output device; and Based on the determination that the electronic device is coupled to the external audio output device, the first response mode is selected.

21. The method according to claim 18, wherein: The context information includes the detection of the voice input that initiated the digital assistant; and Based on the detected voice input, the first response mode is selected.

22. The method of claim 18, wherein: The context information includes whether the electronic device is in a locked state; and Based on the determination that the electronic device is in the locked state, the first response mode is selected.

23. The method of claim 18, wherein: The context information includes whether the display of the electronic device was showing something before the digital assistant was initiated; and Based on the determination that the display of the electronic device was not displaying before the digital assistant was initiated, the first response mode is selected.

24. The method according to claim 1, further comprising: After the digital assistant presents the response packet, a second natural language input in response to the presentation of the response packet is received; In response to the second natural language input, obtain the second response packet; as well as After receiving the second natural language voice input, a second response mode of the digital assistant is selected from the plurality of digital assistant response modes, wherein the second response mode is different from the first response mode; as well as In response to the selection of the second response mode, the digital assistant presents the second response package according to the second response mode.

25. The method of claim 1, wherein after obtaining the response packet, the selection of the first response mode for the digital assistant is performed.

26. An electronic device comprising: monitor; One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 25.

27. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by one or more processors of an electronic device having a display, cause the electronic device to perform the method according to any one of claims 1 to 25.

Citation Information

Patent Citations

  • Intelligent device arbitration and control

    US10089072B2

  • System and method for inferring user intent from speech inputs

    US10176167B2

  • Method and apparatus for integrating manual input

    US20020015024A1

  • Acceleration-based theft detection system for portable electronic devices

    US20050190059A1

  • Methods and apparatuses for operating a portable device based on an accelerometer

    US20060017692A1