Replying events on first device by command using auxiliary device gesture

By detecting event alerts and determining gestures based on motion data, a robust system was developed that allows commands to be executed on the first device using a wireless headset. This solves the problem of low user response efficiency in existing systems and improves the convenience of message processing.

CN121889758APending Publication Date: 2026-04-17APPLE INC
View PDF 24 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
APPLE INC
Filing Date
2024-09-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing systems do not provide a robust framework that allows users to perform actions using wireless headphones, particularly lacking in audible cues for affirmative or negative responses to messages or for canceling long message readings.

Method used

The first electronic device detects event alarms and determines gestures based on motion data from the second electronic device, thereby enabling corresponding outputs and task execution.

Benefits of technology

It provides a robust system that allows users to execute commands on the first device via assistive device gestures, improving message response efficiency and the ease of canceling long message readouts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121889758A_ABST
    Figure CN121889758A_ABST
Patent Text Reader

Abstract

Systems and processes are provided for commands using auxiliary device gestures. In some embodiments, a method includes a first electronic detecting an event alert and causing a message to be provided at a second electronic device, where the message is associated with the event alert. In some implementations, a method includes a first electronic device receiving motion data from a second electronic device corresponding to movement of the second electronic device, and determining a gesture based on the motion data. In some implementations, a method includes a first electronic device causing a first output to be provided at a second electronic device based on a gesture, and performing a first task associated with an event alert.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Patent Application No. 18 / 893,818, filed September 23, 2024, entitled "COMMANDS USING SECONDARY DEVICEGESTURES"; U.S. Provisional Patent Application No. 63 / 670,578, filed July 12, 2024, entitled "COMMANDS USING SECONDARY DEVICE GESTURES"; U.S. Provisional Patent Application No. 63 / 657,828, filed June 8, 2024, entitled "COMMANDS USING SECONDARY DEVICE GESTURES"; and U.S. Provisional Patent Application No. 63 / 541,604, filed September 29, 2023, entitled "COMMANDS USING SECONDARY DEVICEGESTURES". The entire contents of each of these applications are incorporated herein by reference. Technical Field

[0002] This application relates in general to the detection of gestures, and more specifically to commands using gestures from assistive devices. Background Technology

[0003] Electronic devices are used to perform a wide variety of tasks. In some cases, users can use multiple electronic devices interchangeably. For example, a user can use a smartphone connected to a wireless headset to listen to music, send and receive messages, make phone calls, and so on. Event alerts can also be received at the smartphone, so that corresponding information is delivered to the headset. However, conventional systems do not provide a robust framework for allowing users to perform actions using wireless headsets. For example, traditional systems typically do not allow users to perform head gestures to respond positively or negatively to corresponding messages. Furthermore, such systems are inadequate in terms of audible prompts for quickly canceling the reading of long messages. Therefore, there is a need for an improved system for commands using auxiliary device gestures. Summary of the Invention

[0004] Systems and processes are provided for commands to use gestures on assistive devices. In some embodiments, a method includes: a first electronic device detecting an event alarm, and causing a message to be provided at a second electronic device, wherein the message is associated with the event alarm. In some embodiments, the method includes: the first electronic device receiving motion data corresponding to movement of the second electronic device from the second electronic device, and determining a gesture based on the motion data. In some embodiments, the method includes: the first electronic device causing a first output to be provided at the second electronic device based on the gesture, and performing a first task associated with the event alarm. Attached Figure Description

[0005] Figure 1 These are block diagrams illustrating various examples of systems and environments used to implement digital assistants.

[0006] Figure 2A This is a block diagram illustrating a portable multi-functional device that implements the client-side portion of a digital assistant according to various examples.

[0007] Figure 2B This is a block diagram illustrating exemplary components for event handling based on various examples.

[0008] Figure 3 Portable multi-functional devices that implement the client-side portion of a digital assistant, based on various examples, are illustrated.

[0009] Figure 4A This is a block diagram of an exemplary multifunctional device with a display and a touch-sensitive surface, based on various examples.

[0010] Figures 4B to 4G This example demonstrates how to perform an operation using an Application Programming Interface (API).

[0011] Figure 5A Examples of user interfaces for menus on portable multi-functional devices, based on various examples, are shown.

[0012] Figure 5B Exemplary user interfaces of multifunctional devices having a touch-sensitive surface separate from the display are illustrated according to various examples.

[0013] Figure 6A Examples of personal electronic devices are shown, based on various examples.

[0014] Figure 6B This is a block diagram illustrating various examples of personal electronic devices.

[0015] Figure 7A It is a block diagram illustrating a digital assistant system or its server portion according to various examples.

[0016] Figure 7B Examples are shown based on various examples. Figure 7A The digital assistant's functions are shown in the image.

[0017] Figure 7C Examples are provided for a portion of the knowledge ontology based on various examples.

[0018] Figures 8A to 8F The system illustrates commands for using assistive device gestures, based on various examples.

[0019] Figures 9A to 9CThe system illustrates commands for using assistive device gestures, based on various examples.

[0020] Figure 10 The process of using assistive device gestures is illustrated with various examples.

[0021] Figure 11 The process of using assistive device gestures is illustrated with various examples. Detailed Implementation

[0022] The accompanying drawings will be referenced in the following description of the examples, which illustrate specific examples that can be implemented by way of example. It should be understood that other examples may be used and structural changes may be made without departing from the scope of the individual examples.

[0023] Generally, conventional systems associated with gesture-based commands do not provide users with a robust framework for facilitating device interaction. Specifically, systems involving two devices do not provide an efficient mechanism by which a coupled wearable device (such as a headset) executes commands on a first device (such as a smartphone). Such systems simply read messages or notifications to the user without offering the option to enter affirmative or negative responses to questions within the message or to query additional information about the notification content. Furthermore, these systems do not provide an efficient mechanism to cancel the reading of messages currently being presented on the device. Therefore, an improved system for context-sensitive response suggestions is desired.

[0024] Although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the various examples described, a first input may be referred to as a second input, and similarly, a second input may be referred to as a first input. Both the first and second inputs are inputs, and in some cases, they are independent and distinct inputs.

[0025] The terminology used in the description of the various described examples herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various described examples and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context expressly indicates otherwise. It will also be understood that, as used herein, the term “and / or” refers to and covers any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprising” and / or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0026] Depending on the context, the term "if" can be interpreted as meaning "when," "at," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrases "if it is determined that..." or "if [the stated condition or event] is detected" can be interpreted as meaning "when it is determined that..." or "in response to determination that..." or "when [the stated condition or event] is detected," or "in response to detection of [the stated condition or event]."

[0027] 1. System and Environment Figure 1 A block diagram of system 100 according to various examples is illustrated. In some examples, system 100 implements a digital assistant. The terms "digital assistant," "virtual assistant," "intelligent automated assistant," or "automated digital assistant" refer to any information processing system that interprets natural language input in spoken and / or textual form to infer user intent and performs actions based on the inferred user intent. For example, to act on an inferred user intent, the system performs one or more of the following steps: identifying a task flow with steps and parameters designed to achieve the inferred user intent; inputting a specific request into the task flow based on the inferred user intent; executing the task flow by invoking programs, methods, services, APIs, etc.; and generating an output response to the user in an audible (e.g., verbal) and / or visual form.

[0028] Specifically, a digital assistant can accept user requests that are at least partially in the form of natural language commands, requests, statements, narration, and / or inquiries. Typically, user requests seek an informational response or task from the digital assistant. A satisfactory response to a user request includes providing the requested informational response, performing the requested task, or a combination of both. For example, a user asks a digital assistant a question such as, “Where am I now?” Based on the user’s current location, the digital assistant answers, “You are near the west entrance of Central Park.” The user also requests a task, such as, “Please invite my friends to my girlfriend’s birthday party next week.” In response, the digital assistant can confirm the request by saying “Okay, coming right away,” and then send the appropriate calendar invitations to each of the user’s friends listed in the user’s electronic address book. During the performance of the requested task, the digital assistant sometimes interacts with the user in a sustained conversation involving multiple exchanges of information over extended periods. Many other methods exist for interacting with a digital assistant to request information or perform various tasks. In addition to providing verbal responses and taking programmed actions, digital assistants also provide responses in other forms of video or audio, such as text, alerts, music, video, animation, etc.

[0029] like Figure 1 As shown, in some examples, the digital assistant is implemented according to a client-server model. The digital assistant includes a client-side portion 102 (hereinafter referred to as "DA client 102") executing on user device 104 and a server-side portion 106 (hereinafter referred to as "DA server 106") executing on server system 108. DA client 102 communicates with DA server 106 via one or more networks 110. DA client 102 provides client-side functionality, such as user-oriented input and output processing, and communication with DA server 106. DA server 106 provides server-side functionality for any number of DA clients 102, each residing on a corresponding user device 104.

[0030] In some examples, DA server 106 includes a client-facing I / O interface 112, one or more processing modules 114, data and models 116, and an I / O interface 118 to external services. The client-facing I / O interface 112 facilitates client-facing input and output processing of DA server 106. One or more processing modules 114 utilize data and models 116 to process verbal input and determine user intent based on natural language input. Furthermore, one or more processing modules 114 perform task execution based on the inferred user intent. In some examples, DA server 106 communicates with external services 120 via network 110 to complete tasks or collect information. The I / O interface 118 to external services facilitates such communication.

[0031] User equipment 104 can be any suitable electronic device. In some examples, user equipment 104 is a portable multi-functional device (e.g., see reference below). Figure 2A The described device 200), multi-functional device (e.g., see reference below) Figure 4A The described device 400) or personal electronic device (e.g., referred to below) Figures 6A to 6B The described device (600) is a portable multi-functional device, for example, a mobile phone that also includes other functions such as PDA and / or music player functionality. Specific examples of portable multi-functional devices include the Apple Watch from Apple Inc. (Cupertino, California). ® iPhone ® iPod Touch ® and iPad ® Devices. Other examples of portable multifunction devices include, but are not limited to, headphones / headsets, speakers, and laptops or tablets. Additionally, in some examples, user device 104 is a non-portable multifunction device. Specifically, user device 104 is a desktop computer, game console, speaker, television, or set-top box. In some examples, user device 104 includes a touch-sensitive surface (e.g., a touchscreen display and / or touchpad). Furthermore, user device 104 may optionally include one or more other physical user interface devices, such as a physical keyboard, mouse, and / or joystick. Various examples of electronic devices such as multifunction devices are described in more detail below.

[0032] Examples of communication networks 110 include local area networks (LANs) and wide area networks (WANs), such as the Internet. Communication network 110 is implemented using any known network protocol, including various wired or wireless protocols such as Ethernet, Universal Serial Bus (USB), FireWire, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi, Voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.

[0033] Server system 108 is implemented on one or more standalone data processing devices or a distributed computer network. In some examples, server system 108 also utilizes various virtual devices and / or services from third-party service providers (e.g., third-party cloud service providers) to provide potential computing and / or infrastructure resources for server system 108.

[0034] In some examples, user equipment 104 communicates with DA server 106 via a second user equipment 122. The second user equipment 122 is similar to or identical to user equipment 104. For example, the second user equipment 122 is similar to the one described below. Figure 2A , Figure 4A and Figures 6A to 6B The described devices are 200, 400, or 600. User equipment 104 is configured to be communicatively coupled to a second user equipment 122 via a direct communication connection (such as Bluetooth, NFC, BTLE, etc.) or via a wired or wireless network (such as a local Wi-Fi network). In some examples, the second user equipment 122 is configured to act as a proxy between user equipment 104 and DA server 106. For example, a DA client 102 of user equipment 104 is configured to send information (e.g., a user request received at user equipment 104) to DA server 106 via the second user equipment 122. DA server 106 processes the information and returns relevant data (e.g., data content in response to the user request) to user equipment 104 via the second user equipment 122.

[0035] In some examples, user equipment 104 is configured to send a shortened request for data to a second user equipment 122 to reduce the amount of information sent from user equipment 104. The second user equipment 122 is configured to determine supplementary information to add to the shortened request to generate a complete request to be sent to DA server 106. This system architecture can advantageously allow user equipment 104 (e.g., a watch or similar compact electronic device) with limited communication capabilities and / or limited battery power (e.g., a second user equipment 122 with strong communication capabilities and / or battery power, such as a mobile phone, laptop computer, tablet computer, etc.) acting as a proxy to DA server 106 to access the services provided by DA server 106. Although Figure 1 Only two user devices 104 and 122 are shown in this document, but it should be understood that in some examples, system 100 may include any number and type of user devices configured in this agent configuration to communicate with DA server system 106.

[0036] Although Figure 1 The digital assistant illustrated includes both a client-side component (e.g., DA client 102) and a server-side component (e.g., DA server 106), but in some examples, the digital assistant's functionality is implemented as a standalone application installed on the user's device. Furthermore, the functional division between the client and server components of the digital assistant can vary in different implementations. For example, in some examples, the DA client is a thin client that only provides user-facing input and output processing functions and delegates all other functions of the digital assistant to the backend server.

[0037] 2. Electronic devices Now let’s turn our attention to the implementation of electronic devices for the client-side portion of a digital assistant. Figure 2A This is a block diagram illustrating a portable multi-functional device 200 with a touch-sensitive display system 212 according to some embodiments. The touch-sensitive display 212 is sometimes referred to as a “touchscreen” for convenience, and is sometimes referred to as or called a “touch-sensitive display system.” Device 200 includes a memory 202 (which optionally includes one or more computer-readable storage media), a memory controller 222, one or more processing units (CPUs) 220, a peripheral interface 218, RF circuitry 208, audio circuitry 210, a speaker 211, a microphone 213, an input / output (I / O) subsystem 206, other input control devices 216, and an external port 224. Device 200 optionally includes one or more optical sensors 264. Device 200 optionally includes one or more contact strength sensors 265 for detecting the intensity of contact on device 200 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 212 of device 200). Device 200 may optionally include one or more haptic output generators 267 for generating haptic outputs on device 200 (e.g., generating haptic outputs on a haptic surface such as the haptic display system 212 of device 200 or the touchpad 455 of device 400). These components may optionally communicate via one or more communication buses or signal lines 203.

[0038] As used in this specification and claims, the term "intensity" of contact on a tactile surface refers to the force or pressure (force per unit area) of a contact (e.g., finger contact) on a tactile surface, or to a substitute (alternative) for the force or pressure of a contact on a tactile surface. The intensity of contact has a range of values ​​that includes at least four different values ​​and more typically hundreds of different values ​​(e.g., at least 256). The intensity of contact may optionally be determined (or measured) using various methods and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the tactile surface may optionally be used to measure the force at different points on the tactile surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted average) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus may optionally be used to determine the pressure of the stylus on the tactile surface. Alternatively, the size and / or change of the contact area detected on the touch-sensitive surface, the capacitance and / or change of the touch-sensitive surface adjacent to the contact, and / or the resistance and / or change of the touch-sensitive surface adjacent to the contact may be used as substitutes for the force or pressure of the contact on the touch-sensitive surface. In some embodiments, the substitute measurement of the contact force or pressure is used directly to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the substitute measurement result). In some embodiments, the substitute measurement of the contact force or pressure is converted into an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of the contact as an attribute of user input allows users to access additional device functionality that would otherwise be inaccessible to the user on a smaller device with limited physical space, which is used (e.g., on a touch-sensitive display) to display an indication and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls, such as knobs or buttons).

[0039] As used in this specification and claims, the term "tactile output" refers to a physical displacement of the device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., the housing), or a displacement of a component relative to the center of mass of the device, detected by the user using the user's tactile sense. For example, when the device or a component of the device comes into contact with a touch-sensitive surface (e.g., a finger, palm, or other part of the user's hand), the tactile output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or touchpad) may optionally be interpreted by the user as a "press-click" or "release-click" on a physically actuated button. In some cases, the user will feel a tactile sensation, such as a "press-click" or "release-click," even when a physically actuated button associated with the touch-sensitive surface, which has been physically pressed (e.g., displaced) by the user's movement, does not move. As another example, even when the smoothness of the tactile surface remains unchanged, the movement of the tactile surface can optionally be interpreted or sensed by the user as the "roughness" of the tactile surface. While such interpretations of touch by users will be limited by the individualized sensory perceptions of the user, many sensory perceptions of touch are common to most users. Therefore, when a tactile output is described as corresponding to a specific sensory perception of the user (e.g., "release click", "press click", "roughness"), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or its components that will generate the sensory perception described by a typical (or common) user.

[0040] It should be understood that device 200 is merely an example of a portable multifunctional device, and device 200 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of these components. Figure 2A The various components shown are implemented in hardware, software, or a combination of both, including one or more signal processing and / or application-specific integrated circuits.

[0041] Memory 202 includes one or more computer-readable storage media. These computer-readable storage media are, for example, tangible and non-transitory. Memory 202 includes high-speed random access memory and also includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 222 controls other components of device 200 to access memory 202.

[0042] In some examples, the non-transitory computer-readable storage medium of memory 202 is used to store instructions (e.g., aspects of the processes described below) for use by or in conjunction with an instruction execution system, apparatus, or device, such as a computer-based system, a processor-integrated system, or other system from which instructions can be fetched and executed. In other examples, instructions (e.g., aspects of the processes described below) are stored on a non-transitory computer-readable storage medium (not shown) of server system 108, or partitioned between the non-transitory computer-readable storage medium of memory 202 and the non-transitory computer-readable storage medium of server system 108.

[0043] Peripheral interface 218 is used to couple the input and output peripherals of the device to CPU 220 and memory 202. One or more processors 220 run or execute various software programs and / or instruction sets stored in memory 202 to perform various functions of device 200 and process data. In some embodiments, peripheral interface 218, CPU 220, and memory controller 222 are implemented on a single chip, such as chip 204. In some other embodiments, they are implemented on separate chips.

[0044] RF (Radio Frequency) circuit 208 receives and transmits RF signals, also known as electromagnetic signals. RF circuit 208 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 208 optionally includes well-known circuitry for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 208 optionally communicates wirelessly with networks (such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs))) and other devices. RF circuit 208 optionally includes well-known circuitry for detecting near-field communication (NFC) fields, such as via short-range communication radio components. Wireless communication may optionally employ any of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed ​​Downlink Packet Access (HSDPA), High-Speed ​​Uplink Packet Access (HSUPA), Evolution, Pure Data (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), and Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Messaging Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging Processing and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Utilizing Extended Protocol (SIMPLE), Instant Messaging and Presence Service (IMPS)) and / or Short Message Service (SMS), or any other suitable communication protocol, including communication protocols that have not yet been developed as of the date of this document submission.

[0045] Audio circuitry 210, speaker 211, and microphone 213 provide an audio interface between the user and device 200. Audio circuitry 210 receives audio data from peripheral interface 218, converts the audio data into electrical signals, and sends the electrical signals to speaker 211. Speaker 211 converts the electrical signals into sound waves that are audible to humans. Audio circuitry 210 also receives electrical signals converted from sound waves by microphone 213. Audio circuitry 210 converts the electrical signals into audio data and sends the audio data to peripheral interface 218 for processing. The audio data is retrieved from and / or sent to memory 202 and / or RF circuitry 208 via peripheral interface 218. In some embodiments, audio circuitry 210 also includes a headset jack (e.g., ...). Figure 3 (312 in the text). The headset jack provides an interface between the audio circuitry 210 and a removable audio input / output peripheral device, such as an output-only headset or a headset having both an output (e.g., a mono-ear headset or a binaural headset) and an input (e.g., a microphone).

[0046] I / O subsystem 206 couples input / output peripherals (such as touchscreen 212 and other input control devices 216) on device 200 to peripheral interface 218. I / O subsystem 206 optionally includes display controller 256, optical sensor controller 258, intensity sensor controller 259, haptic feedback controller 261, and one or more input controllers 260 for other input or control devices. The one or more input controllers 260 receive electrical signals from / transmit electrical signals to the other input control device 216. Other input control devices 216 optionally include physical buttons (e.g., push-buttons, rocker buttons, etc.), dial pads, slide switches, joysticks, click dials, etc. In some alternative embodiments, input controller 260 may optionally be coupled to (or not coupled to) any of the following: keyboard, infrared port, USB port, and pointing device such as mouse. One or more buttons (e.g., ... Figure 3 Optionally, 308 may include increase / decrease buttons for volume control of speaker 211 and / or microphone 213. These one or more buttons may optionally include a push-button (e.g., Figure 3 (306 in the middle).

[0047] A rapid press of the down button disengages the touchscreen 212 from its lock or initiates a process of unlocking the device using gestures on the touchscreen, as described in U.S. Patent Application 11 / 322,549 (U.S. Patent No. 7,657,849), filed December 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," the entire contents of which are incorporated herein by reference. A longer press of the down button (e.g., 306) powers the device 200 on or off. The user can customize the function of one or more buttons. The touchscreen 212 is used to implement virtual buttons or soft buttons and one or more soft keyboards.

[0048] The touch-sensitive display 212 provides input and output interfaces between the device and the user. The display controller 256 receives electrical signals from and / or transmits electrical signals to the touchscreen 212. The touchscreen 212 displays visual output to the user. Visual output includes graphics, text, icons, video, and any combination thereof (collectively, "graphics"). In some embodiments, some or all of the visual output corresponds to user interface objects.

[0049] Touchscreen 212 has a touch-sensitive surface, sensor, or sensor array that accepts input from a user based on tactile and / or haptic contact. Touchscreen 212 and display controller 256 (along with any associated modules and / or instruction set in memory 202) detect contact on touchscreen 212 (and any movement or interruption of that contact) and translate the detected contact into interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on touchscreen 212. In an exemplary embodiment, the contact point between touchscreen 212 and the user corresponds to the user's finger.

[0050] Touchscreen 212 uses LCD (Liquid Crystal Display) technology, LPD (Light Emitting Polymer Display) technology, or LED (Light Emitting Diode) technology, but other display technologies may be used in other embodiments. Touchscreen 212 and display controller 256 use any of a variety of touch sensing technologies currently known or to be developed thereafter, as well as other proximity sensor arrays or other elements for determining one or more points of contact with touchscreen 212 to detect contact and any movement or interruption. These various touch sensing technologies include, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as in the iPhone from Apple Inc. (Cupertino, California). ® and iPod Touch ® The technology used.

[0051] In some embodiments, the touchscreen 212's touch-sensitive display is similar to the multi-touch panel described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman) and / or U.S. Patent Publication 2002 / 0015024A1, all of which are incorporated herein by reference in their entirety. However, the touchscreen 212 displays visual output from the device 200, while the touch-sensitive panel does not provide visual output.

[0052] The touch-sensitive display in some embodiments of the touchscreen 212 is described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, filed May 2, 2006, “Multipoint Touch Surface Controller”; (2) U.S. Patent Application No. 10 / 840,862, filed May 6, 2004, “Multipoint Touchscreen”; (3) U.S. Patent Application No. 10 / 903,964, filed July 30, 2004, “Gestures For Touch Sensitive Input Devices”; (4) U.S. Patent Application No. 11 / 048,264, filed January 31, 2005, “Gestures For Touch Sensitive Input Devices”; (5) U.S. Patent Application No. 11 / 038,590, filed January 18, 2005, “Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices”; (6) U.S. Patent Application No. 11 / 228,758, filed September 16, 2005, “Virtual Input Device Placement On A Touch Screen User Interface”; (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, “Operation Of A Computer With A Touch Screen Interface”; (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, “Activating Virtual Keys Of A Touch-Screen Virtual Keyboard”; and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, “Multi-Functional Hand-Held Device”. The full text of all these applications is incorporated herein by reference.

[0053] Touchscreen 212 has a video resolution of over 100 dpi, for example. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. The user interacts with touchscreen 212 using any suitable object or accessory such as a stylus, finger, etc. In some embodiments, the user interface is designed to function primarily through finger-based touch and gestures, which may be less precise than stylus-based input due to the larger contact area of ​​a finger on the touchscreen. In some embodiments, the device translates coarse finger-based input into precise pointer / cursor positioning or commands for performing the user-desired actions.

[0054] In some embodiments, in addition to the touchscreen, device 200 also includes a touchpad (not shown) for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of ​​the device that, unlike the touchscreen, does not display visual output. The touchpad is a touch-sensitive surface separate from the touchscreen 212, or an extension of the touch-sensitive surface formed by the touchscreen.

[0055] The device 200 also includes a power system 262 for supplying power to various components. The power system 262 includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in the portable device.

[0056] The device 200 also includes one or more optical sensors 264. Figure 2A An optical sensor 264 is shown coupled to an optical sensor controller 258 in the I / O subsystem 206. The optical sensor 264 includes a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS) phototransistor. The optical sensor 264 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In conjunction with an imaging module 243 (also called a camera module), the optical sensor 264 captures still images or video. In some embodiments, the optical sensor is located at the rear of the device 200, opposite to the touchscreen display 212 at the front of the device, such that the touchscreen display is used as a viewfinder for still image and / or video image acquisition. In some embodiments, the optical sensor is located at the front of the device, such that an image of the user is acquired for use in video conferencing while the user views other video conferencing participants on the touchscreen display. In some embodiments, the positioning of the optical sensor 264 can be changed by the user (e.g., by rotating the lenses and sensors in the device housing), such that a single optical sensor 264 is used in conjunction with the touchscreen display for both video conferencing and still image and / or video image acquisition.

[0057] The device 200 may optionally also include one or more contact strength sensors 265. Figure 2A A contact strength sensor 265 is shown coupled to a strength sensor controller 259 in I / O subsystem 206. The contact strength sensor 265 may optionally include one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric sensors, optical force sensors, capacitive touch-sensitive surfaces, or other strength sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). The contact strength sensor 265 receives contact strength information (e.g., pressure information or a substitute for pressure information) from the environment. In some embodiments, at least one contact strength sensor is co-located or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 212). In some embodiments, at least one contact strength sensor is located on the rear of device 200, opposite to the touchscreen display 212 located on the front of device 200.

[0058] The device 200 also includes one or more proximity sensors 266. Figure 2A A proximity sensor 266 coupled to a peripheral device interface 218 is shown. Alternatively, the proximity sensor 266 is coupled to an input controller 260 in an I / O subsystem 206. The proximity sensor 266 performs as described in the following U.S. patent applications numbered: 11 / 241,839, entitled "Proximity Detector In Handheld Device"; 11 / 240,788, entitled "Proximity Detector In Handheld Device"; 11 / 620,702, entitled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; 11 / 586,862, entitled "Automated Response To And Sensing Of User Activity In Portable Devices"; and 11 / 638,251, entitled "Methods And Systems For Automatic Configuration Of Peripherals", the entire contents of which are incorporated herein by reference. In some implementations, when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call), the proximity sensor is turned off and the touchscreen 212 is disabled.

[0059] The device 200 may optionally also include one or more haptic output generators 267. Figure 2AA haptic output generator coupled to a haptic feedback controller 261 in I / O subsystem 206 is shown. The haptic output generator 267 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components that convert electrical signals into haptic outputs on the device). A contact intensity sensor 265 receives haptic feedback generation instructions from a haptic feedback module 233 and generates a haptic output on device 200 that can be felt by a user of device 200. In some embodiments, at least one haptic output generator is co-located or adjacent to a haptic surface (e.g., haptic display system 212) and optionally generates the haptic output by moving the haptic surface vertically (e.g., in / outward from the surface of device 200) or laterally (e.g., backward and forward in the same plane as the surface of device 200). In some implementations, at least one haptic output generator sensor is located on the rear of the device 200, opposite to the touchscreen display 212 located on the front of the device 200.

[0060] The device 200 also includes one or more accelerometers 268. Figure 2A An accelerometer 268 is shown coupled to a peripheral device interface 218. Alternatively, the accelerometer 268 is coupled to an input controller 260 in an I / O subsystem 206. The accelerometer 268 performs as described in the following U.S. patent publications: U.S. Patent Publication No. 20050190059, “Acceleration-based Theft Detection System for Portable Electronic Devices” and U.S. Patent Publication No. 20060017692, “Methods and Apparatuses For Operating A Portable Device Based On An Accelerometer,” the entire contents of which are incorporated herein by reference. In some embodiments, information is displayed on a touchscreen display in portrait or landscape view based on analysis of data received from one or more accelerometers. The device 200 may optionally include, in addition to the accelerometer 268, a magnetometer (not shown) and a GPS (or GLONASS or other global navigation system) receiver (not shown) for obtaining information about the location and orientation (e.g., portrait or landscape) of the device 200.

[0061] In some embodiments, software components stored in memory 202 include an operating system 226, a communication module (or instruction set) 228, a contact / motion module (or instruction set) 230, a graphics module (or instruction set) 232, a text input module (or instruction set) 234, a Global Positioning System (GPS) module (or instruction set) 235, a digital assistant client module 229, and an application (or instruction set) 236. Furthermore, memory 202 stores data and models, such as user data and models 231. Additionally, in some embodiments, memory 202 ( Figure 2A ) or 470 ( Figure 4A Storage device / global internal state 257, such as Figure 2A and Figure 4A As shown in the diagram. Device / global internal state 257 includes one or more of the following: active application state, which indicates which applications (if any) are currently active; display state, which indicates what applications, views or other information occupy various areas of the touchscreen display 212; sensor state, which includes information obtained from various sensors and input control devices 216 of the device; and position information relating to the device's position and / or orientation.

[0062] The operating system 226 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or embedded operating systems such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware and software components.

[0063] The communication module 228 facilitates communication with other devices via one or more external ports 224 and includes various software components for processing data received by the RF circuitry 208 and / or the external ports 224. The external ports 224 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices or indirectly coupled via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is for use with an iPod. ® (Trademark of Apple Inc.) The same or similar and / or compatible multi-pin (e.g., 30-pin) connectors used in Apple Inc. devices.

[0064] The contact / motion module 230 optionally detects contact with the touchscreen 212 (in conjunction with the display controller 256) and other touch-sensitive devices (e.g., a touchpad or physical click wheel). The contact / motion module 230 includes various software components for performing various operations related to contact detection, such as determining whether a contact has occurred (e.g., detecting a finger press event), determining the contact intensity (e.g., the force or pressure of the contact, or an alternative to force or pressure), determining whether there is movement of the contact and tracking movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or a contact break). The contact / motion module 230 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations can optionally be applied to single-point contact (e.g., single-finger contact) or simultaneous multi-point contact (e.g., "multi-touch" / multiple-finger contact). In some implementations, the contact / motion module 230 and the display controller 256 detect contact on the touchpad.

[0065] In some implementations, the contact / motion module 230 uses a set of one or more intensity thresholds to determine whether an operation has been performed by the user (e.g., determining whether the user has “clicked” an icon). In some implementations, at least a subset of the intensity thresholds is determined based on software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a specific physical actuator and can be adjusted without changing the physical hardware of the device 200). For example, the mouse “click” threshold of a touchpad or touchscreen can be set to any threshold in a wide range of predefined thresholds without changing the touchpad or touchscreen display hardware. Additionally, in some specific implementations, the user of the device is provided with software settings for adjusting one or more intensity thresholds in a set of intensity thresholds (e.g., by adjusting the individual intensity thresholds and / or by adjusting multiple intensity thresholds at once using system-level clicks on the “intensity” parameter).

[0066] The touch / motion module 230 optionally detects gesture input performed by the user. Different gestures on a touch-sensitive surface have different contact patterns (e.g., different movements, timings, and / or intensities of the detected contact). Therefore, gestures can optionally be detected by detecting specific contact patterns. For example, detecting a finger tap gesture includes: detecting a finger press event, and then detecting a finger lift-off (lift-away) event at the same (or substantially the same) location as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on a touch-sensitive surface includes: detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift-off (lift-away) event.

[0067] The graphics module 232 includes various known software components for rendering and displaying graphics on the touchscreen 212 or other displays, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual properties). As used herein, the term "graphics" includes any object that can be displayed to a user, and non-limitingly includes text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.

[0068] In some implementations, the graphics module 232 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 232 receives one or more codes from applications, etc., to specify the graphics to be displayed, and also receives coordinate data and other graphic attribute data if necessary, and then generates screen image data to output to the display controller 256.

[0069] The haptic feedback module 233 includes various software components for generating instructions which are used by the haptic output generator 267 to generate haptic output at one or more locations on the device 200 in response to user interaction with the device 200.

[0070] In some examples, the text input module 234, which is a component of the graphics module 232, provides a soft keyboard for entering text in various applications, such as contacts 237, email 240, IM 241, browser 247, and any other application that requires text input.

[0071] GPS module 235 determines the location of the device and provides that information for use in various applications (e.g., to telephone 238 for use in location-based dialing; to camera 243 as image / video metadata; and to applications that provide location-based services, such as weather widgets, local yellow pages widgets, and map / navigation widgets).

[0072] The digital assistant client module 229 includes various client-side digital assistant commands to provide client-side functionality for the digital assistant. For example, the digital assistant client module 229 can accept voice input (e.g., speech input), text input, touch input, and / or gesture input through various user interfaces of the portable multifunction device 200 (e.g., microphone 213, accelerometer 268, touch-sensitive display system 212, optical sensor 264, other input control devices 216, etc.). The digital assistant client module 229 can also provide audio output (e.g., speech output), visual output, and / or tactile output through various output interfaces of the portable multifunction device 200 (e.g., speaker 211, touch-sensitive display system 212, haptic output generator 267, etc.). For example, output can be provided as voice, sound, alarms, text messages, menus, graphics, video, animation, vibration, and / or combinations of both or more of these. During operation, the digital assistant client module 229 communicates with the DA server 106 using RF circuitry 208.

[0073] User data and models 231 include various data associated with the user (e.g., user-specific vocabulary data, user preference data, user-specified name pronunciation, data from the user's electronic address book, to-do lists, shopping lists, etc.) to provide client-side functionality for the digital assistant. Furthermore, user data and models 231 include various models for processing user input and determining user intent (e.g., speech recognition models, statistical language models, natural language processing models, knowledge ontology, task flow models, service models, etc.).

[0074] In some examples, the digital assistant client module 229 utilizes various sensors, subsystems, and peripherals of the portable multifunction device 200 to collect additional information from the surrounding environment of the portable multifunction device 200 to establish a context associated with the user, the current user interaction, and / or the current user input. In some examples, the digital assistant client module 229 provides the contextual information, or a subset thereof, along with the user input to the DA server 106 to help infer the user's intent. In some examples, the digital assistant also uses the contextual information to determine how to prepare output and deliver it to the user. This contextual information is referred to as contextual data.

[0075] In some examples, the contextual information accompanying user input includes sensor information such as lighting, ambient noise, ambient temperature, and images or videos of the surrounding environment. In some examples, the contextual information may also include the physical state of the device, such as device orientation, device location, device temperature, power level, speed, acceleration, motion mode, and cellular signal strength. In some examples, information related to the software state of the DA server 106, such as the operation of the portable multifunction device 200, installed programs, past and current network activity, background services, error logs, and resource usage, is provided to the DA server 106 as contextual information associated with the user input.

[0076] In some examples, the digital assistant client module 229 selectively provides information (e.g., user data 231) stored on the portable multifunction device 200 in response to a request from the DA server 106. In some examples, the digital assistant client module 229 also elicits additional input from the user via natural language dialogue or other user interfaces when requested by the DA server 106. The digital assistant client module 229 transmits this additional input to the DA server 106 to assist the DA server 106 in intent inference and / or to realize the user intent expressed in the user request.

[0077] The following is for reference. Figures 7A to 7C A more detailed description of the digital assistant follows. It should be understood that the digital assistant client module 229 may include any number of sub-modules of the digital assistant module 726 described below.

[0078] Application 236 includes the following modules (or instruction sets) or subsets or supersets: • Contacts module 237 (sometimes called address book or contact list); • Telephone module 238; • Video conferencing module 239; • Email client module 240; • Instant Messaging (IM) module 241; • Fitness support module 242; • Camera module 243 for still images and / or video images; • Image management module 244; • Video player module; • Music player module; • Browser module 247; • Calendar module 248; • Widget module 249, which in some examples includes one or more of the following: weather widget 249-1, stock market widget 249-2, calculator widget 249-3, alarm clock widget 249-4, dictionary widget 249-5 and other widgets obtained by the user and widgets created by the user 249-6; • Widget creator module 250 for forming user-created widgets 249-6; • Search module 251; • Video and music player module 252, which combines a video player module and a music player module; • Memo module 253; • Map module 254; and / or • Online video module 255.

[0079] Examples of other applications 236 stored in memory 202 include other word processing applications, other image editing applications, drawing applications, rendering applications, Java-enabled applications, encryption, digital access control, speech recognition, and speech duplication.

[0080] In conjunction with touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, contact module 237 manages an address book or contact list (e.g., stored in application internal state 292 of contact module 237 in memory 202 or memory 470), including: adding names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers or email addresses to initiate and / or facilitate communications via telephone 238, video conferencing module 239, email 240, or IM 241; and so on.

[0081] Combining RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touchscreen 212, display controller 256, contact / motion module 230, graphics module 232, and text input module 234, telephone module 238 is used to input character sequences corresponding to telephone numbers, access one or more telephone numbers in contact module 237, modify input telephone numbers, dial corresponding telephone numbers, initiate conversations, and disconnect or hang up when a conversation is completed. As noted above, wireless communication uses any of a variety of communication standards, protocols, and technologies.

[0082] Combining RF circuitry 208, audio circuitry 210, speaker 211, microphone 213, touchscreen 212, display controller 256, optical sensor 264, optical sensor controller 258, contact / motion module 230, graphics module 232, text input module 234, contact module 237, and telephone module 238, video conferencing module 239 includes executable instructions for initiating, conducting, and terminating video conferences between the user and one or more other participants, based on user instructions.

[0083] Incorporating RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, email client module 240 includes executable instructions for creating, sending, receiving, and managing emails in response to user instructions. Combined with image management module 244, email client module 240 makes it very easy to create and send emails containing still images or video images captured by camera module 243.

[0084] In conjunction with RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, the instant messaging module 241 includes executable instructions for performing the following operations: entering a character sequence corresponding to an instant message, modifying previously entered characters, sending a corresponding instant message (e.g., using Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocols for telephone-based instant messaging or using XMPP, SIMPLE, or IMPS for internet-based instant messaging), receiving an instant message, and viewing received instant messages. In some embodiments, the instant messages sent and / or received include graphics, photographs, audio files, video files, and / or other attachments supported by MMS and / or Enhanced Messaging Services (EMS). As used herein, “instant messaging” refers to both telephone-based messages (e.g., messages transmitted using SMS or MMS) and internet-based messages (e.g., messages transmitted using XMPP, SIMPLE, or IMPS).

[0085] Incorporating RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, GPS module 235, map module 254, and music player module, fitness support module 242 includes executable instructions for performing the following operations: creating fitness activities (e.g., with time, distance, and / or calorie burning goals); communicating with fitness sensors (exercise equipment); receiving fitness sensor data; calibrating sensors used to monitor fitness; selecting and playing music for fitness activities; and displaying, storing, and transmitting fitness data.

[0086] In conjunction with the touchscreen 212, display controller 256, optical sensor 264, optical sensor controller 258, contact / motion module 230, graphics module 232, and image management module 244, the camera module 243 includes executable instructions for performing the following operations: capturing still images or videos (including video streams) and storing them in memory 202, modifying the characteristics of still images or videos, or deleting still images or videos from memory 202.

[0087] In conjunction with the touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, and camera module 243, the image management module 244 includes executable instructions for performing operations such as arranging, modifying (e.g., editing) or otherwise manipulating, marking, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or video images.

[0088] Incorporating RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, browser module 247 includes executable instructions for performing the following operations: browsing the Internet according to user instructions, including searching, linking to, receiving, and displaying web pages or portions thereof, as well as links to attachments and other files on web pages.

[0089] Incorporating RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, email client module 240, and browser module 247, calendar module 248 includes executable instructions for creating, displaying, modifying, and storing calendars and associated data (e.g., calendar entries, to-dos, etc.) according to user instructions.

[0090] In conjunction with RF circuitry 208, touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234, and browser module 247, widget module 249 is a micro-application that can be downloaded and used by a user (e.g., weather widget 249-1, stock market widget 249-2, calculator widget 249-3, alarm clock widget 249-4, and dictionary widget 249-5) or a user-created micro-application (e.g., user-created widget 249-6). In some embodiments, the widget includes HTML (Hypertext Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the widget includes XML (Extensible Markup Language) files and JavaScript files (e.g., Yahoo! widgets).

[0091] Combining RF circuit 208, touch screen 212, display controller 256, touch / motion module 230, graphics module 232, text input module 234 and browser module 247, widget creator module 250 is used by the user to create widgets (e.g., to turn user-specified parts of a webpage into widgets).

[0092] In conjunction with the touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, the search module 251 includes executable instructions for performing the following operations: searching the memory 202 for text, music, sound, images, videos, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) according to user instructions.

[0093] Combining touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, and browser module 247, the video and music player module 252 includes executable instructions allowing users to download and play back recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touchscreen 212 or on an external display connected via external port 224). In some embodiments, device 200 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0094] Combining the touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, and text input module 234, the memo module 253 includes executable instructions for creating and managing memos, to-do items, etc., according to user instructions.

[0095] Combining RF circuit 208, touch screen 212, display controller 256, contact / motion module 230, graphics module 232, text input module 234, GPS module 235 and browser module 247, map module 254 is used to receive, display, modify and store maps and data associated with the maps (e.g., driving directions, data related to shops and other points of interest at or near a specific location, and other location-based data) according to user instructions.

[0096] Incorporating touchscreen 212, display controller 256, touch / motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, text input module 234, email client module 240, and browser module 247, the online video module 255 includes instructions for performing the following operations: allowing users to access, browse, receive (e.g., via streaming and / or downloading), play back (e.g., on a touchscreen or on an external display connected via external port 224), send emails with links to specific online videos, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, an instant messaging module 241 is used instead of the email client module 240 to send links to specific online videos. Additional descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” the contents of which are incorporated herein by reference in their entirety.

[0097] Each of the modules and applications described above corresponds to an executable set of instructions for performing one or more functions described above and the methods described in this patent application (e.g., computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) need not be implemented as standalone software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, a video player module can be combined with a music player module into a single module (e.g., Figure 2A (e.g., video and music player module 252). In some embodiments, memory 202 stores a subset of the aforementioned modules and data structures. Additionally, memory 202 stores additional modules and data structures not described above.

[0098] In some implementations, device 200 is a device on which the operation of a predefined set of functions is performed solely via a touchscreen and / or touchpad. By using a touchscreen and / or touchpad as the primary input control device for the operation of device 200, the number of physical input control devices (such as push-buttons, dials, etc.) on device 200 is reduced.

[0099] A predefined set of functions, uniquely performed via a touchscreen and / or touchpad, may optionally include navigation between user interfaces. In some implementations, the touchpad, when touched by a user, navigates device 200 from any user interface displayed on device 200 to the main menu, main desktop menu, or root menu. In such implementations, a "menu button" is implemented using a touchpad. In some other implementations, the menu button is a physical push-button or other physical input control device, rather than a touchpad.

[0100] Figure 2B This is a block diagram illustrating exemplary components for event handling according to some embodiments. In some embodiments, memory 202 ( Figure 2A ) or memory 470 ( Figure 4A This includes an event classifier 270 (e.g., in operating system 226) and a corresponding application 236-1 (e.g., any one of the aforementioned applications 237 to 251, 255, 480 to 490).

[0101] Event classifier 270 receives event information and determines the application 236-1 and application view 291 of application 236-1 to which the event information should be delivered. Event classifier 270 includes event monitor 271 and event dispatcher module 274. In some embodiments, application 236-1 includes application internal state 292, which indicates the current application view displayed on touch-sensitive display 212 when the application is active or running. In some embodiments, device / global internal state 257 is used by event classifier 270 to determine which application(s) is currently active, and application internal state 292 is used by event classifier 270 to determine the application view 291 to which the event information should be delivered.

[0102] In some implementations, the application internal state 292 includes additional information such as one or more of the following: recovery information to be used when the application 236-1 resumes execution, user interface state information indicating that information is being displayed or ready to be displayed by the application 236-1, a state queue for enabling the user to return to the previous state or view of the application 236-1, and a repeat / undo queue for the user's previous actions.

[0103] Event monitor 271 receives event information from peripheral device interface 218. The event information includes information about sub-events, such as user touches on touch-sensitive display 212 as part of a multi-touch gesture. Peripheral device interface 218 transmits information it receives from I / O subsystem 206 or sensors such as proximity sensor 266, accelerometer 268, and / or microphone 213 (via audio circuitry 210). The information received by peripheral device interface 218 from I / O subsystem 206 includes information from touch-sensitive display 212 or touch-sensitive surfaces.

[0104] In some implementations, the event monitor 271 sends requests to the peripheral device interface 218 at predetermined intervals. In response, the peripheral device interface 218 sends event information. In other implementations, the peripheral device interface 218 sends event information only when a significant event occurs (e.g., receiving input above a predetermined noise threshold and / or for a predetermined duration).

[0105] In some implementations, the event classifier 270 also includes a hit view determination module 272 and / or an activity event recognizer determination module 273.

[0106] When the touch-sensitive display 212 displays more than one view, the hit view determination module 272 provides a software process for determining where a sub-event has occurred within one or more views. A view consists of controls and other elements that the user can see on the display.

[0107] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected corresponds to a procedural level within the application's procedural hierarchy or view hierarchy. For example, the lowest-level view in which a touch is detected is called the hit view, and the set of events considered as correct input is determined at least in part based on the hit view of the initial touch that initiates the touch-based gesture.

[0108] The hit view determination module 272 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchical structure, the hit view determination module 272 identifies the hit view as the lowest-level view in the hierarchical structure from which the sub-events should be processed. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events forming an event or potential event) occurs. Once the hit view is identified by the hit view determination module 272, the hit view typically receives all sub-events related to the same touch or input source to which it was identified as the hit view.

[0109] The activity event recognizer determination module 273 determines which views(s) within the view hierarchy should receive a specific sub-event sequence. In some embodiments, the activity event recognizer determination module 273 determines that only the hit view should receive the specific sub-event sequence. In other embodiments, the activity event recognizer determination module 273 determines that all views including the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive the specific sub-event sequence. In other embodiments, even if the touch sub-event is entirely confined to the area associated with a particular view, higher views in the hierarchy will still remain actively participating views.

[0110] Event assigner module 274 assigns event information to event identifiers (e.g., event identifier 280). In embodiments that include active event identifier determination module 273, event assigner module 274 delivers event information to the event identifier determined by active event identifier determination module 273. In some embodiments, event assigner module 274 stores event information in an event queue, which is retrieved by the corresponding event receiver 282.

[0111] In some implementations, operating system 226 includes event classifier 270. Alternatively, application 236-1 includes event classifier 270. In yet another implementation, event classifier 270 is a separate module or part of another module (such as contact / motion module 230) stored in memory 202.

[0112] In some implementations, application 236-1 includes a plurality of event handlers 290 and one or more application views 291, each of which includes instructions for handling touch events occurring within a corresponding view of the application's user interface. Each application view 291 of application 236-1 includes one or more event recognizers 280. Typically, a corresponding application view 291 includes a plurality of event recognizers 280. In other implementations, one or more of the event recognizers 280 are part of a separate module, such as a user interface toolkit (not shown) or a higher-level object from which application 236-1 inherits methods and other properties. In some implementations, the corresponding event handlers 290 include one or more of the following: a data updater 276, an object updater 277, a GUI updater 278, and / or event data 279 received from an event classifier 270. Event handlers 290 utilize or invoke the data updater 276, the object updater 277, or the GUI updater 278 to update the application's internal state 292. Alternatively, one or more application views in application view 291 include one or more corresponding event handlers 290. Additionally, in some embodiments, one or more of data updater 276, object updater 277, and GUI updater 278 are included in the corresponding application view 291.

[0113] The corresponding event identifier 280 receives event information (e.g., event data 279) from the event classifier 270 and identifies the event based on the event information. The event identifier 280 includes an event receiver 282 and an event comparator 284. In some embodiments, the event identifier 280 also includes at least one subset of metadata 283 and event delivery instructions 288 (which includes sub-event delivery instructions).

[0114] Event receiver 282 receives event information from event classifier 270. The event information includes information about sub-events such as touch or touch movement. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves touch movement, the event information also includes the rate and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a lateral orientation, or vice versa), and the event information includes corresponding information about the device's current orientation (also referred to as device pose).

[0115] Event comparator 284 compares event information with predefined event or sub-event definitions and, based on the comparison, determines the event or sub-event, or determines or updates the state of the event or sub-event. In some embodiments, event comparator 284 includes event definition 286. Event definition 286 contains definitions of events (e.g., predefined sequences of sub-events), such as event 1 (287-1), event 2 (287-2), and others. In some embodiments, sub-events in event (287) include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, event 1 (287-1) is defined as a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift-off of a predetermined duration (touch end), a second touch (touch start) of a predetermined duration on the displayed object, and a second lift-off of a predetermined duration (touch end). In another example, event 2 (287-2) is defined as a drag on a displayed object. For example, dragging includes a touch (or contact) of a predetermined duration on the displayed object, movement of the touch on the touch-sensitive display 212, and lifting off the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 290.

[0116] In some implementations, event definition 287 includes definitions of events for corresponding user interface objects. In some implementations, event comparator 284 performs a hit test to determine which user interface object is associated with the sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display 212, when a touch is detected on touch-sensitive display 212, event comparator 284 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 290, the event comparator uses the result of the hit test to determine which event handler 290 should be activated. For example, event comparator 284 selects the event handler associated with the sub-event and the object that triggered the hit test.

[0117] In some implementations, the definition of the corresponding event (287) also includes a delay action that delays the delivery of event information until it has been determined whether the sub-event sequence actually corresponds to or does not correspond to the event type of the event recognizer.

[0118] When the corresponding event recognizer 280 determines that the sub-event sequence does not match any event in event definition 286, the corresponding event recognizer 280 enters an event impossible, event failed, or event ended state, after which subsequent sub-events based on touch gestures are ignored. In this case, other event recognizers (if any) that remain active in the hit view continue to track and process the ongoing sub-events based on touch gestures.

[0119] In some embodiments, the corresponding event recognizer 280 includes metadata 283 having configurable attributes, flags, and / or lists instructing how the event delivery system should perform sub-event delivery to actively participating event recognizers. In some embodiments, the metadata 283 includes configurable attributes, flags, and / or lists instructing how or how event recognizers can interact with each other. In some embodiments, the metadata 283 includes configurable attributes, flags, and / or lists instructing whether sub-events are delivered to different levels in a view or programmatic hierarchy.

[0120] In some implementations, when one or more specific sub-events of an event are identified, the corresponding event recognizer 280 activates the event handler 290 associated with the event. In some implementations, the corresponding event recognizer 280 delivers event information associated with the event to the event handler 290. Activating the event handler 290 is different from delivering (and deferred delivering) the sub-events to the corresponding hit view. In some implementations, the event recognizer 280 throws a flag associated with the identified event, and the event handler 290 associated with the flag retrieves the flag and performs a predefined process.

[0121] In some implementations, event delivery instruction 288 includes a sub-event delivery instruction that delivers event information about a sub-event without activating an event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with the sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or the actively participating view receives the event information and performs a predetermined process.

[0122] In some implementations, data updater 276 creates and updates data used in application 236-1. For example, data updater 276 updates phone numbers used in contact module 237 or stores video files used in video player module. In some implementations, object updater 277 creates and updates objects used in application 236-1. For example, object updater 277 creates new user interface objects or updates the positioning of user interface objects. GUI updater 278 updates the GUI. For example, GUI updater 278 prepares display information and transmits that display information to graphics module 232 for display on a touch-sensitive display.

[0123] In some implementations, event handler 290 includes, or has access to, a data updater 276, an object updater 277, and a GUI updater 278. In some implementations, data updater 276, object updater 277, and GUI updater 278 are included in a single module of the corresponding application 236-1 or application view 291. In other implementations, they are included in two or more software modules.

[0124] It should be understood that the above discussion regarding event handling for user touch on a touch-sensitive display also applies to other forms of user input used to operate the multifunction device 200 using an input device, and not all user input is initiated on a touchscreen. For example, mouse movement and mouse button presses optionally in conjunction with single or multiple keyboard presses or holds; touch movements on the touchpad, such as taps, drags, scrolls, etc.; stylus input; device movement; verbal commands; detected eye movements; biometric input; and / or any combination thereof may optionally be used as input corresponding to sub-events that define the event to be identified.

[0125] Figure 3A portable multifunction device 200 with a touchscreen 212 is illustrated according to some embodiments. The touchscreen optionally displays one or more graphics within a user interface (UI) 300. In this embodiment and other embodiments described below, a user can select one or more graphics by gesturing over the graphics, for example, using one or more fingers 302 (not drawn to scale in the figure) or one or more styluses 303 (not drawn to scale in the figure). In some embodiments, selection of one or more graphics occurs when the user breaks contact with one or more graphics. In some embodiments, gestures optionally include one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or scrolling (from right to left, from left to right, up and / or down) of a finger already in contact with the device 200. In some specific embodiments or in some cases, unintentional contact with a graphic does not select the graphic. For example, a swipe gesture over an application icon may optionally not select the corresponding application when the gesture corresponding to selection is a tap.

[0126] Device 200 also includes one or more physical buttons, such as a "main desktop" or menu button 304. As previously described, menu button 304 is used to navigate to any application 236 of a set of applications running on device 200. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on touchscreen 212.

[0127] In some embodiments, device 200 includes a touchscreen 212, a menu button 304, a push-button 306 for powering on / off and locking the device, a volume control button 308, a SIM card slot 310, a headset jack 312, and a docking / charging external port 224. The push-button 306 may optionally be used to: power on / off the device by pressing the button and holding it in the pressed state for a predefined time interval; lock the device by pressing the button and releasing it before the predefined time interval has elapsed; and / or unlock the device or initiate an unlocking process. In another embodiment, device 200 also accepts verbal input via microphone 213 for activating or deactivating certain functions. Device 200 may also optionally include one or more contact strength sensors 265 for detecting the intensity of contact on the touchscreen 212, and / or one or more haptic output generators 267 for generating haptic output for the user of device 200.

[0128] Figure 4AThis is a block diagram of an exemplary multi-functional device with a display and a touch-sensitive surface according to some embodiments. Device 400 need not be portable. In some embodiments, device 400 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a children's learning toy), gaming system, or control device (e.g., a home controller or industrial controller). Device 400 typically includes one or more processing units (CPUs) 410, one or more network or other communication interfaces 460, memory 470, and one or more communication buses 420 for interconnecting these components. Communication bus 420 optionally includes circuitry (sometimes referred to as a chipset) that interconnects system components and controls communication between system components. Device 400 includes an input / output (I / O) interface 430 with a display 440, which is typically a touchscreen display. I / O interface 430 may also optionally include a keyboard and / or mouse (or other pointing device) 450 and a touchpad 455, and a haptic output generator 457 for generating haptic output on device 400 (e.g., similar to the reference above). Figure 2A The described tactile output generator 267), sensor 459 (e.g., optical sensor, accelerometer, proximity sensor, touch sensor and / or contact intensity sensor (similar to the one described above)) Figure 2A The described contact strength sensor 265). Memory 470 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 470 optionally includes one or more storage devices located remotely from CPU 410. In some embodiments, memory 470 stores data with portable multifunction device 200 (…). Figure 2A The memory 470 stores programs, modules, and data structures similar to those in the memory 202 of the portable multifunction device 200, or subsets thereof. Additionally, the memory 470 may optionally store additional programs, modules, and data structures not present in the memory 202 of the portable multifunction device 200. For example, the memory 470 of the device 400 may optionally store a drawing module 480, a rendering module 482, a word processing module 484, a website creation module 486, a disk editing module 488, and / or a spreadsheet module 490, while the portable multifunction device 200 ( Figure 2A The memory 202 may optionally not store these modules.

[0129] Figure 4AEach of the aforementioned elements is stored in one or more of the previously mentioned memory devices in some examples. Each of the aforementioned modules corresponds to a set of instructions for performing the functions described above. The aforementioned modules or programs (e.g., instruction sets) need not be implemented as standalone software programs, processes, or modules; therefore, various subsets of these modules are combined or otherwise rearranged in various embodiments. In some embodiments, memory 470 stores a subset of the aforementioned modules and data structures. Furthermore, memory 470 stores additional modules and data structures not described above.

[0130] Specific embodiments within the scope of this disclosure may be implemented, in whole or in part, using a tangible computer-readable storage medium (or a plurality of tangible computer-readable storage media of one or more types) that encodes one or more computer-readable instructions. It should be understood that computer-readable instructions may be organized in any format, including applications, widgets, processes, software, and / or components.

[0131] Specific embodiments within the scope of this disclosure include computer-readable storage media that encode instructions organized as an application (e.g., application 3160) that, when executed by one or more processing units, control the execution of an electronic device (e.g., device 3150). Figure 4B Methods Figure 4C The methods and / or one or more other processes and / or methods described herein.

[0132] It should be recognized that the application of 3160 ( Figure 4D The application 3160 (as shown in the diagram) can be any suitable type of application, including one or more of the following: browser applications, applications used as execution environments for plugins, widgets, or other applications, fitness applications, health applications, digital payment applications, media applications, social networking applications, messaging applications, and / or map applications. In some embodiments, application 3160 is an application pre-installed on device 3150 at the time of purchase (e.g., a first-party application). In some embodiments, application 3160 is an application provided to device 3150 via operating system update files (e.g., a first-party or second-party application). In some embodiments, application 3160 is an application provided via an app store. In some embodiments, the app store can be an app store pre-installed on device 3150 at the time of purchase (e.g., a first-party app store). In some embodiments, the app store is a third-party app store (e.g., an app store provided by another app store, downloaded via a network, and / or read from a storage device).

[0133] refer to Figure 4B and Figure 4FApplication 3160 obtains information (e.g., 3010). In some embodiments, at 3010, information is obtained from at least one hardware component of device 3150. In some embodiments, at 3010, information is obtained from at least one software module of device 3150. In some embodiments, at 3010, information is obtained from at least one hardware component external to device 3150 (e.g., peripheral devices, accessory devices, and / or servers). In some embodiments, the information obtained at 3010 includes location information, time information, notification information, user information, environmental information, electronic device status information, weather information, media information, historical information, event information, hardware information, and / or motion information. In some embodiments, in response to obtaining information at 3010 and / or thereafter, application 3160 provides the information to the system (e.g., 3020).

[0134] In some implementations, the system (e.g., Figure 4E The 3110 shown is the operating system hosted on the device 3150. In some implementations, the system (e.g., Figure 4E 3110 shown in the figure is an external device (e.g., a server, peripheral device, accessory and / or personal computing device) that includes an operating system.

[0135] refer to Figure 4C and Figure 4G Application 3160 obtains information (e.g., 3030). In some embodiments, the information obtained at 3030 includes location information, time information, notification information, user information, environmental information, electronic device status information, weather information, media information, historical information, event information, hardware information, and / or motion information. In response to obtaining information at 3030 and / or thereafter, application 3160 performs an operation on the information (e.g., 3040). In some embodiments, the operations performed at 3040 include: providing notifications based on the information, sending messages based on the information, displaying information, controlling the user interface of a fitness application based on the information, controlling the user interface of a health application based on the information, controlling focus mode based on the information, setting reminders based on the information, adding calendar entries based on the information, and / or calling the API of system 3110 based on the information.

[0136] In some implementations, execution is performed in response to a trigger. Figure 4B Methods and / or Figure 4C The method involves one or more steps. In some implementations, triggering includes event detection, notifications received from system 3110, user input, and / or responses to calls to APIs provided by system 3110.

[0137] In some implementations, when the instructions of application 3160 are executed, control device 3150 executes them by calling an application programming interface (API) (e.g., API 3190) provided by system 3110. Figure 4B Methods and / or Figure 4C The method. In some implementations, application 3160 executes without calling API 3190. Figure 4B Methods and / or Figure 4C At least a part of the method.

[0138] In some implementation schemes, Figure 4B Methods and / or Figure 4C One or more steps of the method involve calling the API (e.g., API 3190) using one or more parameters defined by the API. In some implementations, one or more parameters include constants, keys, data structures, objects, object classes, variables, data types, pointers, arrays, lists, or pointers to functions or methods and / or references to data or other items to be passed via the API in another way.

[0139] refer to Figure 4D Example 3150 is shown. In some embodiments, device 3150 is a personal computing device, smartphone, smartwatch, fitness tracker, head-mounted display (HMD) device, media device, public utility, speaker, television, and / or tablet device. Figure 4D As illustrated, device 3150 includes application 3160 and operating system (e.g., Figure 4E The system 3110 shown is an example. Application 3160 includes an application implementation module 3170 and an API call module 3180. System 3110 includes an API 3190 and an implementation module 3100. It should be understood that device 3150, application 3160, and / or system 3110 may include components related to… Figure 4D and Figure 4E The examples illustrate more, fewer, and / or different components.

[0140] In some implementations, application implementation module 3170 includes a set of one or more instructions corresponding to one or more operations performed by application 3160. For example, when application 3160 is a messaging application, application implementation module 3170 may include operations for receiving and transmitting messages. In some implementations, application implementation module 3170 communicates with API calling module 3180 via API 3190 (in... Figure 4E (As shown in the figure) communicates with system 3110.

[0141] In some implementations, API 3190 is a software module (e.g., a set of computer-readable instructions) that provides an interface allowing different modules (e.g., API calling module 3180) to access and / or use one or more functions, methods, procedures, data structures, classes, and / or other services provided by implementation module 3100 of system 3110. For example, API calling module 3180 can access features of implementation module 3100 through one or more API calls or enablements (e.g., embodied by function or method calls) exposed by API 3190 (e.g., software and / or hardware modules capable of receiving, responding to, and / or transmitting API calls), and can pass data and / or control information via API calls or enablements using one or more parameters. In some implementations, API 3190 allows application 3160 to use services provided by a software development kit (SDK) library. In some implementations, application 3160 combines calls to functions or methods provided by the SDK library and API 3190, or uses data types or objects defined in the SDK library and provided by API 3190. In some implementations, API calling module 3180 makes API calls via API 3190 to access and use features of implementation module 3100 specified by API 3190. In such implementations, implementation module 3100 may return a value to API calling module 3180 via API 3190 in response to an API call. This value may report to application 3160 the capabilities or status of hardware components of device 3150, including those capabilities or statuses related to aspects such as input capabilities and status, output capabilities and status, processing capabilities, power status, storage capacity and status, and / or communication capabilities. In some implementations, API 3190 is implemented in part by firmware, microcode, or other low-level logic executed in part on the hardware components.

[0142] In some implementations, API 3190 allows the developer of API calling module 3180 (which may be a third-party developer) to utilize features provided by implementation module 3100. In such implementations, one or more API calling modules (e.g., including API calling module 3180) may exist that communicate with implementation module 3100. In some implementations, API 3190 allows multiple API calling modules written in different programming languages ​​to communicate with implementation module 3100 (e.g., API 3190 may include features for translating calls and returns between implementation module 3100 and API calling module 3180), and API 3190 is implemented in a specific programming language. In some implementations, API calling module 3180 calls APIs from different providers, such as a set of APIs from an OS provider, another set of APIs from a plugin provider, and / or another set of APIs from another provider (e.g., a software library provider) or the creator of another set of APIs.

[0143] Examples of API 3190 may include one or more of the following: pairing API (e.g., for establishing a secure connection, such as with an accessory), device detection API (e.g., for locating nearby devices, such as media devices and / or smartphones), payment API, UIKit API (e.g., for generating user interfaces), location detection API, locator API, map API, health sensor API, sensor API, messaging API, push notification API, streaming API, collaboration API, video conferencing API, app store API, advertising service API, web browser API (e.g., WebKit API), transportation API, networking API, WiFi API, Bluetooth API, NFC API, UWB API, fitness API, smart home API, contact transfer API, photo API, camera API, and / or image processing API. In some implementations, a sensor API is an API for accessing data associated with sensors of device 3150. For example, a sensor API may provide access to raw sensor data. Alternatively, a sensor API may provide data derived (and / or generated) from raw sensor data. In some implementations, sensor data includes temperature data, image data, video data, audio data, heart rate data, IMU (Inertial Measurement Unit) data, lidar data, location data, GPS data, and / or camera data. In some implementations, sensors include one or more of accelerometers, temperature sensors, infrared sensors, optical sensors, heart rate sensors, barometers, gyroscopes, proximity sensors, and / or biometric sensors.

[0144] In some embodiments, implementation module 3100 is a system (e.g., an operating system and / or server system) software module (e.g., a set of computer-readable instructions) configured to perform operations in response to receiving an API call via API 3190. In some embodiments, implementation module 3100 is configured to provide an API response (via API 3190) as a result of processing the API call. For example, implementation module 3100 and API call module 3180 can each be any of an operating system, library, device driver, API, application, or other module. It should be understood that implementation module 3100 and API call module 3180 can be the same or different types of modules. In some embodiments, implementation module 3100 is at least partially embodied in firmware, microcode, or hardware logic.

[0145] In some implementations, implementation module 3100 returns a value via API 3190 in response to an API call from API call module 3180. While API 3190 defines the syntax and results of the API call (e.g., how the API call is enabled and what it does), API 3190 may not reveal how implementation module 3100 performs the functionality specified by the API call. Various API calls are transferred via one or more application programming interfaces between API call module 3180 and implementation module 3100. Transferring API calls may include issuing, initiating, referencing, calling, receiving, returning, and / or responding to function calls or messages. In other words, a transfer may describe the action of either API call module 3180 or implementation module 3100. In some implementations, function calls or other enablements of API 3190 transmit and / or receive one or more parameters via parameter lists or other structures.

[0146] In some implementations, implementation module 3100 provides more than one API, each API providing a different view or aspect of the functionality implemented by implementation module 3100. For example, one API of implementation module 3100 may provide a first set of functions and be exposed to third-party developers, while another API of implementation module 3100 may be hidden (e.g., not exposed) and provide a subset of the first set of functions, and also provide another set of functions, such as test or debug functions not in the first set of functions. In some implementations, implementation module 3100 calls one or more other components via lower-level APIs, thus acting as both an API calling module and an implementation module. It should be recognized that implementation module 3100 may include additional functions, methods, classes, data structures, and / or other features not specified through API 3190 and not available to API calling module 3180. It should also be recognized that API calling module 3180 may be on the same system as implementation module 3100, or may be remotely located and accessed via a network using API 3190. In some implementations, implementation module 3100, API 3190, and / or API calling module 3180 are stored in a machine-readable medium, which includes any means for storing information in a machine-readable (e.g., computer or other data processing system) form. For example, a machine-readable medium may include a magnetic disk, optical disk, random access memory, read-only memory, and / or flash memory devices.

[0147] An Application Programming Interface (API) is an interface between a first software process and a second software process, specifying the format for communication between the two processes. Limited APIs (e.g., private or partner APIs) are APIs accessible to a limited set of software processes (e.g., only software processes within the operating system or only software processes authorized to access the limited API). Public APIs are accessible to a broader set of software processes. Some APIs enable a software process to communicate or set the state of one or more input devices (e.g., one or more touch sensors, proximity sensors, vision sensors, motion / or orientation sensors, pressure sensors, intensity sensors, sound sensors, wireless proximity sensors, biometric sensors, buttons, switches, rotatable elements, and / or external controllers). Some APIs enable a software process to communicate and / or set the state of one or more output generation components (e.g., one or more audio output generation components, one or more display generation components, and / or one or more haptic output generation components). Some APIs enable specific capabilities (e.g., scrolling, handwriting, text input, image editing, and / or image creation) to be accessed, executed, and / or used by a software process (e.g., generating output for use by the software process based on input from the software process). Some APIs enable content from software processes to be inserted into templates and displayed in user interfaces with layouts and / or behaviors specified by the templates.

[0148] Many software platforms include a set of frameworks that provide core objects and behaviors that software developers need to build software applications that can be used on the platform. Software developers use these objects to display content on a screen, interact with that content, and manage interactions with the software platform. The basic behavior of a software application depends on this framework, and this framework provides software developers with numerous ways to customize the application's behavior to match the specific needs of the application. Many of these core objects and behaviors are accessed via APIs. APIs typically specify the format for communication between software processes, including specifying and grouping available variables, functions, and protocols. API calls (sometimes called API requests) are typically passed from a sending software process to a receiving software process as a way to achieve one or more of the following: the sending software process requests information from the receiving software process (e.g., for the sending software process to take an action); the sending software process provides information to the receiving software process (e.g., for the receiving software process to take an action); the sending software process requests an action from the receiving software process; or the sending software process provides information to the receiving software process about the action taken by the sending software process. In some cases, interaction with a device (e.g., using a user interface) will involve transferring and / or receiving one or more API calls (e.g., multiple API calls) between multiple different software processes (e.g., different parts of an operating system, applications and operating systems, or different applications) via one or more APIs (e.g., via multiple different APIs). For example, when input is detected, direct sensor data is frequently processed into one or more input events, which are provided (e.g., via an API) to a receiving software process, which makes some determinations based on the input events and then (e.g., via an API) transmits information to the software process to perform an operation (e.g., change the device state and / or the user interface) based on the determinations. While the determinations and the operations performed in response can be made by the same software process, alternatively, the determinations can be made in a first software process and relayed (e.g., via an API) to a second software process different from the first software process, allowing the operation to be performed by the second software process. Alternatively, the second software process can relay instructions (e.g., via an API) to a third software process different from the first and / or second software processes to perform the operation. It should be understood that some or all user interactions with a computer system may involve one or more API calls within the steps of interacting with the computer system (e.g., between different software components of the computer system or between software components of the computer system and software components of one or more remote computer systems).It should be understood that some or all user interactions with a computer system may involve one or more API calls between steps of interaction with the computer system (e.g., between different software components of the computer system or between software components of the computer system and software components of one or more remote computer systems).

[0149] In some implementations, the application can be any suitable type of application, including one or more of the following: browser applications, applications used as execution environments for plugins, widgets or other applications, fitness applications, health applications, digital payment applications, media applications, social networking applications, messaging applications and / or map applications.

[0150] In some embodiments, the application is an application pre-installed on the first computer system at the time of purchase (e.g., a first-party application). In some embodiments, the application is an application provided to the first computer system via an operating system update file (e.g., a first-party application). In some embodiments, the application is an application provided via an app store. In some embodiments, the app store is pre-installed on the first computer system at the time of purchase (e.g., a first-party app store) and allows the download of one or more applications. In some embodiments, the app store is a third-party app store (e.g., an app store provided by another device, downloaded via a network, and / or read from a storage device). In some embodiments, the application is a third-party application (e.g., an application provided by an app store, downloaded via a network, and / or read from a storage device). In some embodiments, the application controls the first computer system to execute methods 1000 and / or 1100 by calling an application programming interface (API) provided by a system process using one or more parameters. Figure 10 and / or Figure 11 ).

[0151] In some implementations, exemplary APIs provided by system processes include one or more of the following: pairing API (e.g., for establishing a secure connection, such as with an accessory), device detection API (e.g., for locating nearby devices, such as media devices and / or smartphones), payment API, UIKit API (e.g., for generating user interfaces), location detection API, locator API, map API, health sensor API, sensor API, messaging API, push notification API, streaming API, collaboration API, video conferencing API, app store API, advertising service API, web browser API (e.g., WebKit API), transportation API, networking API, WiFi API, Bluetooth API, NFC API, UWB API, fitness API, smart home API, contact transfer API, photo API, camera API, and / or image processing API.

[0152] In some embodiments, at least one API is a software module (e.g., a set of computer-readable instructions) that provides an interface allowing different modules (e.g., an API calling module) to access and use one or more functions, methods, procedures, data structures, classes, and / or other services provided by an implementation module of a system process. The API may define one or more parameters passed between the API calling module and the implementation module. In some embodiments, API 3190 defines a first API call that can be provided by API calling module 3180. An implementation module is a system software module (e.g., a set of computer-readable instructions) configured to perform operations in response to receiving an API call via the API. In some embodiments, the implementation module is configured to provide an API response (via the API) as a result of processing the API call. In some embodiments, the implementation module is included in a device (e.g., 3150) running an application. In some embodiments, the implementation module is included in an electronic device separate from the device running the application.

[0153] Now let’s turn our attention to implementations of user interfaces that can be implemented, for example, on a portable multi-functional device 200.

[0154] Figure 5A An exemplary user interface for an application menu on a portable multifunction device 200 according to some embodiments is illustrated. A similar user interface is implemented on device 400. In some embodiments, user interface 500 includes the following elements or a subset or superset thereof: Signal strength indicator 502 for wireless communications such as cellular signals and Wi-Fi signals; • Time 504; • Bluetooth indicator 505; • Battery status indicator 506; • Tray 508 with icons for frequently used applications, such as: o The telephone module 238 has an icon 516 labeled "telephone", which optionally includes an indicator 514 indicating the number of missed calls or voicemail messages; o An icon 518 labeled "Mail" in the email client module 240, which optionally includes an indicator 510 for the number of unread emails; o The icon 520 labeled "Browser" in browser module 247; and o The video and music player module 252 (also known as the iPod (Apple Inc. trademark) module 252) is marked with an icon 522 labeled "iPod"; and • Icons of other applications, such as: o IM module 241's icon 524 labeled "Message"; o Calendar module 248's icon 526 labeled "Calendar"; o Image management module 244's icon 528 labeled "Photo"; o The icon 530 of camera module 243, labeled "camera"; o The icon 532 of the online video module 255, which is labeled "Online Video"; o Stock Market widget 249-2 with icon 534 labeled "Stock Market"; o Map module 254's icon 536 labeled "Map"; o Weather widget 249-1 with icon 538 labeled "Weather"; o Alarm clock widget 249-4 with icon 540 labeled "Clock"; o The icon 542 labeled "Fitness Support" in the Fitness Support module 242; o The icon 544 labeled "Memo" in the Memo module 253; and o A settings icon 546 labeled "Settings" is used to set up an application or module, which provides access to the settings of the device 200 and its various applications 236.

[0155] It should be pointed out that, Figure 5A The icon labels illustrated herein are merely exemplary. For example, the icon 522 of the video and music player module 252 may optionally be labeled "Music" or "Music Player". Other labels may optionally be used for various application icons. In some embodiments, the label of a particular application icon includes the name of the application corresponding to that particular application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to that particular application icon.

[0156] Figure 5B An example is illustrated having a touch-sensitive surface 551 (e.g., separate from the display 550 (e.g., touchscreen display 212)). Figure 4A A tablet device or touchpad 455) device (e.g., Figure 4A An exemplary user interface on the device 400. The device 400 may also optionally include one or more contact intensity sensors (e.g., one or more of the sensors 459) for detecting the intensity of contact on the tactile surface 551 and / or one or more tactile output generators 457 for generating tactile output for the user of the device 400.

[0157] While some examples of input on a reference touchscreen display 212 (which combines a touch-sensitive surface and a display) are given in the following examples, in some implementations, the device detects input on a touch-sensitive surface separate from the display, such as... Figure 5B As shown in the diagram. In some embodiments, the touch-sensitive surface (e.g., Figure 5B 551) has a spindle (e.g., on the display (e.g., 550) corresponding to the main axis on the display (e.g., Figure 5B The main shaft of 553 (e.g., Figure 5B (552 in the example). According to these embodiments, the device detects the position corresponding to a specific location on the display (e.g., in...). Figure 5B In the middle, contact 560 corresponds to 568 and contact 562 corresponds to 570) at the contact with the touch-sensitive surface 551 (e.g., Figure 5B 560 and 562 in the example). Thus, when the touch-sensitive surface (e.g., Figure 5B 551 in the middle) and the display of a multi-functional device (e.g., Figure 5B When 550 is separated from 560, user input detected by the device on the touch-sensitive surface (e.g., contact with 560 and 562 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods may be used for other user interfaces described herein.

[0158] Additionally, while the examples below are given primarily with reference to finger input (e.g., finger touch, single-finger tap gesture, finger swipe gesture), it should be understood that in some implementations, one or more of these finger inputs may be replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be replaced by a mouse click (e.g., instead of a touch), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the touch). As another example, a tap gesture may optionally be replaced by a mouse click when the cursor is over the location of the tap gesture (e.g., as an alternative to detection of a touch, followed by cessation of touch detection). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice may optionally be used simultaneously, or mouse and finger touch may optionally be used simultaneously.

[0159] Figure 6A An exemplary personal electronic device 600 is illustrated. Device 600 includes a body 602. In some embodiments, device 600 includes devices 200 and 400 (e.g., Figures 2A to 4AThe device 600 may contain some or all of the features described herein. In some embodiments, the device 600 has a touch-sensitive display 604, referred to below as a touchscreen 604. Alternatively, or in addition to the touchscreen 604, the device 600 may also have a display and a touch-sensitive surface. Similar to the cases of devices 200 and 400, in some embodiments, the touchscreen 604 (or touch-sensitive surface) has one or more intensity sensors for detecting the intensity of an applied contact (e.g., a touch). The one or more intensity sensors of the touchscreen 604 (or touch-sensitive surface) provide output data representing the intensity of the touch. The user interface of the device 600 responds to touches based on the touch intensity, meaning that touches of different intensities may invoke different user interface operations on the device 600.

[0160] Techniques for detecting and processing touch intensity may exist, for example, in the following related applications: International Patent Application Serial No. PCT / US2013 / 040061, filed May 8, 2013, entitled “Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application”, and International Patent Application Serial No. PCT / US2013 / 069483, filed November 11, 2013, entitled “Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships”, each of which is incorporated herein by reference in its entirety.

[0161] In some embodiments, device 600 has one or more input mechanisms 606 and 608. Input mechanisms 606 and 608, if included, are physical in form. Examples of physical input mechanisms include push-buttons and rotatable mechanisms. In some embodiments, device 600 has one or more attachment mechanisms. Such attachment mechanisms, if included, allow device 600 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch straps, bangles, trousers, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow a user to wear device 600.

[0162] Figure 6B An exemplary personal electronic device 600 is depicted. In some embodiments, device 600 includes... Figure 2A , Figure 2B and Figure 4ASome or all of the components described. Device 600 has a bus 612 that operatively couples I / O portion 614 to one or more computer processors 616 and memory 618. I / O portion 614 is connected to display 604, which may have touch-sensitive component 622 and optionally also has touch intensity-sensitive component 624. Furthermore, I / O portion 614 is connected to communication unit 630 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular and / or other wireless communication technologies. Device 600 includes input mechanisms 606 and / or 608. For example, input mechanism 606 is a rotatable input device or a pressable input device and a rotatable input device. In some examples, input mechanism 608 is a button.

[0163] In some examples, the input mechanism 608 is a microphone. The personal electronic device 600 includes, for example, various sensors such as a GPS sensor 632, an accelerometer 634, an orientation sensor 640 (e.g., a compass), a gyroscope 636, a motion sensor 638, and / or combinations thereof, all of which are operatively connected to the I / O section 614.

[0164] The memory 618 of the personal electronic device 600 is a non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors 616, cause the computer processors to perform the techniques and processes described above. The computer-executable instructions are also stored and / or transported, for example, in any non-transitory computer-readable storage medium, for use by or in conjunction with an instruction execution system, apparatus, or device, such as a computer-based system, a processor-containing system, or other system capable of retrieving and executing instructions from and from an instruction execution system, apparatus, or device. The personal electronic device 600 is not limited to... Figure 6B It can be the components and configurations, or it can include other components or additional components in a variety of configurations.

[0165] As used herein, the term "power indication" refers, for example, in devices 200, 400, and / or 600 ( Figure 2A , Figure 4A and Figures 6A to 6B A graphical user interface object displayed on a screen. For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each constitute a representation.

[0166] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface with which a user is interacting. In some specific implementations that include a cursor or other positional marker, the cursor acts as a "focus selector," such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the cursor is positioned on a touch-sensitive surface (e.g., a...). Figure 4A The touchpad 455 or Figure 5B When an input (e.g., a press input) is detected on the touch-sensitive surface 551 of the display, the specific user interface element is adjusted according to the detected input. This applies to touchscreen displays (e.g., those capable of direct interaction with user interface elements on a touchscreen display) that enable direct interaction with user interface elements on the touchscreen display. Figure 2A The touch-sensitive display system 212 or Figure 5A In some embodiments of the touchscreen 212, a touch detected on the touchscreen acts as a "focus selector," such that when input (e.g., a press input by touch) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, that particular user interface element is adjusted according to the detected input. In some embodiments, focus moves from one area of ​​the user interface to another without corresponding movement of the cursor or movement of a touch on the touchscreen display (e.g., moving focus from one button to another using tab keys or arrow keys); in these embodiments, the focus selector moves according to the movement of focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is typically a user-controlled user interface element (or a touch on the touchscreen display) that delivers the user-expected interaction with the user interface (e.g., by indicating to the device the element of the user interface that the user expects to interact with). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen), the position of the focus selector (e.g., a cursor, touch, or selection box) above the corresponding button will indicate to the user that they expect to activate the corresponding button (rather than other user interface elements shown on the device's display).

[0167] As used in the specification and claims, the term "characteristic strength" of a contact refers to a characteristic of the contact based on one or more intensities of the contact. In some embodiments, the characteristic strength is based on multiple intensity samples. The characteristic strength may optionally be based on a predefined number of intensity samples or a set of intensity samples collected over a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after contact is detected, before contact is detected to be lifted off, before or after contact begins to move, before contact ends, before or after contact intensity is detected to increase, and / or before or after contact intensity decreases). The characteristic strength of the contact may optionally be based on one or more of the following: the maximum value of the contact intensity, the mean value of the contact intensity, the average value of the contact intensity, the value at the top 10% of the contact intensity, the half maximum value of the contact intensity, the 90% maximum value of the contact intensity, etc. In some embodiments, the duration of the contact is used when determining the characteristic strength (e.g., when the characteristic strength is the average value of the contact intensity over time). In some implementations, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether a user has performed an action. For example, the set of one or more intensity thresholds may include a first intensity threshold and a second intensity threshold. In this example, contact with a characteristic intensity not exceeding the first threshold results in a first action, contact with a characteristic intensity exceeding the first intensity threshold but not exceeding the second intensity threshold results in a second action, and contact with a characteristic intensity exceeding the second threshold results in a third action. In some implementations, the comparison between the characteristic intensity and one or more thresholds is used to determine whether to perform one or more actions (e.g., whether to perform the corresponding action or abort performing the corresponding action), rather than to determine whether to perform the first or second action.

[0168] In some implementations, a portion of the gesture is identified for determining the characteristic strength. For example, a touch-sensitive surface receives a series of swipes that transition from a starting position to an ending position, where the contact strength increases. In this example, the characteristic strength of the contact at the ending position is based only on a portion of the series of swipes, not the entire swipe (e.g., the swipe contact only covers the portion at the ending position). In some implementations, a smoothing algorithm is applied to the strength of the swipe contact before determining its characteristic strength. For example, the smoothing algorithm may optionally include one or more of the following: unweighted moving average smoothing algorithm, triangular smoothing algorithm, median filter smoothing algorithm, and / or exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow spikes or dips in the strength of the swipe contact to achieve the purpose of determining the characteristic strength.

[0169] The intensity of a contact on a touch-sensitive surface is characterized relative to one or more intensity thresholds, such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device performs an operation typically associated with clicking a button on a physical mouse or touchpad. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device performs an operation different from the operation typically associated with clicking a button on a physical mouse or touchpad. In some embodiments, when a contact with an intensity lower than the light press intensity threshold (e.g., and higher than the nominal contact detection intensity threshold, where contacts lower than the nominal contact detection intensity threshold are no longer detected) is detected, the device will move the focus selector based on the movement of the contact on the touch-sensitive surface without performing the operation associated with the light press intensity threshold or the deep press intensity threshold. Generally, unless otherwise stated, these intensity thresholds are consistent across different groups of user interface figures.

[0170] An increase in contact intensity from below a light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold is sometimes referred to as a "light press" input. An increase in contact intensity from below a deep press intensity threshold to an intensity above the deep press intensity threshold is sometimes referred to as a "deep press" input. An increase in contact intensity from below a contact detection intensity threshold to an intensity between the contact detection intensity threshold and the light press intensity threshold is sometimes referred to as detecting a contact on the touch surface. A decrease in contact intensity from above a contact detection intensity threshold to an intensity below the contact detection intensity threshold is sometimes referred to as detecting a contact being lifted off the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.

[0171] In some embodiments described herein, one or more operations are performed in response to detecting a gesture including a corresponding press input or in response to detecting a corresponding press input performed using a corresponding contact (or multiple contacts), wherein the corresponding press input is detected at least in part based on detecting that the intensity of the contact (or multiple contacts) increases to above a press input intensity threshold. In some embodiments, the corresponding operation is performed in response to detecting that the intensity of the corresponding contact increases to above a press input intensity threshold (e.g., a "downward stroke" of the corresponding press input). In some embodiments, the press input includes the intensity of the corresponding contact increasing to above a press input intensity threshold and the intensity of the contact subsequently decreasing to below the press input intensity threshold, and the corresponding operation is performed in response to detecting that the intensity of the corresponding contact subsequently decreases to below the press input threshold (e.g., an "upward stroke" of the corresponding press input).

[0172] In some implementations, the device employs intensity hysteresis to avoid unintended inputs sometimes referred to as "jitter," wherein the device defines or selects a hysteresis intensity threshold that has a predefined relationship with a press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the press input intensity threshold). Therefore, in some implementations, a press input includes an increase in the intensity of the corresponding contact above the press input intensity threshold and a subsequent decrease in the intensity of that contact below the hysteresis intensity threshold corresponding to the press input intensity threshold, and an operation is performed in response to detecting that the intensity of the corresponding contact subsequently decreases below the hysteresis intensity threshold (e.g., the "upstroke" of the corresponding press input). Similarly, in some implementations, a press input is detected only when the device detects that the intensity of the contact increases from an intensity equal to or below a hysteresis intensity threshold to an intensity equal to or above a press input intensity threshold and optionally the intensity of the contact subsequently decreases to an intensity equal to or below the hysteresis intensity, and corresponding operations are performed in response to the detection of a press input (e.g., depending on the environment, the intensity of the contact increases or decreases).

[0173] For ease of explanation, optionally, the description of an operation triggered in response to a press input associated with a press input strength threshold or in response to a gesture including a press input may be provided in response to detecting any of the following conditions: the contact strength increases to above the press input strength threshold, the contact strength increases from below a hysteresis strength threshold to above the press input strength threshold, the contact strength decreases to below the press input strength threshold, and / or the contact strength decreases to below the hysteresis strength threshold corresponding to the press input strength threshold. Additionally, in the example where the operation is described as being performed in response to detecting a decrease in contact strength below the press input strength threshold, the operation may optionally be performed in response to detecting a decrease in contact strength below a hysteresis strength threshold corresponding to and less than the press input strength threshold.

[0174] 3. Digital Assistant System Figure 7A Block diagrams of digital assistant systems 700 according to various examples are illustrated. In some examples, the digital assistant system 700 is implemented on a standalone computer system. In some examples, the digital assistant system 700 is distributed across multiple computers. In some examples, some of the modules and functions of the digital assistant are divided into server and client parts, wherein the client part resides on one or more user devices (e.g., device 104, device 122, device 200, device 400, or device 600) and communicates with the server part (e.g., server system 108) via one or more networks, for example, as... Figure 1 As shown in the image. In some examples, the digital assistant system 700 is... Figure 1The specific implementation of server system 108 (and / or DA server 106) shown herein. It should be noted that digital assistant system 700 is merely one example of a digital assistant system, and digital assistant system 700 may have more or fewer components than shown, may combine two or more components, or may have different configurations or layouts of components. Figure 7A The various components shown are implemented in hardware, software instructions for execution by one or more processors, firmware (including one or more signal processing integrated circuits and / or application-specific integrated circuits), or a combination thereof.

[0175] The digital assistant system 700 includes a memory 702, an input / output (I / O) interface 706, a network communication interface 708, and one or more processors 704. These components can communicate with each other via one or more communication buses or signal lines 710.

[0176] In some examples, memory 702 includes non-transitory computer-readable media, such as high-speed random access memory and / or non-volatile computer-readable storage media (e.g., one or more disk storage devices, flash memory devices or other non-volatile solid-state memory devices).

[0177] In some examples, I / O interface 706 couples input / output devices 716 of digital assistant system 700, such as a display, keyboard, touchscreen, and microphone, to user interface module 722. I / O interface 706, together with user interface module 722, receives user input (e.g., voice input, keyboard input, touch input, etc.) and processes this input accordingly. In some examples, for instance, when the digital assistant is implemented on a standalone user device, digital assistant system 700 includes components related to… Figure 2A , Figure 4A , Figures 6A to 6B Any of the components and I / O communication interfaces described in device 200, device 400, or device 600. In some examples, digital assistant system 700 represents the server portion of a digital assistant implementation and can interact with the user through a client-side portion located on a user device (e.g., device 104, device 200, device 400, or device 600).

[0178] In some examples, the network communication interface 708 includes a wired communication port 712 and / or wireless transmitting and receiving circuitry 714. The wired communication port receives and transmits communication signals via one or more wired interfaces such as Ethernet, Universal Serial Bus (USB), FireWire, etc. The wireless circuitry 714 receives RF signals and / or optical signals from the communication network and other communication devices, and transmits RF signals and / or optical signals to the communication network and other communication devices. Wireless communication uses any of a variety of communication standards, protocols, and technologies, such as GSM, EDGE, CDMA, TDMA, Bluetooth, Wi-Fi, VoIP, Wi-MAX, or any other suitable communication protocol. The network communication interface 708 enables the digital assistant system 700 to communicate with other devices via networks such as the Internet, intranets, and / or wireless networks such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs).

[0179] In some examples, memory 702 or its computer-readable storage medium stores programs, modules, instructions, and data structures, including all or a subset of the following: operating system 718, communication module 720, user interface module 722, one or more applications 724, and digital assistant module 726. Specifically, memory 702 or its computer-readable storage medium stores instructions for performing the processes described below. One or more processors 704 execute these programs, modules, and instructions, and read data from or write data to data structures.

[0180] Operating systems 718 (e.g., Darwin, RTXC, LINUX, UNIX, iOS, OS X, WINDOWS, or embedded operating systems such as VxWorks) include various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware, firmware, and software components.

[0181] The communication module 720 facilitates communication between the digital assistant system 700 and other devices via the network communication interface 708. For example, the communication module 720 communicates with electronic devices (such as those in…) Figure 2A , Figure 4A , Figures 6A to 6B The RF circuit 208 of the devices 200, 400, and 600 shown communicates. The communication module 720 also includes various components for processing data received by the wireless circuit 714 and / or the wired communication port 712.

[0182] The user interface module 722 receives commands and / or input from the user (e.g., from a keyboard, touchscreen, pointing device, controller, and / or microphone) via the I / O interface 706 and generates user interface objects on the display. The user interface module 722 also prepares output (e.g., speech, sound, animation, text, icons, vibration, haptic feedback, lighting, etc.) and transmits it to the user via the I / O interface 706 (e.g., through a display, audio channel, speaker, touchpad, etc.).

[0183] Application 724 includes programs and / or modules configured to be executed by one or more processors 704. For example, if the digital assistant system is implemented on a standalone user device, application 724 includes user applications such as games, calendar applications, navigation applications, or email applications. If the digital assistant system 700 is implemented on a server, application 724 includes, for example, resource management applications, diagnostic applications, or scheduling applications.

[0184] The memory 702 also stores the digital assistant module 726 (or the server portion of the digital assistant). In some examples, the digital assistant module 726 includes the following submodules or subsets or supersets: input / output processing module 728, speech-to-text (STT) processing module 730, natural language processing module 732, dialogue flow processing module 734, task flow processing module 736, service processing module 738, and speech synthesis processing module 740. Each of these modules has access to one or more, or subsets or supersets of, the following systems or data and models of the digital assistant module 726: knowledge ontology 760, vocabulary index 744, user data 748, task flow model 754, service model 756, and ASR system 758.

[0185] In some examples, using the processing modules, data, and models implemented in the digital assistant module 726, the digital assistant can perform at least some of the following: converting verbal input into text; identifying user intent expressed in natural language input received from the user; proactively eliciting and obtaining the information needed to fully infer the user intent (e.g., by disambiguating words, games, intents, etc.); determining a task flow to satisfy the inferred intent; and executing the task flow to satisfy the inferred intent.

[0186] In some examples, such as Figure 7B As shown, the I / O processing module 728 can... Figure 7A The I / O device 716 in the middle interacts with the user or through Figure 7AThe network communication interface 708 interacts with user equipment (e.g., device 104, device 200, device 400, or device 600) to obtain user input (e.g., speech input) and provide a response to the user input (e.g., as speech output). The I / O processing module 728 optionally obtains contextual information associated with the user input from the user equipment along with or shortly after receiving the user input. Contextual information includes user-specific data, vocabulary, and / or preferences associated with the user input. In some examples, the contextual information also includes the software and hardware states of the user equipment at the time the user request is received, and / or information related to the user's surrounding environment at the time the user request is received. In some examples, the I / O processing module 728 also transmits follow-up questions related to the user request to the user and receives answers from the user. When a user request is received by the I / O processing module 728 and the user request includes speech input, the I / O processing module 728 forwards the speech input to the STT processing module 730 (or speech recognizer) for speech-to-text conversion.

[0187] STT processing module 730 includes one or more ASR systems 758. The one or more ASR systems 758 can process speech input received through I / O processing module 728 to produce recognition results. Each ASR system 758 includes a front-end speech preprocessor. The front-end speech preprocessor extracts representative features from the speech input. For example, the front-end speech preprocessor performs a Fourier transform on the speech input to extract spectral features characterizing the speech input as a sequence of representative multidimensional vectors. Furthermore, each ASR system 758 includes one or more speech recognition models (e.g., acoustic models and / or language models) and implements one or more speech recognition engines. Examples of speech recognition models include Hidden Markov Models, Gaussian Mixture Models, Deep Neural Network Models, n-gram grammar language models, and other statistical models. Examples of speech recognition engines include engines based on Dynamic Time Warping (VTW) and engines based on Weighted Finite State Transformers (WFST). One or more speech recognition models and one or more speech recognition engines are used to process the representative features extracted by the front-end speech preprocessor to produce intermediate recognition results (e.g., phonemes, phoneme strings, and sub-words), and finally to produce text recognition results (e.g., words, word strings, or token sequences). In some examples, the speech input is processed at least in part by a third-party service or on the user's device (e.g., device 104, device 200, device 400, or device 600) to produce the recognition results. Once the STT processing module 730 produces the recognition results containing text strings (e.g., words, or sequences of words, or token sequences), the recognition results are passed to the natural language processing module 732 for intent inference. In some examples, the STT processing module 730 produces multiple candidate text representations of the speech input. Each candidate text representation is a sequence of words or tokens corresponding to the speech input. In some examples, each candidate text representation is associated with a speech recognition confidence score. Based on the speech recognition confidence score, the STT processing module 730 ranks the candidate text representations and provides the n best (e.g., the n highest-ranked) candidate text representations to the natural language processing module 732 for intent inference, where n is a predetermined integer greater than zero. For example, in one example, only the highest-ranked (n=1) candidate text representation is delivered to the natural language processing module 732 for intent inference. In another example, the five highest-ranked (n=5) candidate text representations are passed to the natural language processing module 732 for intent inference.

[0188] Further details regarding speech-to-text processing are described in U.S. Utility Model Patent Application Serial No. 13 / 236,942, entitled "Consolidating Speech Recognition Results," filed on September 20, 2011, the entire disclosure of which is incorporated herein by reference.

[0189] In some examples, the STT processing module 730 includes a vocabulary of recognizable words and / or accesses that vocabulary via a phonetic alphabet conversion module 731. Each vocabulary word is associated with one or more candidate pronunciations of a word represented in a speech recognition phonetic alphabet. Specifically, the vocabulary of recognizable words includes words associated with multiple candidate pronunciations. For example, the vocabulary includes the word “tomato” associated with candidate pronunciations of / tə'meɪɾoʊ / and / tə'mɑtoʊ / . Furthermore, the vocabulary words are associated with custom candidate pronunciations based on previous speech input from the user. Such custom candidate pronunciations are stored in the STT processing module 730 and associated with a specific user via a user profile on the device. In some examples, candidate pronunciations of words are determined based on the spelling of the words and one or more linguistic and / or phonetic rules. In some examples, candidate pronunciations are generated manually, for example, based on known standard pronunciations.

[0190] In some examples, candidate pronunciations are ranked based on their prevalence. For example, the candidate pronunciation / tə'meɪɾoʊ / ranks higher than / tə'mɑtoʊ / because the former is a more commonly used pronunciation (e.g., among all users, for users in a specific geographic region, or for any other suitable subset of users). In some examples, candidate pronunciations are ranked based on whether they are custom candidate pronunciations associated with a user. For example, custom candidate pronunciations rank higher than standard candidate pronunciations. This can be used to identify proper nouns with unique pronunciations that deviate from the canonical pronunciation. In some examples, candidate pronunciations are associated with one or more speech characteristics such as geographic origin, country, or ethnicity. For example, the candidate pronunciation / tə'meɪɾoʊ / is associated with the United States, while the candidate pronunciation / tə'mɑtoʊ / is associated with the United Kingdom. Furthermore, the ranking of candidate pronunciations is based on one or more characteristics of a user (e.g., geographic origin, country, ethnicity, etc.) stored in a user profile on the device. For example, it can be determined from the user profile that the user is associated with the United States. Based on the user's association with the United States, the candidate pronunciation / tə'meɪɾoʊ / (associated with the United States) may rank higher than the candidate pronunciation / tə'mɑtoʊ / (associated with the United Kingdom). In some examples, one of the ranked candidate pronunciations may be selected as the predicted pronunciation (e.g., the most likely pronunciation).

[0191] Upon receiving speech input, the STT processing module 730 is used (e.g., using an acoustic model) to determine the phonemes corresponding to the speech input, and then attempts (e.g., using a language model) to determine the word that matches the phonemes. For example, if the STT processing module 730 first identifies the phoneme sequence / tə'meɪɾoʊ / corresponding to a portion of the speech input, then it can subsequently determine, based on the vocabulary index 744, that the sequence corresponds to the word "tomato".

[0192] In some examples, the STT processing module 730 uses fuzzy matching techniques to determine words in a utterance. Thus, for example, the STT processing module 730 determines that the phoneme sequence / tə'meɪɾoʊ / corresponds to the word "tomato," even if that particular phoneme sequence is not a candidate phoneme sequence for that word.

[0193] The digital assistant's natural language processing module 732 ("natural language processor") acquires n best candidate text representations ("word sequences" or "symbol sequences") generated by the STT processing module 730 and attempts to associate each candidate text representation with one or more "executable intentions" recognized by the digital assistant. An "executable intention" (or "user intention") represents a task that can be performed by the digital assistant and may have an associated task flow implemented in the task flow model 754. An associated task flow is a series of programmed actions and steps taken by the digital assistant to perform the task. The capabilities of the digital assistant depend on the number and variety of task flows implemented and stored in the task flow model 754, or in other words, on the number and variety of "executable intentions" recognized by the digital assistant. However, the effectiveness of the digital assistant also depends on its ability to infer the correct "executable intention" from a user request expressed in natural language.

[0194] In some examples, in addition to the sequence of words or symbols obtained from the STT processing module 730, the natural language processing module 732 also receives, for example, contextual information associated with the user request from the I / O processing module 728. The natural language processing module 732 may optionally use the contextual information to clarify, supplement, and / or further define the information contained in the candidate text representation received from the STT processing module 730. Contextual information includes, for example, user preferences, the hardware and / or software state of the user's device, sensor information collected before, during, or shortly after the user request, previous interactions (e.g., conversations) between the digital assistant and the user, etc. As described herein, in some examples, the contextual information is dynamic and varies with the time, location, content, and other factors of the conversation.

[0195] In some examples, natural language processing is based on, for example, a knowledge ontology 760. Knowledge ontology 760 is a hierarchical structure containing many nodes, each node representing an "executable intent" or an "attribute" associated with one or more of the "executable intent" or other "attributes." As noted above, an "executable intent" represents a task that a digital assistant can perform; that is, the task is "executable" or can be done. An "attribute" represents a parameter associated with a sub-aspect of an executable intent or another attribute. The connections between executable intent nodes and attribute nodes in knowledge ontology 760 define how the parameters represented by the attribute nodes are subordinate to the task represented by the executable intent nodes.

[0196] In some examples, the knowledge ontology 760 consists of executable intent nodes and attribute nodes. Within the knowledge ontology 760, each executable intent node is directly connected to or connected to one or more attribute nodes via one or more intermediate attribute nodes. Similarly, each attribute node is directly connected to or connected to one or more executable intent nodes via one or more intermediate attribute nodes. For example, as... Figure 7C As shown, knowledge ontology 760 includes a "Restaurant Reservation" node (i.e., an executable intent node). The attribute nodes "Restaurant", "Date / Time" (for reservations) and "Participant Size" are all directly connected to the executable intent node (i.e., the "Restaurant Reservation" node).

[0197] Furthermore, the attribute nodes "Cuisine," "Price Range," "Phone Number," and "Location" are child nodes of the attribute node "Restaurant," and all are linked to the "Restaurant Reservation" node (i.e., the executable intent node) through the intermediate attribute node "Restaurant." For example, ... Figure 7C As shown, knowledge ontology 760 also includes a "Set Reminder" node (i.e., another executable intent node). The attribute nodes "Date / Time" (for setting reminders) and "Topic" (for reminders) are both connected to the "Set Reminder" node. Since the attribute "Date / Time" is related to both the task of making a restaurant reservation and the task of setting a reminder, the attribute node "Date / Time" is connected to both the "Restaurant Reservation" node and the "Set Reminder" node in knowledge ontology 760.

[0198] An executable intent node, along with its linked attribute nodes, is described as a "domain." In this discussion, each domain is associated with a corresponding executable intent and refers to a set of nodes (and the relationships between these nodes) associated with a particular executable intent. For example, Figure 7CThe knowledge ontology 760 shown includes examples of a restaurant reservation domain 762 and a reminder domain 764 within the knowledge ontology 760. The restaurant reservation domain includes an actionable intent node “Restaurant Reservation”, attribute nodes “Restaurant”, “Date / Time”, and “Participant Size”, and sub-attribute nodes “Cuisine”, “Price Range”, “Phone Number”, and “Location”. The reminder domain 764 includes an actionable intent node “Set Reminder” and attribute nodes “Topic” and “Date / Time”. In some examples, the knowledge ontology 760 consists of multiple domains. Each domain shares one or more attribute nodes with one or more other domains. For example, in addition to the restaurant reservation domain 762 and the reminder domain 764, the “Date / Time” attribute node is associated with many different domains (e.g., itinerary domain, travel booking domain, movie ticket domain, etc.).

[0199] although Figure 7C Two example fields within knowledge ontology 760 are illustrated, but other fields include, for example, "Find a movie," "Initiate a phone call," "Find directions," "Schedule a meeting," "Send a message," and "Provide answers to questions," "Reading lists," "Provide navigation instructions," and "Provide instructions for a task," etc. The "Send a message" field is associated with the "Send a message" executable intent node and also includes attribute nodes such as "Recipient," "Message type," and "Message body." The attribute node "Recipient" is further defined, for example, by sub-attribute nodes such as "Recipient name" and "Message address."

[0200] In some examples, knowledge ontology 760 includes all domains (and thus executable intents) that a digital assistant can understand and act upon. In some examples, knowledge ontology 760 is modified, such as by adding or removing entire domains or nodes, or by modifying the relationships between nodes within knowledge ontology 760.

[0201] In some examples, nodes associated with multiple related executable intents are clustered under a “superdomain” in Knowledge Ontology 760. For example, the “Travel” superdomain includes clusters of travel-related attribute nodes and executable intent nodes. Travel-related executable intent nodes include “Flight Booking,” “Hotel Booking,” “Car Rental,” “Route Planning,” “Find Points of Interest,” and so on. Executable intent nodes under the same superdomain (e.g., the “Travel” superdomain) have multiple shared attribute nodes. For example, executable intent nodes for “Flight Booking,” “Hotel Booking,” “Car Rental,” “Get Route,” and “Find Points of Interest” share one or more of the attribute nodes “Starting Location,” “Destination,” “Departure Date / Time,” “Arrival Date / Time,” and “Number of People in Party.”

[0202] In some examples, each node in the knowledge ontology 760 is associated with a set of words and / or phrases related to the attribute or executable intent represented by the node. The corresponding set of words and / or phrases associated with each node is called the "vocabulary" associated with the node. The corresponding set of words and / or phrases associated with each node is stored in the vocabulary index 744 associated with the attribute or executable intent represented by the node. For example, returning... Figure 7B The vocabulary associated with nodes of the "restaurant" attribute includes words such as "food," "drinks," "cuisine," "hunger," "eat," "pizza," "fast food," and "meals." Similarly, the vocabulary associated with nodes of the "initiate a phone call" action includes words and phrases such as "call," "make a phone call," "dial," "talk to," "call this number," and "make a phone call." The vocabulary index 744 optionally includes words and phrases from different languages.

[0203] Natural Language Processing (NLP) module 732 receives candidate text representations (e.g., text strings or symbol sequences) from STT processing module 730 and, for each candidate representation, determines which nodes the words in the candidate text representation relate to. In some examples, if a word or phrase in a candidate text representation is found to be associated with one or more nodes in knowledge ontology 760 (via lexical index 744), the word or phrase "triggers" or "activates" those nodes. Based on the number and / or relative importance of the activated nodes, NLP module 732 selects one executable intent as the task the user intends the digital assistant to perform. In some examples, the domain with the most "triggered" nodes is selected. In some examples, the domain with the highest confidence (e.g., based on the relative importance of its individual triggered nodes) is selected. In some examples, the domain is selected based on a combination of the number and importance of the triggered nodes. In some examples, additional factors, such as whether the digital assistant has previously correctly interpreted similar requests from the user, are also considered in the node selection process.

[0204] User data 748 includes user-specific information such as user-specific vocabulary, user preferences, user address, user's default second language, user's contact list, and other short- or long-term information for each user. In some examples, the natural language processing module 732 uses user-specific information to supplement the information contained in the user input to further refine the user's intent. For example, in response to a user request "Invite my friends to my birthday party," the natural language processing module 732 can access user data 748 to determine who the "friends" are and when and where the "birthday party" will be held, without requiring the user to explicitly provide such information in their request.

[0205] It should be recognized that, in some examples, the natural language processing module 732 is implemented using one or more machine learning agencies (e.g., neural networks). Specifically, the one or more machine learning agencies are configured to receive candidate text representations and contextual information associated with the candidate text representations. Based on the candidate text representations and the associated contextual information, the one or more machine learning agencies are configured to determine an intent confidence score based on a set of candidate executable intents. The natural language processing module 732 can select one or more candidate executable intents from the set of candidate executable intents based on the determined intent confidence score. In some examples, a knowledge ontology (e.g., knowledge ontology 760) is also utilized to select one or more candidate executable intents from the set of candidate executable intents.

[0206] Further details regarding the search of knowledge ontology based on symbol strings are described in U.S. Utility Model Patent Application Serial No. 12 / 341743, entitled “Method and Apparatus for Searching Using An Active Ontology,” filed on December 22, 2008, the entire disclosure of which is incorporated herein by reference.

[0207] In some examples, once the natural language processing module 732 identifies an executable intent (or domain) based on a user request, it generates a structured query to represent the identified executable intent. In some examples, the structured query includes parameters for one or more nodes within the domain of the executable intent, and at least some of these parameters are populated with specific information and requirements specified in the user request. For example, a user says, "Reserve a table at a sushi restaurant for 7 pm." In this case, the natural language processing module 732 is able to correctly identify the executable intent as "restaurant reservation" based on the user input. According to the knowledge ontology, the structured query for the "restaurant reservation" domain includes parameters such as {cuisine}, {time}, {date}, {number of people}, etc. In some examples, based on verbal input and text derived from the verbal input using the STT processing module 730, the natural language processing module 732 generates a partially structured query for the restaurant reservation domain, where the partially structured query includes the parameters {cuisine = "sushi"} and {time = "7 pm"}. However, in this example, the user's utterance contains insufficient information to complete a structured query associated with the domain. Therefore, based on the currently available information, no other necessary parameters such as {number of people at the party} and {date} are specified in the structured query. In some examples, the natural language processing module 732 uses the received context information to populate some parameters of the structured query. For example, in some examples, if a user requests a "nearby" sushi restaurant, the natural language processing module 732 uses GPS coordinates from the user's device to populate the {location} parameter in the structured query.

[0208] In some examples, the Natural Language Processing (NLP) module 732 identifies multiple candidate executable intents for each candidate text representation received from the STT processing module 730. Furthermore, in some examples, a corresponding structured query (partially or entirely) is generated for each identified candidate executable intent. The NLP module 732 determines an intent confidence score for each candidate executable intent and ranks the candidate executable intents based on the intent confidence scores. In some examples, the NLP module 732 transmits one or more of the generated structured queries (including any completed parameters) to the task flow processing module 736 (“task flow processor”). In some examples, one or more structured queries for the m best (e.g., the m highest-ranked) candidate executable intents are provided to the task flow processing module 736, where m is a predetermined integer greater than zero. In some examples, one or more structured queries for the m best candidate executable intents, along with their corresponding candidate text representations, are provided to the task flow processing module 736.

[0209] Further details regarding the inference of user intent based on multiple candidate executable intents determined from multiple candidate text representations of speech input are described in U.S. Utility Model Patent Application Serial No. 14 / 298,725, filed June 6, 2014, entitled “System and Method for Inferring UserIntent From Speech Inputs,” the entire disclosure of which is incorporated herein by reference.

[0210] Task flow processing module 736 is configured to receive one or more structured queries from natural language processing module 732, complete the structured queries (if necessary), and perform the actions required to "complete" the user's final request. In some examples, the various processes necessary to complete these tasks are provided in task flow model 754. In some examples, task flow model 754 includes processes for obtaining additional information from the user, and task flows for performing actions associated with the executable intent.

[0211] As described above, to complete a structured query, task flow processing module 736 needs to initiate additional dialogue with the user to obtain additional information and / or clarify potentially ambiguous statements. When such interaction is necessary, task flow processing module 736 invokes dialogue flow processing module 734 to participate in the dialogue with the user. In some examples, dialogue flow processing module 734 determines how (and / or when) to request additional information from the user and receives and processes user responses. Questions are presented to the user and answers are received from the user via I / O processing module 728. In some examples, dialogue flow processing module 734 presents dialogue output to the user via audible and / or visual output and receives input from the user via spoken or physical (e.g., click) responses. Continuing with the above example, when task flow processing module 736 invokes dialogue flow processing module 734 to determine the "party size" and "date" information for a structured query associated with the domain "restaurant reservation," dialogue flow processing module 734 generates questions such as "How many people in a row?" and "Which day to book?" and presents them to the user. Once a response is received from the user, the dialogue flow processing module 734 either fills the structured query with the missing information or passes the information to the task flow processing module 736 to complete the missing information based on the structured query.

[0212] Once the task flow processing module 736 has completed the structured query for the executable intent, it begins executing the final task associated with the executable intent. Therefore, the task flow processing module 736 executes the steps and instructions in the task flow model based on the specific parameters contained in the structured query. For example, the task flow model for the executable intent "restaurant reservation" includes steps and instructions for contacting the restaurant and actually requesting a reservation for a specific number of people at a specific time for a specific party. For example, using a structured query such as: {restaurant reservation, restaurant = ABC Cafe, date = 3 / 12 / 2012, time = 7pm, number of people = 5}, the task flow processing module 736 can perform the following steps: (1) log in to ABC Cafe's server or such as OPENTABLE. ® The restaurant reservation system (2) enters the date, time and party number information on the website, (3) submits the form, and (4) creates a calendar entry for the reservation in the user's calendar.

[0213] In some examples, task flow processing module 736, with the assistance of service processing module 738 (“service processing module”), completes the task requested in the user input or provides the informational answer requested in the user input. For example, service processing module 738, on behalf of task flow processing module 736, initiates a phone call, sets a calendar entry, invokes a map search, invokes or interacts with other user applications installed on the user's device, and invokes or interacts with third-party services (e.g., restaurant reservation portals, social networking sites, bank portals, etc.). In some examples, the protocols and application programming interfaces (APIs) required for each service are specified through the corresponding service model in service model 756. Service processing module 738 accesses the appropriate service model for a service and, based on the service model, generates a request for that service according to the protocols and APIs required by that service.

[0214] For example, if a restaurant has enabled an online reservation service, it submits a service model that specifies the necessary parameters for making a reservation and the values ​​of those parameters to be transmitted to the online reservation service's API. When requested by the task flow processing module 736, the service processing module 738 can use the web address stored in the service model to establish a network connection with the online reservation service and transmit the necessary reservation parameters (e.g., time, date, number of party members) to the online reservation interface in a format appropriate to the online reservation service's API.

[0215] In some examples, the natural language processing module 732, the dialogue flow processing module 734, and the task flow processing module 736 are used together and repeatedly to infer and define the user's intent, obtain information to further clarify and refine the user's intent, and ultimately generate a response (i.e., output to the user, or to complete a task) to satisfy the user's intent. The generated response is a dialogue response to the verbal input that at least partially satisfies the user's intent. Furthermore, in some examples, the generated response is output as verbal output. In these examples, the generated response is passed to the speech synthesis processing module 740 (e.g., a speech synthesizer), which processes the generated response to synthesize the dialogue response in verbal form. In still other examples, the generated response is data content related to satisfying the user's request in the verbal input.

[0216] In an example where the task flow processing module 736 receives multiple structured queries from the natural language processing module 732, the task flow processing module 736 first processes a first structured query of the received structured queries to attempt to complete the first structured query and / or execute one or more tasks or actions represented by the first structured query. In some examples, the first structured query corresponds to the highest-ranking executable intent. In other examples, the first structured query is selected from structured queries received based on a combination of a corresponding speech recognition confidence score and a corresponding intent confidence score. In some examples, if the task flow processing module 736 encounters an error during the processing of the first structured query (e.g., due to the inability to determine necessary parameters), the task flow processing module 736 may continue to select and process a second structured query from the received structured queries that corresponds to a lower-ranking executable intent. For example, the second structured query may be selected based on a speech recognition confidence score of a corresponding candidate text representation, an intent confidence score of a corresponding candidate executable intent, missing necessary parameters in the first structured query, or any combination thereof.

[0217] Speech synthesis processing module 740 is configured to synthesize speech output for presentation to a user. Speech synthesis processing module 740 synthesizes speech output based on text provided by a digital assistant. For example, the generated dialogue response is in the form of a text string. Speech synthesis processing module 740 converts the text string into audible speech output. Speech synthesis processing module 740 uses any appropriate speech synthesis techniques to generate speech output from text, including but not limited to: concatenation synthesis, unit selection synthesis, diphone synthesis, domain-specific synthesis, formant synthesis, articulation synthesis, Hidden Markov Model (HMM) based synthesis, and sine wave synthesis. In some examples, speech synthesis processing module 740 is configured to synthesize individual words based on phoneme strings corresponding to those words. For example, phoneme strings are associated with words in the generated dialogue response. Phoneme strings are stored in metadata associated with the words. Speech synthesis processing module 740 is configured to directly process the phoneme strings in the metadata to synthesize words in speech form.

[0218] In some examples, as an alternative to (or as a supplement to) using the speech synthesis processing module 740, speech synthesis is performed on a remote device (e.g., server system 108), and the synthesized speech is transmitted to a user device for output to the user. For example, this could occur in some implementations where the output of a digital assistant is generated at the server system. And since the server system typically has greater processing power or more resources than the user device, it is possible to obtain speech output of higher quality than that achieved through client-side synthesis.

[0219] Additional details regarding digital assistants can be found in U.S. Utility Model Patent Application No. 12 / 987,982, filed January 10, 2011, entitled “Intelligent Automated Assistant,” and U.S. Utility Model Patent Application No. 13 / 251,088, filed September 30, 2011, the entire disclosure of which is incorporated herein by reference.

[0220] As described herein, content is automatically generated by one or more computers in response to a request to generate content. The automatically generated content may optionally be generated on a device (e.g., at least in part by the computer system that received the request to generate content) and / or generated outside the device (e.g., at least in part by one or more nearby computers available via a local network or one or more computers available via the Internet). The automatically generated content may optionally include visual content (e.g., images, graphics, and / or video), audio content, and / or text content.

[0221] In some implementations, novel, automatically generated content produced via one or more artificial intelligence (AI) processes is referred to as generative content (e.g., generative images, generative graphics, generative videos, generative audio, and / or generative text). Generative content is typically generated by an AI process based on prompts provided to that AI process. The AI ​​process typically uses one or more AI models to generate output based on input. The AI ​​process may optionally include one or more preprocessing steps (e.g., adjusting user-provided prompts, creating system-generated prompts, and / or selecting AI models) to adjust the input before it is used by the AI ​​model to generate output. The AI ​​process may optionally include one or more postprocessing steps (e.g., passing the AI ​​model output to different AI models, scaling up, down, cropping, formatting, and / or adding or removing metadata) to adjust the AI ​​model output before it is used for other purposes (e.g., providing it to different software processes for further processing or presenting it to a user, for example, visually or aurally). The AI ​​process that generates generative content is sometimes referred to as a generative AI process.

[0222] Hints used to generate generative content may include one or more of the following: one or more words (e.g., natural language hindrances in written or spoken language), one or more images, one or more pictures, and / or one or more videos. AI processes may include machine learning models, which may include neural networks. Neural networks may include transformer-based deep neural networks, such as Large Language Models (LLMs). A generative pre-trained transformer model is an LLM that can efficiently generate new generative content based on hindrances. Some AI processes use hindrances that include text to generate diverse generative text, generative audio content, and / or generative visual content. Some AI processes use hindrances that include visual content and / or audio content to generate generative text (e.g., transcriptions of audio and / or descriptions of visual content). Some multimodal AI processes use hindrances that include multiple types of content (e.g., text, images, audio, video, and / or other sensor data) to generate generative content. Hints may sometimes also include values ​​for one or more parameters indicating the importance of the various parts of the hindrance. Some tips include a structured set of instructions that the AI ​​process can understand, including wording, specified style, relevant context (e.g., starting content and / or one or more examples) and / or the role of the AI ​​process.

[0223] Generative content is typically based on a cue, but it is not deterministically selected from pre-generated content; rather, it is generated using the cue as a starting point. In some implementations, pre-existing content (e.g., audio, text, and / or visual content) is used as part of the cue for creating the generative content (e.g., pre-existing content is used as a starting point for creating the generative content). For example, the cue may request that a piece of text be summarized or rewritten in a different tone, and the output will be the generative text summarized or written in a different tone. Similarly, the cue may request that visual content be modified to include or exclude content specified by the cue (e.g., removing identified features from the visual content, adding features described in the cue to the visual content, changing the visual style of the visual content, and / or creating additional visual elements based on the visual content beyond its spatial or temporal boundaries). In some implementations, a random or pseudo-random seed is used as part of the cue for creating the generative content (e.g., random or pseudo-random seed content is used as a starting point for creating the generative content). For example, when generating images from a diffusion model, random noise patterns are iteratively denoised based on the cue to generate images based on the cue. While this article has described specific types of AI processes, it should be understood that a variety of different AI processes can be used to generate generative content based on prompts.

[0224] 4. Commands using assistive device gestures Figures 8A to 8E and Figures 9A to 9CA system for issuing commands using assistive device gestures is illustrated. For example, device 802 may include any device described herein, including but not limited to devices 104, 200, 400, and 600. Figure 1 , Figure 2A , Figure 4A and Figures 6A to 6B Therefore, it should be understood that, with Figures 8A to 8E and Figures 9A to 9C The associated device 802 may correspond to a user device, such as a laptop computer, tablet, telephone, etc. Device 804 may correspond to a wearable device, such as, for example, a headset, smart glasses, or head-mounted display. Furthermore, the processes described herein can be performed by a server having information delivered to and from the device, information executed on the device, or a combination thereof. Additionally, as described herein, motion data may be maintained on the assistive device, and audio events may be generated and played accordingly by the second device. In some embodiments, the first and second electronic devices may reside within a single device (e.g., a head-mounted display with a microphone, speakers, etc.).

[0225] refer to Figure 8AGenerally, a user can interact with multiple devices (such as devices 802 and 804) within environment 800 within the context of the gesture response framework. Device 802 may correspond to a smartphone, and device 804 may correspond to multiple headsets or earpieces. Devices 802 and 804 may be further communicatively coupled via one or more wireless connections. In some examples, user 806 may utilize devices 802 and 804 to perform various functions, such as listening to music, participating in communication sessions (e.g., making phone calls, sending and receiving text messages, making video calls, etc.), interacting with digital assistants, etc. For example, device 802 may receive a phone call from another electronic device, such that if user 806 does not answer the call, the caller leaves a voicemail. The caller may correspond to a contact stored on device 802, such as a contact named "John". Once the voicemail is obtained on device 802 (e.g., locally stored and / or stored on a server), device 802 may detect an event alarm corresponding to the received voicemail. In some cases, user 806 may not actively use device 802 (e.g., device 802 is locked and / or placed on a surface far from the user). In such cases, device 802 can then provide a message to user 806 via device 804. Specifically, device 802 may enable the provision of an audible message 808 at device 804, such as “You have a new voicemail from John. Would you like to listen?” In this example, the audible message 808 includes a question or inquiry for the user to respond to. The question may be optional in nature, so that the user does not need to provide a response. For example, if the user does not take any action on the prompt, device 802 may infer that the user does not wish to continue with the corresponding task (e.g., listening to the voicemail), and therefore, device 802 does not take any action. Event alerts and corresponding messages may include a variety of different types of alerts and potential actions, such as incoming calls, incoming text messages, reminders associated with one or more applications (e.g., calendar, social media, lifestyle, productivity), alarms, application notifications, etc.

[0226] In some cases, the message 808 provided to device 804 may include a summary or brief description of the event alert. For example, the message may include something like, “You have a new long text message from Mary. Do you want to listen to it?” In some cases, the message may include the content of the received message (e.g., if the received message is less than a predefined character length) and a prompt to perform a related task, such as, “John said, ‘Are you coming?’ Do you want to respond?” Messages from multiple users (such as predefined groups within a messaging app) may also be received at device 802. In these cases, the message provided to device 804 may include something like, “There are ten new messages in the group chat. Do you want to listen to them?” or “There are 20 new messages in the group chat. Do you want to silence the chat notifications for this group?” Now for reference Figure 8B and Figure 8C Once message 808 is provided to device 804, device 802 can begin receiving motion data from device 804. Specifically, the motion data may correspond to the movement of device 804 as the user makes various movements, such as head gestures corresponding to affirmative actions (e.g., accepting) or negative actions (e.g., refusing). For example, device 804 may include one or more inertial measurement units (IMUs) that can provide various measurements, such as the rotational rate and acceleration rate of device 804. When user 806 makes head movements (such as nodding or shaking), IMU information may reflect the user's head movement via acceleration and rotational rates in the form of, for example, X, Y, and Z coordinate information. The motion data can then be received at device 802, allowing device 802 to process the motion data for classification. Generally, the motion data may be sampled at, for example, a rate of 25 Hz, with a prediction window size of 20, 25, or 30 samples. Motion data can be provided to a classifier that determines labels for the motion data, such as "yes" (e.g., a gesture consistent with a nod or other affirmative gesture), "no" (e.g., a gesture consistent with a head shake or other rejection gesture), or "no gesture," where "no gesture" indicates that the motion data does not correspond to a "yes" or "no" gesture with sufficient confidence. Each set of received motion data can be used to obtain the motion classification probability associated with a specific gesture. For example, if user 806 performs a gesture such as... Figure 8B For the nodding gesture shown, the classifier can assign the label to "yes" motion data with high confidence (e.g., 0.93 or 93%). In this case, the classifier can assign the label to "no" or "no gesture" motion data with low confidence (e.g., 0.05 and 0.02, respectively). For labels exceeding a threshold confidence level (e.g., 85%), the corresponding motion data can be classified based on the corresponding label (e.g., "yes" in this example).

[0227] Now for reference Figure 8DVarious mechanisms can be employed to provide feedback to the user regarding the gesture response framework. For example, an announcement tone 810a can be provided at device 804 before a message is provided. Once the announcement tone 810a is provided, a corresponding message 812a, such as “New long message from Justin. Listen now?”, is provided at device 804. After providing message 812a, a continuous audible sound 814a is provided to device 804 to indicate to the user that devices 802 and / or 804 are waiting and actively detecting whether the user has performed a gesture input. The continuous audible sound 814a can be provided for a predetermined amount of time (e.g., 3 seconds, 5 seconds, 10 seconds, etc.) and can include a soft, continuous tone to make the user aware that the device is in “listening mode” or “waiting loop”. In some instances, devices 802 and / or 804 can begin actively detecting whether a gesture input has been performed before providing the continuous audible sound 814a (e.g., when the announcement tone 810a is started or when message 812a is started). In some examples, device 804 may provide the user with one or more initial audible messages (e.g., during device setup, before or after a message is received) to inform the user that head gestures can be used to respond to various message sending and receiving (e.g., "You can respond to this message by nodding or shaking your head"). Device 802 may also provide various displayed and / or audible messages, including training the user to inform the user about the device's head gesture capabilities.

[0228] During a predetermined period of time during which continuous audible sound 814a is provided, device 802 may receive motion data from device 804 and determine one or more gestures based on the motion data, such as regarding Figures 8B to 8CAs described. A confirmation sound 816a is provided at device 804 to ensure that device 802 determines the corresponding gesture with sufficient confidence. For example, if user 806 performs a gesture consistent with a nod or other confirmation gesture, a positive confirmation sound (e.g., a cheerful chime, ring, or other sound) may be provided. Alternatively, if user 806 performs a gesture consistent with a head-shaking or other rejection gesture, a negative confirmation sound may be provided, such as a tone reversed relative to a positive confirmation sound (e.g., a reversal of pitch, melody, or other acoustic characteristic). Once confirmation sound 816a has been provided, the provision of a continuous audible sound 814a may be stopped to indicate that device 804 is no longer listening for gestural input. In some cases, user 806 may not perform any gesture determined to have a sufficiently high confidence level to be a confirmation or rejection gesture. In such cases, the continuous audible sound 814a may continue to be provided for a predetermined duration. If a predetermined duration expires without detecting any gesture with sufficient confidence, the continuous audible sound 814a may be stopped to instruct the device 804 to no longer listen for gestural input. In some examples, an acknowledgment sound may be provided at the end of the predetermined duration, while in other examples, the continuous audible sound 814a ends without an acknowledgment sound.

[0229] refer to Figure 8E In some examples, dynamic feedback can be provided to indicate the progress of detection regarding the detection of a single gesture. For example, as regarding... Figure 8DAs described, an announcement tone 810b, a message 812b, and a continuous audible sound 814b are provided. While providing the continuous audible sound 814b, gesture feedback 818b is also provided to device 804. Gesture feedback 818b typically includes multiple short, discrete tones (referred to as "ding," "ring," etc.) to indicate the progress of overall gesture detection. For example, user 806 may begin making a "nodding" gesture, causing the user to tilt their head upwards. Therefore, device 802 can detect the first partial gesture (e.g., head movement upwards) based on corresponding partial gesture criteria (such as whether rotational rate and acceleration data correspond to head movement upwards). Furthermore, heuristics based on the speed of head movement are also used to trigger partial gesture feedback. Thus, for example, the speed of head movement consistent with the user moving their head upwards can be used to detect the corresponding partial gesture. Therefore, a first tone can be provided to device 804 to indicate the first initial movement point of the corresponding "nodding" gesture. Similarly, user 806 can then tilt their head downwards, allowing device 802 to detect a second part of the gesture (e.g., head movement downwards) based on corresponding partial gesture criteria (such as whether rotational rate and acceleration data correspond to a downward head movement). Therefore, a second tone is provided to device 804, indicating the second point of movement for the corresponding "nodding" gesture. As the gesture continues, additional short, discrete tones can be provided throughout the gesture's duration (e.g., a third tone for another upward head movement, a fourth tone for another downward head movement, etc.). Each consecutive audible tone can differ from the others. For example, the first audible tone of gesture feedback 818b may be relatively low in terms of volume, pitch, intensity, or other acoustic characteristics. The second audible tone may be slightly higher or moderately higher than the first audible tone in terms of volume, pitch, intensity, or other acoustic characteristics. The third audible tone may be slightly higher or moderately sufficiently higher than the second audible tone in terms of volume, pitch, intensity, or other acoustic characteristics. This pattern can continue for each audible tone of gesture feedback 818b. When a user performs a partial gesture (e.g., the user raises their head) but does not perform any additional partial gestures to complete the gesture, device 802 may eventually reset the listening and thus begin listening for new gestures during the waiting loop period.

[0230] In the event of a detected affirmative or negative gesture corresponding to a message, output is provided at device 804 based on the gesture, and a corresponding task associated with the event alarm and / or message is performed. For example, in response to an event alarm corresponding to a new voicemail from a contact named "John" and a corresponding question about whether the user wants to listen to the voicemail, user 806 may perform a "nodding" gesture. In response to the detected "nodding" gesture, an acceptance confirmation tone is provided at device 804, such as regarding... Figures 8D to 8EThe above is discussed. Furthermore, tasks are performed based on the user's accepted gesture input. Specifically, in this case, device 802 can deliver voicemail audio to device 804. As another example, in response to an event alert corresponding to a new long message from a contact named "Justin" and a corresponding question about whether the user wants to listen to the message, user 806 can perform a back-and-forth head-shaking gesture. In response to a detected "no" gesture, a rejection-type confirmation tone is provided at device 804, such as regarding... Figures 8D to 8E The above is discussed. Furthermore, the corresponding task can be performed based on the user's rejection gesture input. Specifically, in this case, device 802 can provide a brief follow-up confirmation of rejection, such as "okay" or "understood." Alternatively, in response to a rejection gesture, a rejection confirmation tone is provided without any additional audible output, causing device 802 to stop receiving motion data from device 804 and thus stop detecting the corresponding gesture. Various other tasks can be performed based on event alerts and corresponding messages, such as connecting an audio call from device 802 to device 804, sending messages from device 802 (e.g., text messages or emails based on dictation from user 806 at device 804), creating or modifying calendar entries, reading application notifications, etc.

[0231] Generally speaking, gesture response frameworks can be implemented in several ways. For example, as an alternative to (or as a supplement to) using a "listen mode" or "wait loop," one or more discrete tones or other types of indications can be provided to the user to suggest additional content if the user provides a confirming gesture. For example, see [reference]. Figure 8F Before providing message 812c, an announcement tone 810c may be provided at the second electronic device. Message 812c may include a general topic, subject, or other information summarizing the content of the received message. In some examples, a message summary is provided when the message content exceeds a certain length (e.g., word count, character count, etc.). In this example, message 812c may include audible output such as "New message from Ron about lunch plans." Generally, a message summary may be generated based on various natural language processing techniques to identify concepts or other keywords within the text. Once message 812c is provided, a specific tone 820c may be provided to indicate to the user that the device is listening for a confirmation gesture. Tone 820c may include one or more discrete tones varying in volume, intensity, pitch, etc.

[0232] Once tone 820c is provided, the first device can receive motion data corresponding to the movement of the second device, and then determine gestures based on the motion data, as mentioned above. Figures 8A to 8EAs described. With respect to a rejection gesture, the first device may optionally provide an output indicating that a rejection gesture has been detected, and subsequently take no further action (not depicted) on the received message. Alternatively, with respect to a confirmation gesture 822c, the first device may cause to provide output 824c, including additional content corresponding to the full content of the received message.

[0233] In some implementations, the gesture response framework can be used to provide contextual response suggestions based on incoming messages. Additional details regarding providing contextual response suggestions based on incoming messages can be found in U.S. Utility Model Patent Application No. 18 / 373,211, filed September 26, 2023, entitled “Contextual Response Suggestions Using Secondary Electronic Device,” which claims priority to U.S. Provisional Patent Application No. 63 / 462,956, filed April 28, 2023, also entitled “Contextual Response Suggestions Using Secondary Electronic Device.” The disclosures of each of these patent applications are incorporated herein by reference.

[0234] As discussed herein, once an event alarm is detected, a message associated with the event alarm is provided at a second electronic device. In some implementations, the message may include a first part and a second part. The first part of the received message may include the content of an incoming text message (or alternatively, a summary of the text message), such as “John said, ‘When does your flight arrive?’”. The second part may include an indication of a potential response message to be transmitted back to the message sender. Based on one or more queries identified in the received message, a data source can be located to obtain results that satisfy the queries. In this example, the identified query may be parsed by retrieving flight information from an application such as an email application, calendar application, third-party flight application, etc. The current flight status information is then used to prepare a response message to be transmitted back to the message sender “John”. Specifically, the flight status information may include “On time, arriving at 8:20 p.m. tonight.” Therefore, the second part of the message provided to the second electronic device may include the prepared response message “My flight arrives at 8:20 p.m.” Once the first and second parts are provided, the device may begin receiving motion data and determining gestures based on the motion data. Regarding a rejection gesture, the first device may optionally provide an output indicating that a rejection gesture has been detected, and subsequently take no further action on the received message. Alternatively, regarding an acknowledgment gesture, the first device may send a corresponding prepared response message as a message to the message sender (e.g., "on time, arriving at 8:20 tonight"), and provide an output at the second device indicating that the message has been delivered.

[0235] Gesture response frameworks can also be used to confirm one or more options provided by the digital assistant. For example, a user may have dictated a message to be sent to another contact stored on the user's device. Once the message has been dictated, the digital assistant can request confirmation from the user, such as, "Okay, message reply 'I'll be there soon. Send it?'" The user can then respond to the question using confirming or rejecting gestures, as discussed herein. If the device detects a confirming gesture, the message is sent to the appropriate user. If the device detects a rejecting gesture, the message is not sent. Users can also respond via various other modalities such as touch input and / or voice input.

[0236] Now for reference Figure 9A The text describes an environment 900 where the user utilizes an early release framework, allowing audible output to be stopped based on user gestures. Specifically, the user 906 can utilize, as shown in the context of... Figures 8A to 8BDevices 902 and 904 are described. Device 904 may receive a message 908 corresponding to an event alarm detected at device 902. For example, message 908 may include an audible output “A new long message from John. John says, ‘Thanks for contacting me. My initial thought is…’”, explained in more detail below, and message 908 may continue to be read until a rejection gesture is detected from user 906. In some cases, before delivering a message to device 804, an information prompt may be provided to device 804 indicating that message reading may be stopped if the user performs a specific gesture such as a rejection gesture (e.g., “You can dismiss this message by shaking your head”).

[0237] refer to Figures 9B to 9C When a message is provided at device 904, user 906 may perform a gesture such as shaking their head back and forth. Therefore, device 906 can determine that user 906 is performing a rejection gesture based on motion data received from device 904. More specifically, an initial notification tone 908 is provided at device 902, wherein once notification tone 908 is provided, device 904 begins receiving motion data from device 904 and thus begins determining whether user 906 has performed a gesture. The gesture determination at device 902 may occur simultaneously with or otherwise in parallel with the reading of message 910 at device 904. Therefore, in relation to determining a rejection gesture before the reading of message 910 is completed, the corresponding task is performed at device 902. Specifically, device 902 stops providing message 910 to device 904. Furthermore, when message 910 is stopped (e.g., slightly before, simultaneously with, or slightly after stopping message 910), an acknowledgment sound 912 (e.g., a short audible tone) may also be provided at device 904. Alternatively, in the event of other types of user gestures (such as an affirmative nod), device 902 will continue reading message 910 until message reading is complete or device 902 detects a rejection gesture. In some cases, when using an early release frame, a different approach may also be adopted. Figure 8E The discussion provides dynamic feedback to indicate the progress of detection regarding the detection of individual gestures.

[0238] In some cases, early dismissal framing may not be used. For example, some messages provided at device 904 may include questions or other prompts soliciting a positive or negative response from the user. Scenarios involving messages that include interrogative statements or phrases (e.g., “You have ten events scheduled for today. Would you like to hear about them?”) and multiple pre-defined potential user responses (e.g., yes, no) may lead to the omission of early dismissal framing, making alternative framing possible. Figures 8A to 8EThe gesture response framework described herein. For example, this determination can be made at device 904 when an event alarm is detected and a corresponding message for delivery at device 902 is generated. In other words, if the user can hypothetically provide a rejection-type input corresponding to a negative answer to a question prompt (e.g., whether the user wants to hear the reading of today's events), device 902 will not stop reading the corresponding message in response to the detection of a rejection-type gesture, but will instead perform the corresponding task based on the rejection-type gesture (e.g., device 902 will abandon reading today's events). Similarly, in such scenarios, device 902 will also listen for acceptance-type gestures and perform appropriate tasks, such as regarding... Figures 8A to 8E As described. Alternatively, if device 902 determines that the question prompt has not yet been read (e.g., only the initial portion of the message has been audibly delivered to device 904), an early termination framework may be employed. Once device 902 determines that a question prompt requesting a positive or negative response has been provided, the early termination framework may no longer be applicable, making the discussion about... Figures 8A to 8E The described gesture response framework is now applicable.

[0239] Generally speaking, the early disengagement framework discussed above can be used in a variety of use cases. For example, a user can engage in a back-and-forth interaction with a digital assistant, where the user has spoken a command and the digital assistant is in the process of responding to that command, such as by outputting an audible response like, "Okay, I found a few restaurants for you. ABC Diner is the most popular..." During the audible response, the user can begin to shake their head, consistent with a rejection gesture. Based on the detection of the rejection gesture using the techniques disclosed herein, the digital assistant can stop providing audible output.

[0240] Figures 10 to 11Processes 1000 and 1100 for context-responsive suggestions are illustrated according to various examples. For example, processes 1000 and 1100 may be performed using one or more electronic devices implementing a digital assistant. In some examples, processes 1000 and 1100 may be performed using a client-server system (e.g., system 100), and the boxes for processes 1000 and 1100 may be divided in any way between a server (e.g., DA server 106) and client devices. In other examples, the box for process 1100 of process 1000 may be divided between a server and multiple client devices (e.g., mobile phones and smartwatches). Therefore, while portions of processes 1000 and 1100 are described herein as being performed by a specific device of a client-server system, it should be understood that processes 1000 and 1100 are not limited thereto. In other examples, processes 1000 and 1100 may be performed using only client devices (e.g., user device 104) or only multiple client devices. In processes 1000 and 1100, some boxes may be optionally grouped, the order of some boxes may be optionally changed, and some boxes may be optionally omitted. In some examples, processes 1000 and 1100 may be combined to perform additional steps.

[0241] refer to Figure 10At block 1002, in some embodiments, the first electronic device detects an event alarm. In some embodiments, the event alarm corresponds to one of an incoming call, an incoming text message, an alert, and an application notification. At block 1004, the first electronic device causes a message (e.g., automatically generated audio content and / or generated audio content) to be provided at a second electronic device, wherein the message is associated with the event alarm. In some embodiments, the event alarm corresponds to an incoming text message, and the message associated with the event alarm (e.g., automatically generated audio content and / or generated audio content) corresponds to an indication that the incoming text message exceeds a threshold character length. In some embodiments, the incoming text message corresponds to a group text message, and wherein the message associated with the event alarm (e.g., automatically generated audio content and / or generated audio content) includes an option to silence the group text message. In some embodiments, the event alarm corresponds to an incoming text message that does not exceed a threshold character length, and the message associated with the event alarm (e.g., automatically generated audio content and / or generated audio content) includes the content of the incoming text message. In some embodiments, in response to causing (e.g., using an AI process or a generative AI process) to provide a message at a second electronic device, a first electronic device causes to provide continuous audible sound at the second electronic device, wherein the continuous audible sound is provided for a predetermined duration. In some embodiments, an event alarm corresponds to an incoming text message, and the message associated with the event alarm (e.g., automatically generated audio content and / or generative audio content) includes the subject of the incoming text message. In some embodiments, causing to provide a message (e.g., automatically generated audio content and / or generative audio content) at a second electronic device includes a first portion of causing to provide a message (e.g., automatically generated audio content and / or generative audio content) at a second electronic device, and a second portion of causing to provide a message (e.g., automatically generated audio content and / or generative audio content) at a second electronic device, wherein the first portion of the message includes at least a portion of the content of the incoming text message, and wherein the second portion of the message includes a prompt indicating a response to the message subject.

[0242] By providing continuous audible sound during the waiting loop, the system enhances device functionality by offering enhanced feedback to the user. This enhanced feedback makes the device more efficient by focusing the user's response at the appropriate time and thus eliminating undetectable responses. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the device more quickly and efficiently.

[0243] At block 1006, the first electronic device receives motion data corresponding to the movement of the second electronic device from the second electronic device. In some embodiments, the motion data corresponding to the movement of the second electronic device includes at least one or more rotational rates and at least one or more acceleration rates corresponding to the second electronic device. In some embodiments, when causing a continuous audible sound to be provided at the second electronic device, the first electronic device causes a plurality of audible tones to be provided at the second electronic device, wherein the volume associated with the plurality of audible tones increases with each of the plurality of audible tones being provided. In some embodiments, after causing a message to be provided at the second electronic device, at least one tone is provided at the second electronic device, and in response to providing at least one tone at the second electronic device, receiving motion data from the second electronic device is initiated.

[0244] By providing multiple audible tones with increased volume or other features, the system enhances device functionality by notifying the user that a gesture is currently being detected. This enhanced feedback makes the device more efficient by providing user training on new device features and thus eliminating more conventional interaction methods. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the device more quickly and efficiently.

[0245] At box 1008, a first electronic device determines a gesture based on motion data (e.g., using an AI process or a generative AI process). In some embodiments, determining a gesture based on motion data includes obtaining a motion classification probability based on the motion data, and determining a gesture based on a corresponding motion classification if the determined (e.g., using an AI process or a generative AI process) motion classification probability exceeds a motion classification probability threshold. In some embodiments, if the determined (e.g., using an AI process or a generative AI process) gesture meets a predetermined gesture criterion, the first electronic device stops providing continuous audible sound at a second electronic device before the end of a predetermined time period, and after stopping the continuous audible sound, the first electronic device provides an acknowledgment audible tone based on the gesture. In some embodiments, if the determined (e.g., using an AI process or a generative AI process) gesture does not meet a predetermined gesture criterion, the first electronic device continues to provide continuous audible sound at the second electronic device, wherein the continuous audible sound is provided for a predetermined time period, and motion data continues to be received from the second electronic device. In some embodiments, when providing continuous audible sound at a second electronic device, a first electronic device determines a first portion of a gesture based on motion data (e.g., using an AI process or a generative AI process), and in response to determining (e.g., using an AI process or a generative AI process) that the first portion of the gesture meets a first predetermined partial gesture criterion, provides a first audible tone at the second electronic device. In some embodiments, after providing the first audible tone, the first electronic device determines a second portion of a gesture based on motion data (e.g., using an AI process or a generative AI process), and in response to determining (e.g., using an AI process or a generative AI process) that the second portion of the gesture meets a second predetermined partial gesture criterion, provides a second audible tone at the second electronic device. In some implementations, a response message is sent to the sender of the incoming text message based on a determined gesture (e.g., using an AI process or a generative AI process) corresponding to an accept gesture, wherein the response message (e.g., automatically generated text content and / or generative text content) is generated based on the response message subject (e.g., using an AI process or a generative AI process), and a response message (e.g., automatically generated text content and / or generative text content) is not sent to the sender of the incoming text message based on a determined gesture (e.g., using an AI process or a generative AI process) corresponding to a reject gesture.

[0246] By ceasing to provide continuous audible sound once a gesture is detected, the system enhances device functionality by minimizing unnecessary processing time on the device. This feature makes the device efficient by saving processing resources. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the device faster and more efficiently.

[0247] At step 1010, the first electronic device causes a first output to be provided at the second electronic device based on a gesture. In some embodiments, based on determining (e.g., using an AI process or a generative AI process) that the gesture corresponds to an accept gesture, the first electronic device causes an accept tone to be provided at the second electronic device as the first output. In some embodiments, the first electronic device corresponds to one of a smartphone, smartwatch, tablet computer, desktop computer, and laptop computer, and the second electronic device corresponds to one of a headset device and an earphone device.

[0248] By providing different tones corresponding to the detected gesture type, the system enhances device functionality by informing the user that a gesture has been identified and an action will be performed based on that gesture. This enhanced feedback makes the device more efficient by informing the user of the steps required to complete a gesture-based response. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the devices faster and more efficiently.

[0249] At step 1012, the first electronic device performs a first task associated with the event alarm. In some embodiments, based on determining (e.g., using an AI process or a generative AI process) that the gesture corresponds to a rejection gesture, the first electronic device causes a rejection tone to be provided as a first output at the second electronic device, wherein performing the first task includes stopping the reception of motion data. In some embodiments, performing the first task associated with the event alarm includes one of: connecting an incoming audio call to the second electronic device and providing an audible message at the second electronic device. In some embodiments, performing the first task associated with the event alarm includes receiving verbal input from a user from the second electronic device and sending a text message based on the received verbal input. In some embodiments, performing the first task associated with the event alarm includes providing an audible output containing the content of an incoming text message, wherein the subject of the incoming text message is generated based on the content of the incoming text message (e.g., using an AI process or a generative AI process).

[0250] By using gesture responses to support a variety of device tasks, the system enhances device functionality by increasing the methods available for interaction with the host device. Increasing the options available to the user makes the device more efficient by reducing the need for more cumbersome or traditional input methods. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the device faster and more efficiently.

[0251] refer to Figure 11At block 1102, the first electronic device detects an event alarm. In some embodiments, the event alarm corresponds to an incoming text message, and the message associated with the event alarm (e.g., automatically generated audio content and / or generated audio content) corresponds to an indication that the incoming text message exceeds a threshold character length. In some embodiments, the event alarm corresponds to an alert, and the message associated with the event alarm (e.g., automatically generated audio content and / or generated audio content) corresponds to an alert indicating that the alert includes multiple items exceeding an item threshold. At block 1104, the first electronic device causes a message (e.g., automatically generated audio content and / or generated audio content) to be provided to a second electronic device, wherein the message is associated with the event alarm. In some embodiments, an information prompt is provided before causing the message to be provided to the second electronic device, the information prompt indicating that the message provision to the second electronic device can be stopped based on a corresponding gesture using the second electronic device. In some embodiments, the first electronic device corresponds to one of a smartphone, smartwatch, tablet computer, desktop computer, and laptop computer, and the second electronic device corresponds to one of a headset device and an earphone device.

[0252] By providing users with early notifications about message cancellation, the system enhances device functionality by offering additional options for disabling device features. Allowing users to cancel device features using gesture-based input makes the device more efficient, avoiding unnecessary tasks such as message reading. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the device faster and more efficiently.

[0253] At block 1106, the first electronic device receives motion data corresponding to the movement of the second electronic device from the second electronic device. In some embodiments, while a message is being delivered at the second electronic device, the first electronic device determines (e.g., using an AI process or a generative AI process) a first partial gesture based on the motion data, and in response to determining (e.g., using an AI process or a generative AI process) that the first partial gesture meets a first predetermined partial gesture criterion, the first electronic device causes a first audible tone to be delivered at the second electronic device. In some embodiments, while a message is being delivered at the second electronic device, the first electronic device determines (e.g., using an AI process or a generative AI process) a second partial gesture based on the motion data, and in response to determining (e.g., using an AI process or a generative AI process) that the second partial gesture meets a second predetermined partial gesture criterion, a second audible tone is provided at the second electronic device. In some embodiments, in response to causing a message to be delivered at the second electronic device, motion data corresponding to the movement of the second electronic device is received, wherein the motion data is received until the message is no longer being delivered at the second electronic device.

[0254] By providing different tones corresponding to the detected gesture type, the system enhances device functionality by informing the user that a gesture has been identified and an action will be performed based on that gesture. This enhanced feedback makes the device more efficient by informing the user of the steps required to complete a gesture-based response. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the devices faster and more efficiently.

[0255] At box 1108, the first electronic device determines a gesture based on motion data (e.g., using an AI process or a generative AI process). In some embodiments, determining a gesture based on motion data (e.g., using an AI process or a generative AI process) includes obtaining a motion classification probability based on the motion data, and determining a gesture based on a corresponding motion classification if the determined (e.g., using an AI process or a generative AI process) motion classification probability exceeds a motion classification probability threshold. In some embodiments, the motion data corresponding to the movement of the second electronic device includes at least one or more rotational rates and at least one or more acceleration rates corresponding to the second electronic device. In some embodiments, based on the determination (e.g., using an AI process or a generative AI process) of a message provided at the second electronic device associated with multiple predetermined responses and the determination (e.g., using an AI process or a generative AI process) of a gesture corresponding to a rejection gesture, the first electronic device causes a rejection tone to be provided at the second electronic device and stops receiving motion data. In some implementations, based on a message provided at a second electronic device (e.g., using an AI process or a generative AI process) and associated with multiple predetermined responses, and based on a gesture (e.g., using an AI process or a generative AI process) corresponding to an acceptance gesture, the first electronic device causes an acceptance tone to be provided at the second electronic device and stops receiving motion data. In some implementations, the message includes one or more words corresponding to a query, and the multiple predetermined responses include acceptance of the query and rejection of the query.

[0256] By using gesture responses to support a variety of device tasks, the system enhances device functionality by increasing the methods available for interaction with the host device. Increasing the options available to the user makes the device more efficient by reducing the need for more cumbersome or traditional input methods. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the device faster and more efficiently.

[0257] At block 1110, based on the determination (e.g., using an AI process or a generative AI process) that a predetermined criterion is met, the first electronic device stops providing messages at the second electronic device. In some embodiments, based on the determination (e.g., using an AI process or a generative AI process) that a gesture meets the predetermined criterion, the first electronic device causes to provide an audible tone while stopping message provision at the second electronic device. In some embodiments, determining (e.g., using an AI process or a generative AI process) that a gesture meets the predetermined criterion includes associating the message provided at the second electronic device with multiple predetermined responses, and determining (e.g., using an AI process or a generative AI process) that a gesture does not meet the predetermined criterion. In some embodiments, based on the determination (e.g., using an AI process or a generative AI process) that a gesture does not meet the predetermined criterion, the first electronic device continues to provide messages at the second electronic device. In some embodiments, based on the determination (e.g., using an AI process or a generative AI process) that a gesture does not meet the predetermined criterion, the first electronic device continues to receive motion data corresponding to the movement of the second electronic device from the second electronic device and continues to determine gestures based on the motion data.

[0258] By facilitating the early deactivation of device functions, the system enhances device functionality by minimizing unnecessary processing time on the device. This feature makes the device more efficient by saving device processing resources. Therefore, these features improve human-computer interaction by enabling natural gesture-based input for wearable head-mounted devices, allowing users to use the device faster and more efficiently.

[0259] The above references Figures 10 to 11 The described operation can optionally be... Figures 1 to 4A , Figures 6A to 6B and Figures 7A to 7C The components described herein are used for implementation. For example, the operation of process 900 may be implemented by one or more of the following: operating system 718, application module 724, I / O processing module 728, STT processing module 730, natural language processing module 732, vocabulary index 744, task flow processing module 736, service processing module 738, media service 120-1, or processors 220, 410, and 704. Those skilled in the art will clearly understand how to implement... Figures 1 to 4A , Figures 6A to 6B and Figures 7A to 7C The components described herein are used to implement other processes.

[0260] According to some specific embodiments, a computer-readable storage medium (e.g., a non-transitory computer-readable storage medium) is provided that stores one or more programs executable by one or more processors of an electronic device, the one or more programs including instructions for performing any of the methods or processes described herein.

[0261] According to some specific embodiments, an electronic device (e.g., a portable electronic device) is provided, which includes components for performing any of the methods or processes described herein.

[0262] According to some specific embodiments, an electronic device (e.g., a portable electronic device) is provided, the electronic device including a processing unit configured to perform any of the methods or processes described herein.

[0263] According to some specific embodiments, an electronic device (e.g., a portable electronic device) is provided, the electronic device including one or more processors and a memory storing one or more programs for execution by the one or more processors, the one or more programs including instructions for performing any of the methods or processes described herein.

[0264] For purposes of explanation, the foregoing description has been given by reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible based on the teachings above. These embodiments were chosen and described in order to best explain the principles of these techniques and their practical application. Others skilled in the art will thus be able to best utilize these techniques and the various embodiments with various modifications suitable for the particular intended use.

[0265] While this disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. It should be understood that such changes and modifications are considered to be included within the scope of this disclosure and examples as defined by the claims.

[0266] Some implementations described herein may include the use of artificial intelligence and / or machine learning systems (sometimes referred to herein as AI / ML systems). This use may include the collection, processing, labeling, organizing, analyzing, recommending, and / or generating of data. Entities that collect, share, and / or otherwise utilize user data should provide transparency and / or obtain user consent when collecting such data. This disclosure recognizes that the use of data in AI / ML systems can be used to benefit users. For example, the data can be used to train models that can be deployed to improve the performance, accuracy, and / or functionality of applications and / or services. Therefore, the use of data enables AI / ML systems to adapt and / or optimize operations, thereby providing a more personalized, efficient, and / or enhanced user experience. Such adaptation and / or optimization may include customizing content, recommendations, and / or interactions for individual users, as well as simplifying processes and / or implementing more intuitive interfaces. This disclosure also anticipates further beneficial uses of data in AI / ML systems.

[0267] This disclosure anticipates that, in some implementations, the data used by the AI / ML system will include publicly available data. To protect user privacy, data may be anonymized, aggregated, and / or otherwise processed to remove or, where possible, limit any individual identification. As discussed herein, entities that collect, share, and / or otherwise utilize such data should obtain user consent before collecting such data and / or provide transparency in the process of collecting it. Furthermore, this disclosure anticipates that entities responsible for the use of data (including, but not limited to, data used in connection with AI / ML systems) should strive to comply with robust privacy policies and / or privacy measures.

[0268] For example, such entities can implement and consistently follow strategies and practices deemed to meet or exceed industry standards and regulatory requirements for developing and / or training AI / ML systems. In doing so, efforts should be made to ensure that all intellectual property and privacy considerations are maintained. Training should include measures to protect training data (such as personal information) by providing adequate safeguards against misuse or exploitation. Such strategies and practices should cover all phases of AI / ML system development, training, and use, including data collection, data preparation, model training, model evaluation, model deployment, and ongoing monitoring and maintenance. Transparency and measurability should always be maintained. Such policies should be readily accessible to users and should be updated as data collection and / or use change. User data should be collected for the entity's lawful and reasonable use and not shared or sold outside of these lawful uses. Furthermore, such collection and sharing should be conducted in a transparent manner to users and / or with their informed consent. Additionally, such entities should consider taking any necessary steps to defend and safeguard access to such data and ensure that others with access to the data comply with their privacy policies and processes. Furthermore, such entities should be subject to appropriate third-party assessments for transparency purposes to demonstrate their compliance with widely accepted privacy policies and practices. In addition, policies and / or practices should be appropriate for the specific types of data collected and / or accessed, and tailored to specific use cases and applicable laws and standards, including jurisdiction-specific considerations.

[0269] In some implementations, the AI / ML system may utilize a model that can be trained (e.g., in a supervised or unsupervised learning manner) using a variety of training data, including data collected using the user's device. This use of user-collected data may be limited to operations on the user's device. For example, model training may be performed locally on the user's device, so that no part of the data is transmitted to another device. In other implementations, model training may be performed using one or more other devices besides the user's device (e.g., a server), but in a privacy-preserving manner, such as through multi-party computation, encrypted by secretly sharing data, or other means, ensuring that user data is not leaked to these other devices.

[0270] In some implementations, the trained model can be centrally stored on the user device or on multiple devices, as is the case in federated learning. This distributed storage can also be done in a privacy-preserving manner, for example, through encryption, where each piece of data is fragmented so that the data cannot be reassembled or used by a device alone (i.e., only in conjunction with another device) or only by the user device. In this way, user or device behavior patterns are not leaked, while leveraging the increased computing resources of other devices to train and execute the ML model. Therefore, user-collected data can be protected. In some specific implementations, data from multiple devices can be combined in a privacy-preserving manner to train the ML model.

[0271] In some implementations, this disclosure contemplates that data used for AI / ML systems may be kept strictly separate from the platform on which the AI / ML system is deployed and / or used for user interaction and / or data processing. In such implementations, data used for offline training of the AI / ML system may be kept in a secure data repository with restricted access and / or not retained for longer than necessary for the training purpose. In some implementations, the AI / ML system may utilize a local memory cache to temporarily store data during a user session. Local memory caching can be used to improve the performance of the AI / ML system. However, to protect user privacy, data stored in the local memory cache may be erased after the user session ends. Any temporary cached data used for online learning or inference may be quickly erased after processing. All data collection, transfer, and / or storage should be conducted using industry-standard encryption and / or secure communication.

[0272] In some implementations, as noted above, techniques such as federated learning, differential privacy, secure hardware components, homomorphic encryption and / or multi-party computation, and others, can be used to further protect personal information data during the training and / or use of AI / ML systems. AI / ML systems should be monitored for changes in the underlying data distribution, such as concept drift or data skew that may degrade AI / ML system performance over time.

[0273] In some implementations, a combination of offline and online training is used to train the AI / ML system. Offline training can use a carefully selected dataset to establish baseline model performance, while online training allows the AI / ML system to continuously adapt and / or improve. This disclosure recognizes the importance of maintaining strict data management measures throughout this process to ensure user privacy is protected.

[0274] In some implementations, AI / ML systems may be designed with safeguards to maintain adherence to their original intended purpose, even when the AI / ML system is adapted based on new data. Any significant changes to data collection and / or application in the use of the AI / ML system may (and in some cases should) be transparently communicated to affected stakeholders, including (or including) obtaining user consent when the manner of collecting and / or using user data changes.

[0275] Regardless of the foregoing, this disclosure also contemplates implementations that allow users to selectively restrict and / or block the use and / or access to data. That is, this disclosure contemplates that hardware and / or software elements may be provided to prevent or block access to data. For example, with respect to some services, the inventive technology should be configured to allow users to opt-in or opt-out to participate in data collection at any time during or after service registration. In another example, the inventive technology should be configured to allow users to opt out of providing certain data used for training AI / ML systems and / or as input during the inference phase of such systems. In yet another example, the inventive technology should be configured to allow users to choose to limit the length of time data is retained or to completely prohibit AI / ML systems from using their data. In addition to providing "opt-in" and "opt-out" options, this disclosure also contemplates providing notifications related to access to or use of personal information. For example, users may be notified when their data is input into an AI / ML system for training or inference purposes, and / or alerted when the AI / ML system generates output or makes decisions based on their data.

[0276] This disclosure recognizes that AI / ML systems should incorporate explicit limitations and / or oversight to mitigate risks that may still exist even when such systems have been designed, developed, and / or operated in accordance with industry best practices and standards. For example, outputs may be generated that could be considered erroneous, harmful, offensive, and / or biased; such outputs may not necessarily reflect the opinions or positions of the entity that developed or deployed these systems. Furthermore, in some cases, references to third-party products and / or services in these outputs should not be construed as endorsement or affiliation of the third-party products and / or services by the entity providing the AI / ML system. The generated content may be filtered to remove potentially inappropriate or dangerous material before being presented to users, while maintaining the ability for human oversight and / or to cover or correct erroneous or undesirable outputs as a safeguard.

[0277] This disclosure also anticipates that users of the AI / ML system should avoid using the service in any way that infringes upon, misappropriates, or violates the rights of any party. Furthermore, the AI / ML system should not be used for any illegal or unlawful activities, nor should it be used to develop any application or use case that will commit or assist in the committing of a crime or other tortious, illegal, or unlawful act. The AI / ML system should not violate, misappropriate, or infringe upon any party's copyright, trademark, privacy and portrait rights, trade secrets, patents, or other proprietary or statutory rights, and should be appropriately attributed to the content as required. In addition, the AI / ML system should not interfere with any security, digital signature, digital rights management, content protection, verification, or authentication mechanisms. The AI / ML system should not misrepresent machine-generated output as human-generated.

[0278] As described above, one aspect of the present invention involves collecting and using data available from various sources to improve contextual response recommendations. This disclosure contemplates that, in some instances, such collected data may include personal information data that uniquely identifies or can be used to contact or locate specific individuals. This personal information data may include demographic data, location-based data, telephone numbers, email addresses, Twitter IDs, home addresses, data or records related to a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.

[0279] This disclosure recognizes that the use of such personal information data in the techniques of this invention can be beneficial to users. For example, personal information such as contact information can be used to give commands using gestures on assistive devices. Furthermore, this disclosure also anticipates other uses of personal information data that are beneficial to users. For example, health and fitness data can be used to provide insights into a user's overall health status or can be used as positive feedback for individuals using the technology to pursue health goals.

[0280] This disclosure anticipates that entities responsible for the collection, analysis, disclosure, transfer, storage, or other use of such personal information data will comply with robust privacy policies and / or privacy measures. Specifically, such entities should implement and adhere to privacy policies and measures that are recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. Such policies should be easily accessible to users and should be updated as the collection and / or use of data changes. Personal information from users should be collected for legitimate and reasonable entity purposes and should not be shared or sold outside of these legitimate purposes. Furthermore, such collection / sharing should be conducted only after receiving informed consent from users. Additionally, such entities should consider taking any necessary steps to protect and safeguard the right to access such personal information data and ensure that other entities with access to personal information data comply with the privacy policies and procedures of other entities. Furthermore, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and privacy measures. Moreover, policies and measures should be adapted to the specific types of personal information data collected and / or accessed, and to applicable laws and standards, including considerations of specific jurisdictions. For example, in the United States, the collection or acquisition of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); while in other countries, health data may be subject to other regulations and policies and should be handled accordingly. Therefore, different privacy measures should be advocated for different types of personal data in each country.

[0281] Regardless of the foregoing, this disclosure also contemplates implementation schemes for users to selectively block the use or access to personal information data. That is, this disclosure contemplates providing hardware and / or software components to prevent or block access to such personal information data. For example, the inventive technology can be configured to allow users to opt-in or opt-out to participate in the collection of personal information data during or at any time after registering for the service. In another example, users may choose not to provide gesture-based information. In yet another example, users may choose to limit the details provided regarding gesture information, device messages, etc. In addition to providing "opt-in" and "opt-out" options, this disclosure also contemplates providing notifications related to access to or use of personal information. For example, users may be notified when downloading an application that their personal information data will be accessed, and then reminded again just before the application accesses the personal information data.

[0282] Furthermore, the intent of this disclosure is that personal information data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Once data is no longer needed, this risk can be minimized by restricting data collection and deleting data. Additionally, and where applicable, including in certain health-related applications, data deidentification can be used to protect user privacy. Where appropriate, deidentification can be facilitated by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or characteristics of stored data (e.g., collecting location data at the city level rather than address level), controlling how data is stored (e.g., aggregating data among users), and / or other methods.

[0283] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it is also contemplated that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not become inoperable due to the absence of all or part of such personal information data. For example, commands using assistive device gestures can be executed on non-personal information data or a minimal amount of personal information, such as anonymous gesture information, other non-personal information available to gesture tracking systems, or publicly available information.

Claims

1. A computer-implemented method, the computer-implemented method comprising: At a first electronic device having memory and one or more processors: Detect event alerts; This enables a message to be provided at a second electronic device, wherein the message is associated with the event alarm; Receive motion data corresponding to the movement of the second electronic device; The gesture is determined based on the motion data; The gesture is used to enable a first output to be provided at the second electronic device; as well as Perform the first task associated with the event alarm.

2. The method of claim 1, wherein the event alarm corresponds to one of an incoming call, an incoming text message, an alert, and an application notification.

3. The method according to any one of claims 1 to 2, wherein the event alarm corresponds to an incoming text message, and the message associated with the event alarm corresponds to an indication that the incoming text message exceeds a threshold character length.

4. The method of any one of claims 1 to 3, wherein the incoming text message corresponds to a group text message, and wherein the message associated with the event alarm includes an option to silence the group text message.

5. The method according to any one of claims 1 to 4, wherein the event alarm corresponds to an incoming text message of no more than a threshold character length, and the message associated with the event alarm includes the content of the incoming text message.

6. The method according to any one of claims 1 to 5, wherein the motion data corresponding to the movement of the second electronic device includes at least one or more rotational rates corresponding to the second electronic device and at least one or more acceleration rates corresponding to the second electronic device.

7. The method according to any one of claims 1 to 6, wherein determining the gesture based on the motion data comprises: The motion classification probability is obtained based on the motion data. The gesture is determined based on the corresponding motion classification, since the probability of motion classification exceeds the motion classification probability threshold.

8. The method according to any one of claims 1 to 7, the method comprising: In response to provide the message at the second electronic device: This enables the provision of a continuous audible sound at the second electronic device, wherein the continuous audible sound is provided for a predetermined period of time.

9. The method according to claim 8, wherein the method comprises: Based on the determination that the gesture meets a predetermined gesture standard: The continuous audible sound is stopped being provided at the second electronic device before the end of the predetermined time period; as well as After the continuous audible sound is stopped, a confirmed audible tone is provided based on the gesture.

10. The method according to claim 8, wherein the method comprises: Based on the determination that the gesture does not meet the predetermined gesture standard: The continuous audible sound continues to be provided at the second electronic device, wherein the continuous audible sound is provided for a predetermined period of time; and Continue to receive the motion data from the second electronic device.

11. The method according to claim 8, wherein the method comprises: When the continuous audible sound is provided at the second electronic device: The first part of the gesture is determined based on the motion data; In response to determining that the first partial gesture meets a first predetermined partial gesture standard, a first audible tone is provided at the second electronic device.

12. The method according to claim 11, wherein the method comprises: After providing the first audible tone: The second part of the gesture is determined based on the motion data; In response to determining that the second partial gesture meets a second predetermined partial gesture standard, a second audible tone is provided at the second electronic device.

13. The method according to any one of claims 1 to 12, the method comprising: When providing continuous audible sound at the second electronic device, a plurality of audible tones are provided at the second electronic device, wherein the volume associated with the plurality of audible tones increases with each of the plurality of audible tones being provided.

14. The method according to any one of claims 1 to 13, the method comprising: Based on the determination that the gesture corresponds to a rejection gesture, a rejection tone is provided as the first output at the second electronic device, wherein performing the first task includes stopping the reception of the motion data.

15. The method according to any one of claims 1 to 14, the method comprising: Based on the determination that the gesture corresponds to a receiving gesture, a receiving tone is provided at the second electronic device as the first output.

16. The method of any one of claims 1 to 15, wherein performing the first task associated with the event alarm includes one of: connecting an incoming audio call to the second electronic device and providing an audible message at the second electronic device.

17. The method of any one of claims 1 to 16, wherein performing the first task associated with the event alarm comprises: Receive verbal input from the user from the second electronic device; as well as Send text messages based on the received verbal input.

18. The method according to any one of claims 1 to 17, wherein The first electronic device corresponds to one of a smartphone, smartwatch, tablet computer, desktop computer, and laptop computer, and The second electronic device corresponds to either a headset device or an earbud device.

19. The method of any one of claims 1 to 18, wherein the event alarm corresponds to an incoming text message, and the message associated with the event alarm includes the subject of the incoming text message.

20. The method of claim 19, wherein the method comprises: After the message is provided at the second electronic device, at least one tone is provided at the second electronic device; as well as In response to providing the at least one tone at the second electronic device, motion data is received from the second electronic device.

21. The method of claim 19, wherein performing a first task associated with the event alarm includes providing an audible output containing the content of the incoming text message, wherein the topic of the incoming text message is generated based on the content of the incoming text message.

22. The method according to any one of claims 1 to 21, wherein providing the message at the second electronic device comprises: This enables the first portion of the message to be provided at the second electronic device, wherein the first portion of the message includes at least a portion of the content of the incoming text message; as well as This enables the second portion of the message to be provided at the second electronic device, wherein the second portion of the message includes a prompt indicating a response to the message subject.

23. The method according to claim 22, wherein the method comprises: Based on the determination that the gesture corresponds to a receive gesture, a response message is sent to the sender of the incoming text message, wherein the response message is generated based on the response message subject; as well as Based on the determination that the gesture corresponds to a rejection gesture, the transmission of a response message to the sender of the incoming text message is abandoned.

24. An electronic device, the electronic device comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: Detect event alerts; This enables a message to be provided at a second electronic device, wherein the message is associated with the event alarm; Receive motion data corresponding to the movement of the second electronic device; The gesture is determined based on the motion data; The gesture is used to enable a first output to be provided at the second electronic device; as well as Perform the first task associated with the event alarm.

25. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first electronic device, the one or more programs including instructions for: Detect event alerts; This enables a message to be provided at a second electronic device, wherein the message is associated with the event alarm; Receive motion data corresponding to the movement of the second electronic device; The gesture is determined based on the motion data; The gesture is used to enable a first output to be provided at the second electronic device; as well as Perform the first task associated with the event alarm.

26. An electronic device, the electronic device comprising: Components used to detect event alarms; Components for enabling the provision of a message at a second electronic device, wherein the message is associated with the event alarm; A component for receiving motion data corresponding to the movement of the second electronic device from the second electronic device; A component for determining gestures based on the motion data; Components for providing a first output at the second electronic device based on the gesture; and A component used to perform the first task associated with the event alarm.

27. An electronic device, the electronic device comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 23.

28. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 23.

29. An electronic device, the electronic device comprising: Components for performing the method according to any one of claims 1 to 23.

30. A computer-implemented method, the computer-implemented method comprising: At a first electronic device having memory and one or more processors: Detect event alerts; This enables a message to be provided at a second electronic device, wherein the message is associated with the event alarm; Receive motion data corresponding to the movement of the second electronic device; The gesture is determined based on the motion data; as well as If the gesture is determined to meet a predetermined standard, the delivery of the message at the second electronic device is stopped.

31. The method of claim 30, wherein the first electronic device corresponds to one of a smartphone, a smartwatch, a tablet computer, a desktop computer, and a laptop computer.

32. The method according to any one of claims 30 to 31, wherein the second electronic device corresponds to one of a headphone device and an earphone device.

33. The method according to any one of claims 30 to 32, the method comprising: Before the message is provided at the second electronic device, an information prompt is provided indicating that the delivery of the message at the second electronic device can be stopped based on a corresponding gesture using the second electronic device.

34. The method according to any one of claims 30 to 33, the method comprising: While the message is being provided at the second electronic device: The first part of the gesture is determined based on the motion data; In response to determining that the first partial gesture meets a first predetermined partial gesture standard, a first audible tone is provided at the second electronic device.

35. The method according to claim 34, wherein the method comprises: While the message is being provided at the second electronic device: The second part of the gesture is determined based on the motion data; In response to determining that the second partial gesture meets a second predetermined partial gesture standard, a second audible tone is provided at the second electronic device.

36. The method according to any one of claims 30 to 35, the method comprising: In response to providing the message at the second electronic device, motion data corresponding to the movement of the second electronic device is received, wherein the motion data is received until the message is no longer provided at the second electronic device.

37. The method of any one of claims 30 to 36, wherein the event alarm corresponds to an incoming text message, and the message associated with the event alarm corresponds to an indication that the incoming text message exceeds a threshold character length.

38. The method of any one of claims 30 to 37, wherein the event alarm corresponds to an alert, and the message associated with the event alarm corresponds to the alert including an indication that multiple items exceed an item threshold.

39. The method according to any one of claims 30 to 38, wherein the motion data corresponding to the movement of the second electronic device includes at least one or more rotational rates corresponding to the second electronic device and at least one or more acceleration rates corresponding to the second electronic device.

40. The method according to any one of claims 30 to 39, wherein determining the gesture based on the motion data comprises: The motion classification probability is obtained based on the motion data. as well as The gesture is determined based on the corresponding motion classification, since the probability of motion classification exceeds the motion classification probability threshold.

41. The method according to any one of claims 30 to 40, the method comprising: Based on the determination that the gesture meets the predetermined criteria, an audible tone is provided while the message is stopped at the second electronic device.

42. The method according to any one of claims 30 to 41, wherein determining that the gesture satisfies the predetermined criterion comprises: Based on the determination that the message provided at the second electronic device is associated with a plurality of predetermined responses, it is determined that the gesture does not meet the predetermined criteria.

43. The method according to any one of claims 30 to 42, the method comprising: Based on the determination that the message provided at the second electronic device is associated with a plurality of predetermined responses and that the gesture corresponds to a rejection gesture: This enables a rejection tone to be provided at the second electronic device; as well as Stop receiving the motion data.

44. The method according to any one of claims 30 to 43, the method comprising: Based on the determination that the message provided at the second electronic device is associated with a plurality of predetermined responses and that the gesture corresponds to an accept gesture: This enables the receiving tone to be provided at the second electronic device; as well as Stop receiving the motion data.

45. The method according to any one of claims 30 to 44, wherein The message includes one or more words corresponding to the query, and The plurality of predetermined responses include acceptance of the inquiry and rejection of the inquiry.

46. ​​The method according to any one of claims 30 to 45, the method comprising: If it is determined that the gesture does not meet the predetermined criteria, the message continues to be provided at the second electronic device.

47. The method according to any one of claims 30 to 46, the method comprising: Based on the determination that the gesture does not meet the predetermined standard: Continue to receive motion data corresponding to the movement of the second electronic device; as well as The gesture is then determined based on the motion data.

48. A system comprising: A first electronic device includes: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the following operations: Detect event alerts; This enables a message to be provided at a second electronic device, wherein the message is associated with the event alarm; Receive motion data corresponding to the movement of the second electronic device; Determine the gesture based on the motion data; and If the gesture is determined to meet a predetermined standard, the delivery of the message at the second electronic device is stopped.

49. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a first electronic device, the one or more programs including instructions for: Detect event alerts; This enables a message to be provided at a second electronic device, wherein the message is associated with the event alarm; Receive motion data corresponding to the movement of the second electronic device; as well as The gesture is determined based on the motion data; as well as If the gesture is determined to meet a predetermined standard, the delivery of the message at the second electronic device is stopped.

50. A system including a first electronic device, the first electronic device comprising: Components used to detect event alarms; Components for enabling the provision of a message at a second electronic device, wherein the message is associated with the event alarm; A component for receiving motion data corresponding to the movement of the second electronic device from the second electronic device; and A component for determining gestures based on the motion data; and The component used to stop providing the message at the second electronic device is determined to meet a predetermined standard based on the gesture.

51. A system comprising: A first electronic device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for performing the method according to any one of claims 30 to 47.

52. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method according to any one of claims 30 to 47.

53. A system comprising: Components for performing the method according to any one of claims 30 to 47.

Citation Information

Patent Citations

  • System and method for inferring user intent from speech inputs

    US10176167B2

  • Method and apparatus for integrating manual input

    US20020015024A1

  • Acceleration-based theft detection system for portable electronic devices

    US20050190059A1

  • Methods and apparatuses for operating a portable device based on an accelerometer

    US20060017692A1

  • Gestures for touch sensitive input devices

    US20060026521A1