Method and system for interacting with audio events via motion input
By integrating sensors into audio output devices to detect motion input, the complex and inefficient audio interaction problems in existing technologies are solved, providing a faster and more efficient interaction method, improving user experience and device energy efficiency.
Patent Information
- Application Number
- CN202480019974.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-25
- Filing Date
- 2024-03-28
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies for audio output devices to interact with audio data are complex and inefficient, requiring multiple button presses or voice inputs, wasting time and device energy, especially in battery-powered devices.
By integrating sensors into audio output devices, motion input can be detected and related operations can be performed based on the motion detection results within a threshold time period, providing a faster and more efficient interaction method, reducing the cognitive burden on users and saving power.
It achieves a faster and more efficient human-computer interface, reduces redundant input, extends battery life, and improves the efficiency of users responding to audio notifications.
Smart Images

Figure CN120958429A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Patent Application No. 18 / 614,974, filed March 25, 2024, entitled “METHODS AND SYSTEMS FOR INTERACTING WITH AUDIO EVENTS VIA MOTION INPUTS”, and U.S. Provisional Application No. 63 / 456,449, filed March 31, 2023, entitled “METHODS AND SYSTEMS FOR INTERACTING WITH AUDIO EVENTS VIA MOTION INPUTS”. The entire contents of each of these patent applications are incorporated herein by reference. Technical Field
[0003] This disclosure relates generally to audio output devices, and more specifically to techniques for interacting with audio data via motion input. Background Technology
[0004] Electronic devices can provide audio data to audio output devices such as wireless speakers and wireless headphones via wireless connections. Example audio output devices can interact with audio data using various input technologies. Summary of the Invention
[0005] However, some technologies used to interact with audio data using electronic devices and / or audio output devices are often cumbersome and inefficient. For example, some existing technologies are complex, time-consuming, and limiting, potentially requiring voice input and / or multiple keystrokes or button presses. Existing technologies take more time than necessary, resulting in wasted user time and device power. This latter consideration is particularly important in battery-powered devices.
[0006] Therefore, the present invention provides audio output devices with a faster and more efficient method for interacting with audio data. Such methods optionally supplement or replace other methods for interacting with audio data. These methods and interfaces reduce the cognitive burden on the user and result in a more efficient human-computer interface. For battery-powered computing devices, such methods reduce the amount of irrelevant received input, save power, and increase the time interval between battery charges.
[0007] According to some embodiments, a method performed at one or more audio output devices is described. The method includes: outputting a first audio notification; after outputting the first audio notification, detecting motion input based on one or more sensor measurements from one or more sensors in the one or more audio output devices; and in response to the detected motion input and based on determining that a first set of criteria is met, causing a first operation associated with the first audio notification to be performed, wherein the first set of criteria includes a first criterion that is met when motion input is detected within a threshold time period during which the first audio notification is output.
[0008] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors at one or more audio output devices, the one or more programs including instructions for: outputting a first audio notification; after outputting the first audio notification, detecting motion input based on one or more sensor measurements from one or more sensors in one or more audio output devices; and in response to the detected motion input and based on determining that a first set of criteria is met, causing a first operation associated with the first audio notification to be performed, wherein the first set of criteria includes a first criterion that is met when motion input is detected within a threshold time period during which the first audio notification is output.
[0009] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors at one or more audio output devices, the one or more programs including instructions for: outputting a first audio notification; after outputting the first audio notification, detecting motion input based on one or more sensor measurements from one or more sensors in one or more audio output devices; and in response to the detected motion input and based on determining that a first set of criteria is met, causing a first operation associated with the first audio notification to be performed, wherein the first set of criteria includes a first criterion that is met when motion input is detected within a threshold time period during which the first audio notification is output.
[0010] According to some embodiments, one or more audio output devices are described, the one or more audio output devices including: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: outputting a first audio notification; after outputting the first audio notification, detecting motion input based on one or more sensor measurements from one or more sensors in the one or more audio output devices; and in response to the detected motion input and based on determining that a first set of criteria is met, causing a first operation associated with the first audio notification to be performed, wherein the first set of criteria includes a first criterion that is met when motion input is detected within a threshold time period during which the first audio notification is output.
[0011] According to some embodiments, one or more audio output devices are described. The one or more audio output devices include: components for outputting a first audio notification; components for detecting motion input based on one or more sensor measurements from one or more sensors in the one or more audio output devices after the first audio notification is output; and components for performing a first operation associated with the first audio notification in response to the detected motion input and based on determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion that is met when motion input is detected within a threshold time period during which the first audio notification is output.
[0012] According to some embodiments, a computer program product is described, comprising one or more programs configured to be executed by one or more processors at one or more audio output devices. The one or more programs include instructions for: outputting a first audio notification; after outputting the first audio notification, detecting motion input based on one or more sensor measurements from one or more sensors in one or more audio output devices; and in response to the detected motion input and based on determining that a first set of criteria is met, causing a first operation associated with the first audio notification to be performed, wherein the first set of criteria includes a first criterion that is met when motion input is detected within a threshold time period during which the first audio notification is output.
[0013] According to some embodiments, a method performed at one or more audio output devices is described. The method includes: detecting one or more sensor measurements corresponding to the start of a motion posture; providing first audio feedback indicating progress of the motion posture via one or more audio output devices after the detection of the one or more sensor measurements corresponding to the start of the motion posture and while the detection of the one or more sensor measurements is in progress; and after providing the first audio feedback and based on the determination that the motion posture is complete, causing an operation associated with the motion posture to be performed.
[0014] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors at one or more audio output devices, the one or more programs including instructions for: detecting one or more sensor measurements corresponding to the start of a motion posture; providing first audio feedback indicating the progress of the motion posture via one or more audio output devices after detecting the one or more sensor measurements corresponding to the start of the motion posture and while the detection of the one or more sensor measurements is in progress; and, after providing the first audio feedback and based on determining that the motion posture is complete, causing operations associated with the motion posture to be performed.
[0015] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors at one or more audio output devices, the one or more programs including instructions for: detecting one or more sensor measurements corresponding to the start of a motion posture; providing first audio feedback indicating the progress of the motion posture via one or more audio output devices after detecting the one or more sensor measurements corresponding to the start of the motion posture and while the detection of the one or more sensor measurements is in progress; and, after providing the first audio feedback and based on determining that the motion posture is complete, causing operations associated with the motion posture to be performed.
[0016] According to some embodiments, one or more audio output devices are described, the one or more audio output devices comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: detecting one or more sensor measurements corresponding to the start of a motion posture; providing first audio feedback indicating the progress of the motion posture via the one or more audio output devices after the detection of the one or more sensor measurements corresponding to the start of the motion posture and while the detection of the one or more sensor measurements is in progress; and, after providing the first audio feedback and based on the determination that the motion posture is complete, causing the execution of operations associated with the motion posture.
[0017] According to some embodiments, one or more audio output devices are described. The one or more audio output devices include: components for detecting one or more sensor measurements corresponding to the start of a motion posture; components for providing first audio feedback indicating the progress of the motion posture via the one or more audio output devices after detecting the one or more sensor measurements corresponding to the start of the motion posture and while the detection of the one or more sensor measurements is in progress; and components for performing operations associated with the motion posture after providing the first audio feedback and based on determining that the motion posture is complete.
[0018] According to some embodiments, a computer program product is described, comprising one or more programs configured to be executed by one or more processors at one or more audio output devices. The one or more programs include instructions for: detecting one or more sensor measurements corresponding to the start of a motion posture; providing first audio feedback indicating the progress of the motion posture via one or more audio output devices after the detection of the one or more sensor measurements corresponding to the start of the motion posture and while the detection of the one or more sensor measurements is in progress; and, after providing the first audio feedback and based on the determination that the motion posture is complete, causing the execution of operations associated with the motion posture.
[0019] According to some embodiments, a method performed at one or more audio output devices is described. The method includes: detecting one or more sensor measurements of a first movement in a three-dimensional environment corresponding to a corresponding portion of a user in relation to the one or more audio output devices; and in response to detecting the one or more sensor measurements corresponding to the first movement: based on determining that the first movement corresponds to a first positional orientation of the corresponding portion of the user toward the three-dimensional environment, outputting a first sound having an analog spatial position corresponding to the first position in the three-dimensional environment, wherein the first sound corresponds to a first optional option among one or more optional options; and based on determining that the first movement corresponds to a second positional orientation of the corresponding portion of the user toward the three-dimensional environment different from the first position in the three-dimensional environment, outputting a second sound having an analog spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among one or more optional options different from the first optional option.
[0020] According to some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors at one or more audio output devices, the one or more programs including instructions for: detecting one or more sensor measurements corresponding to a first movement of a corresponding part of a user in a three-dimensional environment; and in response to detecting the one or more sensor measurements corresponding to the first movement: based on determining that the first movement corresponds to a first positional orientation of the corresponding part of the user toward a first position in the three-dimensional environment, outputting a first sound having an analog spatial position corresponding to the first position in the three-dimensional environment, wherein the first sound corresponds to a first optional option among one or more optional options; and based on determining that the first movement corresponds to a second positional orientation of the corresponding part of the user toward a second position in the three-dimensional environment different from the first position in the three-dimensional environment, outputting a second sound having an analog spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among one or more optional options different from the first optional option.
[0021] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors at one or more audio output devices, the one or more programs including instructions for: detecting one or more sensor measurements corresponding to a first movement of a corresponding part of a user in a three-dimensional environment; and in response to detecting the one or more sensor measurements corresponding to the first movement: based on determining that the first movement corresponds to a first positional orientation of the corresponding part of the user toward a first position in the three-dimensional environment, outputting a first sound having an analog spatial position corresponding to the first position in the three-dimensional environment, wherein the first sound corresponds to a first optional option among one or more optional options; and based on determining that the first movement corresponds to a second positional orientation of the corresponding part of the user toward a second position in the three-dimensional environment different from the first position in the three-dimensional environment, outputting a second sound having an analog spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among one or more optional options different from the first optional option.
[0022] According to some embodiments, one or more audio output devices are described, the one or more audio output devices comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for: detecting one or more sensor measurements corresponding to a first movement of a corresponding part of a user in a three-dimensional environment; and in response to detecting the one or more sensor measurements corresponding to the first movement: based on determining that the first movement corresponds to a first positional orientation of the corresponding part of the user toward a first position in the three-dimensional environment, outputting a first sound having an analog spatial position corresponding to the first position in the three-dimensional environment, wherein the first sound corresponds to a first optional option among one or more optional options; and based on determining that the first movement corresponds to a second positional orientation of the corresponding part of the user toward a second position in the three-dimensional environment different from the first position in the three-dimensional environment, outputting a second sound having an analog spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among one or more optional options different from the first optional option.
[0023] According to some embodiments, one or more audio output devices are described. The one or more audio output devices include: components for detecting one or more sensor measurements of a first movement of a corresponding part of a user in a three-dimensional environment; and components for performing the following operations in response to detecting the one or more sensor measurements corresponding to the first movement: outputting a first sound having an analog spatial position corresponding to the first position in the three-dimensional environment, wherein the first sound corresponds to a first optional option among one or more optional options, based on determining that the first movement corresponds to a first position orientation of the corresponding part of the user in the three-dimensional environment, and outputting a second sound having an analog spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among one or more optional options, and based on determining that the first movement corresponds to a second position orientation of the corresponding part of the user in the three-dimensional environment, different from the first position in the three-dimensional environment, and based on determining that the first movement corresponds to a second position orientation of the corresponding part of the user in the three-dimensional environment, different from the first optional option ... optional option among one or more optional options, and based on determining that the first movement corresponds to a second position orientation of the corresponding part of the user in the three-dimensional environment, different from the first optional option, and based on determining that the first movement corresponds to a second optional option among one or more optional options, different from the first optional option.
[0024] According to some embodiments, a computer program product is described, comprising one or more programs configured to be executed by one or more processors at one or more audio output devices. The one or more programs include instructions for: detecting one or more sensor measurements corresponding to a first movement of a corresponding part of a user in a three-dimensional environment, corresponding to the one or more audio output devices; and in response to detecting the one or more sensor measurements corresponding to the first movement: based on determining that the first movement corresponds to a first positional orientation of the corresponding part of the user toward a first position in the three-dimensional environment, outputting a first sound having an analog spatial position corresponding to the first position in the three-dimensional environment, wherein the first sound corresponds to a first optional option among one or more optional options; and based on determining that the first movement corresponds to a second positional orientation of the corresponding part of the user toward a second position in the three-dimensional environment different from the first position in the three-dimensional environment, outputting a second sound having an analog spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among one or more optional options different from the first optional option.
[0025] Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0026] Therefore, providing devices with faster and more efficient methods and interfaces for interacting with audio data via motion input improves the effectiveness, efficiency, and user satisfaction of such devices. These methods and interfaces can complement or replace other methods used for interacting with audio data via motion input. Attached Figure Description
[0027] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals in all the drawings indicate corresponding parts.
[0028] Figure 1A This is a block diagram illustrating a portable multi-functional device with a touch-sensitive display according to some implementation schemes.
[0029] Figure 1B This is a block diagram illustrating exemplary components for event handling according to some implementation schemes.
[0030] Figure 2 Examples of portable multi-functional devices with touchscreens according to some implementation schemes are shown.
[0031] Figure 3This is a block diagram of an exemplary multifunctional device having a display and a touch-sensitive surface according to some implementation schemes.
[0032] Figure 4A An exemplary user interface for a menu applied to a portable multi-functional device, according to some implementation schemes, is illustrated.
[0033] Figure 4B An exemplary user interface for a multifunctional device having a touch-sensitive surface separate from the display is illustrated according to some embodiments.
[0034] Figure 5A Examples of personal electronic devices according to some implementation schemes are shown.
[0035] Figure 5B This is a block diagram illustrating a personal electronic device according to some implementation schemes.
[0036] Figures 6A to 6R Example methods for detecting motion input to interact with audio notifications and providing audio feedback for detected motion postures are illustrated according to some implementation schemes.
[0037] Figure 7 This is a block diagram of a method for detecting motion input to interact with audio notifications, based on some implementation schemes.
[0038] Figure 8 This is a block diagram of a method for providing audio feedback for detected motion postures, based on some implementation schemes.
[0039] Figures 9A to 9N An example method for detecting motion input in a spatial audio arrangement, according to the implementation scheme, is illustrated.
[0040] Figure 10 This is a block diagram of a method for detecting motion input in a spatial audio arrangement, according to the implementation plan. Detailed Implementation
[0041] The following description illustrates exemplary methods, parameters, etc. However, it should be understood that such description is not intended to limit the scope of this disclosure, but is provided as a description of exemplary embodiments.
[0042] Electronic devices require efficient methods and interfaces for interacting with audio data. For example, when providing audio notifications from an electronic device to a connected audio output device, users typically must provide a set of voice and / or touch inputs to respond to the audio notification in the desired manner. The technology disclosed in this invention reduces the amount of voice and touch input required for users to respond to audio notifications and provides an additional category of input methods—motion input—that users can use to create more accurate responses to audio notifications. Such technologies can reduce the cognitive burden on users when interacting with audio, thereby improving productivity. Furthermore, such technologies can reduce processor power and battery power that would otherwise be wasted on redundant user inputs.
[0043] under Figures 1A to 1B , Figure 2 , Figure 3 , Figures 4A to 4B and Figures 5A to 5B A description of an exemplary device for performing techniques for interacting with audio data is provided. Figures 6A to 6R Exemplary methods for detecting motion input to interact with audio notifications and providing audio feedback for detected motion postures are illustrated. Figure 7 This is a flowchart illustrating a method for detecting motion input to interact with audio notifications, according to some implementation schemes. Figure 8 This is a flowchart illustrating a method for providing audio feedback for detected motion postures. Figures 6A to 6R Used to illustrate the process described below, including Figure 7 and Figure 8 The process in. Figures 9A to 9N An exemplary method for detecting motion input in a spatial audio arrangement is illustrated. Figure 10 This is a flowchart illustrating a method for detecting motion input in a spatial audio arrangement according to some implementation schemes. Figures 9A to 9N The user interface in the document is used to illustrate the processes described below, including Figure 10 The process in.
[0044] The processes described below enhance device operability and make the user-device interface more efficient through various technologies (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), including providing improved visual feedback to users, reducing the amount of input required to perform operations, providing additional control options without cluttering the user interface with additional displayed controls, performing operations when a set of conditions are met without requiring further user input, and / or additional technologies. These technologies also reduce power consumption and extend device battery life by enabling users to use the device more quickly and efficiently.
[0045] Furthermore, in methods described herein where one or more steps depend on the satisfaction of one or more conditions, it should be understood that the method may be repeated in multiple repetitions such that, during the repetitions, all conditions determining the steps in the method are satisfied in different repetitions of the method. For example, if the method requires performing a first step (if the condition is satisfied) and a second step (if the condition is not satisfied), those skilled in the art will know that the stated steps are repeated until both the conditions are satisfied and not satisfied (in no particular order). Thus, a method described as having one or more steps depending on the satisfaction of one or more conditions can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing discretionary operations based on the satisfaction of the corresponding one or more conditions, and thus to determine whether possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions determining the steps in the method are satisfied. Those skilled in the art will also understand that, similar to methods having discretionary steps, a system or computer-readable storage medium may repeat the steps of the method multiple times as needed to ensure that all discretionary steps have been performed.
[0046] Although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. In some embodiments, these terms are used to distinguish one element from another. For example, a first touch may be referred to as a second touch, and similarly, a second touch may be referred to as a first touch, without departing from the scope of the various described embodiments. In some embodiments, a first touch and a second touch are two separate references to the same touch. In some embodiments, both a first touch and a second touch are touches, but they are not the same touch.
[0047] The terminology used in the description of the various described embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various described embodiments and in the appended claims, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “comprising” and / or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0048] Depending on the context, the term "if" may optionally be interpreted as meaning "when," "in response to," or "in response to detection." Similarly, depending on the context, the phrases "if it is determined..." or "if [the stated condition or event] is detected" may optionally be interpreted as meaning "in response to determining..." or "in response to detecting [the stated condition or event]."
[0049] This document describes implementations of electronic devices, user interfaces for such devices, and associated processes for using such devices. In some implementations, the device is a portable communication device, such as a mobile phone, that also includes other functionalities such as PDA and / or music player functionality. Exemplary implementations of portable multi-functional devices include, but are not limited to, those from Apple Inc. of Cupertino, California. Devices, iPod Equipment and Device. Optionally, other portable electronic devices may be used, such as laptop computers or tablet computers with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). In some embodiments, the electronic device is a computer system that communicates with a display generating component (e.g., via wireless or wired communication). The display generating component is configured to provide visual output, such as display via a CRT display, via an LED display, or via image projection. In some embodiments, the display generating component is integrated with the computer system. In some embodiments, the display generating component is separate from the computer system. As used herein, “display” content includes displaying content (e.g., video data rendered or decoded by display controller 156) by sending data (e.g., image data or video data) to an integrated or external display generating component via a wired or wireless connection to visually generate content.
[0050] In the following discussion, an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device may optionally include one or more other physical user interface devices, such as a physical keyboard, mouse, and / or joystick.
[0051] The device typically supports a variety of applications, such as one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, game applications, telephone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camera applications, web browsing applications, digital music player applications, and / or digital video player applications.
[0052] Various applications running on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the device are optionally adjusted and / or varied for different applications, and / or within the respective applications. In this way, the common physical architecture of the device (such as the touch-sensitive surface) optionally utilizes a user interface that is intuitive and clear to the user to support various applications.
[0053] Now let’s turn our attention to implementation schemes for portable devices with touch-sensitive displays. Figure 1A This is a block diagram illustrating a portable multi-functional device 100 with a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display 112 is sometimes referred to as a “touchscreen” for convenience, and is sometimes referred to as or called a “touch-sensitive display system.” Device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral interface 118, RF circuitry 108, audio circuitry 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact strength sensors 165 for detecting the intensity of contact on device 100 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more haptic output generators 167 for generating haptic output on device 100 (e.g., generating haptic output on a touch-sensitive surface such as the touch-sensitive display system 112 of device 100 or the touchpad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0054] As used in this specification and claims, the term "intensity" of contact on a tactile surface refers to the force or pressure (force per unit area) of a contact (e.g., finger contact) on a tactile surface, or to a substitute (alternative) for the force or pressure of a contact on a tactile surface. The intensity of contact has a range of values, including at least four different values and more typically hundreds of different values (e.g., at least 256). The intensity of contact is optionally determined (or measured) using various methods and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the tactile surface are optionally used to measure the force at different points on the tactile surface. In some embodiments, force measurements from multiple force sensors are combined (e.g., weighted average) to determine the estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the tactile surface. Alternatively, the size and / or change of the contact area detected on the touch-sensitive surface, the capacitance and / or change of the touch-sensitive surface adjacent to the contact, and / or the resistance and / or change of the touch-sensitive surface adjacent to the contact may optionally be used as substitutes for the force or pressure of the contact on the touch-sensitive surface. In some embodiments, the substitute measurement of the contact force or pressure is used directly to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the substitute measurement). In some embodiments, the substitute measurement of the contact force or pressure is converted into an estimated force or pressure, and the estimated force or pressure is used to determine whether an intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of the contact as an attribute of user input allows users to access additional device functionality that would otherwise be inaccessible to the user on a smaller device with limited physical space, which is used (e.g., on a touch-sensitive display) to display an indication and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls, such as knobs or buttons).
[0055] As used in this specification and claims, the term "haptic output" refers to a physical displacement of the device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., the housing), or a displacement of a component relative to the center of mass of the device, which is detected by the user using the user's tactile sense. For example, when the device or a component of the device comes into contact with a touch-sensitive surface (e.g., a finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or touchpad) may optionally be interpreted by the user as a "press-click" or "release-click" on a physically actuated button. In some cases, the user will feel a tactile sensation, such as a "press-click" or "release-click," even when a physically actuated button associated with a touch-sensitive surface that has been physically pressed (e.g., displaced) by the user's movement does not move. As another example, even when the smoothness of the tactile surface remains unchanged, the movement of the tactile surface can optionally be interpreted or perceived by the user as the “roughness” of the tactile surface. While such interpretations of touch by users will be limited by the individualized sensory perceptions of the user, many sensory perceptions of touch are common to most users. Therefore, when a tactile output is described as corresponding to a specific sensory perception of the user (e.g., “release click,” “press click,” “roughness”), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or its components that will generate the sensory perception described by a typical (or average) user.
[0056] It should be understood that device 100 is merely an example of a portable multifunctional device, and device 100 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of these components. Figure 1A The various components shown are implemented in hardware, software, or a combination of both, including one or more signal processing and / or application-specific integrated circuits.
[0057] Memory 102 optionally includes high-speed random access memory, and also optionally includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.
[0058] Peripheral interface 118 can be used to couple the device's input and output peripherals to CPU 120 and memory 102. The one or more processors 120 run or execute various software programs (such as computer programs (e.g., including instructions)) and / or instruction sets stored in memory 102 to perform various functions of device 100 and process data. In some embodiments, peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0059] RF (Radio Frequency) circuit 108 receives and transmits RF signals, also known as electromagnetic signals. RF circuit 108 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 108 optionally includes well-known circuitry for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, subscriber identity module (SIM) cards, memory, etc. RF circuit 108 optionally communicates wirelessly with networks (such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs))) and other devices. RF circuit 108 optionally includes well-known circuitry for detecting near-field communication (NFC) fields, such as via short-range communication radio components. Wireless communication may optionally employ any of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Unit HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), and Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Messaging Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence with Extended Utility (SIMPLE), Instant Messaging and Presence Service (IMPS)) and / or Short Message Service (SMS), or any other suitable communication protocol that has not been developed as of the date of this document submission.
[0060] Audio circuitry 110, speaker 111, and microphone 113 provide an audio interface between the user and device 100. Audio circuitry 110 receives audio data from peripheral interface 118, converts the audio data into electrical signals, and sends the electrical signals to speaker 111. Speaker 111 converts the electrical signals into sound waves that are audible to humans. Audio circuitry 110 also receives electrical signals converted from sound waves by microphone 113. Audio circuitry 110 converts the electrical signals into audio data and sends the audio data to peripheral interface 118 for processing. Audio data is optionally retrieved by peripheral interface 118 from and / or sent to memory 102 and / or RF circuitry 108. In some embodiments, audio circuitry 110 also includes a headset jack (e.g., ...). Figure 2 (212 in the text). The headset jack provides an interface between the audio circuitry 110 and a removable audio input / output peripheral device, such as an output-only headphone or a headset with both outputs (e.g., a single-ear or dual-ear headphone) and inputs (e.g., a microphone).
[0061] I / O subsystem 106 couples input / output peripherals (such as touchscreen 112 and other input control devices 116) on device 100 to peripheral interface 118. I / O subsystem 106 optionally includes display controller 156, optical sensor controller 158, depth camera controller 169, intensity sensor controller 159, haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive electrical signals from / transmit electrical signals to the other input control device 116. Other input control devices 116 optionally include physical buttons (e.g., push-buttons, rocker buttons, etc.), dials, slide switches, joysticks, click dials, etc. In some embodiments, input controller 160 is optionally coupled to (or not coupled to) any of the following: keyboard, infrared port, USB port, and pointing device such as mouse. One or more buttons (e.g., ... Figure 2 Optionally, 208) includes an increase / decrease button for volume control of speaker 111 and / or microphone 113. The one or more buttons optionally include a push-button (e.g., Figure 2(Ref. 206 in the original text). In some embodiments, the electronic device is a computer system that communicates with one or more input devices (e.g., via wireless communication or via wired communication). In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a touchpad, as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system. In some implementations, air gestures are gestures detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and based on the detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another of the user's fingers or a part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0062] A quick press of a push button optionally disengages the touchscreen 112 from its lock or optionally initiates a process of unlocking the device using gestures on the touchscreen, as described in U.S. Patent Application 11 / 322,549 (i.e., U.S. Patent No. 7,657,849), filed December 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image," the entire contents of which are incorporated herein by reference. A long press of a push button (e.g., 206) optionally powers the device 100 on or off. The functionality of one or more of these buttons is optionally user-customizable. The touchscreen 112 is used to implement virtual buttons or soft buttons and one or more soft keyboards.
[0063] The touch-sensitive display 112 provides input and output interfaces between the device and the user. The display controller 156 receives electrical signals from and / or transmits electrical signals to the touchscreen 112. The touchscreen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively, "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.
[0064] Touchscreen 112 has a touch-sensitive surface, sensor, or sensor array that accepts input from a user based on tactile and / or haptic contact. Touchscreen 112 and display controller 156 (along with any associated modules and / or instruction set in memory 102) detect contact on touchscreen 112 (and any movement or interruption of that contact) and translate the detected contact into interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on touchscreen 112. In an exemplary embodiment, the contact point between touchscreen 112 and the user corresponds to the user's finger.
[0065] Touchscreen 112 optionally employs LCD (Liquid Crystal Display) technology, LPD (Light Emitting Polymer Display) technology, or LED (Light Emitting Diode) technology, but other display technologies are used in other embodiments. Touchscreen 112 and display controller 156 optionally employ any of a variety of touch sensing technologies now known or to be developed hereafter, along with other proximity sensor arrays or other elements for determining one or more points of contact with touchscreen 112, to detect contact and any movement or interruption thereof. These various touch sensing technologies include, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that from Apple Inc. (Cupertino, California). and iPod The technology used.
[0066] In some embodiments of the touchscreen 112, the touch-sensitive display optionally resembles a multi-touch-sensitive touchpad described in the following U.S. patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman) and / or U.S. Patent Publication 2002 / 0015024A1, each of which is incorporated herein by reference in its entirety. However, the touchscreen 112 displays visual output from the device 100, while the touch-sensitive touchpad does not provide visual output.
[0067] The touch-sensitive display in some embodiments of the touchscreen 112 is described in the following applications: (1) U.S. Patent Application No. 11 / 381,313, filed May 2, 2006, “Multipoint Touch Surface Controller”; (2) U.S. Patent Application No. 10 / 840,862, filed May 6, 2004, “Multipoint Touchscreen”; (3) U.S. Patent Application No. 10 / 903,964, filed July 30, 2004, “Gestures For Touch Sensitive Input Devices”; (4) U.S. Patent Application No. 11 / 048,264, filed January 31, 2005, “Gestures For Touch Sensitive Input Devices”; and (5) U.S. Patent Application No. 11 / 038,590, filed January 18, 2005, “Mode-Based Graphical User Interfaces For Touch Sensitive Input”. (6) U.S. Patent Application No. 11 / 228,758, filed September 16, 2005, “Virtual Input Device Placement On A Touch Screen User Interface”; (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, “Operation Of A Computer With A Touch Screen Interface”; (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, “Activating Virtual Keys Of ATouch-Screen Virtual Keyboard”; and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, “Multi-Functional Hand-Held Device”. The full text of all these applications is incorporated herein by reference.
[0068] Touchscreen 112 optionally has a video resolution exceeding 100 dpi. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. Users optionally use any suitable object or accessory such as a stylus, finger, etc., to interact with touchscreen 112. In some embodiments, the user interface is designed to operate primarily through finger-based touch and gestures, which may be less precise than stylus-based input due to the larger contact area of a finger on the touchscreen. In some embodiments, the device translates coarse finger-based input into precise pointer / cursor positioning or commands for performing the user-desired actions.
[0069] In some embodiments, in addition to the touchscreen, device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touchscreen, does not display visual output. Optionally, the touchpad is a touch-sensitive surface separate from the touchscreen 112, or an extension of the touch-sensitive surface formed by the touchscreen.
[0070] The device 100 also includes a power system 162 for supplying power to various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in the portable device.
[0071] The device 100 may optionally also include one or more optical sensors 164. Figure 1AAn optical sensor 164 is shown coupled to an optical sensor controller 158 in the I / O subsystem 106. The optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The optical sensor 164 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In conjunction with an imaging module 143 (also referred to as a camera module), the optical sensor 164 optionally captures still images or video. In some embodiments, the optical sensor is located on the rear of the device 100, facing away from a touchscreen display 112 on the front of the device, allowing the touchscreen display to be used as a viewfinder for still image and / or video image acquisition. In some embodiments, the optical sensor is located on the front of the device, allowing images of the user to be optionally acquired for video conferencing while the user views other video conferencing participants on the touchscreen display. In some embodiments, the positioning of the optical sensor 164 can be changed by the user (e.g., by rotating the lenses and sensors within the device housing), allowing a single optical sensor 164 to be used in conjunction with the touchscreen display for both video conferencing and still image and / or video image acquisition.
[0072] The device 100 optionally also includes one or more depth camera sensors 175. Figure 1A A depth camera sensor 175 is shown coupled to a depth camera controller 169 in I / O subsystem 106. The depth camera sensor 175 receives data from the environment to create a 3D model of an object (e.g., a face) within the scene from a viewpoint (e.g., the depth camera sensor). In some embodiments, in conjunction with an imaging module 143 (also referred to as a camera module), the depth camera sensor 175 is optionally used to determine depth maps of different portions of an image captured by the imaging module 143. In some embodiments, the depth camera sensor is located at the front of device 100, such that user images with depth information are optionally acquired for video conferencing while a user views other video conferencing participants on a touchscreen display, and selfies with depth map data are captured. In some embodiments, the depth camera sensor 175 is located at the rear of the device, or at both the rear and front of device 100. In some embodiments, the positioning of the depth camera sensor 175 can be changed by the user (e.g., by rotating a lens and sensor within the device housing), such that the depth camera sensor 175 is used in conjunction with a touchscreen display for both video conferencing and still image and / or video image acquisition.
[0073] The device 100 may optionally also include one or more contact strength sensors 165. Figure 1AA contact strength sensor is shown coupled to a strength sensor controller 159 in I / O subsystem 106. The contact strength sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric sensors, optical force sensors, capacitive touch-sensitive surfaces, or other strength sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). The contact strength sensor 165 receives contact strength information (e.g., pressure information or a substitute for pressure information) from the environment. In some embodiments, at least one contact strength sensor is arranged juxtaposed with or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact strength sensor is located on the rear of device 100, opposite to the touchscreen display 112 located on the front of device 100.
[0074] The device 100 optionally also includes one or more proximity sensors 166. Figure 1A A proximity sensor 166 coupled to a peripheral device interface 118 is shown. Alternatively, the proximity sensor 166 may optionally be coupled to an input controller 160 in an I / O subsystem 106. The proximity sensor 166 may optionally be configured as described in the following U.S. patent applications: 11 / 241,839, entitled "Proximity Detector In Handheld Device"; 11 / 240,788, entitled "Proximity Detector In Handheld Device"; 11 / 620,702, entitled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; 11 / 586,862, entitled "Automated Response To And Sensing Of User Activity In Portable Devices"; and 11 / 638,251, entitled "Methods And Systems For Automatic Configuration Of Peripherals", the entire contents of which are incorporated herein by reference. In some implementations, the proximity sensor is turned off and the touchscreen 112 is disabled when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call).
[0075] The device 100 may optionally also include one or more tactile output generators 167. Figure 1AA haptic output generator coupled to a haptic feedback controller 161 in I / O subsystem 106 is shown. The haptic output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices for converting energy into linear motion such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components for converting electrical signals into haptic outputs on the device). A contact intensity sensor 165 receives haptic feedback generation instructions from a haptic feedback module 133 and generates a haptic output on device 100 that can be felt by a user of device 100. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a haptic surface (e.g., haptic display system 112) and optionally generates the haptic output by moving the haptic surface vertically (e.g., in / outward from the surface of device 100) or laterally (e.g., backward and forward in the same plane as the surface of device 100). In some embodiments, at least one haptic output generator sensor is located on the rear of the device 100, opposite to the touch screen display 112 located on the front of the device 100.
[0076] The device 100 may optionally also include one or more accelerometers 168. Figure 1A An accelerometer 168 is shown coupled to a peripheral device interface 118. Alternatively, the accelerometer 168 may be coupled to an input controller 160 in an I / O subsystem 106. The accelerometer 168 may optionally be configured as described in the following U.S. Patent Publications: 20050190059, entitled "Acceleration-based Theft Detection System for Portable Electronic Devices" and 20060017692, entitled "Methods And Apparatuses For Operating A Portable DeviceBased On An Accelerometer," both of which are incorporated herein by reference in their entirety. In some embodiments, information is displayed on a touchscreen display in portrait or landscape view based on analysis of data received from one or more accelerometers. Device 100 may optionally include, in addition to the accelerometer 168, a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for acquiring information about the location and orientation (e.g., portrait or landscape) of device 100.
[0077] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a contact / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and an application (or instruction set) 136. Furthermore, in some embodiments, memory 102 ( Figure 1A ) or 370 ( Figure 3 Storage device / global internal state 157, such as Figure 1A and Figure 3 As shown in the figure. Device / global internal state 157 includes one or more of the following: active application state, which indicates which applications (if any) are currently active; display state, indicating what applications, views or other information occupy various areas of the touch screen display 112; sensor state, including information obtained from various sensors and input control devices 116 of the device; and position information relating to the device's position and / or orientation.
[0078] The operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or embedded operating systems such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware and software components.
[0079] The communication module 128 facilitates communication with other devices via one or more external ports 124 and includes various software components for processing data received by the RF circuitry 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices or indirectly coupled via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is connected to… (Trademark of Apple Inc.) The same or similar and / or compatible multi-pin (e.g., 30-pin) connectors used in Apple Inc. devices.
[0080] The contact / motion module 130 optionally detects contact with the touchscreen 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., touchpads or physical click-based rotary dials). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether a contact has occurred (e.g., detecting a finger press event), determining the contact intensity (e.g., the force or pressure of the contact, or an alternative to force or pressure), determining whether there is movement of the contact and tracking movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or contact disengagement). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations are optionally applied to single-point contact (e.g., single-finger contact) or multi-point simultaneous contact (e.g., "multi-touch" / multiple-finger contact). In some implementations, the contact / motion module 130 and the display controller 156 detect contact on the touchpad.
[0081] In some implementations, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an operation has been performed by the user (e.g., determining whether the user has “clicked” an icon). In some implementations, at least a subset of the intensity thresholds is determined based on software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a specific physical actuator and can be adjusted without changing the physical hardware of device 100). For example, the mouse “click” threshold of a touchpad or touchscreen can be set to any threshold in a wide range of predefined thresholds without changing the touchpad or touchscreen display hardware. Additionally, in some specific implementations, the user of the device is provided with software settings for adjusting one or more intensity thresholds in a set (e.g., by adjusting the individual intensity thresholds and / or by adjusting multiple intensity thresholds at once using system-level clicks on the “intensity” parameter).
[0082] The touch / motion module 130 optionally detects gesture input performed by the user. Different gestures on a touch-sensitive surface have different contact patterns (e.g., different movements, timings, and / or intensities of the detected contact). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a finger tap gesture includes: detecting a finger press event, and then detecting a finger lift-off (lift-away) event at the same (or substantially the same) location as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on a touch-sensitive surface includes: detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift-off (lift-away) event.
[0083] The graphics module 132 includes various known software components for rendering and displaying graphics on the touchscreen 112 or other displays, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual properties). As used herein, the term "graphics" includes any object that can be displayed to a user, including but not limited to text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0084] In some implementations, the graphics module 132 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes from applications, etc., to specify the graphics to be displayed, and also receives coordinate data and other graphic attribute data if necessary, and then generates screen image data to output to the display controller 156.
[0085] The haptic feedback module 133 includes various software components for generating instructions which are used by the haptic output generator 167 to generate haptic output at one or more locations on the device 100 in response to user interaction with the device 100.
[0086] Optionally, the text input module 134, a component of the graphics module 132, provides a soft keyboard for entering text in various applications (e.g., the contact module 137, the email client module 140, the IM module 141, the browser module 147, and any other application that requires text input).
[0087] GPS module 135 determines the location of the device and provides that information for use in various applications (e.g., to telephone module 138 for use in location-based dialing; to camera module 143 as image / video metadata; and to applications that provide location-based services, such as weather widgets, local yellow pages widgets, and map / navigation widgets).
[0088] Application 136 optionally includes the following modules (or instruction sets) or subsets or supersets thereof:
[0089] • Contacts module 137 (sometimes called address book or contact list);
[0090] • Telephone module 138;
[0091] • Video conferencing module 139;
[0092] • Email client module 140;
[0093] • Instant Messaging (IM) module 141;
[0094] Fitness support module 142;
[0095] • Camera module 143 for still images and / or video images;
[0096] • Image management module 144;
[0097] • Video player module;
[0098] Music player module;
[0099] • Browser module 147;
[0100] • Calendar module 148;
[0101] • Widget module 149, which optionally includes one or more of the following: weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, dictionary widget 149-5 and other widgets acquired by the user, as well as user-created widgets 149-6;
[0102] • Widget creator module 150 for creating user-created widgets 149-6;
[0103] • Search module 151;
[0104] • Video and music player module 152, which combines a video player module and a music player module;
[0105] • Notepad module 153;
[0106] • Map module 154; and / or
[0107] • Online video module 155.
[0108] Examples of other applications 136 that may be optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, rendering applications, Java-enabled applications, encryption, digital rights management, speech recognition, and speech duplication.
[0109] In conjunction with the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the contact module 137 is optionally used to manage an address book or contact list (e.g., in application internal state 192 of the contact module 137 stored in memory 102 or memory 370), including: adding one or more names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers or email addresses to initiate and / or facilitate communications via telephone 138, video conferencing module 139, email 140, or IM 141; and so on.
[0110] Combining RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, telephone module 138 is optionally used to input character sequences corresponding to telephone numbers, access one or more telephone numbers in contact module 137, modify input telephone numbers, dial corresponding telephone numbers, initiate conversations, and disconnect or hang up when a conversation is completed. As described above, wireless communication optionally uses any of a variety of communication standards, protocols, and technologies.
[0111] Combining RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contact module 137, and telephone module 138, video conferencing module 139 includes executable instructions to initiate, conduct, and terminate video conferences between the user and one or more other participants based on user instructions.
[0112] Incorporating RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, email client module 140 includes executable instructions for creating, sending, receiving, and managing emails in response to user commands. Combined with image management module 144, email client module 140 makes it very easy to create and send emails containing still images or video images captured by camera module 143.
[0113] In conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, instant messaging module 141 includes executable instructions for performing the following operations: entering a character sequence corresponding to an instant message, modifying previously entered characters, sending a corresponding instant message (e.g., using Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocols for telephone-based instant messaging or using XMPP, SIMPLE, or IMPS for internet-based instant messaging), receiving an instant message, and viewing received instant messages. In some embodiments, the sent and / or received instant messages optionally include graphics, photographs, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Services (EMS). As used herein, "instant message" means both telephone-based messages (e.g., messages delivered using SMS or MMS) and internet-based messages (e.g., messages delivered using XMPP, SIMPLE, or IMPS).
[0114] Incorporating RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module, fitness support module 142 includes executable instructions for performing the following operations: creating fitness activities (e.g., with time, distance, and / or calorie burning goals); communicating with fitness sensors (exercise equipment); receiving fitness sensor data; calibrating sensors used to monitor fitness; selecting and playing music for fitness activities; and displaying, storing, and transmitting fitness data.
[0115] In conjunction with the touchscreen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144, the camera module 143 includes executable instructions for performing the following operations: capturing still images or videos (including video streams) and storing them in memory 102, modifying the characteristics of still images or videos, or deleting still images or videos from memory 102.
[0116] In conjunction with the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and camera module 143, the image management module 144 includes executable instructions for performing operations such as arranging, modifying (e.g., editing) or otherwise manipulating, marking, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or video images.
[0117] Incorporating RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, browser module 147 includes executable instructions for performing the following operations: browsing the Internet according to user instructions, including searching, linking to, receiving, and displaying web pages or portions thereof, as well as links to attachments and other files on web pages.
[0118] Combining RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, email client module 140, and browser module 147, calendar module 148 includes executable instructions to create, display, modify, and store calendars and associated data (e.g., calendar entries, to-dos, etc.) according to user instructions.
[0119] In conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, widget module 149 is optionally a micro-application downloaded and used by a user (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) or a user-created micro-application (e.g., user-created widget 149-6). In some embodiments, the widget includes HTML (Hypertext Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the widget includes XML (Extensible Markup Language) files and JavaScript files (e.g., Yahoo! widgets).
[0120] Incorporating RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, widget creator module 150 is optionally used by the user to create widgets (e.g., turning user-specified portions of a webpage into widgets).
[0121] In conjunction with the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the search module 151 includes executable instructions for performing the following operations: searching the memory 102 for text, music, sound, images, videos, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) according to user instructions.
[0122] Incorporating touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions allowing users to download and play back recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touchscreen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).
[0123] Combining the touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the notepad module 153 includes executable instructions for creating and managing notes, to-do items, etc., according to user instructions.
[0124] Combining RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 is optionally used to receive, display, modify, and store maps and map-related data (e.g., driving directions, data related to shops and other points of interest at or near a specific location, and other location-based data) according to user instructions.
[0125] Incorporating touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, email client module 140, and browser module 147, the online video module 155 includes instructions for performing the following operations: allowing users to access, browse, receive (e.g., via streaming and / or downloading), play back (e.g., on the touchscreen or on an external display connected via external port 124), send emails with links to specific online videos, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, an instant messaging module 141 is used instead of the email client module 140 to send links to specific online videos. Additional descriptions of the online video application can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” the contents of which are incorporated herein by reference in their entirety.
[0126] Each of the modules and applications described above corresponds to an executable set of instructions for performing one or more functions described above and the methods described in this patent application (e.g., computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules; therefore, various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. For example, a video player module may optionally be combined with a music player module into a single module (e.g., Figure 1A (e.g., video and music player module 152). In some embodiments, memory 102 optionally stores a subset of the above-described modules and data structures. Additionally, memory 102 optionally stores additional modules and data structures not described above.
[0127] In some implementations, device 100 is a device on which the operation of a predefined set of functions is performed solely via a touchscreen and / or touchpad. By using a touchscreen and / or touchpad as the primary input control device for operating device 100, the number of physical input control devices (e.g., push-buttons and dial pads, etc.) on device 100 is optionally reduced.
[0128] A predefined set of functions, uniquely performed via a touchscreen and / or touchpad, optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 from any user interface displayed on device 100 to the main menu, main desktop menu, or root menu. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is a physical push-button or other physical input control device, rather than a touchpad.
[0129] Figure 1B This is a block diagram illustrating exemplary components for event handling according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370 ( Figure 3 This includes an event classifier 170 (e.g., in operating system 126) and a corresponding application 136-1 (e.g., any of the aforementioned applications 137 to 151, 155, 380 to 390).
[0130] Event classifier 170 receives event information and determines the application 136-1 and application view 191 of application 136-1 to which the event information should be delivered. Event classifier 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192, which indicates one or more current application views displayed on touch-sensitive display 112 when the application is active or running. In some embodiments, device / global internal state 157 is used by event classifier 170 to determine which application(s) is currently active, and application internal state 192 is used by event classifier 170 to determine the application view 191 to which the event information should be delivered.
[0131] In some implementations, the application internal state 192 includes additional information such as one or more of the following: recovery information to be used when the application 136-1 resumes execution, user interface state information indicating that information is being displayed or ready to be displayed by the application 136-1, a state queue for enabling the user to return to the previous state or view of the application 136-1, and a repeat / undo queue for the user's previous actions.
[0132] Event monitor 171 receives event information from peripheral interface 118. The event information includes information about sub-events (e.g., user touches on touch-sensitive display 112 as part of a multi-touch gesture). Peripheral interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 168, and / or microphone 113 (via audio circuitry 110). The information received by peripheral interface 118 from I / O subsystem 106 includes information from touch-sensitive display 112 or touch-sensitive surfaces.
[0133] In some implementations, event monitor 171 sends requests to peripheral device interface 118 at predetermined intervals. In response, peripheral device interface 118 sends event information. In other implementations, peripheral device interface 118 sends event information only when a significant event occurs (e.g., receiving input above a predetermined noise threshold and / or receiving input for a predetermined duration).
[0134] In some implementations, the event classifier 170 also includes a hit view determination module 172 and / or an activity event recognizer determination module 173.
[0135] When the touch-sensitive display 112 displays more than one view, the hit view determination module 172 provides a software process for determining where a sub-event has occurred within one or more views. A view consists of controls and other elements that the user can see on the display.
[0136] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected optionally corresponds to a procedural level within the application's procedural or view hierarchy. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events identified as correct input is optionally determined at least in part based on the hit view of the initial touch that initiates the touch-based gesture.
[0137] The hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchical structure, the hit view determination module 172 identifies the hit view as the lowest-level view in the hierarchical structure from which the sub-events should be processed. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events forming an event or potential event) occurs. Once the hit view is identified by the hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source to which it was identified as the hit view.
[0138] The activity event recognizer determination module 173 determines which views(s) within the view hierarchy should receive a specific sub-event sequence. In some embodiments, the activity event recognizer determination module 173 determines that only the hit view should receive the specific sub-event sequence. In other embodiments, the activity event recognizer determination module 173 determines that all views including the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive the specific sub-event sequence. In other embodiments, even if the touch sub-event is entirely confined to the area associated with a particular view, higher views in the hierarchy will still remain actively participating views.
[0139] Event assigner module 174 assigns event information to event identifiers (e.g., event identifier 180). In embodiments that include active event identifier determination module 173, event assigner module 174 delivers event information to the event identifier determined by active event identifier determination module 173. In some embodiments, event assigner module 174 stores event information in an event queue, which is retrieved by the corresponding event receiver 182.
[0140] In some implementations, operating system 126 includes event classifier 170. Alternatively, application 136-1 includes event classifier 170. In yet another implementation, event classifier 170 is a standalone module or part of another module (such as contact / motion module 130) stored in memory 102.
[0141] In some implementations, application 136-1 includes a plurality of event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events occurring within a corresponding view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes a plurality of event recognizers 180. In other implementations, one or more of the event recognizers 180 are part of a separate module, such as a user interface toolkit or a higher-level object from which application 136-1 inherits methods and other properties. In some implementations, a corresponding event handler 190 includes one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from an event classifier 170. Event handlers 190 optionally utilize or invoke the data updater 176, the object updater 177, or the GUI updater 178 to update the application's internal state 192. Alternatively, one or more application views in application view 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included in the corresponding application view 191.
[0142] The corresponding event identifier 180 receives event information (e.g., event data 179) from the event classifier 170 and identifies the event based on the event information. The event identifier 180 includes an event receiver 182 and an event comparator 184. In some embodiments, the event identifier 180 also includes at least one subset of metadata 183 and event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0143] Event receiver 182 receives event information from event classifier 170. The event information includes information about sub-events, such as touch or touch movement. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves touch movement, the event information optionally also includes the speed and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a longitudinal orientation to a lateral orientation, or vice versa), and the event information includes corresponding information about the device's current orientation (also referred to as device orientation).
[0144] Event comparator 184 compares event information with predefined event or sub-event definitions and determines the event or sub-event based on the comparison, or determines or updates the state of the event or sub-event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 contains definitions of events (e.g., predefined sequences of sub-events), such as event 1 (187-1), event 2 (187-2), and others. In some embodiments, sub-events in events (e.g., 187-1 and / or 187-2) include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, event 1 (187-1) is defined as a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift-off of a predetermined duration (touch end), a second touch (touch start) of a predetermined duration on the displayed object, and a second lift-off of a predetermined duration (touch end). In another example, event 2 (187-2) is defined as a drag on a displayed object. For example, dragging includes a touch (or contact) of a predetermined duration on the displayed object, movement of the touch on the touch-sensitive display 112, and lifting the touch (end of touch). In some embodiments, the event also includes information for one or more associated event handlers 190.
[0145] In some implementations, event definition 186 includes definitions of events for corresponding user interface objects. In some implementations, event comparator 184 performs a hit test to determine which user interface object is associated with the sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display 112, when a touch is detected on touch-sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects the event handler associated with the sub-event and the object that triggered the hit test.
[0146] In some implementations, the definition of the corresponding event (187) also includes a delay action that delays the delivery of event information until it has been determined whether the sub-event sequence actually corresponds to or does not correspond to the event type of the event recognizer.
[0147] When the corresponding event recognizer 180 determines that the sub-event sequence does not match any event in event definition 186, the corresponding event recognizer 180 enters an event impossible, event failed, or event ended state, after which subsequent sub-events based on touch gestures are ignored. In this case, other event recognizers (if any) that remain active in the hit view continue to track and process the ongoing sub-events based on touch gestures.
[0148] In some embodiments, the corresponding event recognizer 180 includes metadata 183 having configurable attributes, flags, and / or lists instructing how the event delivery system should perform sub-event delivery to actively participating event recognizers. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists instructing how or how event recognizers can interact with each other. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists instructing whether sub-events are delivered to different levels in a view or programmatic hierarchy.
[0149] In some implementations, when one or more specific sub-events of an event are identified, the corresponding event recognizer 180 activates the event handler 190 associated with the event. In some implementations, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from delivering (and deferred delivering) the sub-events to the corresponding hit view. In some implementations, the event recognizer 180 throws a flag associated with the identified event, and the event handler 190 associated with the flag retrieves the flag and performs a predefined process.
[0150] In some implementations, event delivery instruction 188 includes a sub-event delivery instruction that delivers event information about a sub-event without activating an event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with the sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or the actively participating view receives the event information and performs a predetermined process.
[0151] In some implementations, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137 or stores video files used in video player module. In some implementations, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates the positioning of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and transmits that display information to graphics module 132 for display on a touch-sensitive display.
[0152] In some implementations, event handler 190 includes, or has access to, a data updater 176, an object updater 177, and a GUI updater 178. In some implementations, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other implementations, they are included in two or more software modules.
[0153] It should be understood that the above discussion regarding event handling for user touch on a touch-sensitive display also applies to other forms of user input used to operate the multifunction device 100 using an input device, and not all user input is initiated on the touchscreen. For example, mouse movement and mouse button presses optionally in conjunction with single or multiple keyboard presses or holds; touch movements on the touchpad, such as taps, drags, scrolls, etc.; stylus input; device movement; verbal commands; detected eye movements; biometric input; and / or any combination thereof may optionally be used as input corresponding to sub-events that define the event to be identified.
[0154] Figure 2A portable multifunction device 100 with a touchscreen 112 is illustrated according to some embodiments. The touchscreen optionally displays one or more graphics within a user interface (UI) 200. In this embodiment and other embodiments described below, a user can select one or more graphics by gesturing over the graphics, for example, using one or more fingers 202 (not drawn to scale in the figure) or one or more styluses 203 (not drawn to scale in the figure). In some embodiments, selection of one or more graphics occurs when the user breaks contact with one or more graphics. In some embodiments, gestures optionally include one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or scrolling (from right to left, from left to right, up and / or down) of a finger already in contact with the device 100. In some specific embodiments or in some cases, unintentional contact with a graphic does not select the graphic. For example, a swipe gesture over an application icon optionally does not select the corresponding application when the gesture corresponding to selection is a tap.
[0155] Device 100 optionally also includes one or more physical buttons, such as a "main desktop" or menu button 204. As previously described, menu button 204 is optionally used to navigate to any application 136 of a set of applications optionally executed on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on touchscreen 112.
[0156] In some embodiments, device 100 includes a touchscreen 112, a menu button 204, a push-button 206 for powering on / off and locking the device, one or more volume control buttons 208, a SIM card slot 210, a headphone jack 212, and a docking / charging external port 124. The push-button 206 is optionally used to: power on / off the device by pressing the button and holding it in the pressed state for a predefined time interval; lock the device by pressing the button and releasing it before the predefined time interval has elapsed; and / or unlock the device or initiate an unlocking process. In another embodiment, device 100 also accepts voice input via microphone 113 for activating or deactivating certain functions. Device 100 also optionally includes one or more contact strength sensors 165 for detecting the intensity of contact on the touchscreen 112, and / or one or more haptic output generators 167 for generating haptic outputs for a user of device 100.
[0157] Figure 3This is a block diagram of an exemplary multifunctional device with a display and a touch-sensitive surface according to some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a children's learning toy), gaming system, or control device (e.g., a home controller or industrial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. The communication bus 320 optionally includes circuitry (sometimes referred to as a chipset) that interconnects system components and controls communication between system components. Device 300 includes an input / output (I / O) interface 330 with a display 340, which is typically a touchscreen display. The I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, and a haptic output generator 357 for generating haptic output on device 300 (e.g., similar to the reference above). Figure 1A The described tactile output generator 167) and sensor 359 (e.g., optical sensor, accelerometer, proximity sensor, touch sensor and / or contact intensity sensor (similar to the one described above)) Figure 1A The described contact strength sensor 165). Memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 370 optionally includes one or more storage devices located remotely from CPU 310. In some embodiments, memory 370 stores information related to portable multifunction device 100. Figure 1A The memory 370 stores programs, modules, and data structures similar to those in the memory 102 of the portable multifunction device 100, or subsets thereof. Additionally, the memory 370 optionally stores additional programs, modules, and data structures not present in the memory 102 of the portable multifunction device 100. For example, the memory 370 of the device 300 optionally stores a drawing module 380, a rendering module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while the portable multifunction device 100 (… Figure 1A The memory 102 may optionally not store these modules.
[0158] Figure 3Each of the elements described above is optionally stored in one or more memory devices of the previously mentioned memory devices. Each module described above corresponds to a set of instructions for performing the functions described above. The modules or computer programs described above (e.g., instruction sets or including instructions) need not be implemented as separate software programs (such as computer programs (e.g., including instructions)), processes, or modules, and therefore various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 370 optionally stores a subset of the modules and data structures described above. In addition, memory 370 optionally stores additional modules and data structures not described above.
[0159] Now let’s turn our attention to the implementation of the user interface, which is optionally implemented on, for example, a portable multifunction device 100.
[0160] Figure 4A An exemplary user interface for an application menu on a portable multifunction device 100 according to some embodiments is illustrated. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes elements or a subset or superset thereof:
[0161] • Signal strength indicator 402 for wireless communications such as cellular signals and Wi-Fi signals;
[0162] • Time 404;
[0163] Bluetooth indicator 405;
[0164] • Battery status indicator 406;
[0165] • Tray icon 408 with icons for frequently used applications, such as:
[0166] ○ The telephone module 138 has an icon 416 labeled "telephone", which optionally includes an indicator 414 indicating the number of missed calls or voicemail messages;
[0167] ○ An icon 418 labeled "Mail" in the email client module 140, which optionally includes an indicator 410 for the number of unread emails;
[0168] ○ The icon 420 labeled "Browser" in browser module 147; and
[0169] ○ The video and music player module 152 (also known as the iPod (Apple Inc. trademark) module 152) is marked with an icon 422 labeled "iPod"; and
[0170] • Icons of other applications, such as:
[0171] ○The icon 424 of the IM module 141 marked as "Message";
[0172] ○The calendar module 148 has an icon 426 labeled "Calendar";
[0173] ○ The icon 428 of the image management module 144 is labeled "Photo".
[0174] ○ The icon 430 of camera module 143, which is labeled "camera";
[0175] ○ The icon 432 of the online video module 155, which is labeled "Online Video";
[0176] ○ The icon 434 labeled "Stock Market" in the Stock Market widget 149-2;
[0177] ○The icon 436 of the map module 154 that is labeled "map";
[0178] ○The weather widget 149-1 has icon 438 labeled "weather";
[0179] ○ The alarm clock widget 149-4 has an icon 440 labeled "clock";
[0180] ○ The icon 442 of the fitness support module 142 is labeled "fitness support";
[0181] ○ The icon 444 labeled "Notepad" in Notepad module 153; and
[0182] ○ Set an icon 446 labeled "Settings" for an application or module, which provides access to settings for device 100 and its various applications 136.
[0183] It should be pointed out that, Figure 4A The illustrated icon labels are merely exemplary. For example, icon 422 of video and music player module 152 is labeled "Music" or "Music Player". Other labels may be optionally used for various application icons. In some embodiments, the label of a particular application icon includes the name of the application corresponding to that particular application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to that particular application icon.
[0184] Figure 4B An example is illustrated having a touch-sensitive surface 451 (e.g., separate from the display 450 (e.g., touchscreen display 112)). Figure 3 A tablet device or touchpad 355) device (e.g., Figure 3An exemplary user interface on the device 300. The device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of the sensors 359) for detecting the intensity of contact on the tactile surface 451 and / or one or more tactile output generators 357 for generating tactile outputs for the user of the device 300.
[0185] While some examples of inputs on a touchscreen display 112 (which combines a touch-sensitive surface and a display) are given below, in some implementations, the device detects inputs on a touch-sensitive surface separate from the display, such as... Figure 4B As shown in the diagram. In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451) has a spindle (e.g., on the display (e.g., 450) corresponding to the main axis on the display (e.g., Figure 4B The main shaft of 453 in the middle (e.g., Figure 4B (452 in the example). According to these embodiments, the device detects the position corresponding to a specific location on the display (e.g., in the example). Figure 4B In the middle, 460 corresponds to 468 and 462 corresponds to 470) is in contact with the touch-sensitive surface 451 (e.g., Figure 4B (460 and 462 in the text). Thus, when the touch-sensitive surface (e.g., ...) Figure 4B 451) and the display of a multi-functional device (e.g., Figure 4B When 450 is separated from the touch-sensitive surface, user input detected by the device on that touch-sensitive surface (e.g., touches 460 and 462 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods may be optionally used for other user interfaces described herein.
[0186] Additionally, while the examples below are primarily given with reference to finger input (e.g., finger touch, single-finger tap gesture, finger swipe gesture), it should be understood that in some implementations, one or more of these finger inputs may be replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be replaced by a mouse click (e.g., instead of a touch), followed by movement of the cursor along the swipe path (e.g., instead of movement of the touch). As another example, a tap gesture may optionally be replaced by a mouse click while the cursor is over the location of the tap gesture (e.g., instead of detection of touch, followed by cessation of touch detection). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice may optionally be used simultaneously, or mouse and finger touch may optionally be used simultaneously.
[0187] Figure 5AAn exemplary personal electronic device 500 is illustrated. Device 500 includes a body 502. In some embodiments, device 500 may include components relative to devices 100 and 300 (e.g., Figures 1A to 4B The device 500 may include some or all of the features described herein. In some embodiments, the device 500 has a touch-sensitive display 504, referred to below as a touchscreen 504. Alternatively, or in addition to the touchscreen 504, the device 500 may also have a display and a touch-sensitive surface. Similar to the cases of devices 100 and 300, in some embodiments, the touchscreen 504 (or touch-sensitive surface) optionally includes one or more intensity sensors for detecting the intensity of an applied contact (e.g., a touch). The one or more intensity sensors of the touchscreen 504 (or touch-sensitive surface) may provide output data representing the intensity of the touch. The user interface of the device 500 may respond to touches based on the intensity of the touch, meaning that touches of different intensities may invoke different user interface operations on the device 500.
[0188] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related applications: International Patent Application Serial No. PCT / US2013 / 040061, filed May 8, 2013, entitled “Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application,” published as WIPO Publication No. WO / 2013 / 169849; and International Patent Application Serial No. PCT / US2013 / 069483, filed November 11, 2013, entitled “Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships,” published as WIPO Publication No. WO / 2014 / 105276, each of which is incorporated herein by reference in its entirety.
[0189] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508 (if included) may be physical. Examples of physical input mechanisms include push-buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms (if included) allow device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch straps, bangles, trousers, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow a user to wear device 500.
[0190] Figure 5B An exemplary personal electronic device 500 is depicted. In some embodiments, device 500 may include, relative to... Figure 1A , Figure 1B and Figure 3 Some or all of the components described herein. Device 500 has a bus 512 that operatively couples I / O portion 514 to one or more computer processors 516 and memory 518. I / O portion 514 may be connected to display 504, which may have touch-sensitive component 522 and optionally have intensity sensor 524 (e.g., contact intensity sensor). Furthermore, I / O portion 514 may be connected to communication unit 530 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular and / or other wireless communication technologies. Device 500 may include input mechanisms 506 and / or 508. For example, input mechanism 506 may optionally be a rotatable input device or a pressable and rotatable input device. In some examples, input mechanism 508 may optionally be a button.
[0191] In some examples, the input mechanism 508 is optionally a microphone. The personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which are operatively connected to the I / O section 514.
[0192] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause those computer processors to perform, for example, the techniques described below, including processes 700, 800, and 1000. Figure 7 , Figure 8 and Figure 10 A computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media can include, but is not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state storage such as flash memory, solid-state drives, etc. Personal electronic devices are not limited to... Figure 5B It can be the components and configurations, or it can include other components or additional components in a variety of configurations.
[0193] As used herein, the term "power indication" refers optionally to the power indication in devices 100, 300, and / or 500 ( Figure 1A , Figure 3 and Figures 5A to 5B A user-interactive graphical user interface object displayed on a screen. For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) optionally each constitute a functional representation.
[0194] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface with which a user is interacting. In some specific implementations that include a cursor or other positional marker, the cursor acts as a "focus selector," such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), the cursor is positioned on a touch-sensitive surface (e.g., a...). Figure 3 The touchpad 355 or Figure 4B When an input (e.g., a press input) is detected on the touch-sensitive surface 451 of the display, the specific user interface element is adjusted according to the detected input. This applies to touchscreen displays (e.g., those capable of enabling direct interaction with user interface elements on the touchscreen display) Figure 1A The touch-sensitive display system 112 or Figure 4A In some embodiments of the touchscreen 112, a touch detected on the touchscreen acts as a "focus selector," such that when input (e.g., a press input by touch) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, that particular user interface element is adjusted according to the detected input. In some embodiments, focus moves from one area of the user interface to another without corresponding movement of the cursor or movement of a touch on the touchscreen display (e.g., moving focus from one button to another using tab keys or arrow keys); in these embodiments, the focus selector moves according to the movement of focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is typically a user-controlled user interface element (or a touch on the touchscreen display) that delivers the user-expected interaction with the user interface (e.g., by indicating to the device the element of the user interface that the user expects to interact with). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen), the position of the focus selector (e.g., a cursor, touch, or selection box) above the corresponding button will indicate to the user that they expect to activate the corresponding button (rather than other user interface elements shown on the device's display).
[0195] As used in the specification and claims, the term "characteristic strength" of a contact refers to a characteristic of the contact based on one or more strengths of the contact. In some embodiments, the characteristic strength is based on multiple strength samples. The characteristic strength is optionally based on a predefined number of strength samples or a set of strength samples collected over a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after contact is detected, before contact is detected to lift off, before or after contact begins to move, before contact ends, before or after contact strength increases and / or before or after contact strength decreases). The characteristic strength of the contact is optionally based on one or more of the following: the maximum value of the contact strength, the mean value of the contact strength, the average value of the contact strength, the value at the top 10% of the contact strength, the half maximum value of the contact strength, or the 90% maximum value of the contact strength, etc. In some embodiments, the duration of the contact is used when determining the characteristic strength (e.g., when the characteristic strength is the average value of the contact strength over time). In some implementations, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether a user has performed an action. For example, the set of one or more intensity thresholds may optionally include a first intensity threshold and a second intensity threshold. In this example, contact with a characteristic intensity not exceeding the first threshold results in a first action, contact with a characteristic intensity exceeding the first intensity threshold but not exceeding the second intensity threshold results in a second action, and contact with a characteristic intensity exceeding the second threshold results in a third action. In some implementations, a comparison between the characteristic intensity and one or more thresholds is used to determine whether one or more actions should be performed (e.g., whether to perform the corresponding action or abandon performing the corresponding action) rather than to determine whether to perform the first action or the second action.
[0196] Now let’s turn our attention to the implementation of user interfaces (“UIs”) and associated processes on electronic devices such as portable multifunction devices 100, 300 or 500.
[0197] Figures 6A to 6R Exemplary methods for detecting motion input to interact with audio notifications and providing audio feedback for detected motion postures, according to some embodiments, are illustrated. The user interface in these figures is used to illustrate the processes described below, including... Figure 7 and Figure 8 The process in.
[0198] exist Figures 6A to 6RIn this embodiment, device 600 is a pair of wireless earbuds having integrated sensors (e.g., GPS sensor 532, accelerometer 534, orientation sensor 540 (e.g., compass), gyroscope 536, motion sensor 538 described relative to personal electronic device 500) for detecting motion (e.g., head movement of user 602) and integrated speakers (e.g., speaker 111) for outputting audio to user 602. In some embodiments, device 600 is a head-mounted display device or other wearable device (e.g., a pair of earrings or over-ear headphones). In some embodiments, device 600 connects wirelessly (e.g., via Bluetooth) to a mobile computing device (e.g., a smartphone, smartwatch, or laptop computer) and outputs audio content and notifications generated at the mobile computing device. In such embodiments, device 600 may operate as an input device for the mobile computing device (e.g., for motion, audio, and / or touch input). In some embodiments, device 600 includes one or more features of devices 100, 300, and / or 500.
[0199] Figures 6A to 6R It also includes graphs 620a to 620i depicting timelines of audio-related events (e.g., audio output) at various states (e.g., at various points in time). For example, in Figure 6A In the figure, graph 620a illustrates the complete state of audio-related events, while Figure 6B The graph 640b illustrates the initial state of the event.
[0200] exist Figure 6A In the diagram, user 602 wears a device 600 that outputs audio in user 602's ear. The audio output by device 600 is represented by rows of rectangular bars plotted along the time axis on graph 620a. Each bar represents a specific type of audio output event, the duration of each type of audio output event (e.g., the length of the output for each audio output type), and the timing of each type of audio output event (e.g., the start of each audio output relative to the beginning of the graph). For example, graph 620a depicts various audio output events that will be discussed in more detail below: notification tone 606a (e.g., as...). Figure 6B (as discussed); Notices and announcements 608a (e.g., such as...) Figure 6B (as discussed); waiting for the loop tone 612a (e.g., as Figures 6B to 6E (as discussed); waiting loop 614a (e.g., as discussed) Figures 6B to 6F , Figure 6I , Figure 6K , Figures 6L to 6P as well as Figure 6R (as discussed); posture feedback 616a, 618a, 622a and 624a (e.g., as discussed) Figure 6C , Figure 6D , Figure 6G , Figure 6H and Figure 6K (as discussed); and attitude confirmation 626a (e.g., as discussed). Figure 6E and Figure 6I (As discussed). Graph 620a further depicts... Figure 6L The gesture cancellation line will be discussed in more detail in the middle.
[0201] As mentioned above, in contrast to Figure 6A As described above, graph 620a illustrates a timeline showing the complete state of different types of audio-related events (e.g., audio output). The following description... Figures 6B to 6R Examples of similar timelines depicting various stages of development (e.g., at various points in time) are shown in graphs 620b to 620i. Additionally, graphs 620b to 620i depict timelines relative to... Figure 6A The curve 620a describes output events similar to (if not identical to) audio output events.
[0202] Figures 6B to 6E A graph 620b illustrates an exemplary timeline depicting audio output events related to message exchange between user 602 and user 602's friend Kate. Furthermore, Figures 6B to 6E An example is device 626 (e.g., a smartphone communicating with device 600 and / or associated with user 602), which displays an incoming message 628a from Kate on a user interface 632a (e.g., a text messaging interface with the contact "Kate"). In some embodiments, the incoming message 628a is initially received at device 626 (e.g., before the information is sent to device 600 for notification purposes).
[0203] Figure 6B An exemplary, non-limiting example of a timeline for the initial state of an audio output event is provided. For example, Figure 6B The initial phase of the audio output events related to the message exchange between Kate and user 602 is depicted (e.g., from time 0 to 4 seconds), in which Kate inquires about arranging a meeting for a baseball game. Figure 6BAs depicted in graph 620b, device 600 outputs a notification tone 606b (e.g., a musical tone for a new message), as indicated by the 0-second mark on the time axis 604b of graph 620b, followed by a notification announcement 608b for the incoming message 628a. In this example, notification announcement 608b is an announcement that a new message from Kate (e.g., “New message from Kate; want to read?”) is available to be read aloud (e.g., broadcast by a virtual assistant associated with device 600). In some implementations, device 600 outputs notification announcement 606a instead of broadcasting the entire incoming message 628a because the incoming message 628a exceeds a threshold length (e.g., exceeds the maximum number of characters and / or words).
[0204] In addition, Figure 6B At the same time, device 600 outputs a wait loop start tone 612b (e.g., a low-high humming tone) along with a notification prompt tone 606b. This wait loop start tone indicates the time period during which device 600 is monitoring motion input (e.g., detecting movement of user 602's head that device 600 interprets as a request to interact with incoming message 628a and / or notification 606b) (e.g., the full duration of wait loop 614b). Figure 6B (Partially shown in the middle).
[0205] exist Figure 6C At this point, after outputting notification 606b, device 600 detects the start of movement 634a of the user 602's motion posture. In this example, the start of movement 634a is the beginning of a nodding gesture (e.g., the first nodding movement) of the user 602 attempting to respond affirmatively to notification 606b and request device 600 to broadcast the entire message 628a.
[0206] exist Figure 6C At this point, in response to detecting the start of motion 634a, and while the waiting cycle 614b is in progress, device 600 outputs attitude feedback 616b, which provides discrete sounds (e.g., musical notes and / or "ding-dongs") indicating the start and progress of the motion attitude. Furthermore, Figure 6CAn exemplary, non-limiting example of a discrete sound indicating the onset of a motion posture is provided. In this example, the discrete sound in posture feedback 616b has a low volume (e.g., within a volume range of 1 dB to 30 dB, 5 dB to 25 dB, 10 dB to 40 dB, or another reasonable volume range), which corresponds to an initial confidence level that the start of motion 634a is a motion progressing toward the completion of the nodding posture (e.g., having a minimum confidence level of 1 dB, 5 dB, or 10 dB or another reasonable minimum decibel level within a volume range corresponding to the lowest confidence level, and a maximum confidence level of 25 dB, 30 dB, or 40 dB or another reasonable maximum decibel level within a volume range corresponding to the highest confidence level).
[0207] In addition, Figure 6C At this point, device 600 outputs a modified pitch 636b of the waiting loop pitch 612b, which is modified (e.g., modulation frequency, pitch). In addition, device 600 outputs the modified pitch 636b synchronously with the output of attitude feedback 616b (e.g., both at the 5-second mark) to further indicate the beginning of the progress of the motion posture (e.g., the beginning of a nodding posture towards completion of progress).
[0208] exist Figure 6D At this point, device 600 detects continuous movements 638b, 642b, and 644b (e.g., a series of continuous head nodding movements) from user 602. In response to these continuous movements 638b, 642b, and 644b, and while a waiting cycle 614b is in progress, device 600 outputs posture feedback 618b, 622b, and 624b, respectively. This posture feedback consists of a series of discrete sounds with increasing volume levels (e.g., increasing decibels, each volume level higher than the next), indicating an increased confidence level that the detected head movements continue to progress toward a completed nodding posture. Additionally, posture feedback 618b, 622b, and 624b are output synchronously with modified tones 646b, 648b, and 65b, which further modify the waiting cycle tone 612b and indicate that the detected head movements continue to progress toward a completed nodding posture. In some implementations, the discrete sounds of the posture feedback 616b, 618b, 622b, and 624b are high-pitched sounds with increased volume (e.g., treble notes or "ding-dongs") because they correspond to the progression of a nodding posture rather than a head-shaking posture.
[0209] In addition, such as Figure 6D As shown, the posture feedback 644b is the last discrete sound in a series of discrete sounds with increasing volume levels (e.g., decibel levels 25, 30, 40, or another reasonable maximum decibel level within a volume range) and indicates the completion status of the progress of the motion posture (e.g., the highest confidence level corresponding to the detected motion being a nodding posture).
[0210] exist Figure 6E At this point, device 600 determines that the nodding posture is completed within the time period of waiting loop 614b, and in response, outputs posture confirmation 626b (e.g., a musical note or a ding-dong sound), thereby confirming the success of the nodding posture to user 602. In some embodiments, posture confirmation 626b is a high-pitched sound (e.g., a high-pitched note or a "ding-dong sound") because it corresponds to the completion of the nodding posture rather than the head-shaking posture.
[0211] In addition, Figure 6E At this point, in response to the detection of a completed nodding gesture within the threshold time period of waiting loop 614b (e.g., after detecting movements 634b, 638b, 642b, and 644b), device 600 outputs a complete message, represented by message output 654 (e.g., “I’m going to a baseball game with Eric and Jonathan. I have an extra ticket. Would you like to come with us?”).
[0212] Figures 6F to 6I A graph 620c illustrates an exemplary timeline depicting audio output events related to a response provided by user 602 to a broadcast message from Kate.
[0213] exist Figure 6F At this point, device 600 confirms message output 654 (e.g., Figure 6E The output (shown) contains a yes or no question, and in response, a notification 608c is output. In this example, notification 608c is an inquiry asking the user whether they wish to send a message (e.g., via a virtual assistant associated with device 600) in response to message output 654. After outputting notification 608c, device 600 initiates a response period represented by wait loop 614c, during which device 600 monitors motion input.
[0214] Figure 6G and Figure 6H This example illustrates how device 600 detects a head-shaking gesture from user 602 in response to Kate's message, indicating a "no" response. For instance, in... Figure 6G and Figure 6H In this process, device 600 detects a sequence of head movements from user 602 for head-shaking postures (e.g., the start of movement 634c and subsequent movements 638c, 642c, and 644c), similar to... Figure 6C and Figure 6D The image shows a sequence of head movements used for the nodding posture. Furthermore, in... Figure 6G and Figure 6H In the middle, device 600 provides something similar to... Figure 6C and Figure 6DThe described posture feedback includes posture feedback (e.g., 616c, 618c, 622c, and 624c). However, in this example, the posture feedback is a series of low-high discrete sounds (e.g., low-high notes and / or "ding-dongs") with progressively higher volumes (e.g., increasing confidence indicating head-shaking posture), as they correspond to the progression of head-shaking posture rather than head-nodding posture.
[0215] exist Figure 6I At this point, device 600 determines that the head-shaking posture was completed within the waiting loop 614c time period, and in response, outputs a posture confirmation 626c to user 602 confirming that the head-shaking posture was successfully detected. In some embodiments, posture confirmation 626c is a low-high tone sound (e.g., a low-high note or a "ding-dong") because it corresponds to the completion of the head-shaking posture rather than the nodding posture.
[0216] In addition, Figure 6I In response to the detection of a completed head-shaking gesture within a threshold time period of waiting loop 614c (e.g., after detecting movements 634c, 638c, 642c, and 644c), device 600 causes device 626 to send message 656a as a negative response to Kate's invitation. For example, user interface 632b on device 626 displays message 656a as a "sent" text message to Kate stating "No, thank you" (e.g., sent to Kate's smartphone).
[0217] Figures 6J to 6L An exemplary timeline is illustrated in graph 620d, depicting an audio output event related to motion detection after user 602 receives a subsequent message from Kate and device 600 detects motion after a threshold time period for motion detection has ended, without initiating a response to send a message to Kate.
[0218] like Figure 6J As shown, device 626 receives a follow-up message 658a asking user 602, "How's the movie on Thursday?". In response to (e.g., at device 600 and / or device 626) receiving the follow-up message 658a, device 600 outputs a notification tone 606d, followed by a notification announcement 608d similar to notification 608c, providing an inquiry to user 602 whether they wish to send a message in response (e.g., via a virtual assistant associated with device 600). Furthermore, after outputting notification announcement 606d, device 600 initiates a response period represented by a wait loop 614d, during which device 600 monitors motion input. Figures 6J to 6L An exemplary, non-limiting example of a time period for responding to an audio notification is provided. Figures 6J to 6LIn the example, device 600 monitors a total time period of up to 5 seconds in response to notification 608d (e.g., the maximum duration of the wait loop in this example).
[0219] exist Figure 6K At this point, 5 seconds after the start of the waiting loop 614d, device 600 detects the start of movement 634d from user 602. In this example, user 602 does not intend to respond to Kate, so the start of movement 634d represents a coincidental head movement by user 602 (e.g., an unexpected nodding movement during physical activity such as running or jogging), and device 600 determines that this coincidental head movement is the first movement in a series of movements progressing toward a nodding posture. Furthermore, in response to the detection of start of movement 634d, device 600 outputs posture feedback 616d. Similar to posture 616b, such as relative to... Figure 6C The posture feedback 616d is a discrete sound with a low volume, which corresponds to the start of movement 634d being the initial confidence level of the user 602 toward the completion of the nodding posture.
[0220] exist Figure 6L At this point, the time period for responding (e.g., 5 seconds) expires, as indicated by the end of wait loop 614 at the 6-second mark. After wait loop 614d ends, user 602 performs another series of coincidental nodding movements (e.g., accidental nodding). Device 600 detects this series of nodding movements but does not process the movement as a completed nodding posture (e.g., ignores the accidental nodding movement) because the threshold time period of wait loop 614d for monitoring movement input has expired. In other words, because movement is detected outside of wait loop 614d, the confidence that the head movement is progressing toward the completion of user 602's intentional nodding posture does not increase. Therefore, device 600 does not send a response to subsequent message 658a from Kate and outputs posture cancellation 662. Posture cancellation 662 is a sound indicating that the nodding posture has been cancelled (e.g., a note or "ding" different from posture confirmation 626b).
[0221] Figure 6M and Figure 6N Example of curve 620e, and Figure 6O Graph 620f is illustrated. Graphs 620e and 620f depict corresponding exemplary timelines of audio output events related to the notification that user 602 interrupted the message received in the group message chat between users 602, Eric, and Jonathan. Figures 6M to 6ODevice 626 is also illustrated, which displays message 664b from Eric and message 666b from Jonathan on user interface 632b (e.g., a text messaging interface for group chat with contacts “Eric” and “Jonathan”).
[0222] like Figure 6M As shown, device 626 receives two consecutive messages, messages 664a and 666a, from Eric and Jonathan respectively. Furthermore, Figure 6M An illustrative, non-limiting example of a threshold time period for responding to audio notifications is provided. For example, such as... Figure 6M As shown, in response to the consecutive receipt of messages 664b and 666b at device 626 (e.g., within a threshold time period (e.g., 3 seconds, 4 seconds, 5 seconds, or another reasonable threshold time period)), device 600 outputs a notification tone 606e, followed by a notification announcement 608e. In this example, notification announcement 608e begins with an announcement that user 602 has two messages in the message queue to be played consecutively (e.g., "You have two new messages from Eric and Jonathan").
[0223] exist Figure 6N Meanwhile, device 600 continues to output notification announcement 608e, in which the beginning portion of the first message in the queue (e.g., message 664b from Eric) is broadcast. While the output of notification announcement 608e is in progress, device 600 detects a complete sequence of head movements from user 602 for the head-shaking gesture (e.g., the start of movement 634e and subsequent movements 638e, 642e, and 644e), similar to... Figure 6G and Figure 6H The illustrated head movement sequence for the head-shaking posture.
[0224] In addition, Figure 6N At this point, device 600 determines that the head-shaking gesture is completed within the time period of waiting loop 614e (e.g., a time period matching the duration of the output notification 608e), and in response, interrupts the ongoing output notification of notification 608e and jumps to the notification of the next message in the queue (e.g., message 666b from Jonathan).
[0225] Figure 6O Describing as Figure 6NA similar exemplary timeline of audio output events is shown. In this example, notification 608f is an announcement of the beginning portion of a second message in the queue (e.g., message 666a from Jonathan). Similarly, device 600 detects a complete sequence of head movements for a head-shaking gesture, and in response, device 600 interrupts the ongoing output of notification 608f, thereby ending the announcement of messages in the queue.
[0226] Figure 6P A graph 620g illustrates an exemplary timeline depicting the completion of audio output events related to a request to change the operating mode of device 600 (e.g., via a virtual assistant associated with device 600). For example, in response to detecting multiple consecutive head-shaking gestures (such as...) to skip an audio notification... Figure 6M and Figure 6O As shown), device 600 outputs notification 608g. Notification 608g is a prompt to mute the audio notification of a text message.
[0227] In addition, Figure 6P In this process, device 600 detects a completed nodding posture 644g (e.g., a complete head motion sequence of the nodding posture) within a threshold time period (e.g., wait loop 614g), similar to relative to Figures 6C to 6D The descriptions are as follows. However, device 600 also detects audio input 668 from user 602 during the same threshold time period. In this example, audio input 668 is a negative response to a prompt (e.g., "no"), which conflicts with a completed nodding gesture 644g indicating user 602's intention to give a positive response to the prompt (e.g., "yes"). As a result of detecting conflicting inputs during the same threshold time period, device 600 identifies one of the inputs as the expected input of user 602 based on a set of conflict criteria. In this example, device 600 identifies the completed nodding gesture 644g as reflecting the expected input of user 602 because the beginning of the completed nodding gesture 644g (e.g., the first head movement in a complete nodding sequence) was detected before audio input 668. Therefore, device 600 changes the state of an audio notification mode in which future audio notifications for text messages are muted. In some implementations, in response to a completed nodding gesture, device 600 sends a signal to device 626, which causes device 626 to change the "notification notification" setting (e.g., via toggle key 672 on notification notification interface 674) from an "on" position to an "off" position.
[0228] Figure 6QA graph 620h illustrates a complete exemplary timeline depicting audio output events related to motion input used to change media playback while user 602 is listening to music. For example, device 600 outputs song 676, which is the beginning of a currently playing media file (e.g., "Song 1"). In some embodiments, song 676 is a media file played on device 626, where associated audio is sent to device 600. In some embodiments, such as... Figure 6Q As illustrated, the playback status of song 676 (e.g., "Song 1") is displayed on the media playback interface 678a on device 626.
[0229] also, Figure 6Q Indicative, non-limiting examples are provided for time periods used to change the playback state of media files. For example, such as... Figure 6Q As shown, after the beginning of the output of song 1 and before the playback state of song 1 has reached the 5-second mark, device 600 detects the completed head-shaking posture 644h (e.g., the complete head movement sequence of the head-shaking posture) within a threshold time period (e.g., 5 seconds) of waiting loop 614h, similar to relative to Figures 6G to 6H Those described. In response to detecting a completed head-shaking gesture 644h, device 600 interrupts media playback of song 1 and jumps to the second song (e.g., "song 2" as depicted in media playback interface 678b).
[0230] Figure 6R A graph 620i illustrates a complete exemplary timeline depicting audio output events associated with joining an incoming real-time communication (e.g., an incoming call). For example, device 600 outputs a notification 608i that is a request to answer an incoming call from Kate (e.g., “Kate is calling. Answer?”) via a virtual assistant associated with device 600. In some embodiments, the incoming call is received at device 626, and relevant audio associated with the call (e.g., the join request, the ringtone of the incoming call, and / or audio of the real-time communication session after joining) is sent to device 600.
[0231] In addition, Figure 6R In the above implementation, after outputting notification 608i, device 600 detects a completed head nodding gesture 644i within a threshold time period of waiting loop 614i. In response to the detection of the completed head nodding gesture 644i, device 600 joins the call to initiate a real-time communication session between user 602 and Kate. In some implementations, if device 600 detects a negative response (e.g., a completed head shaking gesture) during the threshold time period of waiting loop 644i, the call will be rejected.
[0232] Figure 7This is a flowchart illustrating a method for detecting motion input to interact with an audio notification using one or more audio output devices, according to some embodiments. Method 700 is performed at one or more audio output devices (e.g., 600 (e.g., speakers, headphones, and / or earphones)). Some operations in method 700 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0233] In some embodiments, one or more audio input devices (e.g., 600) are integrated into a computer system. The computer system optionally communicates (e.g., wired communication, wireless communication) with a display generation component and one or more input devices. The display generation component is configured to provide visual output, such as display via a CRT monitor, an LED monitor, or an image projection display. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. One or more input devices are configured to receive input, such as a touch-sensitive surface that receives user input. In some embodiments, one or more input devices are integrated with the computer system. In some embodiments, one or more input devices are separate from the computer system. Therefore, the computer system can transmit data (e.g., image data or video data) via wired or wireless connections to an integrated or external display generation component to visually generate content (e.g., using a display device), and can receive input from one or more input devices via wired or wireless connections.
[0234] As described below, method 700 provides an intuitive way to interact with audio notifications. This method reduces the cognitive load on users when interacting with audio notifications, thereby creating a more efficient human-computer interface. For battery-powered computing devices, enabling users to interact with audio notifications more quickly and efficiently saves power and increases the time interval between battery charges.
[0235] One or more audio output devices (e.g., 600) (e.g., speakers, headphones, and / or earbuds) (in some embodiments, one or more audio output devices communicate with external electronic devices and / or computer systems (e.g., smartphones, smartwatches, tablets, and / or personal computers)) output (702) a first audio notification (e.g., an audio tone, verbal notification, and / or audio notification broadcast by a virtual assistant associated with one or more audio output devices) (in some embodiments, the first audio notification is a first sub-part of an audio notification in progress) (in some embodiments, the first audio notification is generated by an external electronic device and / or computer system and sent to one or more audio output devices for output).
[0236] After the output of a first audio notification (e.g., 608a, 608b, 608c, 608d, 606e, 606f, 606g, 606h and / or 606i) (in some embodiments, after at least the beginning of the output of the first audio notification) (in some embodiments, after the entire output of the first audio notification), detection of motion input occurs based on one or more sensors from one or more audio output devices (e.g., one or more accelerometers, gyroscopes, magnetometers, inertial measurement units, optical sensors and / or sensors capable of detecting one or more audio output devices in space). (704) Detect motion input (e.g., motion input corresponding to one or more head rotations along a lateral axis (e.g., pitch rotation) indicating a nodding posture or one or more head rotations along a vertical axis (e.g., yaw rotation) indicating a head-shaking posture) based on the processing of sensor measurements at one or more audio output devices and / or based on the processing of sensor measurements at companion devices such as smartphones, smartwatches, tablets, wearable computing devices, laptops and / or desktop computers).
[0237] In response to detected motion input and based on a determination that a first set of criteria is met, wherein the first set of criteria includes when the first audio notification is output within a threshold time period (e.g., a waiting loop (e.g., 614a, 614b, 614c, 614d, 614e, 614f, 614g, 614h and / or 614i) and / or a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.3 seconds, 0.4 seconds, 0.5 seconds, 1 second, 2 seconds, 4 seconds, or 8 seconds)) (e.g., from the start, midpoint, or end of the first audio notification). When a motion input is detected within a predetermined time period, and a first criterion is met, one or more audio output devices cause (706) to perform a first operation associated with the first audio notification (in some embodiments, causing the first operation includes sending commands and / or instructions to an external device (e.g., a companion device) to cause that device to perform the operation associated with the first audio notification) (e.g., the first operation is an audio output operation (e.g., message playback) that can be performed via one or more audio output devices based on head posture (e.g., nodding or shaking)) (e.g., as Figure 6E , Figure 6I , Figure 6N , Figure 6O , Figure 6P , Figure 6Q and Figure 6R(As illustrated). In some embodiments, the first set of criteria includes a second criterion that is met when the first audio notification is a controllable audio notification (e.g., a notification associated with an operation that can be performed based on detected head movement posture) (e.g., the second criterion is not met when the first audio notification is a non-controllable audio notification (e.g., a notification not associated with an operation that can be performed based on detected head movement posture)). In some embodiments, based on determining that the first notification is a first type of controllable audio notification corresponding to a type of motion input (e.g., a head-shaking posture) (e.g., a notification announcement with only one associated operation type (e.g., a skip / interrupt notification)), the first set of criteria includes a third criterion that is met when the detected motion input is of the first type (e.g., the third criterion is not met when the detected motion input is of the second type (e.g., a head-nodding posture)) (e.g., as shown). Figure 6N and Figure 6O (As illustrated). Responding to detected motion input, a first operation associated with a first audio notification is executed. This provides the user with greater control over one or more audio output devices by allowing the user to perform the operation associated with the incoming audio notification without using a visually displayed UI and without requiring the user to provide a response to voice commands. Furthermore, executing the first operation based on a first set of criteria, including a first criterion satisfied when motion input is detected within a threshold time period for outputting the first audio notification, facilitates a precise operation window for motion input. This provides additional control over one or more audio output devices by improving the accuracy of correctly associating detected motion input with the operation associated with the first audio notification and reduces false alarms corresponding to user movement occurring outside the threshold time period. Providing additional control options for one or more audio output devices enhances system operability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the system). This, in turn, reduces power consumption and extends the system's battery life by enabling the user to use the system more quickly and efficiently.
[0238] In some implementations, in response to the detection of motion input and based on the determination that a first set of criteria is not met, one or more audio output devices (e.g., 600) abandon the execution of a first operation associated with the first audio notification (e.g., such as...). Figure 6L(As illustrated). In some implementations, in response to the detection of motion input and based on the determination that a first set of criteria is not met, one or more audio output devices output a second audio notification indicating that the first operation has been cancelled (e.g., cancel tone (e.g., 662)). Abandoning the execution of the first operation based on the determination that the first set of criteria is not met (e.g., no motion input was detected within a threshold time period for which the first audio notification was output) provides additional control over one or more audio output devices by further reducing user error. For example, requiring that motion input be detected within a certain time period after the first audio notification reduces false alarms from coincidental movements by the user that are not intended to correspond to predefined motion inputs. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0239] In some implementations, causing the execution of a first operation associated with the first audio notification (in some implementations, the first audio notification (e.g., a prompt for broadcasting a message) has multiple associated operations of different controllable types (e.g., broadcasting a message and de-announcing a message) includes: determining that the detected motion input is a first type of motion input (e.g., a nodding gesture) (e.g., 634b, 638b, 642b, 644b, 634b, 644g and / or 644i) (e.g., as...) Figure 6C , Figure 6D , Figure 6P and Figure 6R As illustrated), causing the execution of a first type of operation associated with the first audio notification (e.g., broadcasting a message); and determining that the detected motion input is a second type of motion input different from the first type of motion input (e.g., head shaking gesture) (e.g., 634c, 638c, 642c, 644c, 634e, 638e, 643e, 644e and / or 644h) (e.g., as shown). Figure 6G , Figure 6H , Figure 6N and Figure 6Q(As illustrated), this causes the execution of a second type of operation (cancel message notification) associated with the first audio notification, different from the first type. By performing a first type of operation based on determining that the detected motion input is a first type of motion input, and performing a second type of operation different from the first type based on determining that the detected motion input is a second type of motion input, the user is provided with a wider range of control over the operations associated with the first audio notification without using a visual UI and without requiring the user to respond to voice commands. This enhances the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the system), which in turn reduces power consumption and extends the system's battery life by enabling the user to use the system more quickly and efficiently.
[0240] In some implementations, the first audio notification includes a controllable cue (e.g., 608b, 608c, 608d, 608e, 608f, and 608g) (e.g., a query that can be responded to via an operation corresponding to a detected head movement posture (e.g., a query asking a user of one or more audio output devices whether they wish to respond to a received message from an external device)); and a first type of motion input (e.g., a nodding gesture) corresponds to an affirmative response to the controllable cue (the nodding gesture causes an "yes" response sent to an external device in response to the cue) (e.g., as...). Figure 6C , Figure 6D , Figure 6P and Figure 6R (As illustrated). Responding to a positive response to a controllable prompt enables the execution of an action, providing the user with control over a wider range of operations associated with the initial audio notification without using a visually displayed UI or requiring the user to respond to voice commands. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0241] In some implementations, the first audio notification includes a controllable cue (e.g., 608b, 608c, 608d, 608e, 608f and / or 608g) (e.g., an inquiry that can be responded to via an operation corresponding to the detected head movement posture (e.g., an inquiry asking a user of one or more audio output devices whether they wish to respond to a received message from an external device)); and a second type of motion input (e.g., a head-shaking gesture) corresponds to a negative response to the controllable cue (the head-shaking gesture causes the execution of a "no" message sent to an external device in response to the controllable cue) (e.g., as...). Figure 6G , Figure 6H , Figure 6N and Figure 6Q (As illustrated). Responding to a positive response to a controllable prompt enables the execution of an action, providing the user with control over a wider range of operations associated with the initial audio notification without using a visually displayed UI or requiring the user to respond to voice commands. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0242] In some implementations, the first audio notification is a first sub-part of a first in-progress audio notification (e.g., 608e) (e.g., output by one or more audio devices) (e.g., the first part of a complete message being played (e.g., 644b)); motion input (e.g., 634e, 638e, 642e, and / or 644e) is detected during the output of the first in-progress audio notification (e.g., while the complete message is still being played); and causing the execution of the first operation includes interrupting (e.g., pausing and / or stopping) the output of the first in-progress audio notification (e.g., stopping the output of the first in-progress audio notification and not outputting a second sub-part of the first in-progress audio notification). In some implementations, the first operation (e.g., as determined by determining that the motion input is non-gestural (e.g., head shaking) is caused to be performed. Figure 6M (As illustrated). This enables the execution of the first operation for interrupting the output of an in-progress audio notification, allowing the user to quickly and efficiently filter incoming notifications by providing a mechanism for interrupting and / or skipping to the next notification based on detected motion input. This increased efficiency enhances device usability by allowing the user to avoid unwanted audio notifications and skip to important ones without having to independently view the visual representation of the incoming audio notification on the companion device's UI display. This improves device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0243] In some implementations, interrupting the output of a first in-progress audio notification (e.g., 608e) includes: stopping the output of the first in-progress audio notification; and outputting a second audio notification (e.g., 608f) (e.g., broadcasting a second message notification in a message notification queue (e.g., 664b and 666b)). This causes the execution of a first operation for interrupting the output of the first in-progress audio notification, wherein the first operation includes allowing the user to quickly and efficiently filter incoming notifications by providing a mechanism for interrupting and / or skipping to the next notification based on motion input. This increased efficiency enhances device usability by allowing the user to avoid unwanted audio notifications and skip to important audio notifications without having to independently view the visual representation of incoming audio notifications on the UI display of the companion device. This improves device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0244] In some embodiments, the first audio notification includes a prompt (e.g., 608g) indicating a change in the control state of a mode associated with one or more audio output devices (e.g., a programming function (e.g., a function for broadcasting / suppressing audio notifications) (e.g., 674)) (e.g., a toggle on / off (e.g., 672)) (e.g., changing the mode of one or more audio devices and / or changing the mode of one or more companion devices such as a smartphone, smartwatch, tablet, wearable computing device, laptop computer, and / or desktop computer) (in some embodiments, the prompt is a first sub-part of a prompt indicating a change in the mode associated with one or more audio output devices); and causing the execution of the first operation includes changing the mode associated with one or more audio output devices from a first mode associated with one or more audio output devices to a second mode associated with one or more audio output devices that is different from the first mode (e.g., as shown in the image). Figure 6P(As illustrated). In some embodiments, causing the execution of a first operation corresponding to changing the mode associated with one or more audio output devices includes: changing the mode associated with one or more audio output devices (e.g., playing / suppressing audio notifications) based on determining that the motion input is a first type of motion input (e.g., head nodding); and abandoning the change of the mode associated with one or more audio output devices (e.g., deactivating the first audio notification) based on determining that the motion input is a second type of motion input (e.g., head shaking). Causing the execution of the first operation includes changing the mode associated with one or more audio output devices from a first mode associated with one or more audio output devices to a second mode associated with one or more audio output devices. This allows a user to change the operating mode of one or more audio devices and / or one or more companion devices via a single motion input without physically touching one or more audio devices or interacting with the UI of a visual display associated with one or more audio devices, thereby increasing device usability. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0245] In some implementations, a first mode associated with one or more audio output devices is a first notification mode, which includes a first set of notification settings (e.g., settings affecting notification grouping, output, and / or suppression); and a second mode associated with one or more audio output devices is a second notification mode, which includes a second set of notification settings different from the first set of notification settings (e.g., one or more characteristics of the notification are different when in the first mode compared to the second mode) (e.g., such as...). Figure 6P(As illustrated). In some implementations, the first mode and / or the second mode are modes that affect the operation of one or more audio output devices and companion devices (e.g., smartphones, smartwatches, and / or computers communicating with one or more audio output devices) (e.g., grouping notifications in the first mode based on a set of grouping criteria). Performing the first operation involves changing the mode associated with one or more audio output devices from the first mode associated with the one or more audio output devices to the second mode associated with the one or more audio output devices. This allows a user to change the operating mode of one or more audio devices and / or one or more companion devices via a single motion input, without having to physically touch one or more audio devices or interact with the UI of a visual display associated with one or more audio devices, thereby increasing device usability. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0246] In some implementations, a first mode associated with one or more audio output devices is a first audio notification mode, which includes a first set of audio notification output settings (e.g., settings affecting how audio notifications are output (e.g., timing of output, suppression of output, grouping of output, type of output notification)) that influence the output of audio notifications via the one or more audio output devices; and a second mode associated with the one or more audio output devices is a second audio notification mode, which includes a second set of audio notification output settings that influence the output of audio notifications via the one or more audio output devices, wherein the second set of audio notification output settings differs from the first set of audio notification settings (e.g., ...). Figure 6P (As illustrated). Performing the first operation involves changing the mode associated with one or more audio output devices from a first mode associated with the one or more audio output devices to a second mode associated with the one or more audio output devices. This allows a user to change the operating mode of one or more audio devices and / or one or more companion devices via a single motion input, without having to physically touch one or more audio devices or interact with the UI of the visual display associated with the one or more audio devices, thereby increasing device usability. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0247] In some implementations, the first audio notification, including a prompt to change the mode associated with one or more audio output devices, is output after multiple previous motion inputs have been detected (e.g., as shown in the image). Figure 6N and Figure 6O (As illustrated), where multiple prior motion inputs satisfy a second set of criteria. In some embodiments, the second set of criteria includes a first criterion satisfied when multiple motion inputs are of the same type (e.g., head-shaking gesture) and / or multiple prior actions associated with multiple prior motion inputs are of the same type (e.g., and are skip actions). In some embodiments, the second set of criteria includes a second criterion satisfied when multiple prior motion inputs are detected consecutively (e.g., there is no intermediate input of a different type between two prior motion inputs of the same type (e.g., head-nodding gesture)). Performing the first action includes changing the mode associated with one or more audio output devices from a first mode associated with one or more audio output devices to a second mode associated with one or more audio output devices. This allows a user to change the operating mode of one or more audio devices and / or one or more companion devices via a single motion input without physically touching one or more audio devices or interacting with the UI of a visual display associated with one or more audio devices, thereby increasing device usability. This enhances the device's operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0248] In some implementations, performing the first operation includes changing the playback state of the media item (e.g., pausing, unpausing, skipping, and / or restarting) (e.g., as...). Figure 6Q(As illustrated). In some embodiments, the media item is a first media item in a queue of one or more media items (e.g., a song playlist), and the execution of the first operation includes stopping the playback of the first media item and initiating the playback of a second media item in the queue of one or more media items (e.g., skipping to the next song). In some embodiments, the first audio notification is the audio output of a first sub-section of the media item being played (e.g., a media playback file (e.g., 676) (e.g., songs and / or movies) stored on one or more audio output devices and / or a media playback file (e.g., songs and / or movies) stored on one or more companion devices such as a smartphone, smartwatch, tablet, wearable computing device, laptop computer, and / or desktop computer) (in some embodiments, motion input is detected while the media item is playing). In some implementations, changing the playback state of a media item includes: changing the playback state of the media item to a first state (e.g., pausing the media item) based on determining that the motion input is a first type of motion input (e.g., a head nodding gesture); and changing the playback state of the media item to a second state (e.g., stopping and / or skipping the media item and initiating playback of a second media item) based on determining that the motion input is a second type of motion input (e.g., a head shaking gesture). This increases usability and provides better control over one or more audio devices while in media playback operation mode. Specifically, users can efficiently change the playback state of a media item (e.g., pausing, unpausing, skipping, or restarting a song) via a single motion input without having to interact independently with the visual UI of the media item associated with a separate device. This enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0249] In some implementations, the first audio notification is associated with an inquiry (in some implementations, a prompt) to join a live communication session (e.g., 608i) (e.g., a ringtone for an incoming call initiated by a user of an external device and / or an inquiry played by a virtual assistant associated with one or more audio output devices, asking the user whether they wish to answer an incoming call, video chat, or other live communication session); and causes performing the first operation to include joining a live communication session (e.g., answering an incoming call, video chat, or other live communication session) (e.g., as...). Figure 6R(As illustrated). This increases usability by enabling the initial action of joining a live communication session and provides better control over one or more audio devices when in a live communication-related operating mode. Specifically, users can efficiently join or decline live communication sessions via a single motion input, without having to interact independently with the visual UI of media items associated with individual devices. This enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0250] In some implementations, the first audio notification includes a prompt (e.g., 608c) (e.g., as...). Figure 6F (Example) (e.g., an inquiry asking a user of one or more audio output devices whether they wish to respond to a message received from an external device (e.g., 626) (e.g., 628a)) to transmit (e.g., send to the external device) a message (e.g., a text message) (in some embodiments, a first audio notification is output based on determining that the received message contains a question that can be answered with a yes or no response) (in some embodiments, the notification is a first sub-part of a continuous notification of the transmission of the message); and causing the performance of the first operation to include sending the message (e.g., 656a) (e.g., as... Figure 6I(As illustrated) (In some embodiments, the first audio notification is integrated into a computer system including one or more companion devices, and performing the first operation includes displaying a visual representation of the sent message (e.g., display of the sent text message) on one or more companion devices (e.g., 626) (e.g., a smartphone and / or a smartwatch). In some embodiments, transmitting the message includes: transmitting a message indicating a positive response (e.g., a "yes" message) based on determining that the motion input is a first type of motion input (e.g., a nodding gesture); and transmitting a message indicating a negative response (e.g., 656a) (e.g., "no") based on determining that the motion input is a second type of motion input (e.g., 634c, 638c, 643c, and 644c) (e.g., a head shaking gesture). (Messages). This increases usability by enabling the initial operation of sending messages and provides better control over one or more audio devices when in an operating mode related to sending messages to external devices. Specifically, users can efficiently send messages to external devices via a single motion input without having to interact independently with a visual UI associated with a messaging platform on a separate device and without being required to provide verbal commands to generate the message to be sent. This enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0251] In some implementations, performing the first operation includes outputting a notification (e.g., 654) corresponding to the received message (e.g., a description of the content of the received message). Figure 6E(As illustrated). In some embodiments, the first audio notification is a first sub-part of an announcement corresponding to the received message (e.g., 608b) (e.g., an announcement indicating that the received message is longer than a threshold length (e.g., a long message) and / or a description of the content of the first part of the message) (in some embodiments, the first audio notification is output based on determining that the received message is longer than a threshold length to be read in its entirety); and causing the first operation to be performed includes outputting a second sub-part (e.g., the remainder) of the announcement (e.g., 654). In some embodiments, the first operation corresponding to outputting an announcement corresponding to the received message is performed: based on determining that the motion input is a first type of motion input (e.g., 634b, 638b, 642b, and 644b) (e.g., a head nodding gesture), causing the announcement corresponding to the received message to be output; and based on determining that the motion input is a second type (e.g., a head shaking gesture), causing the first audio notification to be deactivated and causing the output of the announcement corresponding to the received message to be abandoned. This increases usability by enabling the execution of the first operation (which includes outputting a notification corresponding to the received message) and provides greater control over one or more audio devices when in an operating mode related to the playback of messages received from external devices. For example, users can provide selective motion input to select which messages have been played in full without having to interact independently with the visual UI of the messaging platform associated with a separate device, thereby filtering received messages efficiently. This enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0252] In some implementations, in response to the detection of motion input, one or more audio output devices provide first audio feedback (e.g., a confirmation tone for a posture (e.g., 626b)) indicating that the motion input has been identified (e.g., the motion input is identified as motion input from a predefined (e.g., pre-programmed) set of one or more motion inputs (e.g., head pose) associated with one or more audio output devices). Figure 6E (as illustrated) or the tone release for no gesture (e.g., 626c) (e.g., as shown) Figure 6I(As illustrated). Providing initial audio feedback indicating that motion input has been recognized offers the user improved audio feedback. Providing feedback to the user indicating when motion input is successful informs the user how to effectively generate the recognized motion input, thereby enabling more efficient control of one or more audio output devices. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0253] In some implementations, providing first audio feedback includes: providing first-type audio feedback (e.g., 626b) (e.g., confirmation tone) based on determining that the detected motion input is a third type of motion input (e.g., 634b, 638b, 642b, and 644b) (e.g., head nodding posture); and providing second-type audio feedback (e.g., 626c) (e.g., de-tone) (e.g., de-tone includes a sound with a different pitch (e.g., higher or lower) than the sound included in the confirmation tone) based on determining that the detected motion input is a third type of motion input (e.g., 634b, 638b, 642b, and 644b) (e.g., head shaking posture) based on determining that the detected motion input is a third type of motion input (e.g., 634c, 638c, 642c, and 644c) (e.g., head shaking posture) based on determining that the detected motion input is a third type of motion input (e.g., a different frequency sound wave)). Providing first-type audio feedback based on determining that the motion input is a first type of motion input and providing second-type audio feedback based on determining that the motion input is a second type of motion input provides the user with improved audio feedback. Specifically, providing feedback based on the type of motion input informs the user how to effectively generate the identified type of motion input, thereby enabling more efficient control of one or more audio output devices. This enhances the device's operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0254] In some implementations, the first operation is performed without any voice input from the user from one or more audio output devices (in some implementations, the first operation is not performed based on and / or in response to voice input from the user) (e.g., as...). Figure 6E , Figure 6I , Figure 6N , Figure 6O , Figure 6P , Figure 6Q and Figure 6R(As illustrated). A first operation associated with a first audio notification is executed in response to detected motion input, wherein the first operation is executed even without voice input from the user of one or more audio devices. By allowing the user to perform an operation associated with an incoming audio notification without requiring the user to provide a response voice command, the user gains greater control over one or more audio output devices. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0255] In some implementations, the first set of criteria includes a second criterion that is met when: based on determining that conflicting voice input (e.g., 668) is detected during a second threshold time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.3 seconds, 0.4 seconds, 0.5 seconds, 1 second, 2 seconds, 4 seconds, or 8 seconds) during which a first audio notification is output (in some implementations, the threshold time period and the second threshold time period are the same (e.g., 614g)), the detected motion input is identified as expected input (e.g., correct input) based on a set of conflict resolution criteria (in some implementations, conflicting voice input is identified as unexpected input) (e.g., such as...). Figure 6P(As illustrated). (In some embodiments, when a conflicting voice input is detected during a second threshold time period and the detected motion input is not identified as expected input (e.g., a conflicting voice input is identified as expected voice input), the first set of criteria is not met, and an operation corresponding to the conflicting voice input is performed. In some embodiments, the set of conflict resolution criteria includes a first conflict criterion that is met when motion input is detected before the conflicting voice input.) In some embodiments, the set of conflict resolution criteria is met when no other motion input is detected during the threshold time period (e.g., motion input will always be identified as expected input rather than conflicting voice input as long as no additional motion input is conflicting). The first operation is performed based on the determination that the first set of criteria is met, wherein the first set of criteria includes a second criterion that is met when the detected motion input is identified as expected input based on a set of conflict resolution criteria, based on the determination that a conflicting voice input is detected during a second threshold time period for outputting the first audio notification. This provides additional control over one or more audio output devices by further reducing user error. For example, requiring the detected motion input to be identified as expected input when there is recently detected conflicting voice input reduces false alarms from coincidental user utterances that are not intended to perform actions associated with audio notifications. This enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0256] It should be noted that the above is relative to method 700 (e.g., Figure 7 The details of the process described herein also apply in a similar manner to the methods described below. For example, methods 800 and 1000 optionally include one or more characteristics of the various methods described above with reference to method 700. For example, using the techniques described in methods 800 and 1000, one or more audio output devices may cause the operations described with respect to method 700 to be performed. For example, method 800 may be used to provide feedback on the progress of a motion posture that causes the first operation to be performed according to method 700. As an additional example, sound having an analog spatial arrangement according to method 1000 may be an audio notification in method 700. For the sake of brevity, these details will not be repeated below.
[0257] Figure 8This is a flowchart illustrating a method for providing audio notification for a detected motion posture using one or more audio output devices, according to some embodiments. Method 800 is performed at one or more audio output devices (e.g., 600) (e.g., speakers, headphones, and / or earphones). Some operations in method 800 may be combined, the order of some operations may be changed, and some operations may be omitted.
[0258] In some embodiments, one or more audio input devices (e.g., 600) are integrated into a computer system. The computer system optionally communicates (e.g., wired communication, wireless communication) with a display generation component and one or more input devices. The display generation component is configured to provide visual output, such as display via a CRT monitor, an LED monitor, or an image projection display. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. One or more input devices are configured to receive input, such as a touch-sensitive surface that receives user input. In some embodiments, one or more input devices are integrated with the computer system. In some embodiments, one or more input devices are separate from the computer system. Therefore, the computer system can transmit data (e.g., image data or video data) via wired or wireless connections to an integrated or external display generation component to visually generate content (e.g., using a display device), and can receive input from one or more input devices via wired or wireless connections.
[0259] As described below, Method 800 provides an intuitive way to interact with audio feedback. This method reduces the cognitive load on the user when interacting with audio feedback, thereby creating a more efficient human-computer interface. For battery-powered computing devices, enabling users to interact with audio feedback more quickly and efficiently saves power and increases the time interval between battery charges.
[0260] One or more audio output devices (e.g., speakers, headphones, and / or earbuds) (in some embodiments, one or more audio output devices communicate with external electronic devices and / or computer systems (e.g., smartphones, smartwatches, tablets, and / or personal computers)) detect (802) one or more sensor measurements corresponding to the start of a motion posture (e.g., 634b and / or 634c) (e.g., based on processing of sensor measurements at one or more audio output devices and / or based on processing of sensor measurements at companion devices such as smartphones, smartwatches, tablets, wearable computing devices, laptops, and / or desktop computers) corresponding to the start of a motion posture (e.g., 634b and / or 634c). These measurements are obtained via one or more sensors (e.g., one or more accelerometers, gyroscopes, magnetometers, inertial measurement units, optical sensors, and / or sensors capable of detecting motion). Other sensors that detect the movement of one or more audio output devices in space (e.g., where the complete motion posture requires detecting multiple sub-parts of a user's action (e.g., head rotation) in the computer system (e.g., initial sub-parts (e.g., 634b and / or 634c), intermediate sub-parts (e.g., 638b and 642b, and / or 638c and 642c) and an ending sub-part (e.g., 644b and / or 644c)), wherein for each sequential sub-part of the detected motion, the computer system has an increasing confidence level (e.g., as indicated by the predefined motion posture) of the detected motion corresponding to a predefined motion posture (e.g., head motion posture (e.g., one or more head rotations along a lateral axis indicating a nodding posture (e.g., pitch rotation) or one or more head rotations along a vertical axis indicating a head-shaking posture (e.g., yaw rotation))). Figure 6C , Figure 6D , Figure 6G and Figure 6H (As illustrated) (In some embodiments, the start of detecting a motion posture corresponds to the initial sub-part of detecting motion (e.g., the start of head rotation), wherein the computer system has an initial threshold confidence level for the detected motion to correspond to a predefined motion posture. (In some embodiments, one or more sensors detect motion that does not meet the initial threshold confidence level for the detected motion to correspond to a predefined motion posture (e.g., slight head movement).)
[0261] After detecting one or more sensor measurements corresponding to the start of a motion posture, and while the detection of one or more sensor measurements (e.g., by one or more audio output devices and / or by companion devices such as smartphones, smartwatches, tablets, wearable computing devices, laptops, and / or desktop computers) is in progress (e.g., detecting intermediate sub-parts of the motion), one or more audio output devices provide (804) (e.g., output) first audio feedback (e.g., 616b, 618b, 622b, 624b, 616c, 618c, 622c, and / or 624c) indicating the progress of the motion posture via one or more audio output devices (e.g., as shown in the image). Figure 6C , Figure 6D , Figure 6G and Figure 6H (Example) (e.g., the state of progress toward completion) (e.g., audio feedback corresponding to the relative confidence level of the detected motion corresponding to a predefined motion posture).
[0262] After providing the first audio feedback and based on the determined (e.g., via one or more audio output devices and / or via an accessory such as a smartphone, smartwatch, tablet, wearable computing device, laptop, and / or desktop computer) motion posture (e.g., the detected motion corresponds to a confirmed determination of the motion posture) (e.g., detecting the end portion of the motion), one or more audio output devices cause (806) to perform an operation associated with the motion posture (e.g., via one or more audio output devices and / or via an accessory such as a smartphone, smartwatch, tablet, wearable computing device, laptop, and / or desktop computer) (e.g., interrupting the audio notification) (e.g., transmitting a response to a message (e.g., a yes or no text response)) (e.g., as Figure 6E , Figure 6I , Figure 6N , Figure 6O , Figure 6P , Figure 6Q and Figure 6R(As illustrated). Providing initial audio feedback indicating the progress of a motion posture while the detection of measurements from one or more sensors is in progress provides the user with improved feedback on the real-time status of the motion posture toward completion. Providing real-time feedback corresponding to the progress status of the motion posture allows the user to effectively assess the amount of additional motion required to successfully execute the expected motion input. Providing the user with improved feedback further enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently. Based on determining that the motion posture is complete, the execution of a first action associated with the motion posture is performed, providing the user with greater control over one or more audio output devices by allowing the user to perform actions associated with the motion posture without using a visually displayed UI and without requiring the user to respond to voice commands. Providing additional control options for one or more audio output devices enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0263] In some implementations, after providing the first audio feedback (e.g., 616d) (e.g., as...) Figure 6K (as illustrated) and based on the determination that the motion posture is not completed (e.g., not all desired sub-parts of the motion (initial sub-part, intermediate sub-part, and / or final sub-part) are detected within a threshold time period), one or more audio output devices abandon the execution of operations associated with the motion posture (e.g., such as...). Figure 6L (As illustrated). Abandoning the execution of an operation based on the determination that a motion posture is incomplete provides additional control over one or more audio output devices by further reducing user error. For example, requiring the detection of multiple sub-parts of motion corresponding to a motion posture within a threshold time period reduces false alarms from coincidental movements by the user that are not intended to correspond to the motion posture. This enhances device operability and makes the user-device interface more efficient (e.g., by assisting the user in providing appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0264] In some implementations, after providing initial audio feedback and based on the determination that the motion posture is not complete, one or more audio output devices provide a first audio output indicating that the operation associated with the motion posture has been cancelled (e.g., 662) (e.g., cancellation tone and / or cancellation notification) (e.g., an indication that a threshold time period (e.g., 614d) for detecting all desired sub-parts of the motion associated with the motion posture has ended (e.g., subsequent sub-parts of the motion will no longer advance the motion posture)) (e.g., as... Figure 6L (As illustrated). Providing a first audio output indicating that an operation has been cancelled provides the user with improved audio feedback. Specifically, the feedback provides a real-time indication that a threshold time period for detecting motion posture has ended, allowing the user to operate one or more audio output devices (e.g., using a second motion input) without unintentionally performing actions associated with the motion posture. Furthermore, providing the user with feedback indicating when a motion input is unsuccessful informs how to effectively generate future motion inputs, thereby enabling more efficient control of one or more audio output devices. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0265] In some implementations, after providing initial audio feedback and upon determining the completion of the motion posture, one or more audio output devices provide audio output indicating successful completion of the motion posture (e.g., 626b and / or 626c) (e.g., acknowledgment tone and / or acknowledgment notification) (e.g., such as...). Figure 6E and Figure 6I (As illustrated). Providing a first audio output indicating successful completion of a motion posture offers the user improved audio feedback. Specifically, the feedback provides a real-time indication that the necessary sub-part of the motion required to complete the motion posture has been detected, allowing the user to stop the motion associated with the posture. This reduces the likelihood of the user accidentally triggering a second motion posture in an attempt to complete the posture. Furthermore, providing the user with feedback indicating when a motion input has been successfully completed informs how to effectively generate future motion inputs, enabling more efficient control of one or more audio output devices. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0266] In some implementations, one or more sensor measurements are detected via one or more sensors (e.g., one or more accelerometers, gyroscopes, magnetometers, inertial measurement units, optical sensors, and / or other sensors capable of detecting the movement of one or more audio output devices in space) of one or more audio output devices (e.g., based on processing of the sensor measurements at one or more audio output devices); and one or more audio output devices are included in one or more wearable devices (e.g., wearable headphones, earbuds, and / or head-mounted displays with integrated audio) (e.g., such as...). Figures 6A to 6R (As illustrated). First audio feedback indicating the progress of a motion posture is provided, wherein one or more sensor measurements are detected at one or more audio output devices, which are part of the wearable device. This provides the user with greater control over one or more audio devices by allowing the user to perform motion posture-related actions without using touch-based input on the one or more audio output devices. This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0267] In some implementations, one or more wearable devices are a set of one or more earbuds or headphones (e.g., such as...). Figures 6A to 6R (As illustrated). First audio feedback indicating the progress of a motion posture is provided, wherein one or more sensor measurements are detected at the same one or more audio output devices, and wherein one or more audio output devices are a set of one or more earbuds or headphones. This provides the user with greater control over one or more audio devices by allowing hands-free operation via head posture (e.g., using touch-based input and not requiring the motion posture to correspond to an arm or hand posture). This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0268] In some implementations, providing first audio feedback indicating the progress of a motion posture includes outputting multiple discrete sounds (e.g., 616b, 618b, 622b, 624b, 616c, 618c, 622c and / or 624c) (e.g., as shown in the image). Figure 6C , Figure 6D , Figure 6G and Figure 6H (As illustrated) (e.g., each of a plurality of sounds includes one or more tones that end after a corresponding amount of time and without further input from the user) (in some embodiments, discrete sounds are distinct sounds (e.g., sounds with different tones, pitches, and / or volumes)). Providing first audio feedback indicating the progress of a motion posture, wherein providing the first audio feedback includes outputting a plurality of discrete sounds, improves the user feedback for one or more audio devices. Providing real-time feedback corresponding to the progress state of the motion posture allows the user to assess the amount of additional motion required to successfully execute the expected motion input. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0269] In some embodiments, the progression of the motion posture includes a first intermediate sub-part of the detected motion posture (e.g., a first intermediate sub-part of the motion required for the complete motion posture (e.g., 638b and / or 638c)) (in some embodiments, the progression of the motion posture includes the start of the detected motion posture, and the first intermediate sub-part of the motion includes an initial sub-part of the motion required for the complete motion posture (e.g., 634b and / or 634c)) and a second intermediate sub-part of the detected motion posture (e.g., 642b and / or 642c); multiple discrete sounds include responses to the detected motion posture. The first discrete sound output by the first intermediate sub-part (e.g., 618b and 618c, and / or 616b and 616c) (e.g., one or more tones with a first pitch / volume that end after a corresponding time amount and without further input from the user); and the plurality of discrete sounds include a second discrete sound output by the second intermediate sub-part in response to the detected motion posture (e.g., 622b and 622c, and / or 618b and 618c) (e.g., one or more tones with a second pitch / volume that end after a corresponding time amount and without further input from the user) (e.g., as... Figure 6C , Figure 6D , Figure 6G and Figure 6H(As illustrated). In some embodiments, each discrete sound is output based on determining that the corresponding sensor measurement in one or more sensor measurements satisfies a corresponding threshold confidence level for the motion posture to progress toward completion. Providing first audio feedback includes outputting a plurality of discrete sounds, wherein the plurality of discrete sounds includes a first discrete sound output in response to a first intermediate sub-part of the detected motion posture; and the plurality of discrete sounds includes a second discrete sound output in response to a second intermediate sub-part of the detected motion posture, which improves feedback to the user of one or more audio devices. Providing real-time feedback corresponding to the progress state of the motion posture allows the user to assess the amount of additional motion required to successfully execute the expected motion input. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0270] In some embodiments, a first discrete sound indicates the beginning of the progress of the motion posture (in some embodiments, the first discrete sound corresponds to a first sensor measurement result among one or more sensor measurements that meets a first threshold confidence level toward completion of the motion posture); and a second discrete sound indicates the continuation of the progress of the motion posture (in some embodiments, the second discrete sound corresponds to a second sensor measurement result among one or more sensor measurements that meets a second threshold confidence level higher than the first threshold confidence level) (e.g., as...). Figure 6C , Figure 6D , Figure 6G and Figure 6H (As illustrated). Providing first audio feedback includes outputting a plurality of discrete sounds, wherein the plurality of discrete sounds includes a first discrete sound output in response to a first intermediate sub-part of the detected motion posture; and the plurality of discrete sounds includes a second discrete sound output in response to a second intermediate sub-part of the detected motion posture, which improves the feedback to a user of one or more audio devices. Providing real-time feedback corresponding to the progress state of the motion posture allows the user to assess the amount of additional motion required to successfully execute the expected motion input. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0271] In some embodiments, the progress of the motion posture includes the final sub-part of the detected motion posture (e.g., 644b and / or 644c); multiple discrete sounds include a third discrete sound (e.g., 624b and / or 624c) output in response to the final sub-part of the detected motion (e.g., one or more tones with a third pitch / volume that end after a corresponding amount of time and without further input from the user); the third discrete sound indicates the completion status of the progress of the motion posture (in some embodiments, the third discrete sound corresponds to a third sensor measurement result in one or more sensor measurements that satisfies a third threshold confidence level above a second threshold confidence level for the motion posture to progress toward completion) (e.g., such as...). Figure 6D and Figure 6H (As illustrated). Providing first audio feedback includes outputting a plurality of discrete sounds, wherein the plurality of discrete sounds includes a first discrete sound output in response to a first intermediate sub-part of the detected motion posture; and the plurality of discrete sounds includes a second discrete sound output in response to a second intermediate sub-part of the detected motion posture, which improves the feedback to a user of one or more audio devices. Providing real-time feedback corresponding to the progress state of the motion posture allows the user to assess the amount of additional motion required to successfully execute the expected motion input. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the battery life of the device by enabling the user to use the device more quickly and efficiently.
[0272] In some embodiments, a first intermediate portion of the motion posture is detected before a second intermediate portion of the motion posture is detected; the first intermediate portion of the motion posture corresponds to (e.g., is determined and / or identified as corresponding to) a first confidence level of the motion posture toward completion progress (in some embodiments, the confidence level of the motion posture toward completion progress increases as a larger portion of the completed posture is detected); the second intermediate portion of the motion posture corresponds to (e.g., is determined and / or identified as corresponding to) a second confidence level of the motion posture toward completion progress that is higher than the first confidence level; and the first discrete sound The first discrete sound has a first value of an audio characteristic within a range of audio characteristic values (e.g., pitch, volume, and / or volume), which corresponds to a first confidence level of progress toward completion of the motion posture (e.g., a volume range from 1 dB to 30 dB, where 1 corresponds to the lowest confidence level and 30 corresponds to the highest confidence level); and the second discrete sound has a second value of an audio characteristic that is further away from the first value of the audio characteristic within the range of audio characteristic values, and the second discrete sound corresponds to a second confidence level of progress toward completion of the motion posture (e.g., the first value is 5 dB and the second value is 10 dB) (e.g., as...). Figure 6C , Figure 6D , Figure 6G and Figure 6H (As illustrated). Providing audio feedback, including discrete sounds with audio characteristics that progress as the confidence level increases towards completion of the motion posture, improves the feedback provided to users of one or more audio devices. This enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0273] In some embodiments, a first intermediate portion of the motion posture is detected before a second intermediate portion of the motion posture is detected; the first intermediate portion of the motion posture corresponds to (e.g., is determined and / or identified as corresponding to) a first confidence level of progress toward completion of the motion posture (in some embodiments, the confidence level of progress toward completion of the motion posture increases as a larger portion of the completed posture is detected); the second intermediate portion of the motion posture corresponds to (e.g., is determined and / or identified as corresponding to) the first confidence level of progress toward completion of the motion posture (e.g., the second intermediate portion does not indicate progress toward completion of the motion posture). A higher confidence level (e.g., a second confidence level); and a first discrete sound having a first value of an audio characteristic within a range of audio characteristics (e.g., pitch, volume, and / or volume) corresponding to a first confidence level of progress toward completion of the motion posture (e.g., a volume range from 1 dB to 30 dB, where 1 corresponds to the lowest confidence level and 30 corresponds to the highest confidence level); and a second discrete sound having a first value of an audio characteristic, and the second discrete sound corresponding to the first confidence level of progress toward completion of the motion posture (e.g., both the first and second discrete sounds output at 5 dB). Providing audio feedback including discrete sounds with audio characteristics that do not progress with the same confidence level toward completion of the motion posture improves the feedback provided to a user of one or more audio devices and can signal to the user that the motion posture should be modified / progressed to complete the motion posture. This enhances the device's operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0274] In some implementations, providing first audio feedback indicating the progress of a movement posture includes: determining that the movement posture is a first type of movement posture (e.g., a "yes" posture (e.g., a nodding posture)). Figure 6C and Figure 6D As illustrated, it provides a first type of audio feedback (e.g., a first pitch / volume sequence corresponding to the progression of a nodding posture); and based on determining that the movement posture is a second type of movement posture different from the first type (e.g., a "no" posture (e.g., a head-shaking posture)) (e.g., as shown in the example), it provides a first type of audio feedback of a first type (e.g., a first pitch / volume sequence corresponding to the progression of a nodding posture); and based on determining that the movement posture is a second type of movement posture different from the first type (e.g., a "no" posture (e.g., a head-shaking posture)). Figure 6G and Figure 6H(As illustrated), a second type of first audio feedback (e.g., a second tone sequence corresponding to the pitch / volume of the head-shaking posture during the progression) is provided. In some embodiments, after providing the first audio feedback and based on the determined motion posture: based on the determination that the motion posture is a first type of motion posture (e.g., a nodding posture), a first type of second audio feedback (e.g., one or more affirmative tones) is provided. In some embodiments, after providing the first audio feedback and based on the determined motion posture: based on the determination that the motion posture is a second type of motion posture (e.g., a head-shaking posture), a first type of second audio feedback (e.g., one or more release tones) is provided. Providing first type of first audio feedback based on the determination that the motion posture is a first type of motion input and providing second type of first audio feedback based on the determination that the posture input is a second type provides the user with improved audio feedback. Specifically, providing feedback based on the type of motion posture informs the user how to effectively generate the identified type of motion posture, thereby enabling more efficient control of one or more audio output devices. This enhances the device's operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0275] In some implementations, before detecting one or more sensor measurements corresponding to the start of a motion posture, one or more audio output devices provide a first portion of the in-process audio effect (e.g., such as...). Figure 6C (Examples) (e.g., continuing and / or repeating a waiting loop tone until a corresponding condition is met (e.g., 612b) (e.g., a condition met when a predetermined time period associated with the waiting loop (e.g., 614b) expires, a condition met when a completed motion posture is detected before the waiting loop expires, a condition met when voice input is detected before the waiting loop expires, a condition met when touch input is detected on one or more audio output devices before the waiting loop expires, and / or a condition met when input is detected on a companion device before the waiting loop expires).
[0276] In some implementations, providing the first audio feedback includes providing a second part of the in-process audio effect (e.g., by modifying one or more audio characteristics of the in-process audio effect (e.g., 636b, 646b, 648b and / or 652b) (e.g., pitch, volume, and / or tone)). Figure 6C and Figure 6D(As illustrated). Based on one or more modifications to the audio effect caused by the first audio feedback, improved audio feedback is provided to the user of one or more audio devices by outputting feedback representing the correlation between the audio effect (e.g., a waiting loop tone) and the progression of head movement posture. Providing feedback representing this correlation helps the user understand the appropriate response of the progressing head movement posture to the controllable cues associated with the waiting loop tone. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0277] In some embodiments, one or more audio output devices communicate with an audio input device (e.g., an integrated and / or connected microphone); and the in-process audio effect indicates a time period (the duration of a waiting loop) during which the one or more audio output devices listen to (e.g., based on recording and / or processing one or more user utterances at one or more audio output devices, and / or based on recording and / or processing one or more user utterances at an accessory device such as a smartphone, smartwatch, tablet, wearable computing device, laptop, and / or desktop computer) one or more audio inputs (e.g., verbal commands to a virtual assistant on one or more output devices and / or accessory devices). In some embodiments, the audio effect indicates a time period during which the audio output devices detect one or more sensor measurements corresponding to motion posture (e.g., the same time period during which one or more audio output devices listen to one or more audio inputs) (e.g., as shown in the image). Figure 6C and Figure 6D (As illustrated). A first part of the audio effect is provided before the detection of one or more sensor measurements. This audio effect indicates the time period during which the audio output device listens to one or more audio inputs. This further improves the feedback provided to the user of the one or more audio output devices by signaling to the user that motion posture is being detected and / or that progress is being made during the specified time period during which the audio output device listens to a particular input. This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0278] In some implementations, providing the first audio feedback includes outputting one or more sounds, wherein a corresponding sound from the one or more sounds is output based on determining that a corresponding sensor measurement among one or more sensor measurements satisfies a corresponding threshold confidence level for the completion of a motion posture; and one or more modifications to the audio effect (e.g., 636b, 646b, 648b, and / or 652b) correspond to the output of the one or more sounds (e.g., synchronized with it), wherein a corresponding modification among the one or more modifications to the audio effect indicates a corresponding state of the progress of the motion posture (e.g., the detected motion corresponds to a predefined relative confidence level of the motion posture). The one or more modifications to the audio effect caused by the first audio feedback, wherein different modifications among the one or more modifications to the audio effect indicate a corresponding state of the progress of the motion posture, improve the user's feedback. Providing real-time feedback corresponding to the progress state of the motion posture allows the user to effectively assess the amount of additional motion required to successfully execute the expected motion input, and simultaneously provides feedback indicating the correlation between the audio effect (e.g., a waiting loop tone) and the progress of the head motion posture. This enhances the device's operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0279] In some implementations, causing the execution of operations associated with a motion posture includes: determining that the motion posture is a first type of motion posture (e.g., a nodding posture), causing the execution of a first type of operation (e.g., such as...). Figure 6C , Figure 6D , Figure 6P and Figure 6R (as illustrated above with respect to method 700) (e.g., transmitting a "yes" message); and based on determining that the motion posture is a second type of motion posture (e.g., head-shaking posture) different from the first type of operation, performing a second type of operation (e.g., as shown in the example) is a second type of motion posture (e.g., head-shaking posture). Figure 6G , Figure 6H , Figure 6N and Figure 6Q(As illustrated above regarding method 700) (e.g., transmitting a "no" message or abandoning the transmission message). Based on determining that the motion input is a second type of motion input different from the first type, the first type of operation, which performs an operation different from the second type, provides the user with control over a wider range of operations associated with one or more audio devices without using a visually displayed UI and without requiring the user to provide a response to voice commands. This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0280] It should be noted that the above is relative to method 800 (e.g., Figure 8 The details of the process described herein also apply in a similar manner to the methods described below / above. For example, methods 700 and 1000 optionally include one or more characteristics of the various methods described above with reference to method 800. For example, using the techniques described in methods 800 and 1000, one or more audio output devices may cause the operations described with respect to method 700 to be performed. For example, method 800 may be used to provide feedback on the progress of a motion posture that causes the first operation to be performed according to method 700. As an additional example, method 800 may be used to provide feedback on the progress of a motion posture that causes the operation on optional options to be performed in a simulated spatial arrangement according to method 1000. For the sake of brevity, these details will not be repeated below.
[0281] Figures 9A to 9N Exemplary methods for detecting motion input in a spatial audio arrangement, according to some embodiments, are illustrated. The user interface in these figures is used to illustrate the process described below, including... Figure 10 The process in.
[0282] Generally, the various techniques described below for providing audio relate to spatial audio (e.g., binaural audio). In some embodiments, spatial audio is manipulated in two audio channels (e.g., left and right) of the headphones to resemble audio of directional sound reaching the ear canal. For example, the headphones may reproduce spatial audio signals simulating the spatial location around a listener (e.g., user 602), which differs from the location of physical speakers on the headphones and is optionally adjusted based on head movement. Effective spatial location simulation renders spatial locations that appear fixed in space (e.g., the listener perceives sound as coming from a fixed location), even if the audio output components themselves move in space (e.g., when the listener's head moves).
[0283] Generally speaking, Figures 9A to 9N Various spatial audio arrangements (e.g., generated via spatial audio experience) are illustrated, in which device 600 outputs simulated sound in the spatial region of the various spatial audio arrangements and detects motion input from user 602 when interacting with the various spatial audio arrangements.
[0284] Figures 9A to 9H An example of a spatial audio arrangement 900 for interacting with a menu of telephone contacts that user 602 can selectively call. Specifically, Figures 9A to 9I The device 600 is described as detecting head movements of user 602 to contact different telephone contact options in a selected spatial area of spatial audio arrangement 900, and outputting spatialized sound options for each option that user 602 interacts with.
[0285] exist Figure 9A At this point, device 600 detects head posture 902a, which is a double head tilt to the left (e.g., a rolling rotation along a longitudinal axis (e.g., an axis pointing in the forward direction of the user 602's face)). Figure 9A Multiple views depict user 602 performing a dual head tilting pose. Figure 9A The top left illustration depicts the front view of user 602, the bottom left illustration shows the rear view of user 602, and the bottom right illustration shows the isometric view of user 602.
[0286] In response to the detection of head pose 902a, and because head pose 902a is a specific type of pose (e.g., a double head tilt to the left rather than the right), device 600 invokes (e.g., via spatial audio experience) spatial audio arrangement 900. Figure 9A The upper right illustration depicts a spatial audio arrangement 900, showing a top view of a user facing forward when the spatial audio arrangement 900 is activated. In some examples, head posture 902a is a double head tilt along a longitudinal axis positioned relative to the body posture of user 602. For example, device 600 records a double head tilt rotation from user 602 as head posture 902a, regardless of whether user 602 is lying down or standing.
[0287] In addition, such as Figure 9AAs illustrated, the spatial audio arrangement 900 includes spatial regions 904a (e.g., to the left of user 602), spatial regions 906a (e.g., in front of the user), and spatial regions 908a (e.g., to the right of the user), which correspond to three alternative options (e.g., 916a, 918a, and 922a) for calling (e.g., initiating a real-time communication session via device 626) contacts “Kate,” “Jonathan,” and “Eric,” respectively. The spatial regions are positioned separately from user 602 such that analog sounds from spatial regions 904a, 906a, and 908a are perceived by user 602 as originating from those corresponding locations, rather than from the location of physical speakers on device 602. Furthermore, the analog sounds generated in spatial regions 904a, 906a, and 908a are fixed relative to the initial forward position of user 602's head, such that even when user 602's head shifts and rotates, user 602 will perceive the sounds in spatial regions 904a, 906a, and 908a as originating from a constant location. For example, as user 602's head from Figure 9A Rotating the initial position depicted in the image 90 degrees to the left, user 602 perceives the simulated sound in spatial region 904a as originating from a position directly in front of user 602's face.
[0288] In some implementation schemes, such as Figure 9B and Figure 9C As illustrated, the spatial regions 904a, 906a and 908a of the spatial audio arrangement 900 can be arranged in various ways such that they occupy different positions relative to the forward position of the user 602 and / or they have different relative sizes and shapes.
[0289] exist Figure 9D In this example, device 600 detects a first head movement 912a, which is a head rotation to the left (e.g., head tilting or head turning) toward spatial region 904a. In response to the detection of head movement 912a, device 600 outputs an analog sound 914a at spatial region 904a. In this example, analog sound 914a is an announcement of an optional option 916a (e.g., via a virtual assistant associated with device 600), which is the option to initiate a phone call with Kate.
[0290] exist Figure 9EIn this example, device 600 detects a first head movement 924a, which is a rightward (e.g., towards spatial region 908a) rather than a leftward head rotation (e.g., head tilting or turning). In response to the detection of a second head movement 924a, device 600 outputs an analog sound 926a at spatial region 908a. In this example, the analog sound 926a is an announcement of an optional option 922a (e.g., via a virtual assistant associated with device 600), which is the option to initiate a phone call with Eric.
[0291] exist Figure 9F In this example, the device detects a second head movement 928a, which is a head rotation (e.g., head tilting or head turning) from space region 908a toward space region 906a (e.g., in front of user 602). In response to the detection of the second head movement 928a, the device 600 outputs an analog sound 918a at space region 906a. In this example, the analog sound 932a is an announcement of the optional option 918a, which is the option to initiate a phone call with Jonathan.
[0292] exist Figure 9G In this process, device 600 detects a third head movement 934a, which is a head rotation (e.g., head tilting or head turning) from spatial region 906a toward spatial region 908a (e.g., to the right of user 602). In response to detecting the third head movement 934a, device 600 repeatedly outputs an analog sound 926a at spatial region 908a (e.g., a repeated announcement to initiate a telephone call with Eric).
[0293] exist Figure 9H In this embodiment, when user 602 orients towards spatial region 908a (e.g., facing that direction), device 600 detects a motion posture 936a, which is a nodding posture. In response to detecting motion posture 936a when user 602 is in a right-facing orientation, device 600 initiates a real-time communication session (e.g., a telephone call) between user 602 and Eric. In some embodiments, device 600 sends a signal to device 626 to initiate the telephone call, and associated audio for the initiated telephone call is output at device 600. Therefore, in Figures 9A to 9H In the illustrated implementation, user 602 can select among different potential call recipients and then initiate a call by providing an appropriate gesture after broadcasting the desired recipient. In some implementations, device 600 is coupled with... Figures 6B to 6E The motion posture 936a is detected in a similar manner to that described in [the previous section]. In some embodiments, the device 600 provides audio feedback in response to 936a, similar to responding to the detection of [another motion posture]. Figures 6B to 6E The audio feedback is provided by the nodding gesture.
[0294] Figures 9I to 9N Examples of various spatial audio arrangements (e.g., 910 and 920) are shown for interaction with a playlist menu of songs selectable by user 602. For example, Figures 9A to 9I The device 600 is depicted detecting head movement of user 602 to navigate from a playlist menu in a spatial audio arrangement (e.g., 910) to a menu of songs within that playlist that user 602 can individually select to play (e.g., 920).
[0295] exist Figure 9I At this point, device 600 detects head posture 902b, which is a double head tilt to the right made by user 602 (e.g., a rolling rotation along a longitudinal axis (e.g., an axis pointing in the forward direction of user 602's face)). In response to detecting head posture 902b, and because head posture 902b is a specific type of posture (e.g., a double head tilt to the right rather than to the left), device 600 invokes (e.g., via spatial audio experience) spatial audio arrangement 910.
[0296] In addition, such as Figure 9I As illustrated, the spatial audio arrangement 900 includes spatial regions 904b (e.g., to the left of user 602), spatial regions 906b (e.g., in front of user), and spatial regions 908b (e.g., to the right of user), which correspond to three optional options (e.g., 916b, 918b, and 922b) for initiating media playback (e.g., starting and / or shuffling a song playlist) for “playlist 1”, “playlist 2”, and “playlist 3”.
[0297] exist Figure 9J In this embodiment, device 600 detects a first head movement 912b, which is a head rotation to the left (e.g., head tilting or head turning) toward spatial region 904b. In response to the detection of head movement 912b, device 600 outputs an analog sound 914b at spatial region 904b. In this example, the analog sound 914b is an announcement of an optional option 916a (e.g., via a virtual assistant associated with device 600), which is the option to initiate media playback of "Playlist 1" (e.g., start playback and / or shuffle a song playlist).
[0298] exist Figure 9KIn this embodiment, when user 602 orients towards spatial region 904b (e.g., facing that direction), device 600 detects a head nodding posture 936b. In response to detecting a head nodding posture 936b when user 602 is facing right, device 600 initiates playback of "Playlist 1". In some embodiments, device 600 sends a signal to device 626 to initiate playback of "Playlist 1", which is a playlist of media files (e.g., songs) stored on device 626, and associated audio from "Playlist 1" is output on device 600. In some embodiments, device 600 is coupled with... Figures 6B to 6E The motion posture 936b is detected in a similar manner to that described in [the previous section]. In some embodiments, device 600 provides audio feedback in response to 936a, similar to responding to the detection of [another motion posture]. Figures 6B to 6E The audio feedback is provided by the nodding gesture.
[0299] Figure 9L express Figure 9K In an alternative implementation, device 600 detects a tilt posture 938 (e.g., rather than a nodding posture 936b) when user 602 is facing spatial region 904b to navigate to a menu of optional options within optional options 916b (e.g., a list of playable songs or "tracks" within playlist 1). In response to detecting a tilt posture 938 when user 602 is facing spatial region 904b, device 600 invokes (e.g., via spatial audio experience) spatial audio arrangement 920 (e.g., a submenu of optional songs).
[0300] like Figure 9L As illustrated, the spatial audio arrangement 920 has a spatial region arrangement different from that of spatial audio arrangements 900 and 910. For example, in this arrangement, there are six spatial regions around the user 602, each with an associated optional option (e.g., tracks 1-6 of playlist 1), and these optional options are grouped more closely together (e.g., with fewer intervals) than the regions associated with spatial audio arrangements 900 and 910. Furthermore, the way the device 600 announces the options to the user 602 in the spatial audio arrangement 920 differs from that in spatial audio arrangements 900 and 910. For example, when the spatial audio arrangement 920 is invoked, the device 600 begins to automatically announce the optional options in its respective spatial region (e.g., in clockwise order, starting with playlist 1) without requiring the user 602 to move their head (e.g., unlike spatial audio arrangements 900 and 910 where the device 600 detects a head rotation of the user 602 before announcing the optional options).
[0301] Go to Figure 9MAfter invoking the spatial audio arrangement 920, device 600 outputs analog sound 948 at spatial region 942. In this example, analog sound 948 is an announcement of optional option 946 (e.g., via a virtual assistant associated with device 600), which is the option to "Play Track 1" in "Playlist 1". If device 600 does not detect any motion from user 602 within a threshold time period of outputting analog sound 948 (e.g., while device 600 is still announcing optional options and / or before announcing the next optional option), device 600 continues to output the next analog sound in the option menu.
[0302] exist Figure 9N At location 944, device 600 outputs analog sound 954. In this example, analog sound 954 is an announcement of optional option 952 (e.g., via a virtual assistant associated with device 600), which is the option to "Play Track 2" in "Playlist 1". While device 600 is announcing the "Play Track 2" option (e.g., before announcing the next option in the menu (e.g., "Play Track 3")), device 600 detects a head nodding gesture 936c. In response to detecting the head nodding gesture 936c during the output of sound 954, device 600 initiates playback of track 2 in playlist 1.
[0303] Figure 10 This is a flowchart illustrating a method for detecting motion input in a spatial audio arrangement using one or more audio output devices, according to some embodiments. Method 1000 is performed at one or more audio output devices (e.g., 600) (e.g., speakers, headphones, and / or earphones). Some operations in method 1000 may be combined, the order of some operations may be changed, and some operations may be omitted.
[0304] In some embodiments, one or more audio input devices (e.g., 600) are integrated into a computer system. The computer system optionally communicates (e.g., wired communication, wireless communication) with a display generation component and one or more input devices. The display generation component is configured to provide visual output, such as display via a CRT monitor, an LED monitor, or an image projection display. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separate from the computer system. One or more input devices are configured to receive input, such as a touch-sensitive surface that receives user input. In some embodiments, one or more input devices are integrated with the computer system. In some embodiments, one or more input devices are separate from the computer system. Therefore, the computer system can transmit data (e.g., image data or video data) via wired or wireless connections to an integrated or external display generation component to visually generate content (e.g., using a display device), and can receive input from one or more input devices via wired or wireless connections.
[0305] As described below, Method 1000 provides an intuitive way to interact with audio data via spatial audio arrangement. This method reduces the cognitive burden on users when interacting with audio data, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to interact with audio data more quickly and efficiently, saving power and increasing the time interval between battery charges.
[0306] One or more audio output devices (e.g., 600) (e.g., speakers, headphones, and / or earbuds) (in some embodiments, one or more audio output devices communicate with external electronic devices and / or computer systems (e.g., smartphones, smartwatches, tablets, and / or personal computers)) detect (1002) (e.g., based on processing of sensor measurements at one or more audio output devices and / or based on processing of sensor measurements at companion devices such as smartphones, smartwatches, tablets, wearable computing devices, laptops, and / or desktop computers) corresponding to a user's respective part (e.g., the user's appendages) of one or more audio output devices. For example, the first movement (e.g., 912a and 912b) of a head, arm and / or leg) and / or its extensions (objects held, worn and / or attached to the user) in a three-dimensional environment (e.g., 900 and 910) (e.g., an extended reality environment, an augmented reality environment and / or a virtual reality environment) (in some embodiments, the three-dimensional environment is not perceptible to the user visually) is measured by one or more sensors (e.g., via one or more sensors such as one or more accelerometers, gyroscopes, magnetometers, inertial measurement units, optical sensors and / or other sensors capable of detecting the movement of one or more audio output devices in space).
[0307] In response to detecting one or more sensor measurements corresponding to the first movement (1004) and determining (e.g., via one or more audio output devices and / or via companion devices such as smartphones, smartwatches, tablets, wearable computing devices, laptops, and / or desktop computers) that the first movement corresponds to the orientation of a corresponding part of the user toward a first position (e.g., 904a and / or 904b) in the three-dimensional environment, one or more audio output devices output (1006) a first sound (e.g., 914a and / or 914b) having a simulated spatial location corresponding to the first position in the three-dimensional environment (e.g., simulating sound in the three-dimensional environment via spatial audio experience, such that the user (e.g., a listener) perceives the sound as coming from a specific direction associated with an optional option), wherein the first sound corresponds to one or more optional options. The first optional option in the item (e.g., 916a and / or 916b) (e.g., the option to initiate a phone call with a first contact in the contact list) (in some embodiments, determining that the first movement corresponds to a corresponding part of the user is determining that a part of the user (e.g., head) is moved (e.g., rotated to the left) (e.g., tilted to the left) such that its orientation angle, position, and / or direction of movement corresponds to the same position (e.g., the first position) associated with the simulated spatial position of the first sound (e.g., a leftward head rotation that causes the user's face to face the first position) (e.g., a leftward head tilt that causes the head to move in a directional direction toward the first position)) (in some embodiments, determining that the first movement is a first type of movement is determining that the first movement is associated with a first predefined motion posture (e.g., a head posture (e.g., a double head tilt)) (e.g., as Figure 9D and Figure 9J exemplified).
[0308] In response to detecting one or more sensor measurements corresponding to the first movement (1004) and determining (e.g., via one or more audio output devices and / or via an accessory such as a smartphone, smartwatch, tablet, wearable computing device, laptop, and / or desktop computer) that the first movement (e.g., 924a) corresponds to the orientation of a corresponding part of the user toward a second position (e.g., 908a) in the three-dimensional environment, which is different from the first position in the three-dimensional environment (e.g., the movement of a part of the user (e.g., head) is shifted (e.g., rotated to the right) (e.g., tilted to the right), such that its orientation, position, and / or movement... The angle of direction corresponds to the same position associated with the simulated spatial position of the second sound (e.g., the second position) (e.g., a leftward head rotation that turns the user's face toward the first position) (e.g., a leftward head tilt that moves the head toward the first position)), and one or more audio output devices output (e.g., 1008) a second sound having a simulated spatial position corresponding to the second position in a three-dimensional environment (e.g., 926a), wherein the second sound corresponds to a second optional option (e.g., 922a) among one or more optional options that is different from the first optional option (e.g., the option to initiate a phone call with a second contact in a contact list) (e.g., such as...). Figure 9E exemplified).
[0309] In some implementations, a spatial audio experience in the headphones is generated by manipulating the sounds in two audio channels (e.g., left and right) to resemble directional sounds reaching the ear canal. For example, the headphones can reproduce spatial audio signals simulating the spatial location around a listener (e.g., a user), which differs from the location of physical speakers on the headphones and is optionally adjusted based on head movement. Effective spatial location simulation renders spatial locations that appear fixed in space (e.g., the listener perceives sound as originating from a fixed location), even if the audio output components themselves move in space (e.g., when the listener's head moves). Outputting a first sound with a simulated spatial location corresponding to the first location in the three-dimensional environment based on determining that the first movement corresponds to a first positional orientation of the user's corresponding part towards a first position in the three-dimensional environment, or outputting a second sound with a simulated spatial location corresponding to the second location in the three-dimensional environment based on determining that the first movement corresponds to a second positional orientation of the user's corresponding part towards a second position in the three-dimensional environment different from the first position in the three-dimensional environment, provides the user with a non-visual user interface for navigating one or more optional options, thereby improving control over one or more audio output devices. For example, providing different simulated spatialized sounds based on the user's orientation towards different locations in the 3D environment allows the user to effectively navigate through a set of options without using a visually displayed UI and without requiring the user to generate voice input. Improved control options for one or more audio output devices enhance device operability and make the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently. Furthermore, outputting sounds with different simulated spatial locations based on the user's orientation towards different locations in the 3D environment improves user feedback. The sound output provides feedback to the user, allowing them to associate a specific simulated spatial location with a specific option, and further provides real-time feedback that the orientation of the corresponding part of the user (e.g., head position / orientation / direction movement) triggers the output of the corresponding sound in that corresponding simulated spatial region of the 3D environment. Providing users with improved feedback further enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0310] In some implementations, after detecting one or more sensor measurements corresponding to the first movement, one or more audio output devices detect one or more sensor measurements corresponding to a second movement (e.g., 928a) of a corresponding part of the user in a three-dimensional environment (in some implementations, the second movement is a continuation of the first movement (e.g., a continuous rotation of the user's head, such that the user's face changes from facing a first or second direction to facing a third direction during the same continuous rotation)).
[0311] In some implementations, in response to detecting one or more sensor measurements corresponding to a second movement and based on determining that the second movement corresponds to a third position orientation of a corresponding part of the user in a three-dimensional environment, which is different from the corresponding position in the three-dimensional environment corresponding to the first movement (e.g., different from the first position or different from the second position), one or more audio output devices output a third sound (e.g., 932a) having an analog spatial position corresponding to the third position (e.g., 906a) in the three-dimensional environment, wherein the third sound corresponds to a third optional option (e.g., 918a) among one or more optional options, which is different from the corresponding optional option corresponding to the first movement (e.g., different from the first optional option or different from the second optional option) (e.g., as shown in the image). Figure 9F(As illustrated). In some embodiments, in response to detecting one or more sensor measurements corresponding to a second movement and based on determining that the second movement corresponds to a fourth position orientation of a corresponding part of the user in a three-dimensional environment, which is different from the corresponding position in the three-dimensional environment corresponding to the first movement (e.g., different from the first position or different from the second position) and different from the third position in the three-dimensional environment, a fourth sound is output having an analog spatial position corresponding to the fourth position in the three-dimensional environment, wherein the fourth sound corresponds to a fourth optional option among one or more optional options, which is different from the corresponding optional option corresponding to the first movement (e.g., different from the first optional option or different from the second optional option) and different from the third optional option. Outputting a third sound having an analog spatial position corresponding to the third position in the three-dimensional environment provides the user with a non-visual user interface for navigating one or more optional options, thereby improving control over one or more audio output devices. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the battery life of the device by enabling the user to use the device more quickly and efficiently. Furthermore, outputting sounds with different simulated spatial locations based on the user's orientation towards different positions in the 3D environment improves user feedback. This enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0312] In some implementations, after detecting one or more sensor measurements corresponding to the second movement, one or more audio output devices detect one or more sensor measurements corresponding to a third movement (e.g., 934a) of a portion of the user in a three-dimensional environment (in some implementations, the third movement is the opposite of the direction of the second movement (e.g., the rotation of the user's head in a direction opposite to the direction of rotation in the second movement (e.g., clockwise versus counterclockwise))).
[0313] In some implementations, in response to detecting one or more sensor measurements corresponding to a third movement and based on determining that the third movement corresponds to a user's corresponding portion oriented toward a fourth position in the three-dimensional environment that is the same as the first position in the three-dimensional environment corresponding to the first movement (e.g., 908a), one or more audio output devices output a fourth sound (e.g., 926a) having an analog spatial position corresponding to the fourth position in the three-dimensional environment (e.g., repeating the first sound), wherein the fourth sound corresponds to a fourth optional option (e.g., 922a) among one or more optional options, which is the same as the first optional option corresponding to the first movement (e.g., as...). Figure 9G (As illustrated). In some embodiments, in response to detecting one or more sensor measurements corresponding to a third movement and based on determining that the third movement corresponds to a fifth position orientation of the user's corresponding part towards the three-dimensional environment, which is the same as a second position in the three-dimensional environment corresponding to the first movement, a fifth sound (e.g., repeating the second sound) with an analog spatial position corresponding to the fifth position in the three-dimensional environment is output, wherein the fifth sound corresponds to a fifth optional option among one or more optional options, which is the same as a second optional option corresponding to the first movement. Outputting a fourth sound with an analog spatial position corresponding to a fourth position in the three-dimensional environment provides the user with a non-visual user interface for navigating one or more optional options, thereby improving control over one or more audio output devices. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the battery life of the device by enabling the user to use the device more quickly and efficiently. Furthermore, outputting sounds with different analog spatial positions based on the user's orientation towards different positions in the three-dimensional environment improves feedback to the user. This enhances the device's operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0314] In some implementations, the simulated spatial position of the first sound is perceptibly fixed in space relative to a corresponding part of the user (e.g., relative to the user's head (e.g., the user's head in a forward position before movement begins)) (simulated such that the user perceives the first sound as coming from a fixed position) (in some implementations, the simulated spatial position is perceptibly fixed in space relative to the user's head's initial position (e.g., default position) such that even during subsequent head movements, the user will perceive the sound corresponding to the simulated space as coming from a constant position); and the simulated spatial position of the second sound is perceptibly fixed in space relative to a corresponding part of the user (e.g., such as...). Figure 9D and Figure 9E (As illustrated). The output has a corresponding simulated spatial position corresponding to a corresponding location in a three-dimensional environment, wherein the corresponding simulated spatial position is perceptibly fixed in space relative to a corresponding part of the user, providing the user with greater control over one or more audio output devices. Making the simulated spatial position perceptibly fixed in space relative to the user facilitates a consistent way for the user to receive audio feedback when interacting with optional options in the three-dimensional environment. For example, the user perceives the first sound as coming from a first simulated spatial position that is always perceptibly fixed to the left side of the user's head (e.g., based on head rotation or head tilting to the left), regardless of the user's current body orientation. Doing so enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0315] In some implementations, the simulated spatial location of the first sound is perceptibly separated from the location of a corresponding part of the user and / or one or more audio devices, wherein the simulated spatial location of the first sound is generated via spatialized audio simulation (e.g., spatial audio experience) (in some implementations, the spatialized audio simulation is generated at one or more audio output devices based on processing performed at one or more audio output devices and / or based on processing performed at one or more external devices such as a smartphone; in some implementations, the spatialized audio simulation is generated at an external device communicating with one or more audio output devices); and the simulated spatial location of the second sound is perceptibly separated from the location of a corresponding part of the user and / or one or more audio devices, wherein the simulated spatial location of the second sound is generated via spatialized audio simulation at one or more audio output devices (e.g., as...). Figure 9E and Figure 9D(As illustrated). The output has a corresponding simulated spatial location corresponding to a corresponding position in a three-dimensional environment, where the corresponding simulated spatial location is perceptually separated from the corresponding part of the user, providing the user with greater control over one or more audio output devices. Making the simulated spatial location perceptually fixed relative to the user in space facilitates a consistent way for the user to receive audio feedback when interacting with optional options in the three-dimensional environment. For example, the user perceives the first sound as coming from a first simulated spatial location that is always perceptually fixed to the left side of the user's head (e.g., based on head rotation or tilting the head to the left), regardless of the user's current body orientation. This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0316] In some implementations, the first position in the three-dimensional environment is a first simulated position located relative to a corresponding orientation (e.g., default orientation and / or forward position) of a corresponding part of the user (e.g., body and / or head); and the second position in the three-dimensional environment is a second simulated position located relative to a corresponding orientation of a corresponding part of the user (e.g., as shown in the image). Figure 9E and Figure 9D (As illustrated). Based on determining that the corresponding movement corresponds to the user's corresponding part's orientation in the three-dimensional environment, the system outputs sound with a corresponding simulated spatial position in the three-dimensional environment corresponding to that position, where the corresponding position is a simulated position positioned relative to the user's corresponding part's orientation. This provides the user with greater control over one or more audio output devices. The simulated position in the three-dimensional environment, positioned relative to the user's orientation, facilitates a consistent interface for the user to interact with, regardless of whether the user's body is in motion. For example, a user can interact with a simulated position that is always located to the left of the user's head (e.g., generating head movement toward that simulated position) (e.g., based on head rotation or tilting the head to the left), regardless of the user's current body orientation. This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0317] In some implementations, the first simulated position is fixed relative to the corresponding orientation of the user's corresponding body part (e.g., the position does not change) (e.g., it is a simulated position always positioned to the left of the user's head (e.g., based on head rotation or head tilt to the left), regardless of the user's current body orientation and / or movement); and the second simulated position is fixed relative to the corresponding orientation of the user's corresponding body part (e.g., as...). Figure 9E and Figure 9D (As illustrated). Based on determining that the corresponding movement corresponds to the user's corresponding part's orientation in the three-dimensional environment, the system outputs sound with a corresponding simulated spatial position in the three-dimensional environment corresponding to that position, where the corresponding position is a simulated position positioned relative to the user's corresponding part's orientation. This provides the user with greater control over one or more audio output devices. The simulated position in the three-dimensional environment, positioned relative to the user's orientation, facilitates a consistent interface for the user to interact with, regardless of whether the user's body is in motion. For example, a user can interact with a simulated position that is always located to the left of the user's head (e.g., generating head movement toward that simulated position) (e.g., based on head rotation or tilting the head to the left), regardless of the user's current body orientation. This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0318] In some implementations, the corresponding part of the user is the user's head, and the first movement is a rotation of the user's head (e.g., as shown in the image). Figure 9E and Figure 9D (As illustrated). A first sound is output based on determining that the first movement corresponds to a first position orientation of the user's corresponding part in the three-dimensional environment, with the user's head being the corresponding part and the first movement causing rotation of the user's head. This provides the user with a non-visual, hands-free user interface for navigating one or more optional options, thereby improving control over one or more audio output devices. For example, providing different analog spatialized sounds based on directional head movement allows the user to effectively navigate among a set of optional options without using a visually displayed UI and without requiring the user to provide hand or arm movement gestures. This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0319] In some implementations, the axis of rotation of the user's head is positioned relative to the user's posture (e.g., body position and / or orientation) (e.g., this axis is always positioned to pass through the user's body line, regardless of whether the user is standing or tilted) (e.g., as... Figure 9E and Figure 9D (As illustrated). By positioning the user's head rotation axis relative to the user's posture, additional control over one or more audio output devices is provided to the user through a consistent interface for interaction, regardless of the user's body position and / or orientation. For example, a user can interact with an option that is always perceptibly fixed to the left side of the user's head (e.g., generating head movement toward that option) (e.g., based on head rotation or tilting the head to the left), whether the user is standing or lying down. This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0320] In some implementations, the first sound is an announcement describing a first optional option among one or more options (e.g., an option to initiate a real-time communication session with a first contact); and the second sound is an announcement describing a second optional option among one or more options (e.g., an option to initiate a real-time communication session with a second contact); wherein the second sound is different from the first sound (e.g., as...). Figure 9E and Figure 9D (As illustrated). The system outputs a first sound with a simulated spatial position corresponding to a first location in the three-dimensional environment, where the first sound is an announcement describing a first optional option among one or more options, and outputs a second sound with a simulated spatial position corresponding to a second location in the three-dimensional environment, where the second sound is an announcement describing a second optional option among one or more options. This improves user feedback. Doing so enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0321] In some embodiments, one or more audio output devices detect one or more sensor measurements corresponding to a first motion posture (in some embodiments, the detection of one or more sensor measurements occurs after the output of a corresponding sound corresponding to a corresponding optional option); and in response to the detection of one or more sensor measurements corresponding to the first motion posture (e.g., 936c) (e.g., a first type of posture (e.g., nodding, shaking, tilting, and / or double tilting of the user's head)) and based on determining that the first motion posture was detected within a threshold time period of the output of the first sound (e.g., 948) (e.g., during the first announcement of the first optional option and / or before the second announcement of the second optional option), the one or more audio output devices cause to perform a first operation associated with the first optional option (e.g., 946).
[0322] In some implementations, in response to detecting one or more sensor measurements corresponding to a first motion posture (e.g., 936c) and based on determining that the first motion posture was detected within a threshold time period for outputting a second sound (e.g., 954) (e.g., during the second announcement of the first optional option and / or before the third announcement of the third optional option), one or more audio output devices cause to perform a first operation associated with the second optional option (e.g., 952) (e.g., as...). Figure 9M (As illustrated). By determining that a first motion posture is detected within a threshold time period for outputting a first sound, a first operation associated with a first optional option is executed; and by determining that a first motion posture is detected within a threshold time period for outputting a second sound, a first operation associated with a second optional option is executed. This provides the user with a non-visual user interface for performing operations corresponding to the respective optional options, thereby improving control over one or more audio output devices. For example, by executing operations corresponding to the respective optional options based on motion posture detected within a threshold time period for outputting the corresponding sound, precise control of the non-visual interface is provided without using a visual UI and without requiring the user to generate voice input. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0323] In some embodiments, before detecting one or more sensor measurements of a first movement of a corresponding part of a user in a three-dimensional environment corresponding to one or more audio output devices, one or more audio output devices detect one or more sensor measurements of a movement of a second part of the user (e.g., 902a and / or 902b) (in some embodiments, the second part of the user is the same as the corresponding part of the user), wherein: the one or more sensor measurements determining the movement of the second part of the user correspond to a first type of movement (e.g., left head tilt or double left head tilt (e.g., 902a)). One or more optional options are a first set of optional options (e.g., 916a, 918a, and / or 922a) (e.g., options in a first menu (e.g., a menu that can be contacted by the user)); and one or more optional options are a second set of optional options (e.g., right head tilt or double right head tilt (e.g., 902b)) that correspond to a movement of a second part of the user, based on one or more sensor measurements. Figure 9A and Figure 9I (As illustrated). One or more sensor measurements corresponding to the movement of the user's second part are detected, wherein the one or more sensor measurements corresponding to the movement of the user's second part correspond to a corresponding type of movement, and one or more optional options are a set of corresponding optional options. This improves control of one or more audio output devices by providing the user with control over a wider range of optional options in a three-dimensional environment. For example, the user can efficiently navigate to different sets of optional options based on the movement of the user's second part and depending on the type of movement corresponding to the movement of the user's second part (e.g., tilting the head to the left or right). This enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0324] In some implementations, one or more optional options (e.g., 916a, 918a and / or 922a) are the first subgroup of multiple options arranged in a hierarchical structure (e.g., options in a hierarchical menu (e.g., a song menu arranged by genre (e.g., level 1), artist (e.g., level 2), album (e.g., level 3), and then song (e.g., level 4))), and one or more optional options correspond to the first level of the hierarchical structure.
[0325] In some embodiments, one or more audio output devices detect (e.g., before or after detecting a first movement of the user's corresponding part) one or more sensor measurements corresponding to the user's third part (e.g., 938) (in some embodiments, the user's third part is the same as the user's corresponding part) (in some embodiments, one or more sensor measurements corresponding to the movement of the user's third part are detected after the corresponding sound (e.g., the first sound or the second sound) is output); and in response to detecting one or more sensor measurements corresponding to the movement of the user's third part, one or more audio output devices cause a hierarchical structure navigation operation to be performed, wherein causing the hierarchical structure navigation operation to be performed includes navigating to a second level of the hierarchical structure (e.g., a higher or lower level in the hierarchical structure), the second level comprising a second subgroup of multiple options that is different from one or more optional options (e.g., 946 and 952) (e.g., one or more optional options correspond to a genre of music, and the second subgroup of multiple options corresponds to an artist within that genre) (e.g., as Figure 9L and Figure 9M (As illustrated). A hierarchical navigation operation is performed in response to one or more sensor measurements detecting movement corresponding to a third part of the user's body. This hierarchical navigation operation includes navigating to a second level of the hierarchy, which comprises a second subgroup of multiple options distinct from one or more optional options. This improves control of one or more audio devices by providing the user with control over a wider range of optional options corresponding to different submenus. For example, when interacting with a menu of optional options, the user can efficiently navigate to a set of optional options within a submenu of the menu without using a visually displayed UI and without requiring the user to respond to voice commands. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0326] In some implementations, one or more optional options corresponding to the first level of the hierarchical structure have a first number of optional options; and a second subgroup of multiple options corresponding to the second level of the hierarchical structure has a second number of optional options (e.g., a larger or smaller number) than the first number of optional options (e.g., such as...). Figure 9L and Figure 9M(As illustrated). A hierarchical navigation operation is performed in response to one or more sensor measurements detecting movement corresponding to a third part of the user's body. This hierarchical navigation operation includes navigating to a second level of the hierarchy, which comprises a second subgroup of multiple options distinct from one or more optional options. This improves control of one or more audio devices by providing the user with control over a wider range of optional options corresponding to different submenus. For example, when interacting with a menu of optional options, the user can efficiently navigate to a set of optional options within a submenu of the menu without using a visually displayed UI and without requiring the user to respond to voice commands. This enhances device operability and makes the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.
[0327] In some implementations, the simulated spatial position corresponding to a first position (e.g., 904b) in the three-dimensional environment (e.g., the position corresponding to a first optional option among one or more optional options) and the simulated spatial position corresponding to a second position (e.g., 906b) in the three-dimensional environment (e.g., the position corresponding to a second optional option among one or more optional options) have a first spatial interval (e.g., a linear distance and / or angular interval separating them) in the three-dimensional environment (e.g., such as...). Figure 9I and Figure 9L (As shown).
[0328] In some implementations, a second subgroup of multiple options corresponding to a second level of the hierarchical structure includes: a first optional option (e.g., 946) in the second level of the hierarchical structure corresponding to a fifth position (e.g., 942) in a three-dimensional environment (e.g., the first second-level optional option corresponds to the sound output at a third position); and a second optional option (e.g., 952) in the second level of the hierarchical structure corresponding to a sixth position (e.g., 944) in a three-dimensional environment, wherein the third and fourth positions in the three-dimensional environment have a second spatial interval (e.g., a larger or smaller spatial interval) in the three-dimensional environment than the first spatial interval (in some implementations, the degree of spatial interval between options at different levels of the hierarchical structure is different (e.g., because the number of options differs between levels)). A hierarchical navigation operation is performed in response to detecting one or more sensor measurements corresponding to movement of a third portion of the user's data. Performing the hierarchical navigation operation includes navigating to a second level of the hierarchical structure, which includes a second subgroup of multiple options different from one or more optional options. This improves control of one or more audio devices by providing the user with control over a wider range of optional options corresponding to different submenus. For example, when interacting with a menu of optional options, users can efficiently navigate to a set of options within a submenu without using a visually displayed UI or requiring the user to respond to voice commands. This enhances device operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0329] In some embodiments, one or more audio output devices detect one or more sensor measurements corresponding to a second motion posture (e.g., 936a) (e.g., head posture (e.g., nodding or shaking)); and in response to detecting one or more sensor measurements corresponding to the second motion posture and based on determining that the second motion posture is detected when the user's corresponding part is oriented toward a first positional orientation in the three-dimensional environment (in some embodiments, the user's current orientation is detected via one or more sensors (e.g., one or more accelerometers, gyroscopes, magnetometers, inertial measurement units, optical sensors, and / or other sensors capable of detecting orientation) located at one or more audio output devices and / or at accompanying devices such as smartphones, smartwatches, tablets, wearable computing devices, laptops, and / or desktop computers), the one or more audio output devices cause the operation associated with the first optional option to be performed.
[0330] In some implementations, in response to detecting one or more sensor measurements corresponding to a second motion posture and based on determining that the second motion posture (e.g., 936a) is detected when the corresponding part of the user is oriented toward a second position (e.g., 908a) in the three-dimensional environment, one or more audio output devices cause the execution of an operation associated with a second optional option (e.g., 922a) (e.g., as...). Figure 9H (As illustrated). In some embodiments, the execution of the operation associated with the first optional option is independent of the output of the first sound (e.g., a user can make a gesture in the direction of a first position in the three-dimensional environment to perform the associated operation without first hearing the first sound). In some embodiments, the operation associated with the first optional option is performed after the first sound is output (e.g., the user must make a first movement to cause the first sound to be output before making a head gesture in the direction of the sound to perform the associated operation). Performing the operation associated with the first optional option based on detecting a second motion posture when the user's corresponding part is oriented towards a first position in the three-dimensional environment, and performing the operation associated with the second optional option based on detecting a second motion posture when the user's corresponding part is oriented towards a second position in the three-dimensional environment, provides the user with a non-visual user interface for performing operations corresponding to the respective optional option, thereby improving control over one or more audio output devices. For example, performing the operation corresponding to the respective optional option based on the direction the user is facing when a motion posture is received provides precise control of the non-visual interface without using a visually displayed UI and without requiring the user to generate voice input. This enhances the device's operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0331] In some implementations, causing the operation associated with the first optional option includes: based on determining that the second motion posture is a third type of motion posture (e.g., 936b) (e.g., a nodding posture), causing the execution of a first type of operation (e.g., shuffling media playback of the first playlist) associated with the first optional option (e.g., an option corresponding to the first playlist) (e.g., such as...). Figure 9K (As illustrated above regarding method 700 and) Figures 6A to 6RThe above); and based on determining that the second motion posture is a fourth type of motion posture (e.g., 938) (e.g., head tilt posture), causing the execution of a second type of operation (e.g., navigating to a submenu for playing a single song within the first playlist) that is different from the first type of operation associated with the first optional option (e.g., the option corresponding to the first playlist) (e.g., as described above); and based on determining that the second motion posture is a fourth type of motion posture (e.g., 938) (e.g., head tilt posture), causing the execution of a second type of operation (e.g., navigating to a submenu for playing a single song within the first playlist) (e.g., as described above); Figure 9L (As illustrated). In some implementations, causing the operation associated with the second optional option includes: causing the execution of a first type of operation (e.g., shuffling media playback of the second playlist) associated with the second optional option (e.g., an option corresponding to the second playlist) based on determining that the second motion posture is a third type of motion posture (e.g., a nodding posture); and causing the execution of a second type of operation (e.g., navigating to a submenu for playing a single song within the second playlist) associated with the second optional option (e.g., an option corresponding to the second playlist) based on determining that the second motion posture is a fourth type of motion posture. This second type of operation, different from the first type, based on determining that the second motion posture is a fourth type of motion posture, provides the user with control over a wider range of operations based on different motion posture types without using a visually displayed UI and without requiring the user to respond to voice commands. This enhances the device's operability and makes the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and extends the device's battery life by enabling users to use the device more quickly and efficiently.
[0332] It should be noted that the above is relative to method 1000 (for example, Figure 10 The details of the process described also apply in a similar manner to the methods described above. For example, methods 700 and 800 optionally include one or more characteristics of the various methods described above with reference to method 1000. For example, using the techniques described in methods 800 and 1000, one or more audio output devices can cause the operations described with respect to method 700 to be performed. For example, method 800 can be used to provide feedback on the progress of a motion posture that causes the operation of optional options to be performed in an analog spatial arrangement according to method 1000. As an additional example, sound having an analog spatial arrangement according to method 1000 can be an audio notification in method 700. For the sake of brevity, these details will not be repeated below.
[0333] For purposes of explanation, the foregoing description has been given by reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible based on the teachings above. These embodiments were chosen and described in order to best explain the principles of these techniques and their practical application. Others skilled in the art will thus be able to best utilize these techniques and the various embodiments with various modifications suitable for the particular intended use.
[0334] While this disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. It should be understood that such changes and modifications are considered to be included within the scope of this disclosure and examples as defined by the claims.
[0335] As described above, one aspect of the present invention is the collection and use of data from various sources to improve the efficiency of interacting with audio data that may be of interest to them. This disclosure contemplates that, in some instances, such collected data may include personal information that uniquely identifies or can be used to contact or locate specific individuals. Such personal information may include demographic data, location-based data, telephone numbers, email addresses, social network IDs, home addresses, data or records related to a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.
[0336] This disclosure recognizes that the use of such personal information data in the techniques of this invention can benefit users. For example, personal information data can be used as input data and / or to determine user intent. Therefore, the use of such personal information data enables users to exercise planned control over the content delivered. Furthermore, this disclosure also contemplates other uses of personal information data that benefit users. For example, health and fitness data can be used to provide insights into a user's overall health status or can be used as positive feedback for individuals using the technology to pursue health goals.
[0337] This disclosure anticipates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will comply with robust privacy policies and / or privacy measures. Specifically, such entities should implement and adhere to privacy policies and measures that are recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. Such policies should be easily accessible to users and should be updated as the collection and / or use of data changes. Personal information from users should be collected for legitimate and reasonable entity purposes and should not be shared or sold outside of these legitimate purposes. Furthermore, such collection / sharing should be conducted only after receiving informed consent from users. Additionally, such entities should consider taking any necessary steps to protect and safeguard the right to access such personal information data and ensure that other entities with access to such personal information data comply with the privacy policies and procedures of those other entities. Furthermore, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and privacy measures. Moreover, policies and measures should be adapted to the specific types of personal information data collected and / or accessed, and to applicable laws and standards, including considerations of specific jurisdictions. For example, in the United States, the collection or acquisition of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); while in other countries, health data may be subject to other regulations and policies and should be handled accordingly. Therefore, different privacy measures should be advocated for different types of personal data in each country.
[0338] Regardless of the foregoing, this disclosure also contemplates implementation schemes for users to selectively block the use or access to personal information data. That is, this disclosure contemplates providing hardware and / or software components to prevent or block access to such personal information data. For example, with regard to determining user intent and / or input, the inventive technology can be configured to allow users to opt-in or opt-out at any time during or after registration for the service to participate in the collection of personal information data. In another example, users may choose not to provide mobility-related data used to determine user intent. In yet another example, users may choose to limit the length of time mobility-related data is retained, or completely prohibit the development of a baseline mobility profile. In addition to providing opt-in and opt-out options, this disclosure also contemplates providing notifications related to access to or use of personal information. For example, users may be notified when downloading an application that their personal information data will be accessed, and then reminded again just before the application accesses the personal information data.
[0339] Furthermore, the intent of this disclosure is that personal information data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Once data is no longer needed, this risk can be minimized by restricting data collection and deleting data. Additionally, and where applicable, including in certain health-related applications, data deidentification can be used to protect user privacy. Deidentification can be facilitated, where appropriate, by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than the address level), controlling how data is stored (e.g., aggregating data among users), and / or other methods.
[0340] Therefore, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, it is also contemplated that various embodiments can be implemented without access to such personal information data. That is, various embodiments of the present invention will not become inoperable due to the absence of all or part of such personal information data. For example, preferences can be inferred based on non-personal information data or a minimal amount of personal information such as content requested by a device associated with a user, other non-personal information available to the system, or publicly available information, thereby selecting content and delivering it to the user via audio data.
Claims
1. A method, the method comprising: At one or more audio output devices: Output the first audio notification; After outputting the first audio notification, motion input is detected based on one or more sensor measurements from one or more sensors in the one or more audio output devices; as well as In response to detected motion input and based on determination that a first set of criteria is met, a first operation associated with the first audio notification is performed, wherein the first set of criteria includes a first criterion that is met when the motion input is detected within a threshold time period during which the first audio notification is output.
2. The method according to claim 1, further comprising: In response to the detection of the motion input and based on the determination that the first set of criteria is not met, the execution of the first operation associated with the first audio notification is abandoned.
3. The method of any one of claims 1 to 2, wherein causing the execution of the first operation associated with the first audio notification comprises: Based on the determination that the detected motion input is a first type of motion input, the first type of operation associated with the first audio notification is executed; as well as If the detected motion input is determined to be a second type of motion input that is different from the first type of motion input, then an operation of a second type, different from the first type, associated with the first audio notification is performed.
4. The method according to claim 3, wherein: The first audio notification includes controllable prompts; and The motion input of the first type corresponds to an affirmative response to the controllable prompt.
5. The method according to claim 3, wherein: The first audio notification includes controllable prompts; and The second type of motion input corresponds to a negative response to the controllable prompt.
6. The method according to any one of claims 1 to 5, wherein: The first audio notification is the first sub-part of the first audio notification in progress; The motion input is detected during the output of the first in-process audio notification; and Performing the first operation includes interrupting the output of the first in-progress audio notification.
7. The method of claim 6, wherein interrupting the output of the first in-process audio notification comprises: Stop the output of the first in-process audio notification; as well as Output a second audio notification.
8. The method according to any one of claims 1 to 7, wherein: The first audio notification includes a prompt indicating a change in the mode associated with the one or more audio output devices; and Performing the first operation includes changing the mode associated with the one or more audio output devices from a first mode associated with the one or more audio output devices to a second mode associated with the one or more audio output devices that is different from the first mode.
9. The method according to claim 8, wherein: The first mode associated with the one or more audio output devices is a first notification mode, which includes a first set of notification settings; and The second mode associated with the one or more audio output devices is a second notification mode, which includes a second set of notification settings that are different from the first set of notification settings.
10. The method according to any one of claims 8 to 9, wherein: The first mode associated with the one or more audio output devices is a first audio notification mode, which includes a first set of audio notification output settings that affect the output of audio notifications via the one or more audio output devices; and The second mode associated with the one or more audio output devices is a second audio notification mode, which includes a second set of audio notification output settings that affect the output of audio notifications via the one or more audio output devices, wherein the second set of audio notification output settings is different from the first set of audio notification settings.
11. The method of any one of claims 8 to 10, wherein the first audio notification of the prompt that changes the mode associated with the one or more audio output devices is output after detecting a plurality of previous motion inputs, wherein the plurality of previous motion inputs satisfy a second set of criteria.
12. The method according to any one of claims 1 to 11, wherein: Performing the first operation includes changing the playback state of the media item.
13. The method according to any one of claims 1 to 12, wherein: The first audio notification is associated with an inquiry to join a live communication session; and This makes performing the first operation include joining the real-time communication session.
14. The method according to any one of claims 1 to 13, wherein: The first audio notification includes a prompt to send a message; and Performing the first operation includes transmitting the message.
15. The method according to any one of claims 1 to 14, wherein: Performing the first operation includes outputting a notification corresponding to the received message.
16. The method according to any one of claims 1 to 15, further comprising: In response to the detection of the motion input, a first audio feedback indicating that the motion input has been recognized is provided.
17. The method of claim 16, wherein providing the first audio feedback comprises: Based on the determination that the detected motion input is a third type of motion input, provide first type of audio feedback; as well as Based on the determination that the detected motion input is a fourth type of motion input, a second type of audio feedback, different from the first type, is provided.
18. The method according to any one of claims 1 to 17, wherein the first operation is performed in the absence of voice input from a user from the one or more audio output devices.
19. The method according to any one of claims 1 to 18, wherein the first set of criteria includes a second criterion satisfied in the following circumstances: Based on the determination that conflicting voice input was detected during a second threshold time period when the first audio notification was output, the detected motion input was identified as expected input based on a set of conflict resolution criteria.
20. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 19.
21. One or more audio output devices, said one or more audio output devices comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 19.
22. One or more audio output devices, said one or more audio output devices comprising: Components for performing the method according to any one of claims 1 to 19.
23. A computer program product comprising one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 19.
24. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs including instructions for: Output the first audio notification; After outputting the first audio notification, motion input is detected based on one or more sensor measurements from one or more sensors in the one or more audio output devices; as well as In response to detected motion input and based on determination that a first set of criteria is met, a first operation associated with the first audio notification is performed, wherein the first set of criteria includes a first criterion that is met when the motion input is detected within a threshold time period during which the first audio notification is output.
25. One or more audio output devices, said one or more audio output devices comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: Output the first audio notification; After outputting the first audio notification, motion input is detected based on one or more sensor measurements from one or more sensors in the one or more audio output devices; as well as In response to detected motion input and based on determination that a first set of criteria is met, a first operation associated with the first audio notification is performed, wherein the first set of criteria includes a first criterion that is met when the motion input is detected within a threshold time period during which the first audio notification is output.
26. One or more audio output devices, said one or more audio output devices comprising: Components used to output the first audio notification; A component for detecting motion input based on one or more sensor measurements from one or more sensors in one or more audio output devices after the first audio notification has been output; and Components for responding to detected motion input and performing a first operation associated with the first audio notification based on determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion that is met when the motion input is detected within a threshold time period during which the first audio notification is output.
27. A computer program product comprising one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs comprising instructions for performing the following operations: Output the first audio notification; After outputting the first audio notification, motion input is detected based on one or more sensor measurements from one or more sensors in the one or more audio output devices; as well as In response to detected motion input and based on determination that a first set of criteria is met, a first operation associated with the first audio notification is performed, wherein the first set of criteria includes a first criterion that is met when the motion input is detected within a threshold time period during which the first audio notification is output.
28. A method, the method comprising: At one or more audio output devices: Detect the results of one or more sensors corresponding to the start of the motion posture; After the detection of the one or more sensor measurements corresponding to the start of the motion posture, and while the detection of the one or more sensor measurements is in progress, a first audio feedback indicating the progress of the motion posture is provided via the one or more audio output devices; as well as After providing the first audio feedback and determining the motion posture, an operation associated with the motion posture is performed.
29. The method according to claim 28, further comprising: After providing the first audio feedback and determining that the motion posture is not completed, the operation associated with the motion posture is abandoned.
30. The method according to claim 29, further comprising: After providing the first audio feedback and based on the determination that the motion posture was not completed, a first audio output indicating that the operation associated with the motion posture was canceled is provided.
31. The method according to any one of claims 28 to 30, the method further comprising: After providing the first audio feedback and determining that the motion posture has been completed, an audio output indicating that the motion posture has been successfully completed is provided.
32. The method according to any one of claims 28 to 31, wherein: The measurement results from the one or more sensors were detected via one or more sensors of the one or more audio output devices; and The one or more audio output devices are included in one or more wearable devices.
33. The method of claim 32, wherein the one or more wearable devices are a group of one or more earbuds or headphones.
34. The method of any one of claims 28 to 33, wherein providing the first audio feedback indicating the progress of the motion posture comprises outputting a plurality of discrete sounds.
35. The method according to claim 34, wherein: The progress of the motion posture includes a first intermediate sub-part of the detected motion posture and a second intermediate sub-part of the detected motion posture; The plurality of discrete sounds include a first discrete sound output in response to the first intermediate sub-part of the detected motion posture; and The plurality of discrete sounds include a second discrete sound output in response to the second intermediate sub-part of the detected motion posture.
36. The method of claim 35, wherein: The first discrete sound indicates the beginning state of the progression of the motion posture; and The second discrete sound indicates the ongoing state of the progress of the motion posture.
37. The method according to any one of claims 35 to 36, wherein: The progression of the motion posture includes the final sub-part of the detected motion posture; The plurality of discrete sounds include a third discrete sound output in response to the final sub-part of the detected motion; The third discrete sound indicates the completion status of the progress of the motion posture.
38. The method according to any one of claims 35 to 37, wherein: The first intermediate sub-part of the motion posture is detected before the second intermediate sub-part of the motion posture is detected; The first intermediate sub-part of the motion posture corresponds to a first confidence level of the motion posture toward completion progress; The second intermediate sub-part of the motion posture corresponds to a second confidence level higher than the first confidence level, indicating that the motion posture is progressing towards completion; and The first discrete sound has a first value of the audio characteristic within the range of its value, the first value corresponding to the first confidence level of the motion posture toward completion of progress; and The second discrete sound has a second value of the audio characteristic that is further away from the first value of the audio characteristic within the value range of the audio characteristic, and the second discrete sound corresponds to the second confidence level of the motion posture toward completion of progress.
39. The method according to any one of claims 35 to 37, wherein: The first intermediate sub-part of the motion posture is detected before the second intermediate sub-part of the motion posture is detected; The first intermediate sub-part of the motion posture corresponds to a first confidence level of the motion posture toward completion progress; The second intermediate sub-part of the motion posture corresponds to the first confidence level of the motion posture toward completion; and The first discrete sound has a first value of the audio characteristic within the range of its value, the first value corresponding to the first confidence level of the motion posture toward completion of progress; and The second discrete sound has the first value of the audio characteristic, and the second discrete sound corresponds to the first confidence level of the motion posture toward completion of progress.
40. The method of any one of claims 28 to 39, wherein providing the first audio feedback indicating the progress of the motion posture comprises: Based on determining that the motion posture is a first type of motion posture, provide first type of first audio feedback; as well as Based on the determination that the motion posture is a second type of motion posture different from the first type, a first type of audio feedback of the second type is provided.
41. The method according to any one of claims 28 to 40, further comprising: Before detecting the sensor measurements corresponding to the start of the motion posture, a first portion of the in-process audio effect is provided via the one or more audio output devices; and Providing the first audio feedback includes providing a second part of the in-process audio effect by modifying one or more audio characteristics of the in-process audio effect.
42. The method according to claim 41, wherein: The one or more audio output devices communicate with the audio input devices; and The in-process audio effect indicates the time period during which the one or more audio output devices listen to one or more audio inputs via the audio input devices.
43. The method according to any one of claims 41 to 42, wherein: Providing the first audio feedback includes outputting one or more sounds, wherein a corresponding sound from the one or more sounds is output based on determining that a corresponding sensor measurement result among the one or more sensor measurements satisfies a corresponding threshold confidence level for the completion of the motion posture orientation; and One or more modifications to the audio effect correspond to the output of the one or more sounds, wherein a corresponding modification in the one or more modifications to the audio effect indicates a corresponding state of the progress of the motion posture.
44. The method of any one of claims 28 to 43, wherein causing the operation associated with the motion posture to be performed comprises: If the motion posture is determined to be a first type of motion posture, then a first type of operation is performed. as well as Based on the determination that the motion posture is a second type of motion posture different from the second type of motion posture, an operation of the second type different from the operation of the first type is performed.
45. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs including instructions for performing the method according to any one of claims 28 to 44.
46. One or more audio output devices, said one or more audio output devices comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 28 to 44.
47. One or more audio output devices, said one or more audio output devices comprising: Components for performing the method according to any one of claims 28 to 44.
48. A computer program product comprising one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs comprising instructions for performing the method according to any one of claims 28 to 44.
49. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors at one or more audio output devices, the one or more programs including instructions for: Detect the results of one or more sensors corresponding to the start of the motion posture; After the detection of the one or more sensor measurements corresponding to the start of the motion posture, and while the detection of the one or more sensor measurements is in progress, a first audio feedback indicating the progress of the motion posture is provided via the one or more audio output devices; as well as After providing the first audio feedback and determining the motion posture, an operation associated with the motion posture is performed.
50. One or more audio output devices, said one or more audio output devices comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: Detect the results of one or more sensors corresponding to the start of the motion posture; After the detection of the one or more sensor measurements corresponding to the start of the motion posture, and while the detection of the one or more sensor measurements is in progress, a first audio feedback indicating the progress of the motion posture is provided via the one or more audio output devices; as well as After providing the first audio feedback and determining the motion posture, an operation associated with the motion posture is performed.
51. One or more audio output devices, said one or more audio output devices comprising: A component used to detect the measurement results of one or more sensors corresponding to the start of a motion posture; A component for providing first audio feedback indicating the progress of the motion posture via the one or more audio output devices after detecting the one or more sensor measurements corresponding to the start of the motion posture and while the detection of the one or more sensor measurements is in progress; and A component for performing an operation associated with the motion posture after providing the first audio feedback and based on the determination of the motion posture.
52. A computer program product comprising one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs comprising instructions for performing the following operations: Detect the results of one or more sensors corresponding to the start of the motion posture; After the detection of the one or more sensor measurements corresponding to the start of the motion posture, and while the detection of the one or more sensor measurements is in progress, a first audio feedback indicating the progress of the motion posture is provided via the one or more audio output devices; as well as After providing the first audio feedback and determining the motion posture, an operation associated with the motion posture is performed.
53. A method, the method comprising: At one or more audio output devices: One or more sensor measurements are used to detect the first movement of a corresponding part of a user in a three-dimensional environment corresponding to one or more audio output devices; as well as In response to detecting the one or more sensor measurements corresponding to the first movement: Based on determining that the first movement corresponds to a first position orientation of the corresponding part of the user toward the three-dimensional environment, a first sound with a simulated spatial position corresponding to the first position in the three-dimensional environment is output, wherein the first sound corresponds to a first optional option among one or more optional options; as well as Based on determining that the first movement corresponds to a second position orientation of the corresponding part of the user toward a different position in the three-dimensional environment than the first position in the three-dimensional environment, a second sound is output having a simulated spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among the one or more optional options that is different from the first optional option.
54. The method according to claim 53, further comprising: After detecting the one or more sensor measurements corresponding to the first movement, detect one or more sensor measurements corresponding to the second movement of the corresponding part of the user in the three-dimensional environment of the one or more audio output devices; In response to detecting the sensor measurements corresponding to the second movement: Based on determining that the second movement corresponds to a third position orientation in the three-dimensional environment that is different from the corresponding position in the three-dimensional environment corresponding to the first movement, a third sound is output having a simulated spatial position corresponding to the third position in the three-dimensional environment, wherein the third sound corresponds to a third optional option among the one or more optional options that is different from the corresponding optional option corresponding to the first movement.
55. The method according to any one of claims 53 to 54, the method further comprising: After detecting the one or more sensor measurements corresponding to the second movement, detect one or more sensor measurements corresponding to the third movement of the user's portion in the three-dimensional environment corresponding to the one or more audio output devices; In response to detecting the one or more sensor measurements corresponding to the third movement: Based on determining that the third movement corresponds to the user's corresponding part facing a fourth position orientation in the three-dimensional environment that is the same as the first position in the three-dimensional environment corresponding to the first movement, a fourth sound is output having a simulated spatial position corresponding to the fourth position in the three-dimensional environment, wherein the fourth sound corresponds to a fourth optional option among the one or more optional options that is the same as the first optional option corresponding to the first movement.
56. The method according to any one of claims 53 to 55, wherein: The simulated spatial position of the first sound is perceptibly fixed in space relative to the corresponding part of the user; and The simulated spatial position of the second sound is perceptually fixed in space relative to the corresponding part of the user.
57. The method according to claim 56, wherein: The simulated spatial location of the first sound is perceptually separated from the location of the corresponding part of the user and / or the location of the one or more audio devices, wherein the simulated spatial location of the first sound is generated via spatialized audio simulation; and The simulated spatial location of the second sound is perceptually located separately from the corresponding part of the user and / or the location of the one or more audio devices, wherein the simulated spatial location of the second sound is generated via the spatialized audio simulation at the one or more audio output devices.
58. The method according to any one of claims 53 to 57, wherein: The first position in the three-dimensional environment is a first simulated position located relative to the corresponding orientation of the corresponding part of the user; and The second position in the three-dimensional environment is a second simulated position located relative to the corresponding orientation of the corresponding part of the user.
59. The method according to claim 58, wherein: The first simulated position is fixed relative to the corresponding orientation of the corresponding part of the user; and The second simulated position is fixed relative to the corresponding orientation of the corresponding part of the user.
60. The method according to any one of claims 53 to 59, wherein the corresponding portion of the user is the user's head and the first movement is a rotation of the user's head.
61. The method of claim 60, wherein the axis of rotation of the user's head is positioned relative to the user's posture.
62. The method according to any one of claims 53 to 61, wherein: The first sound is a notification describing the first optional option among the one or more optional options; and The second sound is a notification describing the second optional option among the one or more optional options; The second sound is different from the first sound.
63. The method according to any one of claims 53 to 62, further comprising: Detect the measurement results of one or more sensors corresponding to the first motion posture; as well as In response to detecting one or more sensor measurements corresponding to the first motion posture: Based on the determination that the first motion posture was detected within a threshold time period for outputting the first sound, a first operation associated with the first optional option is executed; as well as Based on the determination that the first motion posture was detected within a threshold time period for outputting the second sound, a first operation associated with the second optional option is performed.
64. The method according to any one of claims 53 to 63, the method further comprising: Before detecting the first movement of the corresponding part of the user in the three-dimensional environment corresponding to the one or more audio output devices, one or more sensor measurements corresponding to the movement of the second part of the user are detected, wherein: Based on the one or more sensor measurements determining that the movement corresponds to the second part of the user's movement corresponds to a first type of movement, the one or more optional options are a first set of optional options; and The one or more sensor measurements that determine the movement corresponding to the second part of the user correspond to a second type of movement that is different from the first type of movement, and the one or more optional options are a second set of optional options that are different from the first set of optional options.
65. The method of any one of claims 53 to 64, wherein the one or more optional options are a first subgroup of a plurality of options arranged in a hierarchical structure, and the one or more optional options correspond to a first level of the hierarchical structure, the method further comprising: Detect the results of one or more sensors corresponding to the movement of the third part of the user; as well as In response to detecting the movement of the third part of the user, one or more sensor measurements are made to perform a hierarchical navigation operation, wherein performing hierarchical navigation includes navigating to a second level of the hierarchical structure, the second level including a second subgroup of the plurality of options that is different from the one or more optional options.
66. The method according to claim 65, wherein: The one or more optional options corresponding to the first level of the hierarchical structure have a first number of optional options; and The second subgroup of the plurality of options, corresponding to the second level of the hierarchical structure, has a second number of optional options that differ from the first number of optional options.
67. The method according to any one of claims 65 to 66, wherein: The simulated spatial position corresponding to the first position in the three-dimensional environment and the simulated spatial position corresponding to the second position in the three-dimensional environment have a first spatial interval in the three-dimensional environment; and The second subgroup of the plurality of options corresponding to the second level of the hierarchical structure includes: The first optional option in the second level of the hierarchical structure, corresponding to the fifth position in the three-dimensional environment; as well as The second optional option in the second level of the hierarchical structure corresponds to the sixth position in the three-dimensional environment, wherein the fifth position and the sixth position in the three-dimensional environment have a second spatial interval in the three-dimensional environment that is different from the first spatial interval.
68. The method according to any one of claims 53 to 67, further comprising: Detect the measurement results of one or more sensors corresponding to the second motion posture; as well as In response to detecting one or more sensor measurements corresponding to the second motion posture: Based on the determination that the second motion posture is detected when the corresponding part of the user is oriented toward the first position in the three-dimensional environment, the operation associated with the first optional option is executed; as well as Based on the determination that the second motion posture is detected when the corresponding part of the user is oriented toward the second positional orientation in the three-dimensional environment, an operation associated with the second optional option is performed.
69. The method according to claim 68, wherein: Making the operation associated with the first optional option include: Based on determining that the second motion posture is a third type of motion posture, the operation of the first type associated with the first optional option is executed; and Based on the determination that the second motion posture is a fourth type of motion posture, a second type of operation, which is different from the first type of operation, is performed in connection with the first optional option.
70. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs including instructions for performing the method according to any one of claims 53 to 69.
71. One or more audio output devices, said one or more audio output devices comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 53 to 69.
72. One or more audio output devices, said one or more audio output devices comprising: Components for performing the method according to any one of claims 53 to 69.
73. A computer program product comprising one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs comprising instructions for performing the method according to any one of claims 53 to 69.
74. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs including instructions for: One or more sensor measurements are used to detect the first movement of a corresponding part of a user in a three-dimensional environment corresponding to one or more audio output devices; as well as In response to detecting the one or more sensor measurements corresponding to the first movement: Based on determining that the first movement corresponds to a first position orientation of the corresponding part of the user toward the three-dimensional environment, a first sound with a simulated spatial position corresponding to the first position in the three-dimensional environment is output, wherein the first sound corresponds to a first optional option among one or more optional options; as well as Based on determining that the first movement corresponds to a second position orientation of the corresponding part of the user toward a different position in the three-dimensional environment than the first position in the three-dimensional environment, a second sound is output having a simulated spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among the one or more optional options that is different from the first optional option.
75. One or more audio output devices, said one or more audio output devices comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: One or more sensor measurements are used to detect the first movement of a corresponding part of a user in a three-dimensional environment corresponding to one or more audio output devices; as well as In response to detecting the one or more sensor measurements corresponding to the first movement: Based on determining that the first movement corresponds to a first position orientation of the corresponding part of the user toward the three-dimensional environment, a first sound with a simulated spatial position corresponding to the first position in the three-dimensional environment is output, wherein the first sound corresponds to a first optional option among one or more optional options; as well as Based on determining that the first movement corresponds to a second position orientation of the corresponding part of the user toward a different position in the three-dimensional environment than the first position in the three-dimensional environment, a second sound is output having a simulated spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among the one or more optional options that is different from the first optional option.
76. One or more audio output devices, said one or more audio output devices comprising: Components for detecting the first movement of a corresponding part of a user in a three-dimensional environment corresponding to the one or more audio output devices; and Components for performing the following operations in response to detecting the one or more sensor measurements corresponding to the first movement: Based on determining that the first movement corresponds to a first position orientation of the corresponding part of the user toward the three-dimensional environment, a first sound with a simulated spatial position corresponding to the first position in the three-dimensional environment is output, wherein the first sound corresponds to a first optional option among one or more optional options; as well as Based on determining that the first movement corresponds to a second position orientation of the corresponding part of the user toward a different position in the three-dimensional environment than the first position in the three-dimensional environment, a second sound is output having a simulated spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among the one or more optional options that is different from the first optional option.
77. A computer program product comprising one or more programs configured to be executed by one or more processors of one or more audio output devices, the one or more programs comprising instructions for performing the following operations: One or more sensor measurements are used to detect the first movement of a corresponding part of a user in a three-dimensional environment corresponding to one or more audio output devices; as well as In response to detecting the one or more sensor measurements corresponding to the first movement: Based on determining that the first movement corresponds to a first position orientation of the corresponding part of the user toward the three-dimensional environment, a first sound with a simulated spatial position corresponding to the first position in the three-dimensional environment is output, wherein the first sound corresponds to a first optional option among one or more optional options; as well as Based on determining that the first movement corresponds to a second position orientation of the corresponding part of the user toward a different position in the three-dimensional environment than the first position in the three-dimensional environment, a second sound is output having a simulated spatial position corresponding to the second position in the three-dimensional environment, wherein the second sound corresponds to a second optional option among the one or more optional options that is different from the first optional option.
Citation Information
Patent Citations
Method and apparatus for integrating manual input
US20020015024A1
Gestures for touch sensitive input devices
US20060026521A1
Gestures for touch sensitive input devices
US20060026536A1
Virtual input device placement on a touch screen user interface
US20060033724A1
Multipoint touchscreen
US20060097991A1