Devices, methods, and graphical user interfaces for providing feedback related to real-world objects

By integrating sensors into wearable devices to detect user gestures and spatial context, and providing audio and haptic feedback, the problem of existing devices being unable to effectively detect real-world objects and sensor occlusion is solved, improving interaction efficiency, saving energy, and extending device usage time.

CN122111235APending Publication Date: 2026-05-29APPLE INC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
APPLE INC
Filing Date
2024-09-26
Publication Date
2026-05-29

Smart Images

  • Figure CN122111235A_ABST
    Figure CN122111235A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a device, method, and graphical user interface for providing feedback related to real-world objects. An ear-wearable audio output device including one or more sensors and one or more audio output components detects a user gesture via the one or more sensors of the ear-wearable audio output device. In response to detecting the user gesture, the ear-wearable audio output device provides, via the one or more audio output components, first audio feedback corresponding to a first real-world object in accordance with a determination that the user gesture is a first type of gesture and is directed at the first real-world object. In response to detecting the user gesture, the ear-wearable audio output device provides, via the one or more audio output components, second audio feedback corresponding to a second real-world object in accordance with a determination that the user gesture is the first type of gesture and is directed at the second real-world object.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application No. 202480059105.6, filed on September 26, 2024, entitled "Apparatus, Method and Graphical User Interface for Providing Feedback Related to Real-World Objects". Related patent applications This application is a continuation to U.S. Patent Application No. 18 / 895,177, filed September 24, 2024, and claims priority to that U.S. Patent Application No. 63 / 541,751, filed September 29, 2023. Technical Field

[0003] This disclosure generally relates to ear-worn devices such as earbuds and headphones, including but not limited to ear-worn devices that provide feedback related to real-world objects, provide indications of alarm conditions, adjust modifications to ambient sounds, perform actions corresponding to gestures, and / or provide feedback indicating that sensors are obstructed. Background Technology

[0004] Audio output devices, including wearable audio output devices such as headphones and earphones, are widely used to provide audio output to users. However, conventional methods of providing audio output are cumbersome, inefficient, and limited. Additionally, conventional audio output devices cannot detect and provide feedback about real-world objects, or detect and respond to air gestures.

[0005] In some cases, conventional ear-worn audio output devices cannot detect user gestures pointing at real-world objects. Such devices cannot provide audio feedback corresponding to real-world objects.

[0006] In some cases, wearable audio output devices cannot detect alarm conditions that are relevant to the user's spatial context. In other cases, wearable audio output devices do not automatically provide audio feedback about alarm conditions and / or modify the magnitude of ambient sounds from the physical environment (e.g., to make the user better hear sounds relevant to the alarm condition). In some cases, wearable audio output devices cannot detect air gestures (e.g., hand gestures performed near the user's head). Such devices cannot perform actions corresponding to air gestures.

[0007] In some cases, wearable devices cannot determine that a particular sensor is blocked (e.g., completely or partially blocked). In other cases, wearable devices cannot provide feedback to the user that a particular sensor is blocked.

[0008] Furthermore, conventional methods take longer than necessary and require more user interaction, thus wasting energy and providing an inefficient human-computer interface. Conserving device energy is particularly important in battery-powered devices. Summary of the Invention

[0009] Therefore, there is a need for wearable devices (e.g., wearable audio output devices, ear-worn devices, and / or other types of wearable devices) and associated electronics with improved methods and interfaces for control and interaction, such as providing feedback associated with real-world objects, alarm conditions, and / or sensor occlusion, and / or performing actions in response to air gestures. Such methods and interfaces optionally complement or replace conventional methods for controlling the operation of wearable devices. These methods and interfaces reduce the amount, extent, and / or nature of input from the user, resulting in a more efficient human-machine interface. For battery-powered systems and devices, such methods and interfaces conserve power and increase the interval between battery charges.

[0010] The disclosed device reduces or eliminates the aforementioned defects and other problems associated with the user interface of electronic devices (or more generally, computer systems) having touch-sensitive surfaces. In some embodiments, the device is a desktop computer. In some embodiments, the device is portable (e.g., a laptop, tablet, or handheld device). In some embodiments, the device is a personal electronic device (e.g., a wearable electronic device, such as a watch). In some embodiments, the device is a wearable audio output device (e.g., in-ear headphones, earbuds, over-ear headphones, etc.), including the wearable audio output device, and / or communicating with the wearable audio output device. In some embodiments, the device has a touchpad. In some embodiments, the device has a touch-sensitive display (also referred to as a "touchscreen" or "touchscreen display"). In some embodiments, the device has a graphical user interface (GUI), one or more processors, memory, and one or more modules, and a program or instruction set stored in memory for performing multiple functions. In some embodiments, the user interacts with the GUI primarily through stylus and / or finger contact and gestures on the touch-sensitive surface. In some implementations, these functions optionally include image editing, drawing, presentation, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, fitness support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.

[0011] According to some embodiments, a method is performed at an ear-worn audio output device including one or more sensors and one or more audio output components. The method includes: detecting a user gesture via one or more sensors of the ear-worn audio output device. The method further includes: in response to detecting the user gesture, providing first audio feedback corresponding to the first real-world object via one or more audio output components based on determining that the user gesture is a first type of gesture and points to a first real-world object; and providing second audio feedback corresponding to the second real-world object via one or more audio output components based on determining that the user gesture is a first type of gesture and points to a second real-world object.

[0012] According to some implementations, a method is performed at a wearable audio output device. The method includes: when the wearable audio output device is physically positioned relative to a corresponding body part of a user, detecting an alarm condition related to the user's spatial context, in which the magnitude of ambient sound from the physical environment is modified by the wearable audio output device to have a first ambient sound audio level. The method further includes: in response to detecting the alarm condition, providing audio feedback regarding the alarm condition; and changing one or more properties of the wearable audio output device to modify the magnitude of the ambient sound from the physical environment to have a second ambient sound audio level that is louder than the first ambient sound audio level.

[0013] According to some embodiments, a method is performed at a wearable audio output device. The method includes: detecting a gesture performed by a user's hand while outputting audio content. The method further includes: in response to detecting the gesture: performing a first operation corresponding to the gesture based on determining that the gesture is detected at a corresponding distance to the side of the user's head and is a first type of hand gesture determined at least partially based on the shape of the hand during the execution of the gesture; and abandoning the execution of the first operation based on determining that the gesture is not a first type of hand gesture determined at least partially based on the shape of the hand during the execution of the gesture.

[0014] According to some embodiments, a method is performed at a wearable device including one or more sensors. The method includes: detecting the occurrence of one or more events while the wearable device is worn by a user, the one or more events indicating that the device is in a context in which a corresponding sensor among the one or more sensors can be used to perform a corresponding operation. The method further includes: in response to detecting the occurrence of the one or more events: providing feedback to the user indicating that the corresponding sensor is occluded based on determining that the corresponding sensor is occluded; and performing an operation based on information detected by the corresponding sensor among the one or more sensors based on determining that the corresponding sensor is not occluded.

[0015] According to some embodiments, an electronic device (or more generally, a computer system) includes: a display, a touch-sensitive surface, optional one or more sensors for detecting the intensity of contact with the touch-sensitive surface, optional one or more tactile output generators, one or more processors, and a memory storing one or more programs; the one or more programs are configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing the execution of any of the methods described herein. According to some embodiments, a computer-readable storage medium stores instructions therein that, when executed by an electronic device having a display, a touch-sensitive surface, optional one or more sensors for detecting the intensity of contact with the touch-sensitive surface, and optional one or more tactile output generators, cause the device to perform or cause the execution of any of the methods described herein. According to some embodiments, a graphical user interface on an electronic device having a display, a touch-sensitive surface, optional one or more sensors for detecting the intensity of contact with the touch-sensitive surface, optional one or more tactile output generators, a memory, and one or more processors for executing one or more programs stored in the memory includes one or more elements displayed in any of the methods described herein, which update in response to input as described in any of the methods described herein. According to some embodiments, an electronic device includes: a display, a touch-sensitive surface, one or more sensors optionally for detecting the intensity of contact with the touch-sensitive surface, and one or more tactile output generators optionally; and components for performing or causing the performance of any of the methods described herein. According to some embodiments, an information processing device in an electronic device having a display, a touch-sensitive surface, one or more sensors optionally for detecting the intensity of contact with the touch-sensitive surface, and one or more tactile output generators optionally includes components for performing or causing the performance of any of the methods described herein.

[0016] Therefore, electronic devices, including one or more sensors, one or more input devices, and / or one or more audio output devices or communicating with them, are provided with improved methods and interfaces for providing audio / haptic feedback, thereby increasing the effectiveness, efficiency, and user satisfaction of such devices. Such methods and interfaces can complement or replace conventional methods for providing audio / haptic feedback. Attached Figure Description

[0017] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals indicate corresponding parts in all the drawings.

[0018] Figure 1AThis is a block diagram illustrating a portable multi-functional device with a touch-sensitive display according to some implementation schemes.

[0019] Figure 1B This is a block diagram illustrating example components for event handling according to some implementation schemes.

[0020] Figures 1C to 1D Various examples of computer systems according to some implementation schemes are illustrated.

[0021] Figure 2 Examples of portable multi-functional devices with touchscreens according to some implementation schemes are shown.

[0022] Figure 3A This is a block diagram of an example multifunctional device with a display and a touch-sensitive surface, according to some implementation schemes.

[0023] Figure 3B The physical features of an example wearable audio output device according to some implementation schemes are illustrated.

[0024] Figure 3C This is a block diagram of an example wearable audio output device based on some implementation schemes.

[0025] Figure 3D Example audio control performed by a wearable audio output device according to some implementation schemes is illustrated.

[0026] Figure 3E Example audio control performed by another wearable audio output device according to some implementation schemes is illustrated.

[0027] Figure 4A An example user interface for an application menu on a portable multifunction device according to some implementation schemes is shown.

[0028] Figure 4B An example user interface for a multi-functional device having a touch-sensitive surface separate from the display is illustrated according to some implementation schemes.

[0029] Figures 5A to 5R Examples of user interfaces and user interactions involving real-world objects and feedback from wearable devices are illustrated according to some implementation schemes.

[0030] Figures 6A to 6H Examples of user interactions involving real-world objects and feedback from wearable devices are illustrated according to some implementation schemes.

[0031] Figures 7A to 7K Examples of user interactions with wearable devices involving various alarm conditions are illustrated according to some implementation schemes.

[0032] Figures 8A to 8PExamples of user interactions with wearable devices are illustrated according to some implementation schemes.

[0033] Figures 9A to 9K Examples of user interactions with wearable devices are illustrated according to some implementation schemes.

[0034] Figures 10A to 10D It is a flowchart of a process for providing audio feedback related to real-world objects, based on some implementation schemes.

[0035] Figures 11A to 11C This is a flowchart of a process for providing audio feedback on alarm conditions, based on some implementation schemes.

[0036] Figures 12A to 12D It is a flowchart of a process for performing operations in response to air gestures, according to some implementation schemes.

[0037] Figures 13A to 13B This is a flowchart of a process for providing feedback indicating that a sensor is blocked, according to some implementation schemes. Detailed Implementation

[0038] As mentioned above, wearable devices (including wearable audio output devices such as headphones, earbuds, and earphones) are widely used to provide output to users. Many wearable devices cannot detect user gestures (e.g., air gestures and / or gestures pointing at real-world objects) and / or alarm conditions. The methods, systems, and user interfaces / interactions described herein improve the capabilities of wearable devices in a variety of ways. For example, the embodiments disclosed herein describe improved ways to obtain feedback and / or alarm conditions regarding real-world objects at the wearable device and provide an improved user interface for controlling the wearable device.

[0039] The processes described below enhance the capabilities and operability of the device through various technologies, and make the user-device interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device). These include providing users with improved visual, audio, and / or haptic feedback; reducing the amount of input required to perform an operation; providing additional control options without cluttering the user interface with additional displayed controls; and performing an operation when a set of conditions is met without additional user input and / or additional technologies. These technologies also reduce power consumption and extend the device's battery life by enabling users to use the device faster and more efficiently.

[0040] In the following text, Figures 1A to 1D , Figure 2 and Figures 3A to 3E A description of the example device is provided. Figures 4A to 4B , Figures 5A to 5R , Figures 6A to 6H, Figures 7A to 7K , Figures 8A to 8P and Figures 9A to 9K Example user interactions with wearable devices are shown. Figures 10A to 10D A flowchart illustrating a method for providing audio feedback related to real-world objects is provided. Figures 11A to 11C A flowchart illustrating a method for providing audio feedback on alarm conditions is provided. Figures 12A to 12D A flowchart illustrating a method for performing an operation in response to an air gesture is provided. Figures 13A to 13B A flowchart illustrating a method for providing feedback indicating that a sensor is blocked is shown. Figures 5A to 5R , Figures 6A to 6H , Figures 7A to 7K , Figures 8A to 8P and Figures 9A to 9K The user interface in the example is used to demonstrate Figures 10A to 10D , Figures 11A to 11C , Figures 12A to 12D and Figures 13A to 13B The process in.

[0041] Example device Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. Numerous specific details are shown in the following detailed description in order to provide a full understanding of the various described embodiments. However, it will be apparent to those skilled in the art that the various described embodiments can be practiced without these specific details. In other instances, well-known methods, processes, components, circuits, and networks are not described in detail so as not to unnecessarily obscure the various aspects of the embodiments.

[0042] It will also be understood that, although in some cases the terms “first,” “second,” etc., are used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact, without departing from the scope of the various described embodiments. Both the first contact and the second contact are contacts, but they are not the same contact unless the context clearly indicates otherwise.

[0043] The terminology used in the description of the various described embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various described embodiments and in the appended claims, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “comprising” and / or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0044] As used herein, depending on the context, the term “if” is optionally interpreted as meaning “when…” followed by “at…” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrase “if it is determined…” or “if [the stated condition or event] is detected” is optionally interpreted as meaning “when it is determined…” or “in response to determination…” or “when [the stated condition or event] is detected” or “in response to the detection of [the stated condition or event].”

[0045] Implementations of electronic devices (and more generally, computer systems), user interfaces for such devices, and associated processes for using such devices are described. In some implementations, the device is a portable communication device, such as a mobile phone, that also includes other functionalities such as PDA and / or music player functionality. Example implementations of portable multi-functional devices include, but are not limited to, the iPhone from Apple Inc. (Cupertino, California). ® iPod Touch ® and iPad ® Device. Optionally, other portable electronic devices may be used, such as laptops or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0046] In the following discussion, a computer system in the form of an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, mouse, and / or joystick.

[0047] The device typically supports a variety of applications, such as one or more of the following: note-taking applications, drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, game applications, phone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camcorder applications, web browsing applications, digital music player applications, and / or digital video player applications.

[0048] Various applications running on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the device are optionally adjusted and / or varied for different applications, and / or within the respective applications. In this way, the common physical architecture of the device (such as the touch-sensitive surface) optionally utilizes a user interface that is intuitive and clear to the user to support various applications.

[0049] Now let’s turn our attention to implementations of computer systems, such as portable devices with touch-sensitive displays. Figure 1A This is a block diagram illustrating a portable multi-functional device 100 with a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display system 112 is sometimes referred to as a "touchscreen" for convenience, and sometimes simply as a touch-sensitive display. Device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral interface 118, RF circuitry 108, audio circuitry 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input or control devices 116, and an external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more intensity sensors 165 (e.g., touch-sensitive surfaces, such as the touch-sensitive display system 112 of device 100) for detecting the intensity of contact on device 100. Device 100 optionally includes one or more haptic output generators 167 for generating haptic outputs on device 100 (e.g., generating haptic outputs on a touch-sensitive surface such as the touch-sensitive display system 112 of device 100 or the touchpad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.

[0050] As used in this specification and claims, the term "haptic output" refers to a physical displacement of the device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., the housing), or a displacement of a component relative to the center of mass of the device, which is detected by the user using the user's tactile sense. For example, when the device or a component of the device comes into contact with a touch-sensitive surface (e.g., a finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or touchpad) may optionally be interpreted by the user as a "press-click" or "release-click" on a physically actuated button. In some cases, the user will feel a tactile sensation, such as a "press-click" or "release-click," even when a physically actuated button associated with the touch-sensitive surface, which has been physically pressed (e.g., displaced) by the user's movement, does not move. As another example, even when the smoothness of a tactile surface remains unchanged, movement of that surface can optionally be interpreted or sensed by the user as “roughness.” While such interpretations of touch by a user will be limited by the user’s individualized sensory perceptions, many sensory perceptions of touch are common to most users. Therefore, when tactile output is described as corresponding to a user’s specific sensory perception (e.g., “release click,” “press click,” “roughness”), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or its components that will generate the sensory perception described by a typical (or average) user. Providing tactile feedback to the user using tactile output enhances device operability and makes the user device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device). This also further reduces power consumption and extends device battery life by enabling users to use the device more quickly and efficiently.

[0051] In some implementations, the haptic output mode specifies the characteristics of the haptic output, such as the amplitude of the haptic output, the shape of the motion waveform of the haptic output, the frequency of the haptic output, and / or the duration of the haptic output.

[0052] When a device (e.g., one or more haptic output generators that generate haptic output via a movable mass) generates haptic output with different haptic output modes, the haptic output can produce different tactile sensations for a user holding or touching the device. While the user's senses are based on their perception of the haptic output, most users will be able to identify variations in the waveform, frequency, and amplitude of the haptic output generated by the device. Therefore, the waveform, frequency, and amplitude can be adjusted to indicate to the user that different actions have been performed. Thus, haptic output with haptic output modes designed, selected, and / or arranged to simulate the characteristics (e.g., size, material, weight, stiffness, smoothness, etc.); behaviors (e.g., oscillation, displacement, acceleration, rotation, extension, etc.); and / or interactions (e.g., collision, adhesion, repulsion, attraction, friction, etc.) of objects in a given environment (e.g., a user interface including graphical features and objects, a simulated physical environment with virtual boundaries and virtual objects, a real physical environment with physical boundaries and physical objects, and / or any combination thereof), will in some cases provide helpful feedback to the user, reducing input errors and improving the efficiency of the user's operation of the device. Additionally, haptic output can optionally be generated as feedback independent of the simulated physical characteristics, such as input thresholds or object selection. Such haptic output can provide helpful feedback to the user in some cases, reducing input errors and improving the efficiency of the user's operation of the device.

[0053] In some implementations, haptic output with an appropriate haptic output mode serves as a cue for an event of interest occurring in the user interface or behind the screen of the device. Examples of events of interest include activation of a power indication provided on the device or in the user interface (e.g., a real or virtual button, or a toggle switch), success or failure of a requested operation, reaching or crossing a boundary in the user interface, entering a new state, switching input focus between objects, activating a new mode, reaching or crossing an input threshold, detecting or recognizing a type of input or gesture, and so on. In some implementations, haptic output is provided to serve as a warning or cue about an impending event or outcome that will occur unless a change of orientation or interruption of input is detected in a timely manner. Haptic output is also used in other contexts to enrich the user experience, improve the accessibility of the device for users with visual or motor difficulties or other accessibility needs, and / or improve the efficiency and functionality of the user interface and / or the device. Optionally, haptic output can be compared with changes in audio input and / or visual user interface, which further enhances the user experience when interacting with the user interface and / or the device, facilitates better transmission of information about the state of the user interface and / or the device, and reduces input errors and improves the efficiency of user operation of the device.

[0054] It should be understood that device 100 is merely an example of a portable multifunctional device, and device 100 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of these components. Figure 1A The various components shown are implemented in hardware, software, firmware, or any combination thereof, including one or more signal processing circuits and / or application-specific integrated circuits.

[0055] Memory 102 optionally includes high-speed random access memory and also optionally includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Access to memory 102 by other components of device 100, such as CPU 120 and peripheral interface 118, is optionally controlled by memory controller 122.

[0056] Peripheral interface 118 can be used to couple the device's input peripherals and output peripherals to CPU 120 and memory 102. One or more processors 120 run or execute various software programs and / or instruction sets stored in memory 102 to perform various functions of device 100 and process data.

[0057] In some implementations, the peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In other implementations, they are optionally implemented on separate chips.

[0058] RF (Radio Frequency) circuit 108 receives and transmits RF signals, also known as electromagnetic signals. RF circuit 108 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with communication networks and other communication devices via these electromagnetic signals. RF circuit 108 optionally includes well-known circuitry for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, subscriber identity module (SIM) cards, memory, etc. RF circuit 108 optionally communicates wirelessly with networks (such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs))) and other devices. The wireless communication may optionally use any of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed ​​Downlink Packet Access (HSDPA), High-Speed ​​Uplink Packet Access (HSUPA), Evolved Pure Data (EV-DO), HSPA, HSPA+, Dual-Unit HSPA (DC-HSPA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, and Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Messaging Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence with Extended Utility (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol including those not yet developed as of the date of this document submission.

[0059] Audio circuitry 110, speaker 111, and microphone 113 provide an audio interface between the user and device 100. Audio circuitry 110 receives audio data from peripheral interface 118, converts the audio data into electrical signals, and sends the electrical signals to speaker 111. Speaker 111 converts the electrical signals into sound waves that are audible to humans. Audio circuitry 110 also receives electrical signals converted from sound waves by microphone 113. Audio circuitry 110 converts the electrical signals into audio data and sends the audio data to peripheral interface 118 for processing. Audio data is optionally retrieved by peripheral interface 118 from and / or sent to memory 102 and / or RF circuitry 108. In some embodiments, audio circuitry 110 also includes a headset jack (e.g., ...). Figure 2 (212 in the text). The headset jack provides an interface between the audio circuitry 110 and a removable audio input / output peripheral device, such as an output-only headset or a headset with both outputs (e.g., a single-ear or dual-ear headset) and inputs (e.g., a microphone).

[0060] I / O subsystem 106 couples input / output peripherals on device 100, such as touch-sensitive display system 112 and other input or control devices 116, to peripheral interface 118. I / O subsystem 106 optionally includes display controller 156, optical sensor controller 158, intensity sensor controller 159, haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. One or more input controllers 160 receive electrical signals from / transmit electrical signals to other input or control devices 116. Other input control devices 116 optionally include physical buttons (e.g., push-buttons, rocker buttons, etc.), dial pads, slide switches, joysticks, click wheels, etc. In some alternative embodiments, one or more input controllers 160 are optionally coupled to (or not coupled to) any of the following: keyboard, infrared port, USB port, stylus, and / or pointing device such as mouse. The one or more buttons (e.g., Figure 2 Optionally, 208 of the above includes an up / down button (e.g., a single button that shakes in opposite directions, or separate up and down buttons) for volume control of the speaker 111 and / or microphone 113. One or more buttons optionally include a push-down button (e.g., Figure 2 (206 in the middle).

[0061] The touch-sensitive display system 112 provides input and output interfaces between the device and the user. The display controller 156 receives electrical signals from and / or transmits electrical signals to the touch-sensitive display system 112. The touch-sensitive display system 112 displays visual output to the user. Visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively, "graphics"). In some embodiments, some or all of the visual output corresponds to a user interface object. As used herein, the term "enabled representation" refers to a user-interactive graphical user interface object (e.g., a graphical user interface object configured to respond to input directed to it). Examples of user-interactive graphical user interface objects include, but are not limited to, buttons, sliders, icons, selectable menu items, switches, hyperlinks, or other user interface controls.

[0062] The touch-sensitive display system 112 has a touch-sensitive surface, sensor, or sensor array that accepts input from a user based on tactile and / or tactile contact. The touch-sensitive display system 112 and the display controller 156 (along with any associated modules and / or instruction sets in memory 102) detect contact on the touch-sensitive display system 112 (and any movement or interruption of that contact) and translate the detected contact into interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch-sensitive display system 112. In some embodiments, the point of contact between the touch-sensitive display system 112 and the user corresponds to the user's finger or stylus.

[0063] The touch-sensitive display system 112 optionally employs LCD (Liquid Crystal Display) technology, LPD (Light Emitting Polymer Display) technology, or LED (Light Emitting Diode) technology, but other display technologies are used in other embodiments. The touch-sensitive display system 112 and display controller 156 optionally employ any of a variety of touch sensing technologies now known or to be developed hereafter, along with other proximity sensor arrays or other elements for determining one or more points of contact with the touch-sensitive display system 112, to detect contact and any movement or interruption thereof. These various touch sensing technologies include, but are not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies. In some embodiments, projected mutual capacitance sensing technology is used, such as that used in the iPhone from Apple Inc. (Cupertino, California). ® iPod Touch ® and iPad ® The technology discovered in [the text].

[0064] The touch-sensitive display system 112 optionally has a video resolution exceeding 100 dpi. In some embodiments, the touchscreen video resolution exceeds 400 dpi (e.g., 500 dpi, 800 dpi, or greater). The user optionally uses any suitable object or accessory such as a stylus, finger, etc., to interact with the touch-sensitive display system 112. In some embodiments, the user interface is designed to work with finger-based contact and gestures, which may be less precise than stylus-based input due to the larger contact area of ​​a finger on the touchscreen. In some embodiments, the device translates coarse finger-based input into precise pointer / cursor positioning or commands for performing the user-desired actions.

[0065] In some embodiments, in addition to the touchscreen, device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of ​​the device that, unlike the touchscreen, does not display visual output. The touchpad is optionally a touch-sensitive surface separate from the touch-sensitive display system 112, or an extension of the touch-sensitive surface formed by the touchscreen.

[0066] The device 100 also includes a power system 162 for supplying power to various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, power fault detection circuitry, a power converter or inverter, a power status indicator (e.g., light-emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in the portable device.

[0067] The device 100 may also include one or more optical sensors 164 (e.g., as part of one or more cameras). Figure 1A An optical sensor coupled to an optical sensor controller 158 in the I / O subsystem 106 is shown. One or more optical sensors 164 optionally include charge-coupled devices (CCDs) or complementary metal-oxide-semiconductor (CMOS) phototransistors. The one or more optical sensors 164 receive light projected through one or more lenses from the environment and convert the light into data representing an image. In conjunction with an imaging module 143 (also referred to as a camera module), the one or more optical sensors 164 optionally capture still images and / or video. In some embodiments, the optical sensor is located on the rear of the device 100, opposite to the touch-sensitive display system 112 on the front of the device, enabling the touchscreen to be used as a viewfinder for still image and / or video image acquisition. In some embodiments, another optical sensor is located on the front of the device to acquire images of the user (e.g., for selfies, for video conferencing while the user views other video conference participants on the touchscreen, etc.).

[0068] The device 100 may optionally also include one or more contact strength sensors 165. Figure 1A A contact strength sensor coupled to a strength sensor controller 159 in I / O subsystem 106 is shown. One or more contact strength sensors 165 optionally include one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric sensors, optical force sensors, capacitive touch-sensitive surfaces, or other strength sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). One or more contact strength sensors 165 receive contact strength information (e.g., pressure information or a substitute for pressure information) from the environment. In some embodiments, at least one contact strength sensor is arranged juxtaposed with or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact strength sensor is located on the rear of device 100, opposite to the touch-sensitive display system 112 located on the front of device 100.

[0069] The device 100 optionally also includes one or more proximity sensors 166. Figure 1A A proximity sensor 166 coupled to a peripheral device interface 118 is shown. Alternatively, the proximity sensor 166 is coupled to an input controller 160 in an I / O subsystem 106. In some embodiments, the proximity sensor turns off and disables the touch-sensitive display system 112 when the multifunction device is placed near the user's ear (e.g., when the user is making a phone call).

[0070] The device 100 may optionally also include one or more tactile output generators 167. Figure 1AA haptic output generator coupled to a haptic feedback controller 161 in I / O subsystem 106 is shown. In some embodiments, the haptic output generator 167 includes one or more electroacoustic devices such as speakers or other audio components; and / or electromechanical devices such as motors, solenoids, electroactive polymerizers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components for converting electrical signals into haptic outputs on the device). The haptic output generator 167 receives haptic feedback generation instructions from haptic feedback module 133 and generates a haptic output on device 100 that can be felt by a user of device 100. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to a haptic surface (e.g., haptic display system 112) and optionally generates the haptic output by moving the haptic surface vertically (e.g., in / outward from the surface of device 100) or laterally (e.g., backward and forward in the same plane as the surface of device 100). In some embodiments, at least one tactile output generator sensor is located on the rear of the device 100, opposite to the touch-sensitive display system 112 located on the front of the device 100.

[0071] The device 100 may optionally also include one or more accelerometers 168. Figure 1A An accelerometer 168 coupled to a peripheral interface 118 is shown. Alternatively, the accelerometer 168 may be coupled to an input controller 160 in an I / O subsystem 106. In some embodiments, information is displayed on a touchscreen display in portrait or landscape view based on analysis of data received from the one or more accelerometers. The device 100 may optionally include, in addition to the accelerometer 168, a magnetometer and a GPS (or GLONASS or other global navigation system) receiver for acquiring information about the location and orientation (e.g., portrait or landscape) of the device 100.

[0072] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a touch / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a haptic feedback module (or instruction set) 133, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and an application program (or instruction set) 136. Furthermore, in some embodiments, memory 102 stores device / global internal state 157, as shown in Figures 1A and 3. Device / global internal state 157 includes one or more of the following: active application state, indicating which applications (if any) are currently active; display state, indicating what applications, views, or other information occupy the various areas of the touch-sensitive display system 112; sensor state, including information obtained from the device's various sensors and other input or control devices 116; and position and / or location information regarding the device's position and / or orientation.

[0073] Operating system 126 (e.g., iOS, Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or embedded operating systems such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitates communication between various hardware and software components.

[0074] The communication module 128 facilitates communication with other devices via one or more external ports 124 and includes various software components for processing data received by the RF circuitry 108 and / or the external ports 124. The external ports 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to be directly coupled to other devices or indirectly coupled via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is compatible with some iPhones from Apple Inc. (Cupertino, California). ® iPod Touch ® and iPad ® The device uses the same or similar and / or compatible multi-pin (e.g., 30-pin) connectors as the 30-pin connector used in the device. In some implementations, the external port is compatible with some iPhones from Apple Inc. (Cupertino, California). ® iPod Touch ® and iPad ®The device uses the same or similar and / or compatible Lightning connector. In some implementations, the external port is a USB Type-C connector that is the same or similar and / or compatible with the USB Type-C connector used in some electronic devices of Apple Inc. (Cupertino, California).

[0075] The contact / motion module 130 optionally detects contact with the touch-sensitive display system 112 (e.g., in conjunction with display controller 156) and other touch-sensitive devices (e.g., touchpads or physical click wheels). The contact / motion module 130 includes various software components for performing various operations related to contact detection (e.g., via a finger or stylus), such as determining whether contact has occurred (e.g., detecting a finger press event), determining the intensity of the contact (e.g., the force or pressure of the contact, or an alternative to force or pressure), determining whether there is movement of the contact and tracking movement across the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or a contact disconnection). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. These operations can optionally be applied to single-point contact (e.g., single-finger contact or stylus contact) or simultaneous multi-point contact (e.g., "multi-touch" / multi-finger contact). In some embodiments, the contact / motion module 130 and the display controller 156 detect contact on the touchpad.

[0076] The touch / motion module 130 optionally detects gesture input performed by the user. Different gestures on a touch-sensitive surface have different contact patterns (e.g., different movements, timings, and / or intensities of the detected contact). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a single-finger tap gesture includes detecting a finger press event, and then detecting a finger lift-off (lift-away) event at the same (or substantially the same) location as the finger press event (e.g., at the icon location). As another example, detecting a finger swipe gesture on a touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and subsequently detecting a finger lift-off (lift-away) event. Similarly, stylus taps, swipes, drags, and other gestures are optionally detected by detecting specific contact patterns of the stylus.

[0077] In some implementations, detecting a finger tap depends on the length of time between a finger press event and a finger lift event, but is independent of the intensity of finger contact during that time. In some implementations, a tap is detected based on the determination that the length of time between a finger press event and a finger lift event is less than a predetermined value (e.g., less than 0.1 seconds, 0.2 seconds, 0.3 seconds, 0.4 seconds, or 0.5 seconds), regardless of whether the intensity of finger contact during the tap reaches a given intensity threshold (greater than a nominal contact detection intensity threshold), such as a light press or deep press intensity threshold. Therefore, a finger tap can satisfy a specific input criterion that does not require the characteristic intensity of the contact to meet a given intensity threshold to satisfy that specific input criterion. For clarity, finger contact in a tap gesture typically needs to meet a nominal contact detection intensity threshold to detect a finger press event; below this threshold, no contact is detected. Similar analysis applies to detecting tap gestures via a stylus or other contact method. When the device is capable of detecting contact from a finger or stylus hovering above a touch-sensitive surface, the nominal contact detection strength threshold may optionally not correspond to the physical contact between the finger or stylus and the touch-sensitive surface.

[0078] The same concept applies to other types of gestures in a similar manner. For example, swipe gestures, pinch gestures, spread gestures, and / or long press gestures can be optionally detected based on criteria that are independent of the intensity of the contact involved in the gesture or do not require one or more contacts performing the gesture to reach an intensity threshold for recognition. For example, a swipe gesture is detected based on the amount of movement of one or more contacts; a pinch gesture is detected based on the movement of two or more contacts toward each other; a spread gesture is detected based on the movement of two or more contacts away from each other; and a long press gesture is detected based on the duration of contact with less than a threshold amount of movement on a touch-sensitive surface. Therefore, the statement that a particular gesture recognition criterion does not require the contact intensity to meet a corresponding intensity threshold implies that a particular gesture recognition criterion can be met when the contact in the gesture does not reach the corresponding intensity threshold, and also when one or more contacts in the gesture reach or exceed the corresponding intensity threshold. In some implementations, tap gestures are detected based on determining that a finger press event and a finger lift event are detected within a predefined time period, regardless of whether the contact is above or below a corresponding intensity threshold during the predefined time period. Similarly, swipe gestures are detected based on determining that the contact movement is greater than a predefined amount, even if the contact movement ends above a corresponding intensity threshold. Even in specific implementations where gesture detection is affected by the intensity of the contact performing the gesture (e.g., the device detects a long press faster when the contact intensity is above an intensity threshold, or delays the detection of a tap input when the contact intensity is even higher), the detection of these gestures does not require the contact to reach a specific intensity threshold (e.g., even if the amount of time required to recognize the gesture varies).

[0079] In some cases, contact intensity thresholds, duration thresholds, and movement thresholds are combined in various different combinations to create heuristic algorithms that distinguish between two or more different gestures targeting the same input element or region, allowing for a richer set of user interactions and responses from multiple different interactions with the same input element. Statements that a particular set of gesture recognition criteria does not require the intensity of one or more contacts to meet a corresponding intensity threshold in order to satisfy a particular gesture recognition criterion do not preclude the simultaneous evaluation of other intensity-related gesture recognition criteria to identify other gestures that meet criteria when the gesture includes contact with an intensity higher than the corresponding intensity threshold. For example, in some cases, a first gesture recognition criterion for a first gesture (which does not require the contact intensity to meet a corresponding intensity threshold to satisfy the first gesture recognition criterion) competes with a second gesture recognition criterion for a second gesture (which depends on the contact reaching the corresponding intensity threshold). In such competition, if the second gesture recognition criterion for the second gesture is satisfied first, the gesture is optionally not recognized as satisfying the first gesture recognition criterion for the first gesture. For example, if the contact reaches the corresponding intensity threshold before the contact moves a predefined amount of movement, a deep press gesture is detected instead of a swipe gesture. Conversely, if the contact moves a predefined amount of motion before reaching the corresponding intensity threshold, a swipe gesture is detected instead of a deep press gesture. Even in such cases, the first gesture recognition criterion for the first gesture still does not require the contact intensity to meet the corresponding intensity threshold to satisfy the first gesture recognition criterion, because if the contact remains below the corresponding intensity threshold until the gesture ends (e.g., a swipe gesture with a contact intensity that does not increase to above the corresponding intensity threshold), the gesture will be recognized as a swipe gesture by the first gesture recognition criterion. Therefore, a specific gesture recognition criterion that does not require the contact intensity to meet the corresponding intensity threshold to satisfy a specific gesture recognition criterion will (A) in some cases ignore the contact intensity relative to the intensity threshold (e.g., for a tap gesture) and / or (B) in some cases fail to satisfy the specific gesture recognition criterion (e.g., for a long press gesture) if a set of competing intensity-related gesture recognition criteria (e.g., for a deep press gesture) recognize the input as corresponding to an intensity-related gesture before the specific gesture recognition criterion recognizes the gesture corresponding to the input, in this sense, still depend on the contact intensity relative to the intensity threshold (e.g., for a long press gesture that competes with a deep press gesture for recognition).

[0080] The graphics module 132 includes various known software components for rendering and displaying graphics on the touch-sensitive display system 112 or other displays, including components for altering the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual properties). As used herein, the term "graphics" includes any object that can be displayed to a user, and non-limitingly includes text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.

[0081] In some implementations, the graphics module 132 stores data representing the graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes from applications, etc., specifying the graphics to be displayed, and also receives coordinate data and other graphic attribute data if necessary, and then generates screen image data for output to the display controller 156.

[0082] The haptic feedback module 133 includes various software components for generating instructions (e.g., instructions used by the haptic feedback controller 161) to generate haptic output at one or more locations on the device 100 using the haptic output generator 167 in response to user interaction with the device 100.

[0083] The text input module 134 (which is optionally a component of the graphics module 132) provides a soft keyboard for entering text in various applications, such as the contact module 137, email module 140, IM module 141, browser module 147, and any other application that requires text input.

[0084] GPS module 135 determines the location of the device and provides that information for use in various applications (e.g., to telephone module 138 for location-based dialing; to camera module 143 as image / video metadata; and to applications that provide location-based services such as weather widgets, local yellow pages widgets, and map / navigation widgets).

[0085] Application 136 optionally includes the following modules (or instruction sets) or subsets or supersets thereof: • Contacts module 137 (sometimes called address book or contact list); • Telephone module 138; • Video conferencing module 139; • Email client module 140; • Instant Messaging (IM) module 141; • Fitness support module 142; • Camera module 143 for still images and / or video images; • Image management module 144; • Browser module 147; • Calendar module 148; • Widget module 149, which optionally includes one or more of the following: weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, dictionary widget 149-5, and other widgets acquired by the user, and user-created widgets 149-6. • Widget creator module 150 for creating user-created widgets 149-6; • Search module 151; • A video and music player module 152, optionally composed of a video player module and a music player module; • Notes module 153; • Map module 154; and / or • Online video module 155.

[0086] Examples of other applications 136 that may be optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, rendering applications, Java-enabled applications, encryption, digital rights management, speech recognition, and speech duplication.

[0087] In conjunction with the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, the contact module 137 includes executable instructions for managing an address book or contact list (e.g., in the application internal state 192 of the contact module 137 stored in memory 102 or memory 370), including: adding names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers or email addresses to initiate and / or facilitate communications via the telephone module 138, video conferencing module 139, email module 140, or IM module 141; and so on.

[0088] Combining RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, telephone module 138 includes executable instructions for performing the following operations: inputting a character sequence corresponding to a telephone number, accessing one or more telephone numbers in the address book 137, modifying an input telephone number, dialing a corresponding telephone number, initiating a conversation, and disconnecting or hanging up when the conversation is complete. As described above, wireless communication optionally employs any of a variety of communication standards, protocols, and technologies.

[0089] Combining RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch-sensitive display system 112, display controller 156, one or more optical sensors 164, optical sensor controller 158, contact module 130, graphics module 132, text input module 134, contact list 137, and telephone module 138, video conferencing module 139 includes executable instructions to initiate, conduct, and terminate video conferences between the user and one or more other participants based on user instructions.

[0090] Incorporating RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, email client module 140 includes executable instructions for creating, sending, receiving, and managing emails in response to user commands. Combined with image management module 144, email client module 140 makes it very easy to create and send emails containing still images or video images captured by camera module 143.

[0091] In conjunction with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, instant messaging module 141 includes executable instructions for performing the following operations: inputting a character sequence corresponding to an instant message, modifying previously input characters, sending a corresponding instant message (e.g., using Short Message Service (SMS) or Multimedia Messaging Service (MMS) protocols for telephone-based instant messaging or using XMPP, SIMPLE, Apple Push Notification Services (APNs), or IMPS for internet-based instant messaging), receiving an instant message, and viewing received instant messages. In some embodiments, the sent and / or received instant messages optionally include graphics, photographs, audio files, video files, and / or other attachments supported in MMS and / or Enhanced Messaging Services (EMS). As used herein, "instant message" refers to both telephone-based messages (e.g., messages delivered using SMS or MMS) and internet-based messages (e.g., messages delivered using XMPP, SIMPLE, APNs, or IMPS).

[0092] Incorporating RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and video and music player module 152, fitness support module 142 includes executable instructions for creating fitness (e.g., with time, distance, and / or calorie burning goals); communicating with fitness sensors (e.g., in sports equipment and smartwatches); receiving fitness sensor data; calibrating sensors used to monitor fitness; selecting and playing music for fitness; and displaying, storing, and transmitting fitness data.

[0093] In conjunction with the touch-sensitive display system 112, display controller 156, optical sensor 164, optical sensor controller 158, contact module 130, graphics module 132, and image management module 144, camera module 143 includes executable instructions for capturing still images or videos (e.g., including video streams) and storing them in memory 102, modifying the characteristics of still images or videos, and / or deleting still images or videos from memory 102.

[0094] In conjunction with the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, and camera module 143, the image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, tagging, deleting, displaying (e.g., in a digital slideshow or photo album), and storing still images and / or video images.

[0095] Combining RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, browser module 147 includes executable instructions for browsing the Internet (including searching, linking to, receiving, and displaying web pages or portions thereof, as well as attachments and other files linked to web pages) according to user instructions.

[0096] Incorporating RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, email client module 140, and browser module 147, calendar module 148 includes executable instructions for creating, displaying, modifying, and storing calendars and associated data (e.g., calendar entries, to-dos, etc.) according to user instructions.

[0097] In conjunction with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, and browser module 147, widget module 149 is optionally a mini-application downloaded and used by a user (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) or a user-created mini-application (e.g., user-created widget 149-6). In some embodiments, the widget includes HTML (Hypertext Markup Language) files, CSS (Cascading Style Sheets) files, and JavaScript files. In some embodiments, the widget includes XML (Extensible Markup Language) files and JavaScript files (e.g., Yahoo! widgets).

[0098] In conjunction with RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, and browser module 147, widget creator module 150 includes executable instructions for creating widgets (e.g., transferring user-specified portions of a webpage into a widget).

[0099] In conjunction with the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, the search module 151 includes executable instructions for searching the memory 102 for text, music, sound, images, videos, and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) according to user instructions.

[0100] Incorporating the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, the video and music player module 152 includes executable instructions allowing users to download and play back recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing back video (e.g., on the touch-sensitive display system 112, or on an external display connected wirelessly or via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0101] In conjunction with the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, and text input module 134, the note-taking module 153 includes executable instructions for creating and managing notes, to-do items, etc., according to user instructions.

[0102] Combining RF circuitry 108, touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, map module 154 includes executable instructions for receiving, displaying, modifying, and storing maps and map-related data (e.g., driving directions; data on shops and other points of interest at or near specific locations; and other location-based data) according to user instructions.

[0103] In conjunction with the touch-sensitive display system 112, display controller 156, contact module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, email client module 140, and browser module 147, the online video module 155 includes executable instructions that allow users to access, browse, receive (e.g., via streaming and / or downloading), play back (e.g., on the touchscreen 112, or on an external display connected wirelessly or via external port 124), send emails with links to specific online videos, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, an instant messaging module 141 is used instead of the email client module 140 to send links to specific online videos.

[0104] Each module and application identified above corresponds to a set of executable instructions for performing one or more of the functions described above, as well as the methods described in this application (e.g., computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) need not be implemented as standalone software programs, processes, or modules; therefore, various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. In some embodiments, memory 102 optionally stores a subset of the modules and data structures described above. Furthermore, memory 102 optionally stores additional modules and data structures not described above.

[0105] In some implementations, device 100 is a device on which the operation of a predefined set of functions is performed solely via a touchscreen and / or touchpad. By using a touchscreen and / or touchpad as the primary input control device for operating device 100, the number of physical input control devices (such as push-buttons and dial pads) on device 100 is optionally reduced.

[0106] A predefined set of functions, uniquely performed via a touchscreen and / or touchpad, optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 from any user interface displayed on device 100 to the main menu, main desktop menu, or root menu. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is a physical push-button or other physical input control device, rather than a touchpad.

[0107] Figure 1B This is a block diagram illustrating example components for event handling according to some implementation schemes. In some implementations, memory 102 (e.g., as...) Figure 1A (as shown) or 370 (e.g., as shown) Figure 3A (As shown) includes an event classifier 170 (e.g., in operating system 126) and a corresponding application 136-1 (e.g., any one of the aforementioned applications 136, 137-155, 380-390).

[0108] Event classifier 170 receives event information and determines the application 136-1 and application view 191 of application 136-1 to which the event information should be delivered. Event classifier 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192, which indicates one or more current application views displayed on touch-sensitive display system 112 when the application is active or executing. In some embodiments, device / global internal state 157 is used by event classifier 170 to determine which application(s) is currently active, and application internal state 192 is used by event classifier 170 to determine the application view 191 to which the event information should be delivered.

[0109] In some implementations, the application internal state 192 includes additional information such as one or more of the following: recovery information to be used when the application 136-1 resumes execution, user interface state information indicating that information is being displayed or ready to be displayed by the application 136-1, a state queue for enabling the user to return to the previous state or view of the application 136-1, and a repeat / undo queue for the user's previous actions.

[0110] Event monitor 171 receives event information from peripheral device interface 118. The event information includes information about sub-events, such as user touches on touch-sensitive display system 112 as part of a multi-touch gesture. Peripheral device interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 168, and / or microphone 113 (via audio circuitry 110). The information received by peripheral device interface 118 from I / O subsystem 106 includes information from touch-sensitive display system 112 or touch-sensitive surfaces.

[0111] In some implementations, the event monitor 171 sends requests to the peripheral device interface 118 at predetermined intervals. In response, the peripheral device interface 118 sends event information. In other implementations, the peripheral device interface 118 sends event information only when a significant event occurs (e.g., receiving input above a predetermined noise threshold and / or receiving input for a predetermined duration).

[0112] In some implementations, the event classifier 170 also includes a hit view determination module 172 and / or an activity event recognizer determination module 173.

[0113] When the touch-sensitive display system 112 displays more than one view, the hit view determination module 172 provides a software process for determining where a sub-event has occurred within one or more views. A view consists of controls and other elements that the user can see on the display.

[0114] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the corresponding application) in which a touch is detected optionally corresponds to a procedural level within the application's procedural or view hierarchy. For example, the lowest-level view in which a touch is detected is optionally referred to as the hit view, and the set of events identified as correct input is optionally determined, at least in part, based on the hit view of the initial touch that initiates the touch-based gesture.

[0115] The hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchical structure, the hit view determination module 172 identifies the hit view as the lowest-level view in the hierarchical structure from which the sub-events should be processed. In most cases, the hit view is the lowest-level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events forming an event or potential event) occurs. Once the hit view is identified by the hit view determination module, the hit view typically receives all sub-events related to the same touch or input source to which it was identified as the hit view.

[0116] The activity event recognizer determination module 173 determines which views(s) within the view hierarchy should receive a specific sub-event sequence. In some embodiments, the activity event recognizer determination module 173 determines that only the hit view should receive the specific sub-event sequence. In other embodiments, the activity event recognizer determination module 173 determines that all views including the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive the specific sub-event sequence. In other embodiments, even if the touch sub-event is entirely confined to the area associated with a particular view, higher views in the hierarchy will still remain actively participating views.

[0117] Event assigner module 174 assigns event information to event identifiers (e.g., event identifier 180). In embodiments that include active event identifier determination module 173, event assigner module 174 delivers event information to the event identifier determined by active event identifier determination module 173. In some embodiments, event assigner module 174 stores event information in an event queue, which is retrieved by the corresponding event receiver module 182.

[0118] In some embodiments, operating system 126 includes event classifier 170. Alternatively, application 136-1 includes event classifier 170. In yet another embodiment, event classifier 170 is a separate module or part of another module (such as contact / motion module 130) stored in memory 102.

[0119] In some implementations, application 136-1 includes a plurality of event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events occurring within a corresponding view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, a corresponding application view 191 includes a plurality of event recognizers 180. In other implementations, one or more of the event recognizers 180 are part of a separate module, such as a user interface toolkit or a higher-level object from which application 136-1 inherits methods and other properties. In some implementations, a corresponding event handler 190 includes one or more of the following: a data updater 176, an object updater 177, a GUI updater 178, and / or event data 179 received from an event classifier 170. Event handlers 190 optionally utilize or invoke the data updater 176, the object updater 177, or the GUI updater 178 to update the application's internal state 192. Alternatively, one or more application views in application view 191 include one or more corresponding event handlers 190. Additionally, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included in the corresponding application view 191.

[0120] The corresponding event identifier 180 receives event information (e.g., event data 179) from the event classifier 170 and identifies events from the event information. The event identifier 180 includes an event receiver 182 and an event comparator 184. In some embodiments, the event identifier 180 also includes at least one subset of metadata 183 and event delivery instructions 188 (which optionally include sub-event delivery instructions).

[0121] Event receiver 182 receives event information from event classifier 170. The event information includes information about sub-events, such as touch or touch movement. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves touch movement, the event information optionally also includes the speed and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another (e.g., from a portrait orientation to a lateral orientation, or vice versa), and the event information includes corresponding information about the device's current orientation (also referred to as device orientation).

[0122] Event comparator 184 compares event information with predefined event or sub-event definitions and determines the event or sub-event based on the comparison, or determines or updates the state of the event or sub-event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 contains definitions of events (e.g., predefined sequences of sub-events), such as event 1 (187-1), event 2 (187-2), and others. In some embodiments, sub-events in event 187 include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, event 1 (187-1) is defined as a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on the displayed object, a first lift of a predetermined duration (touch end), a second touch (touch start) of a predetermined duration on the displayed object, and a second lift of a predetermined duration (touch end). In another example, event 2 (187-2) is defined as a drag on a displayed object. For example, dragging includes a touch (or contact) on a displayed object for a predetermined duration, movement of the touch on the touch-sensitive display system 112, and lifting off the touch (end of touch). In some embodiments, the event also includes information for one or more associated event handlers 190.

[0123] In some implementations, event definition 187 includes definitions of events for corresponding user interface objects. In some implementations, event comparator 184 performs a hit test to determine which user interface object is associated with the sub-event. For example, in an application view displaying three user interface objects on a touch-sensitive display system 112, when a touch is detected on the touch-sensitive display system 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects the event handler associated with the sub-event and the object that triggered the hit test.

[0124] In some implementations, the definition of the corresponding event 187 also includes delay actions that delay the delivery of event information until it has been determined whether the sub-event sequence actually corresponds to or does not correspond to the event type of the event recognizer.

[0125] When the corresponding event recognizer 180 determines that the sub-event sequence does not match any event in event definition 186, the corresponding event recognizer 180 enters an event impossible, event failed, or event ended state, after which subsequent sub-events based on touch gestures are ignored. In this case, other event recognizers (if any) that remain active in the hit view continue to track and process the ongoing sub-events based on touch gestures.

[0126] In some embodiments, the corresponding event recognizer 180 includes metadata 183 having configurable attributes, flags, and / or lists instructing how the event delivery system should perform sub-event delivery to actively participating event recognizers. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists instructing how or how event recognizers can interact with each other. In some embodiments, the metadata 183 includes configurable attributes, flags, and / or lists instructing whether sub-events are delivered to different levels in a view or programmatic hierarchy.

[0127] In some implementations, when one or more specific sub-events of an event are identified, the corresponding event recognizer 180 activates the event handler 190 associated with the event. In some implementations, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from delivering (and deferred delivering) the sub-events to the corresponding hit view. In some implementations, the event recognizer 180 throws a flag associated with the identified event, and the event handler 190 associated with the flag acquires the flag and performs a predefined process.

[0128] In some implementations, event delivery instruction 188 includes a sub-event delivery instruction that delivers event information about a sub-event without activating an event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with the sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or the actively participating view receives the event information and performs a predetermined process.

[0129] In some implementations, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137, or stores video files used in video or music player module 152. In some implementations, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates the positioning of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and transmits that display information to graphics module 132 for display on a touch-sensitive display.

[0130] In some implementations, event handler 190 includes, or has access to, a data updater 176, an object updater 177, and a GUI updater 178. In some implementations, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other implementations, they are included in two or more software modules.

[0131] It should be understood that the above discussion regarding event handling for user touch on a touch-sensitive display also applies to other forms of user input used to operate the multifunction device 100 using an input device, and not all user input is initiated on the touchscreen. For example, mouse movement and mouse button presses optionally in conjunction with single or multiple keyboard presses or holds; touch movements on the touchpad, such as taps, drags, scrolls, etc.; stylus input; device movement; verbal commands; detected eye movements; biometric input; and / or any combination thereof may optionally be used as input corresponding to sub-events that define the event to be identified.

[0132] Figures 1C to 1DVarious examples of computer systems for performing methods and providing audio, visual, and / or haptic feedback as part of the user interface described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., first display assembly 1-120a and second display assembly 1-120b) for displaying a representation of visual elements and / or physical environment to a user of the computer system, the representation of the visual elements and / or physical environment optionally being generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses to make it easier for a user who would otherwise use glasses or contact lenses to correct their vision to view the user interface, the one or more corrective lenses optionally being removably attached to one or more optical modules in an optical module. While many user interfaces illustrated herein illustrate a single view of the user interface, HMD (e.g., HMD) The user interface in 100b) optionally uses two optical modules (e.g., a first display assembly 1-120a and a second display assembly 1-120b) for display, one optical module for the user's right eye and a different optical module for the user's left eye, and presents slightly different images to the two different eyes to generate the illusion of stereoscopic depth. A single view of the user interface is typically a right-eye view or a left-eye view, and the depth effect is explained in text or using other diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display assembly 1-108) for displaying status information of the computer system to the user of the computer system (when the computer system is not worn) and / or to other persons near the computer system, the status information optionally based on detected events and / or by the computer system. The audio feedback is generated based on detected user input. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, which may optionally be generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device, which may be used (optionally in conjunction with one or more illuminators) to generate digital pass-through images, capture visual media corresponding to the physical environment (e.g., photographs and / or videos), or determine the pose (e.g., positioning and / or orientation) of physical objects and / or surfaces in the physical environment, such that virtual objects can be placed based on the detected pose of the physical objects and / or surfaces.In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand positioning and / or movement, which can be used (optionally in conjunction with one or more illuminators) to determine when one or more air gestures have been performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement, which can be used (optionally in conjunction with one or more lights) to determine attention or gaze positioning and / or gaze movement, which can optionally be used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine user facial expressions and / or hand movements for generating a user avatar or representation, such as an anthropomorphic avatar or representation for real-time communication sessions, wherein the avatar has facial expressions, hand movements, and / or body movements detected by the user based on or similar devices. Gaze and / or attention information may optionally be combined with hand tracking information to determine the interaction between the user and one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first buttons 1-128 and / or second buttons 1-132), knobs (e.g., first buttons 1-128), digital crowns (e.g., pressable and twistable or rotatable first buttons 1-128), touchpads, touchscreens, keyboards, mice, and / or other input devices. One or more buttons (e.g., first buttons 1-128 and / or second buttons 1-132) may optionally be used to perform system operations, such as recentering content in the user-visible 3D environment of the device, displaying the main user interface for launching an application, initiating a real-time communication session, or initiating the display of a virtual 3D background. A knob or digital crown (e.g., a pressable and twistable or rotatable first button 1-128) may optionally be rotatable to adjust parameters of the visual content, such as the level of immersion of the virtual 3D environment (e.g., the extent to which the virtual content occupies the user's viewport in the 3D environment) or other parameters associated with the 3D environment and the virtual content displayed via optical modules (e.g., the first display assembly 1-120a and the second display assembly 1-120b).

[0133] Figure 1CA front top perspective view illustrates an example of a head-mounted display (HMD) device 1-100 configured for wear by a user and to provide virtual and altered / mixed reality (VR / AR) experiences. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strip assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 secured at either end to the electronic strip assembly 1-104. The electronic strip assembly 1-104 and the strap 1-106 may be part of a retention assembly configured to wrap around the user's head to hold the display unit 1-102 against the user's face.

[0134] In at least one example, the strap assembly 1-106 may include a first strap 1-116 configured to wrap around the back of the user's head and a second strap 1-117 configured to extend above the top of the user's head. As shown, the second strap may extend between the first electronic strip 1-105a and the second electronic strip 1-105b of the electronic strip assembly 1-104. The strip assembly 1-104 and the strap assembly 1-106 may be part of a fixing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.

[0135] In at least one example, the fixing mechanism includes a first electronic strip 1-105a, which includes a first proximal end 1-134 coupled to the display unit 1-102 (e.g., the housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite to the first proximal end 1-134. The fixing mechanism may also include a second electronic strip 1-105b, which includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite to the second proximal end 1-138. The fixing mechanism may also include a first strip 1-116 and a second strip 1-117, the first strip including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second strip extending between the first electronic strip 1-105a and the second electronic strip 1-105b. Strips 1-105a-b and strip 1-116 may be coupled via a connecting mechanism or assembly 1-114. In at least one example, the second strip 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between a first proximal end 1-134 and a first distal end 1-136, and a second end 1-148 coupled to the second electronic strip 1-105b between a second proximal end 1-138 and a second distal end 1-140.

[0136] In at least one example, the first and second electronic strips 1-105a-b comprise plastic, metal, or other structural materials forming the substantially rigid shape of the strips 1-105a-b. In at least one example, the first and second strips 1-116, 1-117 are formed of an elastic flexible material (including woven textiles, rubber, etc.). The first strip 1-116 and the second strip 1-117 may be flexible enough to conform to the shape of the user's head when wearing the HMD 1-100.

[0137] In at least one example, one or more of the first and second electronic stripes 1-105a-b may define an inner strip volume and include one or more electronic components disposed within the inner strip volume. In one example, as Figure 1C As shown, the first electronic strip 1-105a may include electronic components 1-112. In one example, electronic components 1-112 may include a speaker. In another example, electronic components 1-112 may include computing components, such as a processor.

[0138] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is located in... Figure 1C The section marked 1-152 with dashed lines is because the display assembly 1-108 is configured to obscure the first opening 1-152 when the HMD 1-100 is assembled, as viewed from above. The housing 1-150 may also define a rearward second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover disposed in or across the front opening 1-152 to obscure the front opening 1-152 and a display screen (shown in other figures). In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 in general, has a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 can be bent as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, wherein the display unit 1-102 is pressed.

[0139] In at least one example, the housing 1-150 may define a first hole 1-126 between a first opening 1-152 and a second opening 1-154, and a second hole 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first hole 1-128, and a second button 1-132 disposed in the second hole 1-130. The first button 1-128 and the second button 1-132 are pressable through their respective holes 1-126 and 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 may be a rotary dial and a pressable button. In at least one example, the first button 1-128 is a pressable and rotary dial button, and the second button 1-132 is a pressable button.

[0140] Figure 1D A rear perspective view of HMD 1-100 is illustrated. HMD 1-100 may include a light seal 1-110 extending rearwardly around the periphery of housing 1-150 of display assembly 1-108, as shown. The light seal 1-110 may be configured to extend from housing 1-150 to the user's face, surrounding the user's eyes, to block external light from being visible. In one example, HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b disposed at or within a rearwardly facing second opening 1-154 defined by housing 1-150 and / or disposed within the internal volume of housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include corresponding display screens 1-122a, 1-122b configured to project light toward the user's eyes in a rearward direction through the second opening 1-154.

[0141] In at least one example, reference Figure 1C and Figure 1D Both, the display assembly 1-108 can be a front-facing display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a-b can be configured to project light in a second rearward direction opposite to the first direction. As described above, the light seal 1-110 can be configured to block light from outside the HMD 1-100 from reaching the user's eyes, including a component made of... Figure 1CThe front perspective view shows the light projected onto the forward-facing display screen of the display assembly 1-108. In at least one example, the HMD 1-100 may also include a curtain 1-124 that blocks the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a-b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.

[0142] Figure 2 Examples are given of devices with touchscreens (e.g., according to some implementation schemes). Figure 1A A portable multi-functional device 100 (a touch-sensitive display system 112). The touchscreen optionally displays one or more graphics within a user interface (UI) 200. In these embodiments and other embodiments described below, a user can select one or more graphics by gesturing over the graphics, for example, using one or more fingers 202 (not drawn to scale in the figures) or one or more styluses 203 (not drawn to scale in the figures). In some embodiments, selection of one or more graphics occurs when the user breaks contact with one or more graphics. In some embodiments, gestures optionally include one or more taps, one or more swipes (from left to right, from right to left, up and / or down), and / or scrolling (from right to left, from left to right, up and / or down) of a finger already in contact with the device 100. In some specific embodiments or in some cases, unintentional contact with a graphic does not select the graphic. For example, a swipe gesture over an application icon optionally does not select the corresponding application when the gesture corresponding to selection is a tap.

[0143] Device 100 optionally also includes one or more physical buttons, such as a "main desktop" or menu button 204. As previously described, menu button 204 is optionally used to navigate to any application 136 of a set of applications optionally executed on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on a touchscreen display, or as a system gesture such as a swipe up the edge.

[0144] In some embodiments, device 100 includes a touchscreen display, a menu button 204 (sometimes referred to as a home button 204), a push-button 206 for powering on / off the device and locking the device, a volume control button 208, a SIM card slot 210, a headset jack 212, and / or a docking / charging external port 124. The push-button 206 is optionally used to: power on / off the device by pressing the button and holding it in the pressed state for a predefined time interval; lock the device by pressing the button and releasing it before the predefined time interval has elapsed; and / or unlock the device or initiate an unlocking process. In some embodiments, device 100 also accepts voice input via microphone 113 for activating or deactivating certain functions. Device 100 also optionally includes one or more contact strength sensors 165 for detecting contact strength on the touch-sensitive display system 112, and / or one or more haptic output generators 167 for generating haptic outputs for a user of device 100.

[0145] Figure 3A This is a block diagram of an example multifunctional device with a display and a touch-sensitive surface according to some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a children's learning toy), gaming system, or control device (e.g., a home controller or industrial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 for interconnecting these components. The communication bus 320 optionally includes circuitry (sometimes referred to as a chipset) that interconnects system components and controls communication between system components. Device 300 includes an input / output (I / O) interface 330 with a display 340, which is typically a touchscreen display. The I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, and a haptic output generator 357 for generating haptic output on device 300 (e.g., similar to the reference above). Figure 1A The described tactile output generator 167) and sensor 359 (e.g., optical or camera sensor, accelerometer, proximity sensor, touch sensor, and / or similar to those described above) are referenced in the reference. Figure 1AThe contact strength sensor 165 described is a contact strength sensor. Memory 370 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 370 optionally includes one or more storage devices located remotely from CPU 310. In some embodiments, memory 370 stores information related to portable multifunction devices 100 (e.g., such as…). Figure 1A The memory 370 stores programs, modules, and data structures similar to those in the memory 102 of the portable multifunction device 100, or subsets thereof. Additionally, the memory 370 optionally stores additional programs, modules, and data structures not present in the memory 102 of the portable multifunction device 100. For example, the memory 370 of the device 300 optionally stores a drawing module 380, a rendering module 382, ​​a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while the portable multifunction device 100 (e.g., as shown) stores additional programs, modules, and data structures not present in the memory 102 of the portable multifunction device 100. Figure 1A The memory 102 (as shown) optionally does not store these modules.

[0146] Wireless interface 381 receives and transmits wireless signals. Wireless interface 381 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with communication networks and other communication devices via electromagnetic signals. Wireless interface 381 optionally includes well-known circuitry for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, CODEC chipsets, subscriber identity module (SIM) cards, memory, etc. Wireless interface 381 optionally communicates wirelessly with networks (such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs))) and other devices.

[0147] Figure 3A Each of the elements identified above is optionally stored in one or more of the previously mentioned memory devices. Each of the modules identified above corresponds to an instruction set for performing the functions described above. The modules or programs identified above (i.e., instruction sets) need not be implemented as standalone software programs, processes, or modules; therefore, various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. In some embodiments, memory 370 optionally stores a subset of the modules and data structures described above. Furthermore, memory 370 optionally stores additional modules and data structures not described above.

[0148] Figure 3B Physical features of an example wearable audio output device 301 according to some embodiments are illustrated. In some embodiments, the wearable audio output device 301 is one or more in-ear headphones, earbuds, over-ear headphones, etc. Figure 3B In the example, the wearable audio output device 301 is an earbud. In some embodiments, the wearable audio output device 301 includes a head portion 303 and a stem portion 305. In some embodiments, the head portion 303 is configured to be inserted into a user's ear. In some embodiments, the stem portion 305 physically extends from the head portion 303 (e.g., is an elongated portion extending from the head portion 303). For example, when the head portion 303 is inserted into the user's ear, the head portion 303 physically extends downward, in front of and / or past the user's earlobe.

[0149] In some embodiments, the wearable audio output device 301 includes one or more audio speakers 306 (e.g., in the head portion 303) for providing audio output (e.g., to a user's ear). In some embodiments, the wearable audio output device 301 includes one or more placement sensors 304 (e.g., placement sensors 304-1 and 304-2 in the head portion 303) to detect the positioning or placement of the wearable audio output device 301 relative to the user's ear, such as detecting placement of the wearable audio output device 301 in the user's ear.

[0150] In some embodiments, the wearable audio output device 301 includes one or more microphones 302 for receiving audio input. In some embodiments, one or more microphones 302 (e.g., microphone 302-1) are included in a head portion 303. In some embodiments, one or more microphones 302 (e.g., microphone 302-2) are included in a stem portion 305. In some embodiments, the microphones 302 detect voice from a user wearing the wearable audio output device 301 and / or ambient noise around the wearable audio output device 301. In some embodiments, multiple microphones of the microphones 302 are positioned at different locations on the wearable audio output device 301 to measure voice and / or ambient noise at different locations around the wearable audio output device 301.

[0151] In some embodiments, the wearable audio output device 301 includes one or more input devices 308 (e.g., in the handle portion 305). In some embodiments, the input device 308 includes a pressure-sensitive (e.g., intensity-sensitive) input device. In some embodiments, the pressure-sensitive input device detects input from the user in response to the user squeezing the input device (e.g., by pinching the handle portion 305 of the wearable audio output device 301 between two fingers). In some embodiments, the input device 308 includes a touch-sensitive surface (e.g., a capacitive sensor) for detecting touch input, an accelerometer and / or a posture sensor (e.g., for determining the posture of the wearable audio output device 301 relative to the physical environment and / or changes in device posture) and / or other input devices through which the user can interact with and provide input to the wearable audio output device 301. In some embodiments, the input device 308 includes one or more capacitive sensors, one or more force sensors, one or more motion sensors, and / or one or more orientation sensors. Figure 3B An input device 308 is shown at a location within the handle portion 305; however, in some embodiments, one or more of the input devices 308 are located at other locations within the wearable audio output device 301 (e.g., other locations within the handle portion 305 and / or the head portion 303). In some embodiments, the wearable audio output device 301 includes a housing having one or more physical differentiating portions 307 at locations corresponding to the input devices 308 (e.g., to assist the user in locating and / or interacting with the input devices 308). In some embodiments, the physical differentiating portions 307 include indentations, protrusions, and / or portions with different textures. In some embodiments, the physical differentiating portions 307 include a single differentiating portion spanning multiple input devices 308. For example, input devices 308 include a set of touch sensors configured to detect swipe gestures, and a single differentiating portion (e.g., a recess or groove) spans the set of touch sensors. In some embodiments, the physical differentiating portions 307 include a corresponding differentiating portion for each input device of the input devices 308.

[0152] In some embodiments, the wearable audio output device 301 includes one or more sensors 311 (e.g., sensors 311-1 and 311-2 in the handle portion 305). In some embodiments, the one or more sensors 311 include one or more image sensors or cameras. In some embodiments, the sensors 311 include a forward-facing sensor (e.g., sensor 311-1) when the wearable audio output device 301 is worn by a user. In some embodiments, the sensors 311 include a rear-facing sensor (e.g., sensor 311-2) when the wearable audio output device 301 is worn by a user. In some embodiments, the sensors 311 consist of a single sensor (e.g., having substantially the same field of view as the wearer of the wearable audio output device 301). In some embodiments, the sensors 311 include three or more sensors (e.g., each sensor has a different field of view). In some embodiments, one or more sensors of the sensors 311 are arranged in relation to... Figure 3B The different locations are shown. For example, one of the sensors in sensor 311 may be located on the head portion 303. As another example, one of the sensors in sensor 311 may be located near the middle or top of the handle portion 305.

[0153] Figure 3CThis is a block diagram of an example wearable audio output device 301 according to some embodiments. In some embodiments, the wearable audio output device 301 is one or more in-ear headphones, earbuds, over-ear headphones, etc. In some examples, the wearable audio output device 301 includes a pair of headphones or earbuds (e.g., one headphone or earbud for each ear in a user's ear). In some examples, the wearable audio output device 301 includes over-ear headphones (e.g., headphones with two over-ear earcups for placement over a user's ears and optionally connected via a headband). In some embodiments, the wearable audio output device 301 includes one or more audio speakers 306 for providing audio output (e.g., to a user's ears). In some embodiments, the wearable audio output device 301 includes one or more placement sensors 304 to detect the positioning or placement of the wearable audio output device 301 relative to a user's ear, such as detecting placement of the wearable audio output device 301 in a user's ear. In some embodiments, the wearable audio output device 301 conditionally outputs audio based on whether it is in or near the user's ear (e.g., abandoning audio output when it is not in the user's ear to reduce power consumption). In some embodiments where the wearable audio output device 301 includes multiple (e.g., a pair) wearable audio output components (e.g., headphones, earbuds, or earmuffs), each component includes one or more corresponding placement sensors, and the wearable audio output device 301 conditionally outputs audio based on whether one or both components are in or near the user's ear, as described herein. In some embodiments, the wearable audio output device 301 also includes an internal rechargeable battery 309 for providing power to the various components of the wearable audio output device 301.

[0154] In some embodiments, the wearable audio output device 301 includes an audio I / O logic component 312 that determines the positioning or placement of the wearable audio output device 301 relative to a user's ear based on information received from a self-placement sensor 304, and in some embodiments, the audio I / O logic component 312 controls the resulting conditional audio output. In some embodiments, the wearable audio output device 301 includes features for communication with one or more multi-functional devices (such as device 100, e.g., such as...). Figure 1A (as shown) or device 300 (e.g., such as Figure 3A Interface 315 (e.g., a wireless interface) for communication with devices such as device 100 (e.g., as shown). In some embodiments, interface 315 includes an interface for communication with multi-functional devices such as device 100 (e.g., as shown). Figure 1A (as shown) or device 300 (e.g., such as Figure 3AThe wearable audio output device 301 is connected via a wired interface (e.g., via a headphone jack or other audio port). In some embodiments, a user can interact with and provide input to the wearable audio output device 301 via interface 315 (e.g., remotely). In some embodiments, the wearable audio output device 301 communicates with multiple devices (e.g., multiple multifunction devices and / or audio output device housings), and the audio I / O logic unit 312 determines from which multifunction device to receive instructions for outputting audio.

[0155] In some embodiments, the wearable audio output device 301 includes one or more microphones 302 for receiving audio input. In some embodiments where the wearable audio output device 301 includes multiple (e.g., a pair) wearable audio output components (e.g., headphones or earbuds), each component includes one or more corresponding microphones. In some embodiments, the audio I / O logic component 312 detects or identifies speech or ambient noise based on information received from the microphones 302.

[0156] In some embodiments, the wearable audio output device 301 includes one or more input devices 308. In some embodiments where the wearable audio output device 301 includes multiple (e.g., a pair) wearable audio output components (e.g., headphones, earbuds, or earmuffs), each component includes one or more corresponding input devices. In some embodiments, the input device 308 includes one or more volume control hardware elements (e.g., up / down buttons for volume control, or as referenced herein) for (e.g., local) volume control of the wearable audio output device 301. Figure 1A The increase button and the separate decrease button are mentioned. In some embodiments, input provided via one or more input devices 308 is processed by the audio I / O logic unit 312. In some embodiments, the audio I / O logic unit 312 is associated with a separate device (e.g., Figure 1A Equipment 100 or Figure 3A The device 300 communicates with the independent device, which provides instructions or content for audio output and optionally receives and processes input (or information about the input) provided via microphone 302, placement sensor 304 and / or input device 308, or via one or more input devices of a separate device. In some embodiments, audio I / O logic component 312 is located in device 100 (e.g., as part of the device 100). Figure 1A (part of peripheral device interface 118) or device 300 (e.g., as part .... Figure 3A The I / O interface 330 is located in a portion of the device 100, rather than in the device 301, or alternatively partially in the device 100 and partially in the device 301, or partially in the device 300 and partially in the device 301.

[0157] Figure 3D Example audio control performed by a wearable audio output device 301 according to some embodiments is illustrated. While the following examples are explained with respect to specific embodiments of wearable audio output devices including earplugs with attachable replaceable ear tips (sometimes referred to as silicone ear tips or silicone seals), the methods, devices, and user interfaces described herein are equally applicable to specific embodiments in which the wearable audio output device does not have ear tips but each has a portion of a body shaped for insertion into a user's ear. In some embodiments, when a wearable audio output device with ear tips capable of attaching replaceable ear tips is worn in a user's ear, the ear tips and ear tips together act as a physical barrier, blocking at least some ambient sounds from the surrounding physical environment from reaching the user's ear. For example, in Figure 3D In this embodiment, the user wears a wearable audio output device 301 such that the head portion 303 and the earplug 314 are in the user's left ear. The earplug 314 extends at least partially into the user's ear canal. Preferably, when the head portion 303 and the earplug 314 are inserted into the user's ear, a seal is formed between the earplug 314 and the user's ear to isolate the user's ear canal from the surrounding physical environment. However, in some embodiments, the head portion 303 and the earplug 314 together block some, but not necessarily all, ambient sounds from the surrounding physical environment from reaching the user's ear. Therefore, in some embodiments, a first microphone (or, in some embodiments, a first group of one or more microphones) 302-1 (and optionally a third microphone 302-3) is located on the wearable audio output device 301 to detect ambient sounds in region 316 of the physical environment surrounding the head portion 303 (e.g., outside the earplug), represented by waveform 322. In some embodiments, (e.g., Figure 3C A second microphone (or, in some embodiments, a second set of one or more microphones) 302-2 is located on the wearable audio output device 301 to detect any ambient sounds, represented by waveform 324, that are not completely blocked by the head portion 303 and the earplug 314 and can be heard in the region 318 within the user's ear canal. Therefore, in some cases where the wearable audio output device 301 does not generate a noise-cancelling (also known as "inverting") audio signal to cancel (e.g., attenuate) ambient sounds from the surrounding physical environment (as indicated by waveform 326-1), the ambient sound waveform 324 can be perceived by the user (as indicated by waveform 328-1). In some cases where the wearable audio output device 301 generates an inverted audio signal to cancel ambient sounds (as indicated by waveform 326-2), the ambient sound waveform 324 is not perceived by the user (as shown by waveform 328-2).

[0158] In some implementations, the ambient sound waveform 322 is compared with the attenuated ambient sound waveform 324 (e.g., via the wearable audio output device 301 or components of the wearable audio output device 301 such as audio I / O logic unit 312, or via an electronic device communicating with the wearable audio output device 301) to determine the passive attenuation provided by the wearable audio output device 301. In some implementations, the amount of passive attenuation provided by the wearable audio output device 301 is taken into account when an inverted audio signal is provided to cancel ambient sound from the surrounding physical environment. For example, the inverted audio signal waveform 326-2 is configured to cancel the attenuated ambient sound waveform 324, rather than the unattenuated ambient sound waveform 322.

[0159] In some implementations, the wearable audio output device 301 is configured to operate in one of a variety of available audio output modes, such as an active noise control audio output mode, an active pass-through audio output mode, and a bypass audio output mode (sometimes also referred to as a noise control off audio output mode). In the active noise control mode (also known as “ANC”), the wearable audio output device 301 outputs one or more audio cancellation audio components (e.g., one or more inverted audio signals, also referred to as “audio cancellation audio components”) to at least partially cancel ambient sounds from the surrounding physical environment that would otherwise be perceived by the user. In the active pass-through audio output mode, the wearable audio output device 301 outputs one or more pass-through audio components (e.g., playing at least a portion of ambient sounds received by, for example, microphone 302-1, from outside the user’s ear), allowing the user to hear a greater amount of ambient sound from the surrounding physical environment than would otherwise be perceived by the user (e.g., a greater amount of ambient sound than would be heard using the passive attenuation of the wearable audio output device 301 placed in the user’s ear). In bypass mode, active noise management is turned off, so that the wearable audio output device 301 outputs neither any audio cancellation audio component nor any pass-through audio component (e.g., so that any amount of ambient sound perceived by the user is due to the physical attenuation of the wearable audio output device 301).

[0160] Figure 3E The physical features of an example wearable audio output device 301 according to some embodiments are illustrated. Figure 3E In one example, wearable audio output device 301 includes over-ear earmuffs worn on a user's ears, which act as a physical barrier blocking at least some ambient sounds from the surrounding physical environment from reaching the user's ears. For example, in Figure 3E In this embodiment, the wearable audio output device 301 is worn by the user, such that the earcup 317 is positioned on the user's left ear. In some implementations, (e.g., Figure 3BThe first microphone (or, in some embodiments, a first group of one or more microphones) 302-1 is located on the wearable audio output device 301 to detect ambient sound in an area 316 of the physical environment surrounding the earcups 317 (e.g., outside the earcups). In some embodiments, the earcups 317 block some, but not necessarily all, ambient sound from the surrounding physical environment from reaching the user's ears. In some embodiments, (e.g., Figure 3B A second microphone (or, in some embodiments, a second group of one or more microphones) 302-2 is located on the wearable audio output device 301 to detect any ambient sounds that are not completely blocked by the earcups 317 and can be heard in the area 318 within the earcups 317. Therefore, in some cases where the wearable audio output device 301 does not generate a noise-cancelling (also known as "inverting") audio signal to cancel (e.g., attenuate) ambient sounds from the surrounding physical environment, the ambient sound waveform 324 may be perceived by the user. In some cases where the wearable audio output device 301 generates an inverted audio signal to cancel ambient sounds, the ambient sound waveform 324 may not be perceived by the user.

[0161] Now let’s turn our attention to the implementation of the user interface (“UI”) optionally implemented on the portable multifunction device 100.

[0162] Figure 4A An example user interface for an application menu on a portable multifunction device 100 according to some embodiments is shown. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements or a subset or superset thereof: • One or more signal strength indicators for one or more wireless communications, such as cellular signals and Wi-Fi signals; • time; • Bluetooth indicator; • Battery status indicator; • Tray 408 with icons for frequently used applications, such as: ○ The telephone module 138 has an icon 416 labeled "telephone", which optionally includes an indicator 414 indicating the number of missed calls or voicemail messages; ○ An icon 418 labeled "Mail" in the email client module 140, which optionally includes an indicator 410 for the number of unread emails; ○ The icon 420 labeled "Browser" in browser module 147; and ○ The icon 422 labeled "Music" in the video and music player module 152; and • Icons of other applications, such as: ○ Icon 424 of IM module 141 marked as "Message"; ○ The icon 426 labeled "Calendar" in calendar module 148; ○ The icon 428 of the image management module 144 that is labeled "Photo"; ○ The icon 430 of camera module 143, which is labeled "camera"; ○ Icon 432 of the online video module 155, labeled "Online Video"; ○ The icon 434 labeled "Stock Market" in the Stock Market widget 149-2; ○ Map module 154's icon 436, labeled "Map"; ○ The weather widget 149-1 with icon 438 labeled "weather"; ○ The alarm clock widget 149-4 has an icon 440 labeled "clock"; ○ Icon 442 of the fitness support module 142, which is labeled "fitness support"; ○ Icon 444 labeled "Notes" in Notes module 153; and ○ Icon 446 for setting applications or modules, which provides access to settings of device 100 and its various applications 136.

[0163] It should be noted that Figure 4A The icon labels illustrated are merely exemplary. Other labels are optionally used for various application icons, for example. In some embodiments, the label of a particular application icon includes the name of the application corresponding to that particular application icon. In some embodiments, the label of a specific application icon is different from the name of the application corresponding to that particular application icon.

[0164] Figure 4B An example user interface is illustrated on a device (e.g., device 300 in FIG. 3) having a touch-sensitive surface 451 separate from the display 450 (e.g., a tablet or touchpad 355 in FIG. 3). Although many subsequent examples are given with reference to input on a touchscreen display 112 (which combines a touch-sensitive surface and a display), in some embodiments the device detects input on a touch-sensitive surface separate from the display, such as... Figure 4B As shown in the diagram. In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451) has a main axis (e.g., on the display (e.g., 450) that is aligned with the main axis on the display (e.g., 451). Figure 4B The principal axis corresponding to 453 in the middle (e.g., Figure 4B(452 in the middle). According to these embodiments, the device detects contact with the touch-sensitive surface 451 at a position corresponding to the corresponding position on the display (e.g., Figure 4B (460 and 462 in the example) Figure 4B In the diagram, 460 corresponds to 468 and 462 corresponds to 470. Thus, on a touch-sensitive surface (e.g., Figure 4B 451 in the middle) and the display of a multi-functional device (e.g., Figure 4B When 450 is separated, user input detected by the device on the touch-sensitive surface (e.g., touches 460 and 462 and their movement) is used by the device to manipulate the user interface on the display. It should be understood that similar methods may be optionally used for other user interfaces described herein.

[0165] Additionally, while the examples below are given primarily with reference to finger input (e.g., finger touch, single-finger tap gesture, finger swipe gesture, etc.), it should be understood that in some implementations, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a touch), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the touch). As another example, a tap gesture is optionally replaced by a mouse click when the cursor is over the location of the tap gesture (e.g., instead of detection of touch, followed by cessation of touch detection). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are optionally used simultaneously, or mouse and finger touch are optionally used simultaneously.

[0166] As used in the specification and claims, the term "characteristic intensity" of a contact refers to a characteristic of the contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on multiple intensity samples. The characteristic intensity is optionally based on a predefined number of intensity samples or a set of intensity samples collected over a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after contact is detected, before contact is detected to lift off, before or after contact begins to move, before contact ends, before or after an increase in contact intensity is detected, and / or before or after a decrease in contact intensity is detected). The characteristic intensity of the contact is optionally based on one or more of the following: the maximum value of the contact intensity, the mean value of the contact intensity, the average value of the contact intensity, the value at the top 10% of the contact intensity, the half maximum value of the contact intensity, the 90% maximum value of the contact intensity, a value generated by low-pass filtering the contact intensity over or from a predefined time period, etc. In some implementations, the duration of contact is used when determining the characteristic intensity (e.g., when the characteristic intensity is the average intensity of the contact over time). In some implementations, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether a user has performed an action. For example, the set of one or more intensity thresholds may include a first intensity threshold and a second intensity threshold. In this example, contact with a characteristic intensity not exceeding the first intensity threshold results in a first action, contact with a characteristic intensity exceeding the first intensity threshold but not exceeding the second intensity threshold results in a second action, and contact with a characteristic intensity exceeding the second intensity threshold results in a third action. In some implementations, a comparison between the characteristic intensity and one or more intensity thresholds is used to determine whether to perform one or more actions (e.g., whether to select an appropriate option or abandon the appropriate action), rather than to determine whether to perform the first or second action.

[0167] In some implementations, a portion of the gesture is identified for determining the characteristic intensity. For example, a touch-sensitive surface may receive a series of swipe contacts that transition from a starting position to an ending position (e.g., a drag gesture), where the intensity of the contact increases. In this example, the characteristic intensity of the contact at the ending position may be based only on a portion of the series of swipe contacts, rather than the entire swipe contact (e.g., only a portion of the swipe contact at the ending position). In some implementations, a smoothing algorithm may be applied to the intensity of the swipe gesture before determining the characteristic intensity of the contact. For example, the smoothing algorithm may optionally include one or more of the following: unweighted moving average smoothing algorithm, triangular smoothing algorithm, median filter smoothing algorithm, and / or exponential smoothing algorithm. In some cases, these smoothing algorithms eliminate narrow spikes or dips in the intensity of the swipe contact to achieve the purpose of determining the characteristic intensity.

[0168] In some implementations, the device's response to input detected by the device depends on a criterion based on the contact intensity during the input. For example, for some "light press" inputs, a first response is triggered by the intensity of contact exceeding a first intensity threshold during the input. In some implementations, the device's response to input detected by the device depends on a criterion that includes both the contact intensity during the input and a time-based criterion. For example, for some "deep press" inputs, a second response is triggered by the intensity of contact exceeding the second intensity threshold (greater than the first light press threshold) during the input, provided a delay time has elapsed between satisfying the first intensity threshold and satisfying the second intensity threshold. The duration of this delay time is typically less than 200 ms (e.g., 40 ms, 100 ms, or 120 ms, depending on the magnitude of the second intensity threshold, where the delay time increases as the second intensity threshold increases). This delay time helps avoid unintentionally identifying deep press inputs. As another example, for some "deep press" inputs, there is a period of decreased sensitivity after the first intensity threshold is reached. During this period of decreased sensitivity, the second intensity threshold increases. This temporary increase in the second intensity threshold also helps to prevent accidental deep press inputs. For other deep press inputs, the response to the detected deep press input does not depend on time-based criteria.

[0169] User interface and related processes Now let’s turn our attention to implementation schemes for user interfaces (“UI”), user interactions and associated processes that can be implemented on electronic devices such as portable multifunction devices 100, 300 and / or wearable audio output devices 301.

[0170] Figures 5A to 5R Examples of user interfaces and user interactions involving real-world objects and feedback from wearable devices are shown. Figures 6A to 6H Examples of user interactions involving real-world objects and feedback from wearable devices are shown. Figures 7A to 7K Examples of user interactions with wearable devices are shown, involving various alarm conditions. Figures 8A to 8P and Figures 9A to 9K Examples of user interaction with wearable devices are illustrated. The user interface in these figures is used to illustrate the processes described below, including... Figures 10A to 10D , Figures 11A to 11C , Figures 12A to 12D and Figures 13A to 13B The process in.

[0171] Figures 5A to 5D Examples of user interfaces and user interactions involving real-world objects and feedback from wearable devices are illustrated according to some implementation schemes. Figure 5AThe illustration shows a user 502 wearing a wearable audio output device 301 (e.g., earbuds) and a head-mounted display (HMD) 100b (e.g., Figure 1C HMD 1-100). In some implementations, user 502 does not wear HMD 100b (e.g., only wears wearable audio output device 301). Figure 5A Real-world objects 508-1 (e.g., a box) and 508-2 (e.g., a clock) near user 502 are also shown. Figure 5A The illustration also shows a user 502 performing a gesture 506 pointing at (e.g., pointing at) a real-world object 508-1. The real-world object 508-1 includes a barcode 509. In some implementations, the real-world object includes different types of machine-readable codes (e.g., QR codes or App Clip codes).

[0172] Figure 5B Audio feedback 510 from wearable audio output device 301 is shown. Audio feedback 510 is in response to gesture 506 and corresponds to real-world object 508-1. According to some embodiments, wearable audio output device 301 uses the context of user 502 to generate audio feedback 510. According to some embodiments, wearable audio output device 301 uses information from one or more sensors (e.g., sensor 311) to detect gesture 506 and / or identify real-world object 508-1. In some embodiments, wearable audio output device 301 uses information from one or more sensors (e.g., sensor 311) to determine that gesture 506 points to real-world object 508-1, and not to another object (such as real-world object 508-2). In some embodiments, wearable audio output device 301 scans barcode 509 (e.g., using sensor 311) to obtain data about real-world object 508-1. For example, audio feedback 510 indicates that real-world object 508-1 is a package sent by the user's mother the previous day. In some embodiments, the wearable audio output device 301 receives information from one or more other devices (e.g., portable multifunction device 100 and / or HMD 100b) and uses the received information to generate audio feedback 510. In some embodiments, the wearable audio output device 301 provides haptic feedback in response to a gesture 506 (e.g., as an alternative to or supplement to the audio feedback 510). In some embodiments, the HMD 100b provides audio, haptic, and / or visual feedback in response to a gesture 506.

[0173] Figure 5C The illustration shows a user 502 wearing a wearable audio output device 301 (e.g., earbuds) and an HMD 100b. Figure 5CA real-world object 508-2 (e.g., a digital clock) near user 502 is also shown. Figure 5C The illustration also shows user 502 listening to music 520 output by wearable audio output device 301. Wearable audio output device 301 is communicatively coupled to portable multifunction device 100, as indicated by arrow line 514. In some embodiments, music 520 corresponds to an application (e.g., a music application) running on portable multifunction device 100. Figure 5C A user interface 516 corresponding to music 520 is shown. In some embodiments, the user interface 516 is displayed on the portable multifunction device 100. The user interface 516 includes a playback control 518 (e.g., a pause element). Selecting the playback control 518 stops the output of music 520 (e.g., pauses music 520). In some embodiments, the user interface 516 is not displayed (e.g., the wearable audio output device 301 and / or the portable multifunction device 100 control the playback of music 520 in response to user input without displaying the user interface 516). In some embodiments, the wearable audio output device 301 controls the playback of music 520 in response to user input (e.g., via input device 308). Figure 5C The illustration also shows user 502 performing a gesture 512 pointing at (e.g., pointing at) a real-world object 508-2. In some embodiments, other gestures are used to indicate the real-world object (e.g., tapping gesture, nodding gesture, circling gesture, and / or other types of gestures). The real-world object 508-2 includes a display of the current time (e.g., the real-world object 508-2 is a digital clock).

[0174] Figure 5DAudio feedback 522 from wearable audio output device 301 is shown. Audio feedback 522 is responsive to gesture 512 and corresponds to real-world object 508-2. According to some embodiments, wearable audio output device 301 uses the context of user 502 to generate audio feedback 522 (e.g., to determine that real-world object 508-2 is user 502's bedroom clock). According to some embodiments, wearable audio output device 301 uses information from one or more sensors (e.g., sensor 311) to detect gesture 512 and / or identify real-world object 508-2. In some embodiments, wearable audio output device 301 uses information from one or more sensors (e.g., sensor 311) to determine that gesture 512 is pointing at real-world object 508-2. In some embodiments, wearable audio output device 301 analyzes the display of real-world object 508-2 to obtain data about real-world object 508-2. For example, audio feedback 522 indicates the current time (9:17 AM) obtained from real-world object 508-2. In some embodiments, the wearable audio output device 301 receives information from one or more other devices (e.g., portable multifunction device 100 and / or HMD 100b) and uses the received information to generate audio feedback 522. For example, the wearable audio output device 301 may use location information from the portable multifunction device 100 to determine that a real-world object 508-2 is located in the bedroom of user 502. In some embodiments, the wearable audio output device 301 provides haptic feedback in response to a gesture 512 (e.g., as an alternative to or supplement to audio feedback 510). In some embodiments, the HMD 100b provides audio, haptic, and / or visual feedback in response to a gesture 512.

[0175] Figure 5E The illustration shows a user 502 wearing a wearable audio output device 301 (e.g., headphones) and an HMD 100b. In some embodiments, the user 502 does not wear the HMD 100b (e.g., only wears the wearable audio output device 301). Figure 5E It also shows real-world objects 508-3 (e.g., a calendar) near user 502. Figure 5E It also shows user 502 performing a gesture 528 pointing to a real-world object 508-3 (e.g., pointing to a specific day shown on a calendar).

[0176] Figure 5FAudio feedback 530 from wearable audio output device 301 is shown. Audio feedback 530 is responsive to gesture 528 and corresponds to real-world object 508-3. According to some embodiments, wearable audio output device 301 uses the context of user 502 to generate audio feedback 530 (e.g., to determine if user 502 is available on a particular date). According to some embodiments, wearable audio output device 301 uses information from one or more sensors (e.g., sensor 311) to detect gesture 528 and / or identify real-world object 508-3. In some embodiments, wearable audio output device 301 uses information from one or more sensors (e.g., sensor 311) to determine that gesture 528 points to real-world object 508-3 (e.g., to a specific portion of real-world object 508-3). In some embodiments, wearable audio output device 301 analyzes real-world object 508-3 to obtain data about text on real-world object 508-3. For example, audio feedback 530 indicates a date (October 1st) obtained from real-world object 508-3. In some embodiments, the wearable audio output device 301 receives information from one or more other devices (e.g., portable multifunction device 100 and / or HMD 100b) and uses the received information to generate audio feedback 530. For example, the wearable audio output device 301 may use calendar information from the portable multifunction device 100 to determine that the user 502 is scheduled to meet with a contact named John on October 1st. In some embodiments, the wearable audio output device 301 provides haptic feedback in response to a gesture 512 (e.g., as an alternative to or supplement to the audio feedback 510). In some embodiments, the HMD 100b provides audio, haptic, and / or visual feedback in response to a gesture 512. Therefore, Figures 5A to 5F Examples of user interfaces and user interactions are illustrated according to some implementations for making gestures at real-world objects (e.g., detected via wearable audio output device 301) and obtaining audio feedback (e.g., provided by wearable audio output device 301) related to the indicated object.

[0177] Figure 5G The illustration shows a user 502 wearing a wearable audio output device 301 (e.g., headphones) and an HMD 100b. In some embodiments, the user 502 does not wear the HMD 100b (e.g., only wears the wearable audio output device 301). Figure 5G It also shows a notepad 534 (e.g., a real-world object) near the user 502 at the first moment. Figure 5G It also shows a user 502 holding a writing instrument 532 (e.g., a pen or pencil). Figure 5H Notebook 534 shows a second timeframe following the first. Figure 5H In this context, the notebook 534 includes writing content 535 (e.g., notes) on it. For example, user 502 has already written the writing content 535 using writing tool 532. In some embodiments, the writing content 535 includes virtual writing content displayed via HMD 100b (e.g., corresponding to writing by user 502 with their finger and / or stylus).

[0178] Figure 5I This illustrates user 502 performing gesture 536 (e.g., circling text on notepad 534). Figure 5I In this scenario, user 502 is using a writing tool to perform gesture 536. In some implementations, user 502 uses their fingers to perform gesture 536 (e.g., pointing, circling, and / or otherwise indicating the text "Call John"). Figure 5J A portion 538 on the notebook 534 corresponding to the selection of gesture 536 (e.g., the result of gesture 536) is shown. In some embodiments, gesture 536 selects text within the selected portion 538. In some embodiments, the selected portion 538 includes visible lines (e.g., drawn with pencil or ink). In some embodiments, the selected portion 538 has invisible boundaries (e.g., drawn by user 502 using their finger or stylus). In some embodiments, the selected portion 538 has virtual boundary lines (e.g., displayed via HMD 100b). Although Figure 5I The selected portion 538 is shown to have a circular shape, but in some embodiments, the selected portion 538 has a non-circular shape (e.g., an elliptical shape, a rectangular shape, or an irregular shape). In some embodiments, the selected portion 538 is defined and / or indicated by an underline in the text.

[0179] Figure 5K The illustration shows a user 502 wearing a wearable audio output device 301 (e.g., headphones) and an HMD 100b. In some embodiments, the user 502 does not wear the HMD 100b (e.g., only wears the wearable audio output device 301). Figure 5K The illustration also shows user 502 performing a gesture 540 pointing to a selected portion 538 containing the text "Call John". In some embodiments, gesture 540 is a pointing gesture (e.g., where the user's finger does not touch the notepad 534). In some embodiments, gesture 540 is a tapping gesture (e.g., where the user's finger touches the selected portion 538 of the notepad 534). Figure 5K Portable multi-functional device 100 (e.g., a companion device to wearable audio output device 301) is also shown.

[0180] Figure 5LThe illustration shows audio feedback 542 provided to user 502 by wearable audio output device 301 in response to gesture 540. Audio feedback 542 instructs user 502 to hold gesture 540 for two seconds to perform an action corresponding to selected portion 538 (e.g., initiating a phone call). Figure 5L The user 502 is also shown continuing gesture 540. In some embodiments, tactile and / or visual feedback is provided in response to gesture 540 (e.g., as a supplement to or alternative to audio feedback 542). In some embodiments, tactile feedback is provided by wearable audio output device 301 and / or HMD 100b. In some embodiments, tactile and / or visual feedback is provided by HMD 100b.

[0181] Figure 5M This illustrates feedback 544 provided to user 502 by wearable audio output device 301 in response to user 502 maintaining gesture 540. For example, Figure 5M Can correspond to Figure 5L The time point one second after the time shown. In some embodiments, feedback 544 includes audio and / or haptic feedback. In some embodiments, feedback 544 is progressive feedback that changes over time as user 502 continues gesture 540. In some embodiments, feedback 544 indicates how long user 502 has held gesture 540 (e.g., indicating that user 502 has held gesture for at least a first threshold amount of time). In some embodiments, audio, visual, and / or haptic feedback is provided to user 502 via HMD 100b (e.g., as a supplement to or alternative to feedback 544). Figure 5M It also shows user 502 continuing gesture 540.

[0182] Figure 5N This illustrates how a wearable audio output device 301 provides audio feedback 546 to user 502 in response to user 502 continuing to hold gesture 540. For example, Figure 5N Can correspond to Figure 5M The time point one second after the time shown in the image. Audio feedback 546 indicates that a call is being initiated (via portable multifunction device 100) to the mobile number of a contact named John. In response to user 502 maintaining gesture 540 for at least a second threshold amount of time (e.g., by...), Figure 5L The audio feedback 542 indicates 2 seconds) while Figure 5N Initiate a call. Figure 5NA user interface 548 on the portable multifunction device 100 is also shown, indicating that the portable multifunction device 100 is initiating a call to John's mobile phone number. In some embodiments, the wearable audio output device 301 initiates the call in response to the user 502 continuing to hold the gesture 540 for at least a second threshold time. In some embodiments, the wearable audio output device 301 transmits a command to the portable multifunction device 100 to initiate the call in response to the user 502 continuing to hold the gesture 540 for at least a second threshold time. In some embodiments, the portable multifunction device 100 initiates the call without displaying the user interface 548. In some embodiments, the wearable audio output device 301 initiates the call without providing audio feedback 546 (e.g., using the audio of the call (such as a ringtone) to notify the user 502 that a call has been initiated). Therefore, Figure 5G to Figure 5N Examples of user interfaces and user interactions are illustrated according to some implementations for writing text, selecting a portion of the written text, and initiating a call in response to a gesture toward the selected portion (e.g., detected via a wearable audio output device 301).

[0183] Figures 50 to 5Q This illustrates user 502 performing a gesture 550 (e.g., a drawing gesture) pointing to portion 554 of notepad 534. Figure 5O In this scenario, user 502 has already initiated gesture 550 with their finger at location 550-a. In some embodiments, gesture 550 is an air gesture performed without touching the notepad 534. In some embodiments, gesture 550 is performed on the surface of the notepad 534. Figures 50 to 5Q In the example, gesture 550 is performed using the user's finger. In some implementations, gesture 550 is performed using a writing instrument (e.g., pen, pencil, stylus, or another type of writing instrument). Figure 5P In the process, user 502 has continued with gesture 550, where their finger has moved to positioning 550-b. In some implementations, HMD 100b displays an indication of the progress of gesture 550 (e.g., a line). Figure 5P In the example, wearable audio output device 301 provides feedback 552 corresponding to gesture 550. In some embodiments, feedback 552 includes audio and / or haptic feedback. In some embodiments, feedback 552 indicates the progress of gesture 550. Figure 5QIn this embodiment, user 502 has already completed gesture 550 with their finger at location 550-c (e.g., has already drawn an ellipse around portion 554 on notepad 534). In some embodiments, wearable audio output device 301 and / or HMD 100b provide audio and / or haptic feedback indicating that user 502 has completed gesture 550. In some embodiments, HMD 100b provides visual feedback that user 502 has completed gesture 550.

[0184] Figure 5R A wearable audio output device 301 is shown providing audio feedback 556 in response to a user 502 completing a gesture 550. The audio feedback 556 indicates to the user 502 that text in section 554 has been added to an active notebook 558. According to some embodiments, notebook 558 corresponds to a note-taking application running on a portable multifunction device 100. Figure 5R A notebook 558 displayed on a portable multifunction device 100 is also shown, where text 560 is copied from portion 554 of a notepad 534. In some embodiments, a wearable audio output device 301 detects the completion of a gesture 550 and analyzes the text in portion 554 of the notepad 534. In some embodiments, the wearable audio output device 301 sends the text to the portable multifunction device 100 for storage and association with the notebook 558. In some embodiments, the notebook 558 and / or text 560 are displayed via an HMD 100b (e.g., as a supplement or alternative to what is displayed on the portable multifunction device 100). Therefore, Figures 5O to 5R Examples of user interfaces and user interactions for adding text to a notebook in response to gestures (e.g., detected via a wearable audio output device 301) are illustrated according to some implementations.

[0185] Figures 6A to 6H Examples of user interactions involving real-world objects and feedback from wearable devices are illustrated according to some implementation schemes. Figure 6A The illustration shows user 602 wearing wearable audio output devices 301-1 and 301-2 (e.g., earbuds) and HMD 100b. In some embodiments, user 602 does not wear HMD 100b (e.g., only wears wearable audio output device 301). Figure 6AUser 602 is outside the geofence boundary 604. In some embodiments, the geofence boundary 604 corresponds to a location previously geofenced by user 602. In some embodiments, the geofence boundary 604 corresponds to user 602's home or office. In some embodiments, the geofence boundary 604 has been previously defined by user 602. In some embodiments, the geofence boundary 604 is defined by a different user and shared with wearable audio output device 301 (e.g., shared via network connection and / or companion devices).

[0186] Figure 6B The illustration shows user 602 crossing geofence boundary 604 and receiving feedback 606 from wearable audio output device 301 in response. In some embodiments, feedback 606 includes audio and / or haptic feedback. For example, feedback 606 includes a beep or ringtone to indicate to user 602 that they have entered a geofenced location. In some embodiments, wearable audio output device 301 provides a first type of feedback in response to user entering a geofenced location and a second type of feedback in response to user leaving a geofenced location. In some embodiments, HMD 100b provides audio, visual, and / or haptic feedback (e.g., as a supplement to or alternative to feedback 606) in response to user 602 crossing geofence boundary 604. In some embodiments, wearable audio output device 301 determines that user 602 has crossed geofence boundary 604 based on sensor data (e.g., from sensor 311) generated at wearable audio output device 301 and / or received from other devices (e.g., companion devices, such as portable multifunction device 100). Figure 6B Real-world objects 610-1, 610-2, and 610-3 within the geofence boundary 604 are also shown.

[0187] Figure 6C The illustration shows a user 602 performing a gesture 612 toward (e.g., pointing at) a real-world object 610-1 and asking a question 614 (“What is that?”). In some embodiments, the wearable audio output device 301 and / or HMD 100b detects the question 614 via one or more microphones (e.g., microphone 302). In some embodiments, the wearable audio output device 301 and / or HMD 100b detects the gesture 612 via one or more sensors (e.g., sensor 311). Although Figure 6C The illustration shows user 602 asking question 614, but in some implementations, gesture 612 is performed when there is no verbal question (e.g., gesture to indicate the question).

[0188] Figure 6DAudio feedback 616 is shown provided in response to gesture 612 and question 614. Audio feedback 616 includes information about real-world object 610-1 (e.g., obtained from analysis of an image of real-world object 610-1). In some embodiments, the audio feedback indicates the object type of the real-world object (e.g., it is a box). In some embodiments, the audio feedback includes information about the appearance of the real-world object (e.g., size, color, material, texture, and / or other appearance information). In some embodiments, the audio feedback includes contextual information (e.g., based on the user's context). Audio feedback 616 indicates that real-world object 610-1 is... Figure 6D The cardboard box in the example. According to some implementations, audio feedback 616 is provided as spatial feedback corresponding to the location of object 610-1 (e.g., appearing to emanate from that location). Spatialized audio feedback simulates a more realistic listening experience, where the audio seems to originate from a sound source within a specific frame of reference (such as the physical environment surrounding the user). For example, audio feedback 616 is provided as spatial audio from a simulated location corresponding to the position of the real-world object 610-1 relative to the wearable audio output device 301. In some implementations, audio feedback 616 is not provided if the user 602 is outside the geofence boundary 604 while performing gesture 612 and / or asking question 614. As an example, when spatial audio is enabled, the audio output from the ear-worn audio output device (e.g., earbuds) sounds as if the corresponding audio for each real-world object originates from a different simulated spatial location (which may change over time) within a frame of reference (such as the physical environment) (e.g., a surround sound effect). The positioning of the real-world object (simulated spatial location) is independent of the movement of the earbuds relative to the frame of reference.

[0189] Figure 6E The illustration shows user 602 performing a gesture 618 toward (e.g., pointing at) real-world objects 610-3 and asking question 620 (“What is that?”). In some embodiments, wearable audio output device 301 and / or HMD 100b detects question 620 via one or more microphones (e.g., microphone 302). In some embodiments, wearable audio output device 301 and / or HMD 100b detects gesture 618 via one or more sensors (e.g., sensor 311). In some embodiments, gesture 618 and question 620 are detected concurrently. In some embodiments, gesture 618 is detected before or after question 620 (e.g., gesture 618 and question 620 are correlated based on whether they are detected within a threshold amount of time to each other). Figure 6E In the example, question 620 is spoken in a normal tone of voice (e.g., neither shouted nor whispered).

[0190] Figure 6FAudio feedback 622 is shown provided in response to gesture 618 and question 620. Audio feedback 622 includes information about the real-world object 610-3 (e.g., obtained from analysis of an image of the real-world object 610-3). In some embodiments, the audio feedback indicates the object type of the real-world object (e.g., it is a traffic cone). In some embodiments, the audio feedback includes information about the appearance of the real-world object (e.g., the real-world object 610-3 is orange). In some embodiments, the audio feedback includes contextual information (e.g., the real-world object 610-3 was present at that location the previous day). According to some embodiments, audio feedback 622 is provided as spatial feedback corresponding to the relative position of object 610-3. Figure 6F Audio feedback 622 is also shown provided at a volume level 626-1 as shown on volume indicator 624. In some embodiments, volume level 626-1 is based on the volume of problem 620. In some embodiments, volume level 626-1 is based on the volume setting of wearable audio output device 301 and / or the ambient sound level of the physical environment in which wearable audio output device 301 is located.

[0191] Figure 6G The illustration shows user 602 performing a gesture 618 toward (e.g., pointing at) real-world objects 610-3 and whispering the question 630 (“What’s that?”). In some embodiments, wearable audio output devices 301 and / or HMD 100b detect the question 630 via one or more microphones (e.g., microphone 302). Figure 6G In the example, question 630 is related to Figure 6E Question 620 involves speaking with different intonations (e.g., being whispered rather than spoken). Questions 620 and 630 involve the same words but are spoken differently.

[0192] Figure 6H Audio feedback 632 is shown provided in response to gesture 618 and question 630. Audio feedback 632 includes information about real-world object 610-3 (e.g., obtained from analysis of an image of real-world object 610-3). Figure 6HAudio feedback 632 is also shown provided at a volume level 626-2 (lower than volume level 626-1) as shown on volume indicator 624. In some embodiments, volume level 626-2 is based on the volume of question 630. In some embodiments, volume level 626-2 is based on the volume setting of wearable audio output device 301 and / or the ambient sound level of the physical environment in which wearable audio output device 301 is located. According to some embodiments, audio feedback 632 includes information that differs from audio feedback 622 based on the difference in the tone of voice of user 602 when asking questions 620 and 630 (and / or the different volume levels of questions 620 and 630). According to some embodiments, audio feedback 632 is provided as spatial feedback corresponding to the location of object 610-3 (e.g., appearing to be emanating from that location). Therefore, Figures 6A to 6H An example user interaction is illustrated according to some implementations for gesturing at a real-world object (e.g., detected via wearable audio output device 301) within a geofenced location and asking the real-world object and obtaining audio feedback (e.g., provided by wearable audio output device 301) related to the indicated object.

[0193] Figures 7A to 7K Examples of user interactions with wearable devices involving various alarm conditions are illustrated according to some implementation schemes. Figure 7A Includes perspective view 701-1 and corresponding top view 701-2, and shows a user 702 wearing a wearable audio output device 301 (e.g., a headset). Figure 7A In the middle, wearable audio output device 301 is outputting music 704. Figure 7A It also indicates that active noise cancellation (ANC) mode is enabled at level 710-a, and music 704 has a corresponding media volume level 712-a. Figure 7A Also shown is a potentially dangerous puddle 706 in the street in front of user 702. According to some embodiments, the wearable audio output device 301 determines that the puddle 706 is more than a threshold distance from user 702 and / or determines that user 702 is not moving toward the puddle 706. Figure 7A No feedback (notification) is provided for puddle 706. Top view 701-2 also shows the field of view 711 of user 702.

[0194] Figure 7BThe diagram includes a perspective view and a corresponding top view, showing user 702 crossing the street toward puddle 706, and in response, wearable audio output device 301 detects an alarm condition. In some embodiments, the alarm condition is based on the distance between user 702 and puddle 706. In some embodiments, wearable audio output device 301 uses data from one or more sensors (e.g., sensor 311) to detect puddle 706. In some embodiments, the alarm condition is based on user 702 approaching puddle 706. In some embodiments, the alarm condition is determined based on an assessed probability that user 702 will slip or fall due to stepping into puddle 706. In some embodiments, the alarm condition is based on one or more user preferences (e.g., avoiding stepping into puddles). In some embodiments, the alarm condition is based on user 702's line of sight (e.g., an assessment of whether user 702 is likely to notice the puddle). Based on the detected alarm condition, Figure 7B The ANC in the range is from level 710-a (e.g., as...) Figure 7A (As shown) is reduced to level 710-b, and the media volume is reduced from volume level 712-a (e.g., as shown) Figure 7A (As shown) The volume is reduced to level 712-b. In some embodiments, the ANC is reduced based on the detection of an alarm condition, and the media volume remains unchanged. In some embodiments, the media volume is reduced (and / or media playback is paused) based on the detection of an alarm condition, and the ANC remains unchanged. In some embodiments, whether the media is paused or the media volume is reduced is based on the type of media (e.g., for music, the volume is reduced; and for audio content from other sources, playback is paused).

[0195] Figure 7C Includes a perspective view and a corresponding top view, and shows the user 702 wearing a wearable audio output device 301. Figure 7C In the middle, the wearable audio output device 301 is playing back media, as indicated by the media playback indicator 718. Figure 7C It also indicates that active noise cancellation (ANC) mode is enabled at level 716-a. Figure 7C A ball 720 is also shown approaching user 702 from the front (e.g., within field of view 711) and indicating a potential hazard. In some embodiments, wearable audio output device 301 uses data from one or more sensors (e.g., sensor 311) to detect ball 720. According to some embodiments, wearable audio output device 301 detects ball 720 based on determining that ball 720 is more than a threshold distance from user 702 and / or determining that ball 720 is within field of view 711. Figure 7C No feedback (notification) is provided for the ball 720. In some implementations, the wearable audio output device 301 determines that the ball 720 is unlikely to come into contact with the user 702. Figure 7CFeedback (notifications) for Ball720 are not provided in the Chinese version. Figure 7D The image shows ball 720 behind user 702 (e.g., having passed user 702). Figure 7D In this process, the wearable audio output device 301 determines that the ball 720 is moving away from the user 702 without providing feedback (e.g., a notification) to the ball 720. Figure 7C and Figure 7D In the meantime, ANC remains at level 716-a, and media continues to play, for example, based on the condition that no alarm is detected.

[0196] Figure 7E Includes a perspective view and a corresponding top view, and shows the user 702 wearing a wearable audio output device 301. Figure 7E In the middle, the wearable audio output device 301 is playing back media, as indicated by the media playback indicator 718, and ANC mode is enabled in level 716-a. Figure 7E A ball 720 is also shown behind the user 702 (e.g., outside the field of view 711) and indicating a potential hazard. According to some embodiments, the wearable audio output device 301 determines that the ball 720 is more than a threshold distance from the user 702 and / or determines that the ball 720 has not moved any closer to the user 702. Figure 7E The system does not provide feedback (e.g., notifications) for Ball 720. Figure 7E In the meantime, ANC remains at level 716-a, and media continues to play, for example, based on the condition that no alarm is detected.

[0197] Figure 7F Including a perspective view and a corresponding top view, ball 720 approaches user 702 from behind (e.g., outside user 702's field of view 711), and in response, wearable audio output device 301 detects an alarm condition and provides audio feedback 722. In some embodiments, the alarm condition is based on the distance between user 702 and ball 720. In some embodiments, the alarm condition is based on ball 720 approaching user 702 from a position outside the field of view 711. In some embodiments, the alarm condition is determined based on an assessed probability of collision between ball 720 and user 702. In some embodiments, the alarm condition is based on one or more user preferences (e.g., a preference for being more cautious about conditions occurring behind the user). In some embodiments, the alarm condition is based on user 702's line of sight (e.g., an assessment of whether user 702 is likely to notice the ball). Depending on the detected alarm condition, Figure 7F The ANC in the range is from level 716-a (e.g., as...) Figure 7EThe ANC is lowered to level 716-b (as shown), and media playback is paused, as indicated by media playback indicator 718. In some embodiments, the ANC is lowered based on the detection of an alarm condition, and media playback remains unchanged. In some embodiments, media playback is paused based on the detection of an alarm condition, and the ANC remains unchanged.

[0198] Figure 7G Includes a perspective view and a corresponding top view, and shows the user 702 wearing a wearable audio output device 301. Figure 7G In the middle, user 702 is waiting to cross the road. (Regarding...) Figure 7G The wearable audio output device 301 in the middle enables an ANC mode with a corresponding ANC level 716-a. Figure 7G In this context, the active transparency mode of the wearable audio output device 301 is disabled, as indicated by level 726-a, and the dialogue enhancement mode is also disabled, as indicated by enhancement indicator 728. Figure 7G It also shows person 730 approaching user 702 from outside user 702's field of view 711. Figure 7G Person 730 is greeting user 702, as instructed by statement 732 (e.g., “Hey John, wait a minute!”).

[0199] Figure 7H Includes perspective views and corresponding top views, and shows the wearable audio output device 301 responding to a person 730 in... Figure 7G An alarm condition is detected when user 702 is called upon. In some embodiments, the alarm condition is based on the distance between user 702 and person 730. In some embodiments, wearable audio output device 301 uses data from one or more sensors (e.g., sensor 311 and / or microphone 302) to detect person 730. In some embodiments, the alarm condition is based on person 730 being close to user 702 and / or outside user 702's field of vision 711. In some embodiments, the alarm condition is determined based on an assessed probability that user 702 will notice person 730. In some embodiments, the alarm condition is based on one or more user preferences (e.g., to notify the user when called upon by someone). In some embodiments, the alarm condition is based on user 702's line of sight (e.g., an assessment of whether user 702 is likely to notice person 730). Based on the detected alarm condition, Figure 7H The ANC in the range is from level 716-a (e.g., as...) Figure 7G (As shown) Reduce to level 716-b (e.g., disable ANC), active transparency from volume level 726-a (e.g., as shown) Figure 7GThe volume level 726-b is increased to (e.g., active transparency is enabled), and conversation enhancement mode is enabled as indicated by indicator 728. In some embodiments, ANC is disabled in response to an alarm condition, and active transparency and / or conversation enhancement remain unchanged. In some embodiments, ANC, active transparency, and / or conversation enhancement are adjusted in response to an alarm condition, depending on whether the alarm condition includes audio components (e.g., a user being greeted, a vehicle horn, an impact sound, or other types of audio characteristics). In some embodiments, the level 726-b of active transparency is based on the volume of statement 732 (e.g., a lower volume of statement 732 corresponds to a higher level of active transparency). In some embodiments, conversation enhancement mode is enabled based on the volume of statement 732 being below a threshold volume level. Figure 7H Also shown is a wearable audio output device 301 providing audio feedback 736 in response to the detection of an alarm condition (e.g., "Your friend Kacie is trying to catch up with you"). According to some embodiments, the audio feedback 736 includes contextual information about the user 702 (e.g., identifying Kacie as the user's friend).

[0200] Figure 7I Includes a perspective view and a corresponding top view, and shows the user 702 wearing a wearable audio output device 301. Figure 7I In the middle, user 702 is waiting to cross the road. (Regarding...) Figure 7I The wearable audio output device 301 in the middle enables an ANC mode with a corresponding ANC level 716-a. Figure 7I In this context, the active transparency mode of the wearable audio output device 301 is disabled, as indicated by level 726-a, and the dialogue enhancement mode is also disabled, as indicated by enhancement indicator 728. Figure 7I Persons 740 and 742 are also shown outside the field of view 711 of user 702. Figure 7J The image shows people 740 and 742 arguing outside the field of vision 711 of user 702, as indicated by a scream 744. In some embodiments, wearable audio output device 301 determines that people are arguing based on audio and / or body movement.

[0201] Figure 7KA wearable audio output device 301 is shown detecting an alarm condition in response to people 740 and 742 arguing. In some embodiments, the alarm condition is based on the distance between user 702 and people 740 and 742. In some embodiments, the wearable audio output device 301 uses data from one or more sensors (e.g., sensor 311 and / or microphone 302) to detect people 740 and 742 arguing. In some embodiments, the alarm condition is based on people 740 and 742 being outside the user 702's field of vision 711. In some embodiments, the alarm condition is determined based on an assessed probability that user 702 has noticed people 740 and 742 arguing. In some embodiments, the alarm condition is based on one or more user preferences (e.g., to notify the user when someone is arguing or fighting nearby). In some embodiments, the alarm condition is based on the user 702's line of sight (e.g., an assessment of whether user 702 is likely to notice the argument). Depending on the detected alarm condition, Figure 7K The ANC in the range is from level 716-a (e.g., as...) Figure 7J (As shown) Reduce to level 716-b (e.g., disable ANC), active transparency from volume level 726-a (e.g., as shown) Figure 7J The volume of the alarm 746 is increased to level 726-c (e.g., active transparency is enabled), and the dialogue enhancement mode remains unchanged, as indicated by indicator 728. In some embodiments, ANC is disabled in response to an alarm condition, and active transparency and / or dialogue enhancement remain unchanged. In some embodiments, ANC, active transparency, and / or dialogue enhancement are adjusted in response to an alarm condition, based on whether the alarm condition includes audio components (e.g., arguing or other types of audio characteristics). In some embodiments, the level 726-c of active transparency is based on the volume of the alarm 746. In some embodiments, the dialogue enhancement mode is enabled based on the relative volume of the alarm 746 compared to other noises in the physical environment. Figure 7K Also shown is a wearable audio output device 301 that provides audio feedback 748 in response to the detection of an alarm condition (e.g., "Attention: There may be a fight behind you"). According to some embodiments, the audio feedback 748 includes information about the relative location of the alarm condition (e.g., behind the user). Figure 7KAn alarm 750 is also shown displayed on a portable multifunction device 100 in response to the detection of an alarm condition. Alarm 750 includes an image 752 of people 740 and 742 arguing (e.g., an image captured by a wearable audio output device 301). In some embodiments, alarm 750 is generated based on user preferences (e.g., user preferences for obtaining audio and visual information about the alarm condition). In some embodiments, alarm 750 is generated based on determining that the wearable audio output device 301 is coupled to a display generation component (e.g., the display generation component of the portable multifunction device 100). In some embodiments, a visual alarm is generated if the alarm condition is a first type of alarm condition (e.g., a fight nearby), and no visual alarm is generated if the alarm condition is a second type of alarm condition (e.g., a puddle in the user's path). Therefore, Figures 7A to 7K Example user interfaces and user interactions for generating audio feedback in response to alarm conditions (e.g., detected via wearable audio output device 301) are illustrated according to some implementation schemes.

[0202] Figures 8A to 8P Examples of user interactions with wearable devices are illustrated according to some implementation schemes. Figure 8A The illustration shows a user 802 wearing a wearable audio output device 301 (e.g., earbuds) and an HMD 100b. In some embodiments, the user 802 does not wear the HMD 100b (e.g., only wears the wearable audio output device 301). Figure 8A It also instructs that active noise cancellation (ANC) mode be enabled for wearable audio output device 301 at level 810-a, active transparency mode is at level 808-a (e.g., disabled), and dialogue enhancement mode is disabled, as indicated by enhancement indicator 812.

[0203] Figure 8BThe illustration shows user 802 performing an air gesture 804 within detection area 806. In some embodiments, detection area 806 is defined relative to wearable audio output device 301. In some embodiments, detection area 806 corresponds to the sensor detection range of a sensor (e.g., sensor 311) of wearable audio output device 301. In some embodiments, detection area 806 is a three-dimensional region (e.g., having a predetermined size). For example, detection area 806 may have dimensions of 1 ft × 1 ft × 1 ft, 6 inches × 6 inches × 8 inches, 5 inches × 8 inches × 4 inches, a radius of 1 ft (e.g., a hemispherical radius), or other dimensions. In some embodiments, wearable audio output device 301 detects only gestures performed within detection area 806. In some embodiments, wearable audio output device 301 ignores and / or disregards gestures performed outside detection area 806. In some embodiments, air gesture 804 includes the user's hand forming a C-shaped cupping gesture near the user's ear. In response to the detection of an air gesture 804, the wearable audio output device 301 performs operations including: Figure 8B The ANC in the range is from level 808-a (e.g., as...) Figure 8A Adjust the volume level (as shown) to 808-b (e.g., disable ANC), and change the active transparency from volume level 810-a (e.g., as shown). Figure 8A The volume is increased to level 810-b (e.g., enabling active transparency), and conversation enhancement mode is enabled, as indicated by indicator 812. In some implementations, only a subset of ANC, active transparency, and conversation enhancement are adjusted in response to air gesture 804. For example, ANC may be disabled in response to air gesture 804, and active transparency may not be enabled.

[0204] Figure 8C This demonstrates user 802 performing an air gesture 804 outside detection area 806. In response to the user's hand moving outside detection area 806, wearable audio output device 301 will... Figure 8C The ANC in the range is from level 808-b (e.g., as...). Figure 8B (As shown) Adjust to level 808-A (e.g., re-enable ANC), and adjust active transparency from volume level 810-b (e.g., as shown). Figure 8B (As shown) Reduce the volume to level 810-a (e.g., disable active transparency) and disable conversation enhancement mode, as indicated by indicator 812 (e.g., wearable audio output device 301 stops performing). Figure 8B (Operation).

[0205] Figure 8DThis illustrates user 802 performing an air gesture 804 within detection area 806 (e.g., the user's hand re-enters detection 806 after leaving the detection area 806 to perform the air gesture 804). Figure 8C (As shown). In response to the detection of an air gesture 804 in the detection area 806, the wearable audio output device 301 performs an operation (e.g., resume). Figure 8B The operation performed in the middle), the operation includes: Figure 8D The ANC in the range is from level 808-a (e.g., as...) Figure 8C Adjust the volume level (as shown) to 808-b (e.g., disable ANC), and change the active transparency from volume level 810-a (e.g., as shown). Figure 8C (As shown) Increase to volume level 810-b (e.g., enable active transparency), and enable conversation enhancement mode, as indicated by indicator 812.

[0206] Figure 8E The illustration shows a user 802 wearing a wearable audio output device 301 (e.g., earbuds) and an HMD 100b. In some embodiments, the user 802 does not wear the HMD 100b (e.g., only wears the wearable audio output device 301). Figure 8A The wearable audio output device 301 is also shown outputting music 822. Figure 8E The wearable audio output device 301 is communicatively coupled to the portable multifunction device 100, as indicated by arrow line 814. In some embodiments, music 822 corresponds to an application (e.g., a music application) executed on the portable multifunction device 100. Figure 8E A user interface 816 corresponding to music 822 is shown. In some embodiments, the user interface 816 is displayed on a portable multifunction device 100. The user interface 816 includes a playback control 818 (e.g., a pause element) and a media item indicator 820 (e.g., indicating that "track 1" is currently playing). Selecting the playback control 818 stops the output of music 822 (e.g., pauses music 822). In some embodiments, the user interface 816 is not displayed (e.g., the wearable audio output device 301 and / or the portable multifunction device 100 controls the playback of music 822 in response to user input without displaying the user interface 816). In some embodiments, the wearable audio output device 301 controls the playback of music 822 in response to user input (e.g., via input device 308).

[0207] Figure 8F This shows user 802's hand 824 entering detection area 806. Figure 8FAlso shown is feedback 826 provided by the wearable audio output device 301 in response to the detection of a hand 824 entering the detection area 806. In some embodiments, the wearable audio output device 301 uses one or more sensors (e.g., sensor 311) to detect the hand 824. In some embodiments, feedback 826 includes audio and / or haptic feedback. In some embodiments, the HMD 110b provides audio, visual, and / or haptic feedback (e.g., as a supplement to or alternative to feedback 826) in response to the hand 824 entering the detection area 806.

[0208] Figure 8G The image shows user 802 performing an air gesture 828 (e.g., a pinch gesture) within detection area 806. Figure 8G Music 822 (from) was also shown. Figure 8E Playback has stopped (e.g., music 822 is paused), as indicated by user interface 816, which includes playback controls 830 (e.g., a playback element) in place of... Figure 8E The playback control 818 is shown in the figure. In some embodiments, the playback of music 822 is paused for a preset amount of time (e.g., 5 seconds, 10 seconds, 20 seconds, or 30 seconds). In some embodiments, the playback of music 822 is paused until the user 802 issues a command, input, and / or gesture to resume playback. For example, in response to the user 802 selecting the playback control 830, the playback of music 822 resumes.

[0209] Figure 8H This shows that user 802's hand 824 leaves the detection area 806. Figure 8H Also shown is feedback 834 provided by wearable audio output device 301 in response to detecting hand 824 leaving detection area 806. In some embodiments, wearable audio output device 301 provides feedback 834 in response to stopping detection of hand 824 (e.g., due to hand 824 leaving detection area 806). In some embodiments, wearable audio output device 301 uses one or more sensors (e.g., sensor 311) to detect hand 824. In some embodiments, feedback 834 includes audio and / or haptic feedback. In some embodiments, HMD 110b provides audio, visual, and / or haptic feedback (e.g., as a supplement to or alternative to feedback 834) in response to hand 824 leaving detection area 806. In some embodiments, Figure 8H Feedback 834 in the middle has the same Figure 8F Feedback 826 may have one or more different attributes. For example, feedback 834 may include a single beep, while feedback 826 may include two or more beeps.

[0210] Figure 8IThe image shows user 802 with their hands placed on either side. In some embodiments, one or more sensors of the wearable audio output device 301 are disabled (e.g., detection area 806 is deactivated) or placed in a low-power mode (e.g., using a reduced sampling or scanning rate to detect hand gestures or other events compared to a normal power mode using a normal sampling or scanning rate), while the hand is not present in that area. Figure 8I It also shows that media playback is paused, as indicated by the playback control 830 of the user interface 816. Figure 8J The illustration shows user 802 performing an air gesture 836 (e.g., a double pinch gesture) within detection area 806. In some embodiments, the air gesture 836 is detected by wearable audio output device 301 (e.g., via sensor 311). Figure 8J It also shows the currently playing media item responding to air gesture 836 from... Figure 8I The “track 1” in the text has been changed to Figure 8J Track 2, as indicated by Media Item Indicator 820.

[0211] Figure 8K The image shows user 802 performing an air gesture 840 (e.g., a pinch gesture) within detection area 806. Figure 8K Music 822 (from) was also shown. Figure 8E Playback has resumed (e.g., music 822 is no longer paused), as indicated by user interface 816, which includes playback controls 818 (e.g., a pause element) in place of... Figure 8J The playback control 830 is shown in the image. Figure 8K As shown, due to the detection Figure 8J The aerial gesture 836 and music 822 in the middle are restored with "track 2".

[0212] Figure 8L The illustration shows a user 802 wearing a wearable audio output device 301 (e.g., earbuds) and an HMD 100b and performing a head gesture 841 (e.g., shaking, nodding, or other gestures). In some embodiments, the head gesture 841 is detected by the wearable audio output device 301 (e.g., via sensor 311). For example, the head gesture 841 is detected by an accelerometer in the wearable audio output device 301.

[0213] Figure 8M The detection area 806 is shown to be activated in response to a head gesture 841. Figure 8M Also shown is a portion 842-a (e.g., an initial portion) (e.g., a pinching portion of the gesture) within the detection area 806 where an air gesture 842 is detected. According to some embodiments, the air gesture 842 corresponds to a volume adjustment operation of audio output by the wearable audio output device 301. Figure 8M In the middle, the wearable audio output device 301 has an audio output volume level at level 844-a.

[0214] Figure 8N Part 842-b (e.g., a follow-up or second part) shows user 802 performing an air gesture 842 (e.g., a pinch and twist gesture). Figure 8N In the example, user 802 twists counterclockwise, which corresponds to the output volume changing from level 844-a (e.g., as shown in the image). Figure 8M (As shown) the volume is reduced to level 844-b. In some embodiments, the amount of volume adjustment is based on the speed and / or amount of movement in the air gesture 842. In some embodiments, the wearable audio output device 301 performs the volume adjustment operation (e.g., in response to a twisting motion) while the pinched portion of the gesture is held.

[0215] Figure 8O Part 842-c (e.g., another part) shows user 802 performing an air gesture 842 (e.g., a pinch and twist gesture). Figure 8O In the example, user 802 twists clockwise, which corresponds to the output volume changing from level 844-b (e.g., as shown in the image). Figure 8N (As shown) increases to level 844-c. In some embodiments, the amount of volume adjustment is based on the speed and / or amount of movement in the air gesture 842. In some embodiments, the wearable audio output device 301 performs the volume adjustment operation (e.g., in response to a twisting motion) while the pinched portion of the gesture is held.

[0216] Figure 8P This illustrates that user 802 ceases performing air gesture 842 (e.g., performing a fist gesture 846 within detection area 806). In response to user 802 ceasing to perform air gesture 842, wearable audio output device 301 ceases performing volume adjustment operations. In some embodiments, gesture execution ceases when the hand shape changes (e.g., ceasing to perform a pinch gesture) and / or when the hand leaves detection area 806. In some embodiments, air gesture 846 corresponds to an operation different from the operation corresponding to air gesture 842 (e.g., volume adjustment) (e.g., a power-off operation). Therefore, Figures 8A to 8P Examples of user air gestures (e.g., detected via wearable audio output device 301) and corresponding operations (e.g., performed by wearable audio output device 301 and / or HMD 100b) according to some implementations are illustrated.

[0217] Figures 9A to 9K Examples of user interactions with wearable devices are illustrated according to some implementation schemes. Figures 9A to 9KIn the example, user 802 wears wearable audio output device 301 and HMD 100b; however, in some embodiments, user 802 wears different types of wearable devices (e.g., different head-mounted devices, ear-worn devices, or other types of wearable devices). In some embodiments, user 802 does not wear HMD 100b (e.g., only wears wearable audio output device 301).

[0218] Figure 9A The illustration shows a user 802 wearing a wearable audio output device 301 and performing a gesture 906 (e.g., a pointing gesture). Figure 9A It is also shown that user 802 issues command 908 (e.g., verbal command) to turn on light 910. In some embodiments, turning on light 910 includes discrete operations (e.g., where a single output is a command to activate light 910). In some embodiments, wearable audio output device 301 detects command 908 via one or more microphones (e.g., microphone 302). In some embodiments, wearable audio output device 301 detects gesture 906 via one or more sensors (e.g., sensor 311). Figure 9A In the example, user 802 tucks his hair behind his ears so that the field of view of wearable audio output device 301 is not obstructed by the user's hair. Figure 9B The lamp 910 is shown to respond to Figure 9A The light is turned on by gesture 906 and command 908, as indicated by the lighting line 912. In some embodiments, the wearable audio output device 301 determines that the gesture 906 is directed at the light 910 (e.g., using data from sensor 311) and sends a command to the light 910 (and / or the controller of the light 910) to turn it on.

[0219] Figure 9C The illustration shows a user 802 wearing a wearable audio output device 301 and performing a gesture 916 (e.g., a pointing gesture). Figure 9C It also shows user 802 issuing command 914 to turn on light 910. In Figure 9C In one example, user 802's hair obscures wearable audio output device 301 and one or more of its sensors (e.g., sensor 311). In some embodiments, the obscured sensor is used to perform an operation requested by command 914 (e.g., turning on a light 910). For example, the obscured sensor includes an image sensor for determining what object a gesture 916 is pointing at.

[0220] Figure 9D A wearable audio output device 301 is shown in response to Figure 9CFeedback 918 is provided in response to command 914. In some embodiments, feedback 918 is provided based on the determination that a sensor (e.g., sensor 311) of the wearable audio output device 301 is obstructed. In some embodiments, feedback 918 includes audio and / or haptic feedback. In some embodiments, feedback 918 includes an indication of the reason for the failure to perform the operation requested by command 914. Figure 9E The image shows user 802 using gesture 920 to tuck his hair behind his ears to resolve the obstruction of the sensors of wearable audio output device 301 (e.g., to make the field of view of wearable audio output device 301 clear).

[0221] Figure 9F This demonstrates that the wearable audio output device 301 is no longer obscured by the hair of the user 802. Figure 9F It also shows a wearable audio output device 301 responding to user 802 in Figure 9E Feedback 922 is provided by moving the hair. In some embodiments, feedback 922 is provided based on determining that the sensors of the wearable audio output device 301 are no longer obstructed. In some embodiments, feedback 922 includes audio and / or haptic feedback. In some embodiments, feedback 922 includes one or more attributes different from feedback 918. For example, feedback 922 includes a single beep, while feedback 918 includes two or more beeps.

[0222] Figure 9G The illustration shows a user 802 wearing a wearable audio output device 301 and issuing a command 924 to activate a navigation assistance mode. In some embodiments, the wearable audio output device 301 detects the command 924 via one or more microphones (e.g., microphone 302). In some embodiments, the wearable audio output device 301 uses data from one or more sensors (e.g., sensor 311) to execute the navigation assistance mode. Figure 9G In one example, user 802 tucks their hair behind their ears so that the field of view of wearable audio output device 301 is not obstructed by the user's hair. In some implementations, the navigation assistance mode includes continuous operation (e.g., scanning the user's physical environment and providing navigation assistance).

[0223] Figure 9HA wearable audio output device 301 is shown providing audio feedback 926 based on navigation assist mode activity. In some embodiments, feedback 926 includes instructions for user 802 to follow. In some embodiments, feedback 926 includes haptic feedback (e.g., providing haptic vibration when user 802 should turn). In some embodiments, the navigation assist mode uses data from one or more sensors of the wearable audio output device 301 (e.g., location data from a geospatial sensor and / or image data from one or more image sensors). In some embodiments, the navigation assist mode uses data obtained from one or more other devices (e.g., portable multifunction device 100).

[0224] Figure 9I The illustration shows a user 802's hair obscuring a wearable audio output device 301, with the wearable audio output device 301 providing audio feedback 928. In some embodiments, the wearable audio output device 301 provides audio feedback 928 based on the determination that a sensor used for navigation assistance mode (e.g., an image sensor) is obstructed. In some embodiments, the wearable audio output device 301 provides haptic feedback as a supplement to or alternative to providing feedback 928. According to some embodiments, feedback 928 includes suggestions for actions that the user 802 can perform to attempt to resolve an error (e.g., remove the obstruction). Figure 9I In the example, feedback 928 included a suggestion from user 802 to tuck his hair behind his ears. Figure 9J The image shows user 892 using gesture 930 to tuck his hair behind his ears to resolve the obstruction of the sensors of wearable audio output device 301 (e.g., to make the field of view of wearable audio output device 301 clear).

[0225] Figure 9K This demonstrates that the wearable audio output device 301 is no longer obscured by the hair of the user 802. Figure 9K It also shows a wearable audio output device 301 responding to user 802 in Figure 9J Feedback 932 is provided by moving the hair. In some embodiments, feedback 932 is provided based on determining that the sensors of the wearable audio output device 301 are no longer obstructed. In some embodiments, feedback 932 includes audio and / or haptic feedback. In some embodiments, feedback 932 includes an indication that the sensors are no longer obstructed and / or an indication to re-enable navigation assistance mode.

[0226] Figures 10A to 10DThis is a flowchart illustrating a method 1000 for providing audio feedback related to real-world objects according to some embodiments. Method 1000 is performed at an ear-worn audio output device (e.g., a wearable audio output device 301, such as earbuds or headphones), which includes one or more sensors (e.g., sensor 311, such as image sensors, motion sensors, and / or other types of sensors) and one or more audio output components (e.g., speakers 306). In some embodiments, the ear-worn audio output device does not obstruct the user's eyes. In some embodiments, the ear-worn audio output device does not extend across the user's face. In some embodiments, the ear-worn audio output device does not affect the user's vision. In some embodiments, the ear-worn audio output device is an in-ear device (e.g., earbuds). In some embodiments, the ear-worn audio output device is an over-ear device (e.g., headphones). In some embodiments, the ear-worn audio output device is mounted to the user's ear (e.g., mounted to the ear canal, such as using earbuds; mounted to the earlobe, such as using earrings; and / or mounted to the helix of the ear, such as using hearing aids). Some operations in method 1000 may be optionally combined, and / or the order of some operations may be optionally changed.

[0227] As described below, method 1000 provides an improved interface for controlling a wearable audio output device by providing audio feedback related to real-world objects in response to user gestures. Detecting and responding to user gestures reduces the amount of input required to provide audio feedback and makes the user-device interface more efficient (e.g., by helping the user achieve expected results and reducing user errors when operating / interacting with the audio output device), which reduces power consumption and extends battery life (e.g., by mitigating the need to power the graphical user interface). Additionally, detecting and responding to user gestures allows the user to avoid directly manipulating the wearable audio output device, which enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0228] The ear-worn audio output device transmits audio data via one or more sensors (e.g., ...). Figure 3B Sensor 311 detects (1002) user gestures. For example, Figures 5A to 5B An example is shown where user 502 performs gesture 506, which is detected by wearable audio output device 301. Example user gestures include pointing gestures, tapping gestures, and tracing gestures. In some embodiments, the user gesture is an air gesture; in other embodiments, the user gesture is performed on the surface of a real-world object.

[0229] In some embodiments, the ear-worn audio output device includes (1004) an audio playback device. For example, the ear-worn audio output device includes one or more speakers (e.g., speaker 306). In some embodiments, the ear-worn audio output device is communicatively coupled to an accessory device (e.g., portable multifunction device 100, device 300, and / or HMD 100b) to play back audio provided by the accessory device.

[0230] In some implementations, user gestures are detected (1006) while audio content is being played back via one or more audio output components. For example, the audio content includes music, audio from audiobooks or podcasts, and / or other types of audio content. As an example, Figures 5C to 5D An example is illustrated where a wearable audio output device 301 detects a gesture 512 performed by a user 502 while music 520 is playing back. Detecting user gestures while providing playback of audio content reduces the amount of input required (e.g., the user does not need to manually stop playback) and allows the device to perform gesture detection automatically.

[0231] In some implementations, an accessory device communicatively coupled to the ear-worn audio output device receives (1008) audio content. For example, the accessory device is a telephone, smartwatch, music player, or other type of device. As an example, audio is received from a portable multifunction device 100. Figure 5C Music 520, as indicated by arrow line 514. Receiving audio content from the companion device reduces the memory required for operations performed by and / or by the ear-worn audio output device, which enhances the operability of the ear-worn audio output device, reduces power consumption, and extends the battery life of the ear-worn audio output device.

[0232] In some implementations, the accompanying equipment includes (1010) playback controls for controlling the playback of audio content at the ear-worn audio output device. For example, Figures 5C to 5DA user interface 516 including playback controls 518 and 524 is illustrated. In some embodiments, the companion device displays a user interface with audio playback controls. For example, the user interface includes multiple controls for the companion device and / or the earphone audio output device (e.g., play, pause, stop, skip forward, rewind, change audio source, change media item, change volume, and / or other controls). In some embodiments, the earphone audio output device detects user input pointing to a first control among the multiple controls (e.g., a pause or stop control) (e.g., occurring at the first control or when attention is drawn to the first control), and stops the playback of audio content at the earphone audio output device in response to detecting user input pointing to the first control. In some embodiments, the earphone audio output device detects user input pointing to or otherwise corresponding to a second control among the multiple controls (e.g., a skip forward, rewind, or change media item control), and changes that portion of the audio content being played back at the earphone audio output device in response to detecting user input pointing to the second control. In some implementations, the ear-worn audio output device detects user input pointing to a volume control among multiple controls, and in response to detecting user input pointing to the volume control, adjusts the output volume of the audio content being played back at the ear-worn audio output device. Presenting playback controls at the companion device enhances the operability of both the companion device and the ear-worn audio output device (e.g., provides flexibility) and makes the user-device interface more efficient.

[0233] In some implementations, the user gesture includes (1012) a boundary gesture outlining a portion of a first real-world object, and the first audio feedback includes audio feedback regarding real-world content on that portion of the first real-world object. As an example, Figures 5I to 5NAn example is illustrated where user 502 performs gesture 536, which outlines text in portion 538 of notebook 534, and wearable audio output device 301 provides feedback 542 regarding the text within portion 538. For example, a boundary gesture is a finger circling gesture. In some embodiments, a boundary gesture describes the boundary surrounding that portion of a first real-world object. In some embodiments, the ear-worn audio output device detects a second user gesture via one or more sensors of the ear-worn audio output device; and in response to detecting the second user gesture: based on determining that the second user gesture is a boundary-type gesture pointing to a portion of the first real-world object, the ear-worn audio output device provides third audio feedback corresponding to the real-world content on that portion of the first real-world object. In some embodiments, the real-world content on that portion of the first real-world object includes musical notes or a sequence of musical notes, and the audio feedback includes playback of the musical notes or the sequence of musical notes. Detecting and responding to different types of gestures in different ways enhances the operability of the device (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to access that flexibility) and makes the user device interface more efficient.

[0234] In some implementations, the real-world content on that portion of the first real-world object includes (1014) text, and the ear-worn audio output device provides an indication (e.g., a visual or audio indication) that the text from the first real-world object has been selected (e.g., by displaying a copy of the text on a companion device's display and / or outputting audio corresponding to the text at the ear-worn device). As an example, Figures 5O to 5R An example is illustrated where text in portion 554 is added to an active notebook in response to gesture 550, and wearable audio output device 301 provides feedback 556 indicating that the text has been selected. For example, a copy of the text is stored at the ear-worn audio output device and / or an accessory device communicating with the ear-worn audio output device. In some embodiments, audio feedback regarding the real-world content on this portion of a first real-world object includes an indication that text has been copied to the ear-worn audio output device and / or the accessory device (e.g., copied to the device's virtual clipboard). Providing an indication that text from a real-world object has been selected provides improved feedback regarding the state of the ear-worn audio output device.

[0235] In some implementations, the ear-worn audio output device appends (1016) text to a document associated with the ear-worn audio output device and / or a companion device. As an example, Figures 5Q to 5RThis example illustrates text from section 554 being added to notebook 558. For instance, text is added to an active notes document. In some implementations, user gestures indicate which document the text should be added to. For example, user gestures may include voice commands, and these voice commands may indicate which document the text should be added to. Adding text to a document in response to gestures allows text manipulation to be performed without displaying additional controls and reduces the amount of input required to perform text manipulation.

[0236] In some implementations, in response to detecting at least a portion of a user's gesture, the ear-worn audio output device provides (1020) feedback corresponding to the user's gesture. For example, Figures 50 to 5P Example 550-a illustrates a wearable audio output device 301 detecting a gesture 550 and, in response, providing feedback 552 regarding the gesture 550. In some embodiments, the feedback corresponding to the user's gesture includes audio and / or haptic feedback. As an example, the feedback corresponding to the user's gesture includes indications of the start, pause, and / or end of the gesture. For instance, the user gesture includes a touch and hold gesture, and the feedback corresponding to the gesture includes a sound at the start of the touch and hold gesture and / or a sound at the end of the touch and hold gesture. In some embodiments, the feedback corresponding to the user's gesture includes indications of the action corresponding to the gesture. For example, the user gesture points to a musical object (e.g., a music album or sheet music), and the feedback corresponding to the gesture includes an indication that music (e.g., a music preview) corresponding to the musical object will be played in response to the completion of the gesture. As another example, the user gesture points to an object with text in a language different from the language assigned to the user, and the feedback corresponding to the gesture indicates that a translation of the text will be played in response to the completion of the gesture. Providing feedback corresponding to the user's gesture provides improved feedback regarding the state of the ear-worn audio output device.

[0237] In some implementations, the feedback corresponding to the user's gesture includes (1022) feedback indicating the pause of the user's gesture. As an example, Figure 5MA wearable audio output device 301 is shown providing feedback 544 regarding a gesture 540, which can indicate the duration of the gesture 540. For example, the gesture is a touch and hold gesture, and the feedback indicating the duration of the user's gesture is feedback indicating that the touch and hold gesture has been held for a threshold amount of time. In some embodiments, the user is provided with dwell feedback based on the duration of the user's gesture meeting the threshold amount of time; and no dwell feedback is provided based on the duration of the user's gesture not meeting the threshold amount of time. For example, if the user's gesture ends before the threshold amount of time, no dwell feedback is provided. As another example, the gesture is a touch and drag gesture, and the feedback indicating the duration of the user's gesture is feedback indicating that the dragging portion of the gesture is being performed. Indicating the duration of the user's gesture provides improved feedback regarding the state of the ear-worn audio output device.

[0238] In some implementations, the feedback corresponding to the user's gesture includes (1024) feedback indicating the progress of the user's gesture. As an example, Figure 5M A wearable audio output device 301 is shown providing feedback 544 regarding a gesture 540, which can indicate the progress of the gesture 540. For example, feedback corresponding to a user gesture includes progressive feedback indicating the progress of the interaction toward an input threshold (e.g., an input threshold of 0.5 seconds, 1 second, 1.5 seconds, or 2 seconds). In some embodiments, one or more properties of the progressive feedback change over time (e.g., to indicate progress toward the input threshold). Example properties of progressive feedback include pitch, amplitude, rhythm, frequency, and / or other audio (and / or haptic) properties. Providing feedback indicating the progress of a user's gesture provides improved feedback on the state of the ear-worn audio output device.

[0239] In response to detecting a user gesture (1018) and determining that the user gesture is a first type of gesture (e.g., an index finger pointing gesture) and points to a first real-world object, the ear-worn audio output device provides (1026) a first audio feedback corresponding to the first real-world object via one or more audio output components. For example, in Figures 5A to 5B In the process, user 502 performs a gesture 506 pointing at real-world object 508-1, and in response, wearable audio output device 301 provides feedback 510 about real-world object 508-1.

[0240] In some embodiments, the ear-worn audio output device provides (1028) first audio feedback without providing visual feedback. In some embodiments, the ear-worn audio output device does not include a display and does not provide visual feedback (e.g., only audio is provided, and optionally haptic feedback is provided). For example, it may provide... Figure 5BThe feedback 510 is provided without any corresponding visual feedback (e.g., in an implementation where the user 502 is not wearing the HMD 100b). Forgoing visual feedback reduces the number of components required in the ear-worn audio output device (e.g., no need for display generation components), which reduces power consumption and extends the battery life of the ear-worn audio output device.

[0241] In some implementations, the first audio feedback corresponding to the first real-world object includes (1030) a description of the first real-world object. For example, Figure 5B The feedback 510 includes a description of the real-world object 508-1. In some embodiments, the first audio feedback corresponding to the second real-world object includes a description of the second real-world object (e.g., size, shape, color, purpose, and / or other descriptive information). In some embodiments, the first audio feedback includes the relative position of the first real-world object (e.g., text describing the relative position or information corresponding to the relative position). Providing a description of the real-world object in response to user gestures enhances the operability of the ear-worn audio output device (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to achieve this flexibility) and makes the user-device interface more efficient.

[0242] In some implementations, the first audio feedback corresponding to the first real-world object includes (1032) an indication of operable data associated with the first real-world object. For example, Figure 5A Real-world objects 508-1 in the code include barcode 509. Figure 5D The real-world object 508-2 in the text includes the current time, and Figure 5F The real-world object 508-3 in the context includes a date. For example, operable data includes time, date, phone number, email address, and / or machine-readable code (e.g., QR code or App Clip code). In some embodiments, the operable data is displayed on the surface of and / or by the first real-world object. In some embodiments, the ear-worn audio output device initiates interaction based on input of a machine-readable code, indicated by gestures (e.g., pointing or touching). Providing indications of operable data associated with the first real-world object in response to user gestures enhances the operability of the ear-worn audio output device (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to achieve this flexibility) and makes the user-device interface more efficient.

[0243] In some implementations, the first audio feedback corresponding to the first real-world object includes (1034) contextual information about the user of the ear-worn audio output device. For example, Figure 5B Feedback 510 includes contextual information about real-world objects 508-1 from the user's mother. Figure 5D Feedback 522 includes contextual information about real-world object 508-2 being the bedroom clock of user 502, and Figure 5F Feedback 530 includes contextual information about user 502's appointment on a given date. Example contextual information includes information about the user's schedule (e.g., whether they are available at a given time), the user's past experiences with a first real-world object, the user's experiences with similar real-world objects, information about the user's preferences, and / or other types of contextual information.

[0244] In some implementations, the first audio feedback corresponding to the first real-world object includes an indication of a future time period associated with the first real-world object, and the contextual information includes information about whether the user has free time during the future time period (e.g., as provided by...). Figure 5F (As indicated by feedback 530 in the audio). For example, the first audio feedback includes an indication of a specific date and an indication of whether the user is available on that date. Providing indications of future time periods associated with real-world objects in response to user gestures enhances the operability of the ear-worn audio output device (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to achieve that flexibility) and makes the user-device interface more efficient.

[0245] In some implementations, the first audio feedback is spatialized (1038) to a first position based on the first real-world object being in a first position, and spatialized to a second position based on the first real-world object being in a second position. For example, Figures 6C to 6FExamples include providing audio feedback 616 at a location corresponding to real-world object 610-1 and audio feedback 622 at a location corresponding to real-world object 610-3. For example, spatializing the first audio feedback to a first location includes indicating the direction and distance between the ear-worn audio output device and the first location via the first audio feedback, and spatializing the first audio feedback to a second location includes indicating the direction and distance between the ear-worn audio output device and the second location via the first audio feedback. Spatialized audio feedback simulates a more realistic listening experience, where the audio appears to originate from a sound source in a specific frame of reference, such as the physical environment surrounding the user. For example, the first audio feedback is provided as spatial audio from a simulated location corresponding to the relative position of the first real-world object to the ear-worn audio output device. Providing spatial feedback enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping the user achieve the expected results and reducing user errors when operating / interacting with the device), which additionally reduces power consumption and extends the device's battery life by enabling the user to use the device more quickly and efficiently.

[0246] As an example, when spatial audio is enabled, the audio output from an ear-worn audio output device (e.g., earbuds) sounds as if the corresponding audio for each real-world object comes from a different simulated spatial location (which may change over time) in a frame of reference (such as the physical environment) (e.g., surround sound effect). The positioning of the real-world object (e.g., simulated spatial location) is independent of the movement of the earbuds relative to the frame of reference.

[0247] As an example, the simulated spatial position of one or more real-world objects is fixed relative to a reference frame when stationary, and moves relative to the reference frame when moving. For example, in the case where the reference frame is the physical environment, one or more real-world objects have corresponding simulated spatial positions in the physical environment. As an earphone audio output device moves in the physical environment, the audio output from the earphone audio output device is automatically adjusted due to user adjustments, so that the audio continues to sound as if it comes from one or more real-world objects at their respective spatial positions in the physical environment. As one or more real-world objects move through a series of spatial positions in the physical environment, the audio output from the earphone audio output device is adjusted so that the audio continues to sound as if it comes from one or more real-world objects at that series of spatial positions in the physical environment. Such adjustments for moving sound sources also take into account any movement of the earphone audio output device relative to the physical environment. For example, if the earphone audio output device moves relative to the physical environment along a path similar to that of a moving real-world object in order to maintain a constant spatial relationship with the real-world object, audio will be output such that the sound appears not to have moved relative to the earphone audio output device.

[0248] In some implementations, the first real-world object includes a (1040) hand-drawn picture, and the first audio feedback includes information indicated by the hand-drawn picture. For example, Figures 5I to 5L An example is illustrated with a notebook 534 containing handwritten text and feedback 542 indicating a portion 538 of the handwritten text. For example, the handwritten image includes one or more items, and the first audio feedback includes one or more items or a description of one or more items. In some embodiments, the handwritten image indicates an action to be performed by an ear-worn audio output device and / or an accompanying device communicating with the ear-worn audio output device. In some embodiments, the first audio feedback includes an instruction for the action. In some embodiments, the ear-worn audio output device performs the action in response to detecting a user gesture. For example, the handwritten image is a labeled button (e.g., a call button, a pause button, a mute button, or other type of button), and the action corresponds to the button (e.g., initiating a phone call, pausing audio content, or muting the microphone of the ear-worn audio output device). In some embodiments, the handwritten image includes one or more instructions for the ear-worn audio output device and / or the accompanying device, and the ear-worn audio output device performs one or more actions based on one or more instructions. For example, one or more instructions include “Call Sally,” and one or more actions include initiating a phone call to a contact labeled “Sally” in the user’s contact list. In some implementations, the user gesture includes pointing to or tapping a hand-drawn image. In some implementations, the first real-world object includes multiple hand-drawn images, the user gesture points to the first hand-drawn image among the multiple hand-drawn images, and the first audio feedback includes information associated with the first hand-drawn image. In some implementations, the first real-world object includes multiple hand-drawn images, the user gesture points to a second hand-drawn image among the multiple hand-drawn images, and the first audio feedback includes information associated with the second hand-drawn image. In some implementations, providing the first audio feedback includes providing information about a first portion of the hand-drawn image based on determining that the user gesture points to a second portion of the hand-drawn image, and providing the first audio feedback includes providing information about a second portion of the hand-drawn image based on determining that the user gesture points to a second portion of the hand-drawn image. Providing information indicated by a hand-drawn image in response to a user gesture enhances the operability of the ear-worn audio output device and makes the user-device interface more efficient.

[0249] In response to detecting a user gesture (1018) and determining that the user gesture is a first type of gesture and points to a second real-world object, the ear-worn audio output device provides (1042) second audio feedback corresponding to the second real-world object via one or more audio output components. For example, Figures 5A to 5B This example illustrates a user 502 making a gesture toward a real-world object 508-1 (e.g., a first real-world object), and Figures 5C to 5DThe example shows user 502 making a gesture toward real-world object 508-2 (e.g., a second real-world object). Figure 5D The wearable audio output device 301 is also shown providing feedback 522 regarding real-world objects 508-2. In some embodiments, the ear-worn audio output device is a set of earbuds, and provides first and second audio feedback at a subset of these earbuds (e.g., audio feedback is output only at the right earbud). In some embodiments, audio feedback is provided at specific earbuds based on preference settings. In some embodiments, the ear-worn audio output device provides second audio feedback but not visual feedback.

[0250] In some embodiments, the ear-worn audio output device detects (1044) a user voice command corresponding to a user gesture, wherein at least one parameter of the first audio feedback is based on the user voice command. For example, the user points to a first real-world object and asks a question or gives a verbal command. In some embodiments, the user voice command is detected simultaneously with the user gesture. In some embodiments, the user voice command is detected within a predefined threshold time before or after the user gesture is detected. In some embodiments, if the user voice command is detected within the threshold time of the user gesture, it is determined that the user voice command corresponds to the user gesture. In some embodiments, a first type of gesture includes (or optionally includes) a voice command component. In some embodiments, a second type of gesture does not include a voice command component. In some embodiments, the first audio feedback includes first information depending on whether the user voice command is a first type of voice command, and includes second information different from the first information depending on whether the user voice command is a second type of voice command. For example, a user voice command “What is that?” causes the first audio feedback to include a description of the first real-world object, and a user voice command “What is written on that?” causes the first audio feedback to include a textual description of the first real-world object. In some implementations, in response to a third type of voice command, the ear-worn audio output device performs an action different from providing first or second audio feedback. Detecting voice commands and providing feedback based on those commands enhances the operability of the ear-worn audio output device and makes the user-device interface more efficient.

[0251] In some implementations, at least one parameter of the first audio feedback is based on the (1046) tone of the user's voice command. For example, Figures 6E to 6H Example of user 602 in Figure 6E The question was raised in section 620 and... Figure 6G The problem is with the soft voice channel 630. Figures 6E to 6HThe volume level 626 and detail level in the feedback also vary in response to a whispered question and a spoken question. In some embodiments, the volume of the first audio feedback is based on the tone of the voice. For example, a whispered voice command results in a lower volume of the first audio feedback compared to a louder voice command. In some embodiments, the level of detail in the first audio feedback is based on the tone of the voice. For example, a whispered voice command results in a more concise first audio feedback compared to a louder voice command. In some embodiments, the type of information in the first audio feedback is based on the tone of the voice. As an example, a whispered user voice command elicits different audio feedback than a spoken or shouted user voice command. In some embodiments, parameters of the first audio feedback (e.g., volume, speed, and / or other types of parameters) are set to a first value based on the user voice command having a first tone, and parameters of the first audio feedback are set to a second value different from the first value based on the user voice command having a second tone different from the first tone. In some implementations, a first audio feedback is provided based on a user's voice command with a first tone, and an operation different from providing the first or second audio feedback is performed based on a user's voice command with a third tone. Providing parameterized feedback based on voice tone enhances the operability of ear-worn audio output devices and makes the user-device interface more efficient.

[0252] In some implementations, the ear-worn audio output device detects (1048) a second user gesture via one or more sensors of the ear-worn audio output device; and in response to detecting the second user gesture: based on determining that the ear-worn audio output device is within a predefined area and determining that the user gesture is a first type of gesture and points to a first real-world object, the ear-worn audio output device provides second audio feedback corresponding to the first real-world object via one or more audio output components; and based on determining that the ear-worn audio output device is not within the predefined area, the ear-worn audio output device abandons providing the first audio feedback (e.g., regardless of whether the user gesture is a first type of gesture and points to the first real-world object). For example... Figures 6A to 6HAn example is illustrated where user 602 enters a geofence boundary and gestures toward object 610, receiving feedback about object 610. In some embodiments, first audio feedback is provided when the ear-worn audio output device is within a predefined area; and the ear-worn audio output device abandons providing the first audio feedback if it is determined that the ear-worn audio output device is not within the predefined area and that the user's gesture is a first type of gesture and points to a first real-world object. In some embodiments, the ear-worn audio output device provides instructions (e.g., audio and / or haptic feedback) to the user when the user enters and / or leaves the geofence area. In some embodiments, the ear-worn audio output device provides audio feedback corresponding to a real-world object only when the ear-worn audio output device and / or the real-world object are within the geofence location. Providing feedback within a predefined area enhances the operability of the ear-worn audio output device and makes the user-device interface more efficient, and reduces power consumption and extends the battery life of the ear-worn audio output device by being able to disable components (e.g., sensors) when outside the predefined area.

[0253] It should be understood that, Figures 10A to 10D The specific order in which the operations are described herein is merely an example and is not intended to indicate that this order is the only possible order in which these operations can be performed. Those skilled in the art will conceive of various ways to reorder the operations described herein. Furthermore, it should be noted that the details of other processes described herein in conjunction with other methods (e.g., methods 1100, 1200, and 1300) are similarly applicable to the above-described combinations. Figure 10A-10D The method 1000. For example, the inputs, gestures, functions, and feedback described above with reference to method 1000 optionally have one or more of the characteristics of inputs, gestures, functions, and feedback described herein with reference to other methods described herein (e.g., methods 1100, 1200, and 1300). For the sake of brevity, these details will not be repeated here.

[0254] Figures 11A to 11C This is a flowchart illustrating method 1100 for providing audio feedback on alarm conditions according to some embodiments. Method 1100 is performed at a wearable audio output device (e.g., wearable audio output device 301, such as earbuds or headphones), which includes one or more audio output components (e.g., speaker 306) and optionally includes one or more sensors (e.g., sensor 311, such as image sensors, motion sensors, and / or other types of sensors). Some operations in method 1100 are optionally combined, and / or the order of some operations is optionally changed.

[0255] As described below, method 1100 provides audio output in an intuitive and efficient manner by offering audio feedback and adjusting modifications to ambient sounds. For example, when an alarm condition is detected in the surrounding physical environment, the ambient sound modifications are automatically adjusted to allow the user to hear more of the surrounding sounds. In this way, the wearable audio output device better adapts to the current state of the surrounding physical environment without requiring additional input from the user. Providing an adaptive and more intuitive user experience while reducing the amount of input required to achieve this enhances the operability of the device and makes the user-device interface more efficient (e.g., by helping users achieve expected results and reducing user errors when operating / interacting with the device), thereby further reducing power consumption and extending the device's battery life by enabling users to use the device more quickly and efficiently.

[0256] When the wearable audio output device is physically positioned relative to a corresponding part of the user's body, the wearable audio output device detects (1102) an alarm condition related to the user's spatial context, in which the magnitude of ambient sound from the physical environment is modified by the wearable audio output device (e.g., via active noise cancellation, passive noise cancellation, and / or active audio transparency) to have a first ambient sound level. For example, Figures 7A to 7B An example is illustrated where user 702 approaches puddle 706 while active at level 710-a in ANC mode (e.g., representing an alarm condition). In some embodiments, the alarm condition is detected via one or more sensors of the wearable audio output device (e.g., one or more image sensors, one or more audio sensors, and / or one or more other types of sensors). In some embodiments, the wearable audio output device includes in-ear and / or over-ear components (e.g., speakers) configured to provide audio feedback. For example, the wearable audio output device is a head-mounted device, such as a headset (e.g., an augmented reality headset), headphones, earbuds, glasses, or earrings. In some embodiments, the wearable audio output device includes one or more audio output components (e.g., speakers).

[0257] In some embodiments, the wearable audio output device is (1104) an ear-worn audio output device (e.g., wearable audio output device 301, such as earbuds or headphones), and alarm conditions are detected via one or more sensors of the ear-worn audio output device. In some embodiments, the ear-worn audio output device does not obstruct the user's eyes. In some embodiments, the ear-worn audio output device does not extend across the user's face. In some embodiments, the ear-worn audio output device does not impair the user's vision. In some embodiments, the ear-worn audio output device is an in-ear device (e.g., earbuds). In some embodiments, the ear-worn audio output device is an over-ear device (e.g., headphones). In some embodiments, the ear-worn audio output device is mounted on the user's ear. Detecting alarm conditions via sensors of the ear-worn audio output device enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0258] In some implementations, the wearable audio output device provides playback of (1106) audio content at the wearable audio output device and detects alarm conditions related to the user's spatial context while providing playback of the audio content. For example, Figures 7A to 7B An example is illustrated where a wearable audio output device 301 detects an alarm condition corresponding to a puddle 706 while music 704 is being provided. The audio content may include, for example, music, spoken audio (e.g., podcasts or audiobooks), and / or other types of audio content. In some embodiments, the audio content is received from a companion device that communicates with the wearable audio output device. Detecting the alarm condition while the audio content is being played reduces the amount of input required (e.g., the user does not need to manually stop playback) and allows the device to perform alarm condition detection automatically.

[0259] In some implementations, in response to the detection of an alarm condition, the wearable audio output device pauses playback of audio content (and / or reduces its volume). For example, Figures 7E to 7FAn example is illustrated where a wearable audio output device 301 detects a ball 720 approaching from behind a user 702 (e.g., indicating an alarm condition) and pauses media playback, as indicated by a media playback indicator 718. In some embodiments, playback of the audio content is paused upon providing audio feedback. In some embodiments, playback of the audio content is paused until the alarm condition ends. In some embodiments, playback of the audio content is paused for a threshold amount of time. In some embodiments, playback of the audio content is paused until input to resume playback is received from the user. In some embodiments, for example, playback of the audio content automatically resumes (e.g., without additional user input) after providing audio feedback or after a threshold amount of time following the provision of audio feedback. Automatically pausing playback of the audio content reduces the amount of input required (e.g., the user does not need to manually pause) and allows the device to automatically perform alarm condition detection.

[0260] In some implementations, the alarm condition is based on the user of the (1108) wearable audio output device being within a threshold distance of a potential hazard. For example, in Figures 7A to 7B In the example, an alarm is issued to user 702 regarding puddle 706 after approaching the puddle (e.g., based on being within a threshold distance of the puddle). In some embodiments, the wearable audio output device determines that the user is within a threshold distance of a potential hazard (e.g., using a camera and / or other types of sensors). In some embodiments, the audio feedback changes based on the user's proximity to the potential hazard. For example, a more urgent warning is given if the user is closer to the potential hazard. In some embodiments, one or more attributes of the audio feedback (e.g., pitch, frequency, and / or other audio attributes) are based on the user's proximity to the potential hazard. In some embodiments, the audio feedback indicates the proximity of the potential hazard (e.g., using verbal feedback). In some embodiments, the wearable audio output device waives the provision of audio feedback regarding the potential hazard if it is determined that the wearable audio output device is not within a threshold distance of the potential hazard. For example, an alarm condition is not met when the wearable audio output device is greater than a threshold distance from the potential hazard. In some embodiments, the wearable audio output device identifies alarm conditions from a predefined list of alarm conditions. In some implementations, the wearable audio output device (and / or the device communicating with it) dynamically (e.g., in real time) determines the presence of alarm conditions based on one or more hazard detection criteria (e.g., comparing the user's spatial context against one or more hazard detection criteria). Potential hazards include pits, holes, sharp or other dangerous objects, high-speed objects, and heavy machinery. Detecting alarm conditions based on threshold distance enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0261] In some implementations, the alarm condition is based on the user of a (1110) wearable audio output device moving toward a potential hazard. For example, in Figures 7A to 7B In the example, from a greater distance (e.g., at) Figure 7A (middle) After approaching the puddle, in Figure 7B An alarm is issued to user 702 regarding puddle 706. In some embodiments, the wearable audio output device determines that the user is moving toward a potential hazard (e.g., using a camera and / or other types of sensors). In some embodiments, the alarm condition is based on the user's direction of travel (e.g., using a gyroscope to determine this). In some embodiments, the alarm condition is based on the speed at which the user approaches the potential hazard. In some embodiments, the alarm condition is based on an estimated amount of time before the user is in the potential hazard. In some embodiments, the alarm condition is not met if the user of the wearable audio output device is not moving toward the potential hazard (e.g., maintaining distance from or moving away from the potential hazard). Detecting alarm conditions based on user movement enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0262] In some implementations, the alarm condition is based on (1112) a real-world entity approaching the user of the wearable audio output device. For example, in Figures 7E to 7F In the example, an alarm is issued to user 702 regarding ball 720 after the ball begins to approach user 702 from behind (e.g., based on ball 720 approaching user 702). For example, the real-world entity is a car, bicycle, or other vehicle. In some embodiments, the wearable audio output device determines that a real-world entity is approaching the user (e.g., using a camera and / or other types of sensors). In some embodiments, the alarm condition is based on the speed at which the real-world entity approaches the user. In some embodiments, the alarm condition is based on the type of real-world entity. In some embodiments, the audio feedback indicates the type of real-world entity, the speed of the real-world entity, and / or the relative position of the real-world entity relative to the user. In some embodiments, the alarm condition is not met when the real-world entity is not approaching the user (e.g., maintaining distance from the user or moving away from the user). Detecting alarm conditions based on entity movement enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0263] In some implementations, the real-world entity is a (1114) person, and the audio feedback indicates that the person is approaching the user of the wearable audio output device. For example, Figures 7G to 7HAn example is illustrated where a person 730 approaches a user 702, and a wearable audio output device 301 provides feedback 736 indicating that the person 730 is approaching. In some embodiments, the wearable audio output device determines that the real-world entity is a person. In some embodiments, the wearable audio output device determines that the person is attempting to attract the user's attention (e.g., the person waves and / or calls to the user), and the audio feedback indicates this situation. In some embodiments, the wearable audio output device determines that the person is moving in a direction of travel where a collision with the user is possible, and the audio feedback indicates this situation. In some embodiments, the person is a known contact of the user, and the audio feedback indicates the person's identity. In some embodiments, an alarm condition is not met (e.g., no audio feedback is provided) as the person moves away from the user of the wearable audio output device. Providing an indication that a person is approaching enhances the operability of the wearable audio output device (e.g., provides flexibility without cluttering the user interface with additional display controls and reduces the amount of input required to achieve this flexibility) and makes the user-device interface more efficient.

[0264] In some implementations, alarm conditions are based on the spatial context following the (1116) user. For example, Figures 7C to 7F An alarm condition based on the approach of ball 720 from behind user 702 is illustrated. In some embodiments, the alarm condition is based on spatial context outside the user's field of view. In some embodiments, audio feedback indicates the spatial context. In some embodiments, the alarm condition is detected based on the satisfaction of one or more criteria (e.g., one or more hazard detection criteria described above) of the spatial context behind the user, and the alarm condition is not satisfied (e.g., no audio feedback is provided) if one or more criteria are not satisfied of the spatial context behind the user. Detecting alarm conditions based on the spatial context behind the user enhances the operability of wearable audio output devices and makes the user-device interface more efficient.

[0265] In some implementations, the alarm condition includes (1118) one or more parameters, and the one or more parameters vary based on whether the alarm condition is within the estimated or detected field of view of the user of the wearable audio output device. For example, Figures 7C to 7FAn example is given where an alarm condition is triggered when ball 720 approaches from behind user 702, but not when ball 720 approaches from in front of user 702. For example, different criteria are used for the alarm condition when the event occurs behind the user rather than in front of the user. For instance, a car approaching from in front of the user while the user is walking along a curb may not trigger an alarm condition, but a car approaching from behind the user may trigger an alarm condition (e.g., to warn the user not to drift out of the lane). In some implementations, an alarm condition is met based on a first set of values ​​for one or more parameters, and is not met based on a second set of values ​​for one or more parameters. Differently detecting alarm conditions based on the user's estimated or detected field of view enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0266] In some implementations, the alarm conditions are based (1120) on audio from the physical environment in which the wearable audio output device is operating. For example, Figures 7I to 7K An example is illustrated where an alarm condition is triggered in response to people 740 and 742 arguing with each other. In some embodiments, the audio is detected by a wearable audio output device. In some embodiments, the audio is received via one or more microphones communicating with the wearable audio output device (e.g., externally or internally). For example, the audio includes sirens, alarms, shouts, and / or other types of audio. Detecting alarm conditions based on audio from the physical environment enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0267] In some implementations, alarm conditions are based (1122) on the rhythm of audio from the physical environment in which the wearable audio output device is operating. For example, the rhythm of the audio can indicate that an object is approaching the user. As another example, the rhythm of the audio can indicate the urgency of the situation. In some implementations, alarm conditions are based on changes in the rhythm of the audio. In some implementations, alarm conditions are based on the frequency of the audio. Detecting alarm conditions based on the rhythm of audio from the physical environment enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0268] In some implementations, the alarm condition is based (1124) on the tone of the audio from the physical environment in which the wearable audio output device is operating. Figures 7I to 7K An example is given of an alarm condition triggered in response to people 740 and 742 arguing with each other. For example, the tone of the conversation detected in the physical environment in which the wearable audio output device is operating indicates an argument. Detecting alarm conditions based on the tone of audio from the physical environment enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0269] In response to the detection of an alarm condition (1126), the wearable audio output device provides (1128) audio feedback regarding the alarm condition. For example, in Figure 7B In this embodiment, the wearable audio output device provides feedback 714 regarding alarm conditions involving the puddle 706. In some embodiments, the audio feedback includes a description of the alarm conditions. In some embodiments, the audio feedback indicates a potential hazard associated with the alarm conditions. In some embodiments, the audio feedback is provided as spatial audio indicating the relative location of the source of the alarm conditions. In some embodiments, the audio feedback indicates the relative location of the alarm conditions. In some embodiments, the audio feedback includes suggestions for avoiding the potential hazards.

[0270] In response to the detection of an alarm condition (1126), the wearable audio output device changes (1130) (e.g., automatically changes in the absence of additional input from the user) one or more attributes of the wearable audio output device to modify the magnitude of ambient sound from the physical environment (e.g., by reducing the degree of active noise cancellation and / or increasing the degree of active transparency) to a second ambient sound audio level that is louder than a first ambient sound audio level. For example, Figures 7A to 7B An example is illustrated where a wearable audio output device 301 raises the ANC level from [previous setting] in response to detecting an alarm condition involving a puddle 706. Figure 7A The level 710-a in the middle decreased to Figure 7B Level 710-b.

[0271] In some implementations, altering one or more properties of the wearable audio output device to modify the magnitude of ambient sound from the physical environment includes (1132) reducing the degree of active noise cancellation and / or increasing the degree of active transparency. For example, Figures 7G to 7H An example is illustrated where a wearable audio output device 301 raises the ANC level from [previous setting] in response to an alarm condition that detects a person 730. Figure 7G The level 716-a in the middle decreased to Figure 7H Level 716-b. Figures 7G to 7J H also exemplifies how a wearable audio output device 301, in response to an alarm condition detecting an individual 730, increases the active transparency level from... Figure 7G Level 726-a increased to Figure 7HLevel 726-b. In some embodiments, changing one or more properties of the wearable audio output device to modify the magnitude of ambient sound from the physical environment includes switching from a noise cancellation mode to an active transparency mode. In some embodiments, changing the degree of active noise cancellation (ANC) includes disabling the ANC mode. In some embodiments, changing the ANC degree includes reducing the ANC from a percentage above a threshold (e.g., 90%, 80%, 75%, or 50%) to a percentage below a threshold. For example, before an alarm condition is detected, the ANC reduces the ambient sound to 20%, 10%, or 5% of the ambient sound audio level, and in response to the detection of an alarm condition, the ANC reduces the ambient sound to 80%, 75%, 70%, or 50% of the ambient sound audio level. In some embodiments, changing the degree of active transparency includes enabling an active transparency mode. In some embodiments, changing the degree of active transparency includes increasing the active transparency from a percentage below a threshold (e.g., 15%, 20%, 30%, or 50%) to a percentage above a threshold. For example, before an alarm condition is detected, active transparency increases the ambient sound level to 10%, 20%, or 35%, and in response to the detection of an alarm condition, it increases the ambient sound level to 80%, 75%, 70%, or 50%. In some embodiments, the wearable audio output device changes one or more properties of the wearable audio output device to modify the magnitude of ambient sound from the physical environment based on determining that the wearable audio output device is operating in a first state (e.g., where ANC is enabled). In some embodiments, the wearable audio output device abandons changing one or more properties of the wearable audio output device based on determining that the wearable audio output device is operating in a second state (e.g., where ANC is disabled). Reducing the degree of active noise cancellation and / or increasing the degree of active transparency based on alarm conditions enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0272] In some implementations, altering one or more properties of the wearable audio output device to modify the magnitude of ambient sound from the physical environment includes (1134) amplifying the ambient sound to a volume level higher than the volume level of ambient sound in the physical environment. For example, Figures 7G to 7HAn example is illustrated where a wearable audio output device 301 activates a dialogue enhancement mode in response to the detection of an alarm condition involving a person 730, as indicated by enhancement indicator 728. For example, ambient sounds are amplified to assist users with hearing difficulties or to emphasize relatively quiet alarm conditions. As an example, a warning shouted from a distance in a noisy environment can be amplified so that the user of the wearable audio output device can hear and understand the warning amidst the noise. Amplifying ambient sounds based on alarm conditions enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0273] In some implementations, the wearable audio output device enables (1136) indication of alarm conditions to be provided at an accessory device communicatively coupled to the wearable audio output device. For example, Figure 7K An example is illustrated where an alarm 750 is provided at a portable multifunction device 100 in response to the detection of alarm conditions involving persons 740 and 742. For example, a notification is displayed on the companion device's display. In some embodiments, the indication of the alarm condition includes visual, audio, and / or tactile feedback. In some embodiments, the indication of the alarm condition includes an image of a potential hazard associated with the alarm condition. Providing indication of the alarm condition at the companion device enhances the operability of the wearable audio output device and provides improved feedback regarding the alarm condition.

[0274] It should be understood that, Figures 11A to 11C The specific order in which the operations are described herein is merely exemplary and not intended to indicate that this order is the only possible order in which these operations can be performed. Those skilled in the art will conceive of various ways to reorder the operations described herein. Furthermore, it should be noted that the details of other processes described herein in conjunction with other methods (e.g., methods 1000, 1200, and 1300) are similarly applicable to the above-described combinations. Figure 11A-11C The method 1100. For example, the inputs, gestures, functions, and feedback described above with reference to method 1100 optionally have one or more of the characteristics of the inputs, gestures, functions, and feedback described herein with reference to other methods described herein (e.g., methods 1000, 1200, and 1300). For the sake of brevity, these details will not be repeated here.

[0275] Figures 12A to 12DThis is a flowchart illustrating a method 1200 for performing operations in response to air gestures, according to some embodiments. Method 1200 is performed at a wearable audio output device (e.g., wearable audio output device 301, such as earbuds or headphones), which includes one or more audio output components (e.g., speaker 306) and optionally includes one or more sensors (e.g., sensor 311, such as image sensors, motion sensors, and / or other types of sensors). Some operations in method 1200 are optionally combined, and / or the order of some operations is optionally changed.

[0276] As described below, method 1200 provides an improved interface for controlling a wearable audio output device by performing actions in response to a user's hand gestures. Detecting and responding to user hand gestures reduces the amount of input required to provide audio feedback and makes the user-device interface more efficient (e.g., by helping the user achieve expected results and reducing user errors when operating / interacting with the audio output device), which reduces power consumption and extends battery life (e.g., by mitigating the need to power the graphical user interface). Additionally, detecting and responding to user gestures allows the user to avoid directly manipulating the wearable audio output device, which enhances the operability of the wearable audio output device and makes the user-device interface more efficient.

[0277] When outputting audio content, the wearable audio output device detects (1202) gestures performed by the user's hand. For example, Figure 8B A wearable audio output device 301 is shown detecting a gesture 804. In some embodiments, the wearable audio output device includes in-ear and / or over-ear components (e.g., speakers) configured to provide audio feedback. For example, the wearable audio output device is a head-mounted device, such as a headset (e.g., an augmented reality headset), headphones, earbuds, glasses, or earrings. In some embodiments, the wearable audio output device includes one or more audio output components (e.g., speakers). In some embodiments, a first operation changes the operating state of the wearable audio output device (e.g., from ANC mode activity to Active Transparency mode activity).

[0278] In some implementations, the gestures include (1204) air gestures. For example, Figure 8GGesture 828 in the document is an air gesture. For example, an air gesture does not involve contact with any surface or part of the user's body. As an example, an air gesture is a cupping gesture performed near the user's ear. An air gesture is a gesture detected without the user touching (or without regard to) an input element that is part of the device and based on the detected movement of a part of the user's body (e.g., head, one or two arms, one or two hands, one or more fingers, and / or one or two legs) through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture including movement of the hand in a predetermined amount and / or speed in a predetermined pose, or a shaking gesture including rotation of a part of the user's body at a predetermined speed or amount)). Detecting air gestures enhances the operability of wearable audio output devices and makes the user-device interface more efficient.

[0279] In some implementations, pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other, optionally followed by an immediate (e.g., within 0.01 seconds to 1 second) interruption of contact. A long pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before an interruption of contact is detected. For example, a long pinch gesture includes the user holding a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some implementations, a double pinch gesture as an air gesture includes two (e.g., more) pinch inputs (e.g., performed by the same hand) that are detected consecutively with each other immediately (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts the contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or 2 seconds). Detecting and responding to different types of gestures in different ways enhances the operability of the device (e.g., providing flexibility without cluttering the user interface with additional display controls, and reducing the amount of input required to access that flexibility) and makes the user device interface more efficient.

[0280] In some embodiments, pinch-and-drag gestures as air gestures (e.g., air drag gestures or air swipe gestures) include pinch gestures (e.g., pinch gestures or long pinch gestures) performed in conjunction with (e.g., following) a drag input that changes the user's hand position from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the drag). In some embodiments, the user holds the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opens their hand, separating the two or more fingers that formed the pinch gesture) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together to touch each other and uses the drag gesture to move the same hand to the second position in the air). In some implementations, pinch input is performed by the user's first hand, and drag input is performed by the user's second hand (e.g., while the user continues pinch input with the user's first hand, the user's second hand moves in the air from a first position to a second position). In some implementations, input gestures as air gestures include inputs performed using both of the user's hands (e.g., pinch and / or tap inputs). For example, input gestures include two (e.g., more) pinch inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., pinch input, long pinch input, or pinch and drag input) is performed using the user's first hand, and a second pinch input is performed using the other hand (e.g., the second hand in both of the user's hands). In some implementations, movement between the user's two hands is performed (e.g., increasing and / or decreasing the distance or relative orientation between the user's two hands).

[0281] In some implementations, a tap input performed as an air gesture (e.g., pointing at a user interface element) includes movement of a user's finger toward the user interface element, movement of a user's hand toward the user interface element (optionally, the user's finger extends toward the user interface element), downward movement of a user's finger (e.g., mimicking a mouse click or a tap on a touchscreen), or other predefined movements of the user's hand. In some implementations, the tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tap gesture movement, which is the finger or hand moving away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by the end of the movement. In some implementations, the end of the movement is detected based on changes in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of finger or hand movement, and / or a reversal of the acceleration direction of finger or hand movement).

[0282] In some implementations, gestures are detected via one or more image sensors (e.g., sensor 311) (1206). For example, gestures are detected by one or more cameras configured to capture images in the visible, infrared, and / or near-infrared spectra. Detecting gestures via one or more image sensors enhances device operability and makes the user-device interface more efficient.

[0283] In some embodiments, the gesture is detected by a wearable audio output device (e.g., via sensor 311) (1208). For example, the gesture is detected via one or more sensors of the wearable audio output device. In some embodiments, the gesture is detected from data captured from the wearable audio output device and another device communicating with the wearable audio output device.

[0284] In some embodiments, the wearable audio output device is (1210) an ear-worn device. In some embodiments, the ear-worn audio output device does not obstruct the user's eyes. In some embodiments, the ear-worn audio output device does not extend across the user's face. In some embodiments, the ear-worn audio output device does not affect the user's vision. In some embodiments, the ear-worn audio output device is an in-ear device. In some embodiments, the ear-worn audio output device is an over-ear device. In some embodiments, the ear-worn audio output device is mounted on the user's ear. The ability to detect gestures via the ear-worn device simplifies the user-device interface.

[0285] In some implementations, the gesture includes (1212) a cupping gesture (e.g., Figure 8B Gestures 804). For example, a cupping gesture involves the user's hand bending into a "C" shape or a backward "C" shape. In some implementations, the cupping gesture is performed near the user's ear. Detecting and responding to different types of gestures in different ways enhances device operability (e.g., providing flexibility without cluttering the user interface with additional display controls, and reducing the amount of input required to access that flexibility) and makes the user device interface more efficient.

[0286] In some implementations, the gestures include (1214) air pinch gestures (e.g., Figure 8G Gestures (828). For example, gestures include a pinching gesture performed by a user using their thumb and forefinger or their thumb and middle finger. In some embodiments, a pinching gesture is a pinching and holding gesture. Detecting and responding to different types of gestures in different ways enhances the operability of the device (e.g., providing flexibility without cluttering the user interface with additional display controls, and reducing the amount of input required to access that flexibility) and makes the user device interface more efficient.

[0287] In some implementations, the gesture includes (1216) a double air pinch gesture (e.g., Figure 8J Gesture 836). For example, a double pinch gesture involves a user performing two pinches consecutively (e.g., within a threshold amount of time each). In some embodiments, a double pinch gesture includes a user pinching their thumb and index (or middle) finger together twice. In some embodiments, a double pinch gesture includes a first pinch using a first finger and a second pinch using a second finger (e.g., not used during the first pinch). Detecting and responding to different types of gestures in different ways enhances device operability (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to access that flexibility) and makes the user device interface more efficient.

[0288] In some implementations, the gestures include (1218) pinching and twisting gestures in the air (e.g., Figures 8M to 8O Gesture 842). For example, a pinch and twist gesture involves a user pinching their thumb and forefinger together and then rotating their wrist (e.g., simulating turning a dial pad) beyond a threshold rotation amount (e.g., more than 5, 10, or 20 degrees). In some implementations, one or more parameters of the first operation vary based on the magnitude, speed, and / or direction of the rotation. For example, if the first operation is an adjustment to an output parameter (e.g., volume, contrast, brightness, or other type of parameter), the degree of adjustment may be based on the magnitude, speed, and / or direction of the rotation. Detecting and responding to different types of gestures in different ways enhances device operability (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to access that flexibility) and makes the user device interface more efficient.

[0289] In some implementations, gestures are detected (1220) within a gesture region (e.g., gesture region 806). For example, the gesture region is relative to the wearable audio output device. In some implementations, the gesture region is based on the field of view of the wearable audio output device. In some implementations, the gesture region is based on the distance from the side of the user's head. In some implementations, the gesture region is a three-dimensional region (e.g., defined relative to the user's head). For example, the gesture region is a 6-inch × 6-inch × 6-inch cube region, a 10-inch × 10-inch × 10-inch cube region, a 10-inch × 8-inch × 6-inch region, or other three-dimensional regions of different sizes. Detecting gestures within the detection region improves device operation by reducing false alarms caused by gestures performed outside the gesture region, and also reduces power consumption and extends the battery life of the wearable audio output device.

[0290] In response to the detection of a gesture (1222), based on a first type of hand gesture determined to be detected at a corresponding distance to the side of the user's head and at least partially based on the shape of the hand during the execution of the gesture, the wearable audio output device performs (1224) a first operation corresponding to the gesture. For example, Figures 8I to 8J An example is shown where a wearable audio output device 301 adjusts the media content playback positioning in response to a gesture 836.

[0291] In some implementations, the first operation includes (1226) adjusting the volume of the audio output at the wearable audio output device. For example, Figures 8M to 8O An example is illustrated where a wearable audio output device 301 adjusts the volume level 844 in response to a gesture 842. For example, the audio output corresponds to playback of media content. As another example, the audio output corresponds to ambient sounds in the physical environment (e.g., audio enhancement). As yet another example, the audio output corresponds to audio received from another device (e.g., a telephone call). In some embodiments, the volume decreases based on movement in a first direction, and increases based on movement in a second direction (e.g., opposite or different from the first direction). In some embodiments, the magnitude of the volume adjustment is based on the magnitude and / or speed of the gesture. Adjusting the volume of the audio output in response to a gesture performed within a gesture area enhances the operability of the device (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to achieve this flexibility) and makes the user-device interface more efficient.

[0292] In some implementations, the first operation includes (1228) adjusting the magnitude of ambient sound from the physical environment. For example, Figures 8A to 8B An example is illustrated where a wearable audio output device 301 adjusts the levels of ANC, active transparency, and conversation enhancement in response to a detected gesture 804. For example, adjusting the magnitude of ambient sound includes adjusting the degree of ANC and / or the degree of active transparency. In some embodiments, adjusting the magnitude of ambient sound includes enabling or disabling ANC and / or active transparency modes. In some embodiments, the magnitude of ambient sound is decreased based on movement in a first direction, and increased based on movement in a second direction (e.g., opposite or different from the first direction). In some embodiments, the adjustment is based on the magnitude and / or speed of the gesture. Adjusting the magnitude of ambient sound in response to a gesture performed in a gesture area enhances device operability (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to achieve this flexibility) and makes the user-device interface more efficient.

[0293] In some implementations, the first operation includes (1230) adjusting the playback of the media content. For example, Figures 8E to 8G An example is illustrated where a wearable audio output device 301 pauses playback of media content in response to a detected gesture 828. For example, a first operation includes pausing or resuming the media content. As another example, a first operation includes fast-forwarding or rewinding the media content. As yet another example, a first operation includes switching media content (e.g., switching tracks, switching songs, switching radio stations, or switching media sources). In some embodiments, adjustments to the playback of the media content include a rewind operation based on movement in a first direction, and adjustments to the playback of the media content include a forward skip operation based on movement in a second direction (e.g., opposite or different from the first direction). In some embodiments, the amount of adjustment is based on the magnitude and / or speed of the gesture. Adjusting media playback in response to gestures performed in a gesture area enhances the operability of the device (e.g., providing flexibility without cluttering the user interface with additional display controls and reducing the amount of input required to achieve this flexibility) and makes the user-device interface more efficient.

[0294] In response to the detection of a gesture (1222), and based on the determination that the gesture is not a first type of hand gesture determined at least in part based on the shape of the hand during the execution of the gesture, the wearable audio output device abandons (1232) the execution of the first operation. For example, Figures 80 to 8P An example is illustrated where a wearable audio output device 301 stops performing volume adjustment in response to a change in hand shape. For example, the first operation is not performed, regardless of whether the hand gesture is within a corresponding distance to the side of the user's head. For example, the corresponding distance could be 2 inches, 5 inches, 10 inches, or 1 foot. Abandoning operation based on hand shape reduces false gesture detection and enhances device operability, making the user-device interface more efficient (e.g., by helping users achieve expected results and reducing user errors when operating / interacting with the device). This additionally reduces power consumption and extends device battery life by enabling users to use the device more quickly and efficiently.

[0295] In some implementations, in response to the detection of a gesture (1222), based on determining that the detected gesture is beyond the side of the user's head at a corresponding distance, the wearable audio output device abandons (1234) performing the first operation (e.g., regardless of whether the hand gesture is a first type of gesture). For example, Figures 8B to 8D An example is a wearable audio output device 301 that responds to the user's hand movements. Figure 8CANC, transparency, and enhancement modifications are stopped when the gesture exceeds a certain distance from the side of the user's head (e.g., outside the gesture area). Abandoning gestures based on a certain distance from the side of the user's head enhances device operability and makes the user-device interface more efficient (e.g., by helping users achieve expected results and reducing user errors when operating / interacting with the device). This additionally reduces power consumption and extends device battery life by enabling users to use the device more quickly and efficiently.

[0296] In some implementations, in response to the detection of a gesture (1222), based on determining that the detected gesture is at a corresponding distance beyond the side of the user's head and that the gesture is a second type of hand gesture determined at least in part based on the shape of the hand during the execution of the gesture, the wearable audio output device performs (1236) a second operation corresponding to the gesture, wherein the second operation differs from the first operation. For example, Figure 8B Examples of adjustments to ANC, opacity, and enhancement in response to the detection of gesture 804 (e.g., a cupping gesture) are shown, and Figure 8G An example of media playback adjustment in response to the detection of gesture 828 (e.g., a pinch gesture) is illustrated. Detecting and responding to different types of gestures in different ways enhances the operability of the device (e.g., providing flexibility without cluttering the user interface with additional display controls, and reducing the amount of input required to access that flexibility) and makes the user device interface more efficient.

[0297] In some implementations, in response to the detection of a gesture (1222), based on the determination that the gesture is not a second type of hand gesture determined at least in part based on the shape of the hand during the execution of the gesture, the wearable audio output device abandons (1238) the execution of the second operation (e.g., regardless of whether the hand gesture is within a corresponding distance to the side of the user's head). For example, Figures 80 to 8P An example is a wearable audio output device 301 that stops performing volume adjustment in response to a change in hand shape. Abandoning operations based on gesture type enhances device operability and makes the user-device interface more efficient, which additionally reduces power consumption and extends device battery life by allowing users to use the device more quickly and efficiently.

[0298] In some implementations, in response to the detection of a gesture (1222), based on determining that the detected gesture is beyond the side of the user's head at a corresponding distance, the wearable audio output device abandons (1240) performing a second operation (e.g., regardless of whether the hand gesture is a second type of gesture). For example, Figures 8B to 8D An example is a wearable audio output device 301 that responds to the user's hand movements. Figure 8CANC, transparency, and enhancement modifications are stopped when the gesture is outside the gesture area. Giving up gestures based on a corresponding distance beyond the side of the user's head enhances device operability and makes the user-device interface more efficient. This, in turn, reduces power consumption and extends device battery life by allowing users to use the device more quickly and efficiently.

[0299] In some implementations, the wearable audio output device detects the end of the gesture (1242) and, in response to detecting the end of the gesture, stops performing the first operation. For example, Figures 80 to 8P An example is illustrated where a wearable audio output device 301 stops performing volume adjustment in response to a change in hand shape. For example, the gesture is a pinch and hold gesture, and a first operation is performed while the gesture is held. In some embodiments, the first operation modifies ambient sound. For example, disabling ANC (or alternatively, reducing the level of ANC) upon detecting a gesture. In some embodiments, the first operation changes the operating state of the wearable audio output device from a first state to a second state, and the wearable audio output device operates in the second state while the gesture is being performed (e.g., the wearable audio output device transitions back to the first state when the gesture is no longer detected). In some embodiments, the wearable audio output device operates in the second state for a corresponding amount of time after the gesture is detected. In some embodiments, the wearable audio output device operates in the second state until a command to change the operating state from the second state is received. Stopping operation in response to the detection of the end of a user gesture enhances device operability and makes the user-device interface more efficient, which additionally reduces power consumption and extends device battery life by enabling users to use the device more quickly and efficiently.

[0300] In some implementations, detecting the end of a gesture includes (1244) detecting a change in the shape of the hand. For example, Figures 80 to 8P An example is illustrated where a wearable audio output device 301 stops adjusting volume in response to a change in hand shape from a pinched shape to a fist shape. For example, the user stops cupping their hand. Detecting the end of a gesture based on a change in hand shape enhances device operability and makes the user-device interface more efficient.

[0301] In some implementations, the end of gesture detection includes (1246) detecting that the hand is further away from the side of the user's head. For example, a wearable audio output device detects that the user's hand has moved away from the side of the user's head. Figures 8B to 8D An example is a wearable audio output device 301 that responds to the user's hand movements. Figure 8CWhen the user's hand is outside the gesture area, ANC, transparency, and enhancement modifications are stopped. Detecting the end of a gesture based on the user's hand being further away than the corresponding distance enhances device operability and makes the user-device interface more efficient.

[0302] In some implementations, the wearable audio output device detects (1248) the user's hand entering the gesture area and, in response to detecting the user's hand entering the gesture area, provides initial feedback indicating that the hand has entered the gesture area. For example, Figure 8F The illustration shows user 802's hand 824 entering gesture area 806, and wearable audio output device 301 providing feedback 826 in response. In some embodiments, the first feedback includes audio and / or haptic feedback. Providing feedback indicating that the user's hand has entered the gesture area provides improved feedback regarding the status of the wearable audio output device.

[0303] In some implementations, the wearable audio output device detects (1250) that the user's hand has left the gesture area, and in response to detecting that the user's hand has left the gesture area, provides second feedback indicating that the hand has left the gesture area. For example, Figure 8H The illustration shows user 802's hand 824 leaving gesture area 806, and wearable audio output device 301 providing feedback 834 in response. In some embodiments, the second feedback includes audio and / or haptic feedback. In some embodiments, the second feedback is of a different type than the first feedback. In some embodiments, the first and second feedbacks are of the same type. In some embodiments, one or more attributes of the first feedback differ from one or more attributes of the second feedback (e.g., different pitch, frequency, and / or other attributes). Providing feedback indicating that the user's hand has left the gesture area provides improved feedback regarding the status of the wearable audio output device.

[0304] In some implementations, the wearable audio output device detects (1252) a user's hand entering a gesture area, and in response to detecting the user's hand entering the gesture area, activates a gesture detection state of the wearable audio output device, wherein a gesture is detected while the gesture detection state is active. For example, in response to detecting hand 824 entering gesture area 806, in Figure 8FThe gesture detection state is activated in some embodiments. In some embodiments, the wearable audio output device detects that the user's hand has left the gesture area and, in response, disables the gesture detection state of the wearable audio output device. In some embodiments, the gesture detection state is activated for a threshold time amount. In some embodiments, activating the gesture detection state includes enabling one or more sensors of the wearable audio output device (e.g., image sensors, audio sensors, and / or other types of sensors). In some embodiments, hand entry into the gesture area is detected via a first type of sensor (e.g., a capacitive sensor, such as the capacitive sensor of the wearable audio output device), and the gesture is detected via a second type of sensor (e.g., an image sensor, such as the image sensor of the wearable audio output device). In some embodiments, the first type of sensor consumes less power than the second type of sensor. Activating the gesture detection state in response to detecting hand entry into the gesture area enhances the operability of the wearable audio output device and makes the user-device interface more efficient, and reduces power consumption and extends the battery life of the wearable audio output device by being able to disable components (e.g., sensors) when outside a predefined area.

[0305] In some implementations, the wearable audio output device detects (1254) a second gesture performed by the user's head, and in response to detecting the second gesture, activates a gesture detection state of the wearable audio output device, wherein the gesture is detected while the gesture detection state is active. For example, Figures 8L to 8M Examples are shown in Figure 8L The head gesture 841 is executed, and the gesture detection state is in Figure 8M In some embodiments, the wearable audio output device detects another gesture performed by the user's head and, in response, disables the gesture detection state of the wearable audio output device. In some embodiments, the gesture detection state is activated for a threshold time amount. In some embodiments, the second gesture includes tilting the head and / or shaking the head. In some embodiments, the second gesture is detected by a first type of sensor (e.g., an accelerometer, such as the accelerometer of the wearable audio output device), and the gesture is detected by a second type of sensor (e.g., an image sensor, such as the image sensor of the wearable audio output device). In some embodiments, the first type of sensor consumes less power than the second type of sensor. Activating the gesture detection state in response to head gestures enhances the operability of the wearable audio output device and makes the user-device interface more efficient, and reduces power consumption and extends the battery life of the wearable audio output device by being able to disable components (e.g., sensors) outside a predefined area.

[0306] It should be understood that, Figures 12A to 12DThe specific order in which the operations are described herein is merely exemplary and is not intended to indicate that the order is the only possible order in which these operations can be performed. Those skilled in the art will conceive of various ways to reorder the operations described herein. Furthermore, it should be noted that the details of other processes described herein in conjunction with other methods (e.g., methods 1000, 1100, and 1300) are similarly applicable to the above-described combinations. Figure 12A-12D The method 1200. For example, the inputs, gestures, functions, and feedback described above with reference to method 1200 optionally have one or more of the characteristics of the inputs, gestures, functions, and feedback described herein with reference to other methods described herein (e.g., methods 1000, 1100, and 1300). For the sake of brevity, these details will not be repeated here.

[0307] Figures 13A to 13B This is a flowchart illustrating method 1300 for providing feedback indicating sensor occlusion according to some embodiments. Method 1300 is performed at a wearable device (e.g., a wearable audio output device 301, such as earbuds, headphones, headsets, necklaces, or glasses), which includes one or more sensors (e.g., sensor 311, such as image sensors, motion sensors, and / or other types of sensors) and optionally includes one or more audio output components (e.g., speaker 306). In some embodiments, the wearable device is a headset, earbuds, necklace, headphones, glasses, or other type of wearable device. Some operations in method 1300 are optionally combined, and / or the order of some operations is optionally changed.

[0308] As described below, method 1300 provides improved feedback on the state of the wearable device. Identifying sensor occlusion and reporting it to the user allows users to achieve desired results and reduces errors when operating / interacting with the wearable device. Reduced errors decrease power consumption and extend battery life, and make the user-device interface more efficient.

[0309] When the wearable device is worn by a user, the wearable device detects (1302) the occurrence of one or more events that indicate that the device is in a context in which the corresponding sensor among one or more of its sensors is available to perform the corresponding operation. For example, Figure 9AThe illustration shows a user 802 performing a gesture 906 and uttering a command 908, corresponding to the occurrence of an event in which sensors of the wearable audio output device 301 can be used (e.g., to turn on a light 910). In some embodiments, the occurrence of one or more events is detected by a corresponding sensor. In some embodiments, the occurrence of one or more events is detected at least in part by one or more sensors other than the corresponding sensor. In some embodiments, the occurrence of one or more events includes input corresponding to a request from the user to perform the corresponding action (e.g., an explicit or implicit request). In some embodiments, the occurrence of one or more events includes detecting a context of the device that might be helpful in performing the corresponding action (e.g., based on information detected by the corresponding sensor). In some embodiments, the context in which the corresponding sensor of one or more sensors can be used to perform the corresponding action includes a scanning state of the wearable device in which the wearable device uses one or more sensors to scan the physical environment in which the wearable audio output device is operating (e.g., to provide information and / or context-based suggestions about the physical environment in which the wearable audio output device is operating).

[0310] In some implementations, the corresponding operation is a (1304) discrete operation (e.g., the operation of turning on a light, such as...). Figures 9A to 9B (As illustrated). In some implementations, discrete operations involve a single output (e.g., audio and / or haptic notification) in response to a user command (e.g., a query and response). For example, a discrete operation is a single (one-time) scan of the physical environment in which the wearable audio output device is operating. As another example, a discrete operation is a "one-off" or bounded operation. As yet another example, a discrete operation is providing information about a real-world object (e.g., the object the user is pointing at). Determining sensor occlusion based on events in which discrete operations can be performed and reporting occlusion to the user allows users to achieve the desired results and reduces errors when operating / interacting with the wearable device. Reducing errors reduces power consumption and extends battery life, and makes the user-device interface more efficient.

[0311] In some implementations, the corresponding operation is (1306) a continuous operation (e.g., Figures 9G to 9K(Examples of navigation-assisted operation are shown below). In some embodiments, continuous operation involves two or more outputs (e.g., a sequence of notifications and / or instructions) in response to a user command. In some embodiments, continuous operation involves activating one or more sensors and providing feedback based on data from one or more of the activated sensors meeting one or more criteria. In some embodiments, performing continuous operation includes enabling a specific state, wherein the state is enabled for a predefined amount of time, and / or the state is enabled until it is disabled (e.g., disabled by the user via a second command). For example, continuous operation is a scanning mode in which the wearable device identifies real-world objects and provides information. As another example, continuous operation is a navigation-assisted mode in which the wearable device provides navigation directions to the user as the user moves. Determining sensor occlusion and reporting occlusion to the user during continuous operation allows users to achieve the desired results and reduces errors when operating / interacting with the wearable device. Reducing errors reduces power consumption and extends battery life, and makes the user-device interface more efficient.

[0312] In some embodiments, the corresponding sensor is (1308) an image sensor (e.g., sensor 311). For example, the corresponding sensor is a camera or other type of vision sensor for the wearable device. As an example, the image sensor may be located in the stem of an earbud or in the frame of a pair of glasses. As another example, the image sensor may be mounted to the housing of a pair of headphones. In some embodiments, the corresponding sensor is a low-resolution sensor and / or a low-fidelity sensor. In some embodiments, the wearable device does not retain or transmit images from the image sensor (e.g., analyzes the image and then discards it). In some embodiments, the image sensor and / or the wearable device is configured to extract information about the physical environment in which the wearable audio output device is operating (e.g., to determine the context in which the corresponding sensor is available to perform the corresponding operation) without storing images. Determining that the image sensor is occluded and reporting the occlusion to the user allows users to achieve the desired results and reduces errors when operating / interacting with the wearable device. Reducing errors reduces power consumption and extends battery life, and makes the user-device interface more efficient.

[0313] In some implementations, the wearable device is (1310) an ear-worn device (e.g., a wearable audio output device 301). For example, the wearable device is an ear-worn audio output device, such as earbuds or headphones. The ability to detect gestures via an ear-worn device simplifies the user-device interface.

[0314] In response to the detection of one or more events (1312), based on the determination that the corresponding sensor is occluded, the wearable device provides (1314) feedback to the user indicating that the corresponding sensor is occluded (e.g., partially or completely occluded). For example, Figures 9C to 9D An example is illustrated where a wearable audio output device 301 provides feedback 918 based on the wearable audio output device 301 being obscured by the hair of a user 802. For example, the corresponding sensor may be at least partially obscured by the user's clothing and / or hair. In some embodiments, the wearable device determines that the corresponding sensor is obscured based on data from the corresponding sensor and / or data from another sensor among one or more sensors.

[0315] In some implementations, depending on whether the corresponding operation is a discrete operation, providing feedback includes (1316) providing a first type of feedback (e....

Claims

1. A method, the method comprising: At an ear-worn audio output device that includes one or more sensors and one or more audio output components: User gestures and corresponding voice commands are detected via one or more sensors in the ear-worn audio output device; as well as In response to detecting the user gesture and the corresponding voice command: Based on determining that the user gesture is a first type of gesture and points to a first real-world object, and determining that the corresponding voice command includes a request for feedback about the first real-world object, a first audio feedback corresponding to the first real-world object is provided via the one or more audio output components, wherein at least one parameter of the first audio feedback is based on the corresponding voice command; as well as Based on determining that the user gesture is a gesture of the first type and points to a second real-world object, and determining that the corresponding voice command includes a request for feedback regarding the second real-world object, a second audio feedback corresponding to the second real-world object is provided via the one or more audio output components, wherein at least one parameter of the second audio feedback is based on the corresponding voice command.

2. The method according to claim 1, wherein the ear-worn audio output device includes an audio playback device.

3. The method of claim 1, wherein the ear-worn audio output device provides the first audio feedback but not visual feedback.

4. The method of claim 1, further comprising providing playback of audio content via the one or more audio output components, wherein the user gesture is detected while providing the playback.

5. The method of claim 4, wherein the audio content is received from an accessory device communicatively coupled to the ear-worn audio output device.

6. The method of claim 5, wherein the accessory device includes a playback control for controlling the playback of the audio content at the earphone audio output device.

7. The method according to any one of claims 1 to 6, wherein the first audio feedback corresponding to the first real-world object includes a description of the first real-world object.

8. The method of any one of claims 1 to 6, wherein the first audio feedback corresponding to the first real-world object includes an indication of operable data associated with the first real-world object.

9. The method according to any one of claims 1 to 6, wherein the first audio feedback corresponding to the first real-world object includes contextual information about the user of the ear-worn audio output device.

10. The method of claim 9, wherein the first audio feedback corresponding to the first real-world object includes an indication of a future time period associated with the first real-world object, and wherein the contextual information includes information about the availability of the user during the future time period.

11. The method of any one of claims 1 to 6, further comprising providing feedback corresponding to the user gesture in response to detecting at least a portion of the user gesture.

12. The method of claim 11, wherein the feedback corresponding to the user gesture includes feedback indicating the pause of the user gesture.

13. The method of claim 11, wherein the feedback corresponding to the user gesture includes feedback indicating the progress of the user gesture.

14. The method according to any one of claims 1 to 6, further comprising: The second user's gesture is detected via one or more sensors of the ear-worn audio output device; as well as In response to the detection of the second user gesture: Based on determining that the ear-worn audio output device is within a predefined area and that the user gesture is the first type of gesture and points to the first real-world object, a third audio feedback corresponding to the first real-world object is provided via the one or more audio output components; as well as If it is determined that the ear-worn audio output device is not within the predefined area, the provision of the third audio feedback is abandoned.

15. The method according to any one of claims 1 to 6, wherein: Based on the first real-world object being located at the first position, the first audio feedback is spatialized to the first position; as well as Based on the first real-world object being located in the second position, the first audio feedback is spatialized to the second position.

16. The method of any one of claims 1 to 6, wherein the first real-world object comprises a hand-drawn picture, and wherein the first audio feedback comprises information indicated by the hand-drawn picture.

17. The method according to any one of claims 1 to 6, wherein at least one parameter of the first audio feedback is based on the tone of the user's voice command.

18. The method of claim 17, wherein the volume of the first audio feedback is based on the tone of the user's voice command.

19. The method of any one of claims 1 to 6, wherein the user gesture includes a boundary gesture outlining a portion of the first real-world object, and wherein the first audio feedback includes audio feedback regarding real-world content on the portion of the first real-world object.

20. The method of claim 19, wherein the real-world content on the portion of the first real-world object comprises text, and the method further comprises providing an indication from the first real-world object that the text has been selected.

21. The method of claim 20, further comprising adding the text to a document associated with the ear-worn audio output device and / or companion device.

22. An ear-worn audio output device, the ear-worn audio output device comprising: One or more sensors; One or more audio output components; One or more processors; as well as A memory storing one or more programs configured to be executed by one or more processors, the one or more programs including instructions for performing the following operations: User gestures and corresponding voice commands are detected via one or more sensors of the ear-worn audio output device; and In response to detecting the user gesture and the corresponding voice command: Based on determining that the user gesture is a first type of gesture and points to a first real-world object, and determining that the corresponding voice command includes a request for feedback about the first real-world object, a first audio feedback corresponding to the first real-world object is provided via the one or more audio output components, wherein at least one parameter of the first audio feedback is based on the corresponding voice command; as well as Based on determining that the user gesture is a gesture of the first type and points to a second real-world object, and determining that the corresponding voice command includes a request for feedback regarding the second real-world object, a second audio feedback corresponding to the second real-world object is provided via the one or more audio output components, wherein at least one parameter of the second audio feedback is based on the corresponding voice command.

23. The ear-worn audio output device of claim 22, wherein the one or more programs include instructions for performing any one of the methods according to claims 2 to 21.

24. A computer-readable storage medium storing one or more programs, said one or more programs including instructions that, when executed by an ear-worn audio output device including one or more sensors and one or more audio output components, cause the ear-worn audio output device to: User gestures and corresponding voice commands are detected via one or more sensors in the ear-worn audio output device; as well as In response to detecting the user gesture and the corresponding voice command: Based on determining that the user gesture is a first type of gesture and points to a first real-world object, and determining that the corresponding voice command includes a request for feedback about the first real-world object, a first audio feedback corresponding to the first real-world object is provided via the one or more audio output components, wherein at least one parameter of the first audio feedback is based on the corresponding voice command; as well as Based on determining that the user gesture is a gesture of the first type and points to a second real-world object, and determining that the corresponding voice command includes a request for feedback regarding the second real-world object, a second audio feedback corresponding to the second real-world object is provided via the one or more audio output components, wherein at least one parameter of the second audio feedback is based on the corresponding voice command.

25. The computer-readable storage medium of claim 24, wherein the one or more programs include instructions that, when executed by an ear-worn audio output device, cause the ear-worn audio output device to perform any one of the methods of claims 2 to 21.

26. A computer program product comprising one or more programs, which, when executed by a processor, cause the processor to perform any one of the methods according to claims 1 to 21.