Devices, methods, and graphical user interfaces for content applications

By improving computer system interfaces and methods, reducing user input, providing visual feedback and animated objects, the problem of low interaction efficiency in virtual/augmented reality environments is solved, improving user experience and device efficiency, especially saving power and extending battery life in battery-powered devices.

CN122018689APending Publication Date: 2026-05-12APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
APPLE INC
Filing Date
2024-05-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for interacting with virtual/augmented reality environments are inefficient, complex, and error-prone, leading to wasted computer system resources and cognitive burden on users, especially in battery-powered devices.

Method used

Improved computer system interfaces and methods reduce the amount and nature of user input, provide visual feedback, display animated objects and controls, respond to user gaze and gestures, reduce resource consumption, save power, and improve interaction efficiency.

Benefits of technology

It improves user interaction efficiency, reduces computer system resource consumption, extends battery life, enhances device operability and user experience, and saves power and reduces heat generation, especially in battery-powered devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018689A_ABST
    Figure CN122018689A_ABST
Patent Text Reader

Abstract

The disclosure relates to devices, methods, and graphical user interfaces for content applications. In some embodiments, a computer system generates a virtual lighting effect while presenting a content item. In some embodiments, a computer system generates an animated three-dimensional object while presenting a content item. In some embodiments, the computer system displays a reduced user interface instead of an extended user interface in response to different inputs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application No. 202480036952.0, filed on May 31, 2024, entitled "Apparatus, Method and Graphical User Interface for Content Application". Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 506,072, filed June 3, 2023, the contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field

[0003] This disclosure relates in its entirety to computer systems that provide computer-generated experiences, including but not limited to electronic devices that provide a user interface via a display for presenting and browsing content. Background Technology

[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices used in computer systems and other electronic computing devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) are used to interact with the virtual / augmented reality environment. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention

[0005] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems where manipulating virtual objects is complex, tedious, and error-prone, impose a significant cognitive burden on users and detract from the immersive experience of virtual / augmented reality environments. Furthermore, these methods take longer than necessary, thus wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.

[0006] Therefore, computer systems with improved methods and interfaces are needed to provide users with computer-generated experiences, making user interaction with the computer system more efficient and intuitive. Such methods and interfaces can optionally supplement or replace conventional methods for providing users with extended reality experiences. By helping users understand the relationship between the input provided and the device's response to that input, such methods and interfaces reduce the quantity, extent, and / or nature of user input, thus creating a more efficient human-computer interface.

[0007] The disclosed system reduces or eliminates the aforementioned defects and other problems associated with the user interface of a computer system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop, tablet, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has (e.g., includes or communicates with) a display generation component (e.g., a display device such as a head-mounted device (HMD), a monitor, a projector, a touch-sensitive display (also referred to as a "touchscreen" or "touchscreen display"), or other device or component that presents visual content to a user, such as visual content generated on or in the display generation component itself or from the display generation component and visible elsewhere). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to the display generation component, the computer system also has one or more output devices, including one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules. A program or set of instructions stored in memory for performing multiple functions. In some embodiments, a user interacts with the GUI through touch and gestures of a stylus and / or fingers on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (such as captured by a camera and other motion sensors), and / or voice input (such as captured by one or more audio input devices). In some embodiments, the functions performed through interaction may optionally include image editing, drawing, presentation, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions may optionally be included in transient and / or non-transitory computer-readable storage media or other computer program products configured for execution by one or more processors.

[0008] Electronic devices with improved methods and interfaces are needed to interact with 3D environments. Such methods and interfaces can complement or replace conventional methods for interacting with 3D environments. They reduce the amount, extent, and / or nature of user input, resulting in more efficient human-computer interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging.

[0009] In some implementations, the computer system displays a set of controls (e.g., transmission controls and / or other types of controls) associated with controlling the playback of media content in response to the detection of a user's gaze and / or gesture. In some implementations, the computer system initially displays a first set of controls in a desalience state (e.g., with reduced visual salience) in response to the detection of a first input, and then displays a second set of controls (which may optionally include additional controls) in an increased salience state in response to the detection of a second input. In this way, the computer system optionally provides feedback to the user that the display of controls has begun to invoke the display without unduly distracting the user from the content (e.g., by initially displaying the controls in a less visually salience manner), and then displays the controls in a more visually salience manner based on the detection of user input indicating that the user wishes to interact further with the controls, to allow for easier and more accurate interaction with the computer system.

[0010] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not exhaustive; in particular, many additional features and advantages will be apparent to those skilled in the art from the accompanying drawings, description, and claims. Furthermore, it should be pointed out that the language used in this specification has been chosen in principle for readability and instruction purposes, and such choice may not be necessary to depict or define the subject matter of the invention. Attached Figure Description

[0011] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals indicate corresponding parts in all the drawings.

[0012] Figure 1A This is a block diagram illustrating the operating environment of a computer system for providing XR experiences according to some implementation schemes.

[0013] Figures 1B to 1P It is used in Figure 1A Examples of computer systems that provide XR experiences in the operating environment.

[0014] Figure 2This is a block diagram illustrating a computer system configured to manage and coordinate an XR experience for a user, according to some implementation schemes.

[0015] Figure 3 This is a block diagram illustrating a display generation component of a computer system configured to provide an XR experience to a user, according to some implementation schemes.

[0016] Figure 4 This is a block diagram illustrating a hand tracking unit of a computer system configured to capture user gesture input according to some implementation schemes.

[0017] Figure 5 This is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input according to some implementation schemes.

[0018] Figure 6 This is a flowchart illustrating a flare-assisted gaze tracking pipeline according to some implementation schemes.

[0019] Figures 7A to 7H Examples are shown of how a computer system, according to some implementation schemes, generates virtual lighting effects when rendering content items.

[0020] Figure 8 This is a flowchart illustrating how a computer system, according to some implementation schemes, generates virtual lighting effects when presenting content items.

[0021] Figures 9A to 9E Examples are shown of how a computer system, according to some implementation schemes, generates animated 3D objects when rendering content items.

[0022] Figure 10 This is a flowchart illustrating how a computer system, according to some implementation schemes, generates animated 3D objects when presenting content items.

[0023] Figures 11A to 11N Examples are given of how computer systems, according to some implementation schemes, display a simplified user interface instead of an extended user interface in response to different inputs.

[0024] Figure 12 This is a flowchart illustrating how a computer system, according to some implementation schemes, displays a simplified user interface instead of an extended user interface in response to different inputs. Detailed Implementation

[0025] According to some implementations, this disclosure relates to a user interface for providing extended reality (XR) experiences to users.

[0026] The systems, methods, and GUIs described in this paper improve user interface interactions with virtual / augmented reality environments in a variety of ways.

[0027] In some implementations, the computer system displays the user interface of an application (such as a content playback application) in a three-dimensional environment. In some implementations, while displaying the user interface and in response to receiving a first input corresponding to a request to initiate playback of the corresponding content, the computer system initiates playback of the corresponding content and displays the user interface with simulated lighting effects. In some implementations, the simulated lighting effects have one or more characteristics based on the playback of the corresponding content. Displaying the user interface with simulated lighting effects based on the playback of the corresponding content reduces the resources required to display simulated lighting effects when the corresponding content is not playing, and reduces the need for manual input to manually enable and / or disable simulated lighting effects.

[0028] In some implementations, the computer system displays the user interface of an application (such as a content playback application) in a three-dimensional environment. In some implementations, while displaying the user interface and in response to receiving a first input corresponding to a request to initiate playback of the corresponding content, the computer system initiates playback of the corresponding content and displays the user interface as having animated objects. In some implementations, the animated objects have one or more characteristics based on the playback of the corresponding content. Displaying the user interface as having animated objects based on the playback of the corresponding content reduces the resources required to display simulated lighting effects when the corresponding content is not playing, and reduces the need for manual input to manually enable and / or disable simulated lighting effects.

[0029] In some implementations, the computer system displays an extended user interface (e.g., a content playback application) that includes selectable options for initiating playback of a second content item. In some implementations, the computer system displays a second selectable option in response to a first input, which can be selected to display a simplified user interface of the application in a three-dimensional environment. In some implementations, the extended user interface is displayed in a first location, the second selectable option in a second location, and the simplified user interface in a third location. Displaying the simplified user interface instead of the extended user interface reduces visual distraction for the user and reduces clutter in the three-dimensional environment, thereby reducing errors in interaction with the computer system.

[0030] Figure 1 to Figure 6 Descriptions of example computer systems for providing XR experiences to users are provided (such as those described below with respect to methods 800, 1000 and / or 1200). Figures 7A to 7H Examples are shown of how a computer system, according to some implementation schemes, generates virtual lighting effects when rendering content items. Figure 8 This is a flowchart illustrating how a computer system, according to some implementation schemes, generates virtual lighting effects when presenting content items. Figures 7A to 7H The user interface in the example is used to demonstrate Figure 8 The process in. Figures 9A to 9E Examples are shown of how a computer system, according to some implementation schemes, generates animated 3D objects when rendering content items. Figure 10 This is a flowchart illustrating how a computer system, according to some implementation schemes, generates animated 3D objects when presenting content items. Figures 9A to 9E The user interface in the example is used to demonstrate Figure 10 The process in. Figures 11A to 11N Examples are given of how computer systems, according to some implementation schemes, display a simplified user interface instead of an extended user interface in response to different inputs. Figure 12 This is a flowchart illustrating how a computer system, according to some implementation schemes, displays a simplified user interface instead of an extended user interface in response to different inputs. Figures 11A to 11N The user interface in the example is used to demonstrate Figure 12 The process in.

[0031] The processes described below enhance device operability and make the user-device interface more efficient through various technologies (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device). These technologies include providing users with improved visual feedback, reducing the amount of input required to perform operations, providing additional control options without cluttering the user interface with additional display controls, performing operations without further user input when a set of conditions are met, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or additional technologies. These technologies also reduce power consumption and extend device battery life by enabling users to use the device faster and more efficiently. Saving battery power, and thus weight, improves the ergonomics of the device. These technologies also enable real-time communication, allowing the use of fewer and / or less precise sensors, resulting in more compact, lighter, and cheaper devices, and enabling the device to be used in a variety of lighting conditions. These technologies reduce energy consumption, thereby reducing the heat emitted by the device, which is particularly important for wearable devices, where excessive heat generated by a device within the operating parameters of its components can make wearing the device uncomfortable for the user.

[0032] Furthermore, in methods described herein where one or more steps depend on the satisfaction of one or more conditions, it should be understood that the method may be repeated in multiple repetitions such that, during the repetitions, all conditions determining the steps in the method are satisfied in different repetitions of the method. For example, if the method requires performing a first step (if the condition is satisfied) and a second step (if the condition is not satisfied), those skilled in the art will know that the stated steps are repeated until both the conditions are satisfied and not satisfied (in no particular order). Thus, a method described as having one or more steps depending on the satisfaction of one or more conditions can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing discretionary operations based on the satisfaction of the corresponding one or more conditions, and thus to determine whether possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions determining the steps in the method are satisfied. Those skilled in the art will also understand that, similar to methods having discretionary steps, a system or computer-readable storage medium may repeat the steps of the method multiple times as needed to ensure that all discretionary steps have been performed.

[0033] In some implementation schemes, such as Figure 1A As shown, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted display (HMD), a monitor, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., image sensors, light sensors, depth sensors, haptic sensors, orientation sensors, proximity sensors, temperature sensors, position sensors, motion sensors, speed sensors, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted or handheld device).

[0034] In describing XR experiences, various terms are used to distinguish several related but different environments that a user can sense and / or interact with (e.g., interacting with inputs detected by the computer system 101 that generates the XR experience, causing the computer system to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101). The following is a subset of these terms: Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.

[0035] Extended Reality: Conversely, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In XR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the properties of virtual objects in the XR environment can be done in response to a representation of physical motion (e.g., a voice command). A person can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. For example, audio objects can enable audio transparency, which selectively introduces ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, people can sense and / or interact only with audio objects.

[0036] Examples of XR include virtual reality and mixed reality.

[0037] Virtual Reality: A virtual reality (VR) environment is a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing human avatars are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through the simulation of a person's presence within the computer-generated environment and / or through the simulation of a subset of a person's physical movements within the computer-generated environment.

[0038] Mixed Reality: Compared to VR environments, which are designed to be entirely based on computer-generated sensory input, mixed reality (MR) environments refer to simulated environments designed to incorporate sensory input from the physical environment, or its representations, in addition to computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any state between, but not limited to, a purely physical environment as one end and a virtual reality environment as the other. In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to present an MR environment can track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical objects or their representations from the physical environment). For example, a system can cause motion so that virtual trees appear stationary relative to the physical ground.

[0039] Examples of mixed reality include augmented reality and augmented virtual reality.

[0040] Augmented Reality (AR): An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are overlaid on a physical environment or a representation of the physical environment. For example, an electronic system for presenting an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or semi-transparent display, allowing a person to perceive the virtual objects overlaid on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via the images or videos of the physical environment and perceives the virtual objects overlaid on the physical environment. As used herein, video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects onto the physical environment, such as as a hologram or onto a physical surface, allowing a person to perceive the virtual objects superimposed on the physical environment. Augmented reality environments also refer to simulated environments in which the representation of the physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, the system can transform one or more sensor images to apply a selected viewpoint (e.g., viewpoint) different from the viewpoint captured by the imaging sensor. As another example, the representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions of it, such that the modified portions can be representative but not realistic versions of the original captured image. Furthermore, the representation of the physical environment can be transformed by graphically removing or blurring portions of it.

[0041] Augmented Virtual: An augmented virtual (AV) environment is a simulated environment in which a virtual or computer-generated environment combines one or more sensory inputs from a physical environment. Sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park could have virtual trees and virtual buildings, but a person's face could be realistically reproduced from an image taken of a physical person. Similarly, virtual objects could adopt the shape or color of a physical object imaged by one or more imaging sensors. Furthermore, virtual objects could adopt shadows that correspond to the sun's position within the physical environment.

[0042] In augmented reality, mixed reality, or virtual reality environments, a view of the three-dimensional environment is visible to the user. This view is typically visible to the user via a virtual viewport through one or more display generating components (e.g., a display providing stereoscopic content to different eyes of the same user), which has a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generating components. In some embodiments, the area defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). The viewport and viewport boundary typically move with the movement of one or more display generating components (e.g., with the user's head for head-mounted devices, or with the user's hand for handheld devices such as tablets or smartphones). The user's viewpoint determines what is visible within the viewport. The viewpoint typically specifies the position and orientation relative to the 3D environment, and as the viewpoint moves, the view of the 3D environment also moves within the viewport. For head-mounted devices, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptibly accurate view of the 3D environment that offers an immersive experience when the user is using the head-mounted device. For handheld or fixed devices, the viewpoint shifts with the movement of the handheld or fixed device and / or with changes in the user's positioning relative to the handheld or fixed device (e.g., the user moves towards, away from, up, down, right, and / or left). For a device including a display generation component with virtual pass-through, portions of the physical environment visible (e.g., displayed and / or projected) via one or more display generation components are based on the field of view of one or more cameras communicating with the display generation component, which typically move with the movement of the display generation component (e.g., for a head-mounted device, it moves with the movement of the user's head, or for a handheld device such as a tablet or smartphone, it moves with the movement of the user's hand), because the user's viewpoint moves with the movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual objects are updated based on the movement of the user's viewpoint)).For a display generating component with optical transparency, portions of the physical environment visible through one or more display generating components (e.g., optically visible through one or more portions or fully transparent portions of the display generating component) are based on the user's field of view through the portion or fully transparent portion of the display generating component (e.g., for a head-mounted device, it moves with the movement of the user's head, or for a handheld device such as a tablet or smartphone, it moves with the movement of the user's hand), because the user's viewpoint moves with the movement of the user's field of view through the portion or fully transparent portion of the display generating component (and the appearance of one or more virtual objects is updated based on the user's viewpoint).

[0043] In some embodiments, the representation of the physical environment (e.g., displayed via virtual passthrough or optical passthrough) may be partially or completely occluded by the virtual environment. In some embodiments, the amount of virtual environment displayed (e.g., the amount of physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level may optionally result in more virtual environment being displayed, replacing and / or occluding more physical environment, and decreasing the immersion level may optionally result in less virtual environment being displayed, thereby revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, displayed with increased transparency), and one or more third background objects are de-emphasized. In some embodiments, the level of immersion includes the associated degree to which virtual content (e.g., a virtual environment and / or virtual content) displayed by the computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual environment. Optionally, this includes the number of items in the displayed background content and / or the displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed via the display generating component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed via the display generating component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included within a background on which the virtual content is displayed (e.g., background content in a representation of the physical environment). In some implementations, background content includes user interfaces (e.g., user interfaces corresponding to applications generated by a computer system), virtual objects not associated with or included in the virtual environment and / or virtual content (e.g., files generated by the computer system or representations of other users), and / or real objects (e.g., transparent objects representing real objects in the user's surrounding physical environment, visible such that they are displayed via display generation components and / or via transparent or semi-transparent components of the display generation components, because the computer system does not obscure / impede their visibility through the display generation components). In some implementations, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real objects are displayed in an unobstructed manner. For example, a virtual environment with a low immersion level may optionally be displayed simultaneously with background content, which may optionally be displayed at full brightness, color, and / or semi-transparency.In some implementations, at higher immersion levels (e.g., a second immersion level above the first immersion level), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full-screen or fully immersive mode). Alternatively, a virtual environment displayed at a medium immersion level is displayed simultaneously with darkened, blurred, or otherwise de-emphasized background content. In some implementations, the visual characteristics of background objects differ between background objects. For example, at a particular immersion level, one or more first background objects are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, and / or displayed with increased transparency), and one or more third background objects are stopped from displaying. In some implementations, zero immersion or a zero immersion level corresponds to a virtual environment that is stopped from displaying, and instead, a representation of the physical environment (optionally having one or more virtual objects, such as an application, window, or virtual 3D object) is displayed, and the representation of the physical environment is not occluded by the virtual environment. Using physical input elements to adjust immersion levels provides a quick and efficient way to adjust immersion, which enhances the operability of computer systems and makes user-device interfaces more efficient.

[0044] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same location and / or position within the user's viewpoint, the virtual object remains viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the direction forward of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); therefore, the user's viewpoint remains fixed even when the user's gaze shifts without moving the user's head. In embodiments where the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the computer system's display generating component. For example, a viewpoint-locked virtual object displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint, even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation of a viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In an implementation where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so the virtual object is also referred to as a "head-locked virtual object".

[0045] Environment-locked visual objects: When a computer system displays a virtual object at a location and / or position within the user's viewpoint, the virtual object is environment-locked (or, "world-locked"), the location and / or position being based on a location and / or object within a three-dimensional environment (e.g., a physical or virtual environment) (e.g., selected and / or anchored to that location and / or object with reference to it). As the user's viewpoint moves, the location and / or object in the environment relative to the user's viewpoint changes, causing the environment-locked virtual object to appear at different locations and / or positions within the user's viewpoint. For example, an environment-locked virtual object locked to a tree immediately in front of the user appears at the center of the user's viewpoint. When the user's viewpoint shifts to the right (e.g., the user's head turns to the right) so that the tree is now centered to the left in the user's viewpoint (e.g., the tree's position shifts in the user's viewpoint), the environment-locked virtual object locked to the tree appears centered to the left in the user's viewpoint. In other words, the position and / or orientation of an environment-locked virtual object displayed in the user's viewpoint depends on the position to which the virtual object is locked and / or the orientation and / or orientation of the object within the environment. In some implementations, the computer system uses a stationary frame of reference (e.g., a coordinate system anchored to a fixed position and / or object in the physical environment) to determine the position of the environment-locked virtual object displayed in the user's viewpoint. An environment-locked virtual object may be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object), or it may be locked to a movable part of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body that moves independently of the user's viewpoint, such as a hand, wrist, arm, or foot), causing the virtual object to move with the viewpoint or that part of the environment to maintain a fixed relationship between the virtual object and that part of the environment.

[0046] In some implementations, environment-locked or viewpoint-locked virtual objects exhibit lazy following behavior, reducing or delaying their movement relative to the movement of a reference point they are following. In some implementations, when exhibiting lazy following behavior, the computer system intentionally delays the movement of the virtual object when movement of the reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) is detected. For example, when the reference point (e.g., that portion of the environment or the viewpoint) moves at a first rate, the virtual object is moved by the device to remain locked to the reference point, but moves at a second rate that is slower than the first rate (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up). In some implementations, when the virtual object exhibits lazy following behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold amount, such as 0 to 5 degrees or 0 cm to 50 cm). For example, when the reference point (e.g., the part of the environment or viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a viewpoint or part of the environment to which the virtual object is locked), and when the reference point (e.g., the part of the environment or viewpoint to which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a viewpoint or part of the environment to which the virtual object is locked), and then decreases as the amount of movement of the reference point increases above a threshold (e.g., a "lazy following" threshold) because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some implementations, maintaining a substantially fixed position of the virtual object relative to a reference point includes displaying the virtual object within a threshold distance (e.g., 1cm, 2cm, 3cm, 5cm, 15cm, 20cm, 50cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the reference point).

[0047] Hardware: Many different types of electronic systems enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet devices, and desktop / laptop computers. Head-mounted systems may have one or more speakers and an integrated opaque display. Alternatively, head-mounted systems may be configured to receive an external opaque display (e.g., a smartphone). Head-mounted systems may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. Head-mounted systems may have transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may have a medium through which light representing the image is directed to the person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display can be configured to selectively become opaque. Projection-based systems can employ retinal projection techniques that project graphic images onto a person's retina. Projection systems can also be configured to project virtual objects into a physical environment, such as as holograms or onto a physical surface. In some embodiments, controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 2The controller 110 is described in more detail. In some embodiments, the controller 110 is a computing device located locally or remotely relative to scene 105 (e.g., physical environment). For example, the controller 110 is a local server located within scene 105. Alternatively, the controller 110 is a remote server (e.g., a cloud server, central server, etc.) located outside scene 105. In some embodiments, the controller 110 is communicatively coupled to display generation components 120 (e.g., HMD, monitor, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., physical enclosure) of display generation component 120 (e.g., HMD or portable electronic device including display and one or more processors), one or more input devices in input device 125, one or more output devices in output device 155, one or more sensors in sensor 190, and / or one or more peripheral devices in peripheral device 195, or shares the same physical housing or support structure with one or more of the aforementioned devices.

[0048] In some implementations, display generation component 120 is configured to provide a user with an XR experience (e.g., at least the visual components of the XR experience). In some implementations, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 3 The display generation component 120 is described in more detail. In some embodiments, the functionality of the controller 110 is provided by and / or combined with the display generation component 120.

[0049] According to some implementation schemes, when a user is virtually and / or physically present within scene 105, display generation component 120 provides the user with an XR experience.

[0050] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, his / her hand, etc.). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device having a display facing the user's field of view and a camera facing scene 105. In some embodiments, the handheld device is optionally placed within a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content, wherein the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware used for displaying XR content (e.g., a handheld device or a tripod-mounted device) can be implemented on another type of hardware used for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface illustrating interaction with XR content triggered by an interaction occurring in the space in front of a handheld device or tripod-mounted device can be similarly implemented using an HMD, where the interaction occurs in the space in front of the HMD and the response to the XR content is displayed via the HMD. Similarly, a user interface illustrating interaction with XR content triggered by movement of a handheld device or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented using an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).

[0051] Despite Figure 1A The relevant features of the operating environment 100 are illustrated herein, but those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the exemplary embodiments disclosed herein.

[0052] Figures 1A to 1PVarious examples of computer systems for performing methods and providing audio, visual, and / or haptic feedback as part of the user interface described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b) for displaying to a user of the computer system a representation of virtual elements and / or a physical environment optionally generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216 to make it easier for a user who would otherwise use glasses or contact lenses to correct their vision to view the user interface, the one or more corrective lenses optionally being removably attached to one or more optical modules in the optical modules. While many user interfaces illustrated herein represent a single view of the user interface, user interfaces in an HMD may optionally use two optical modules (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b) for display, one optical module for the user's right eye and a different optical module for the user's left eye, presenting slightly different images to the two different eyes to generate the illusion of stereoscopic depth. A single view of the user interface is typically a right-eye view or a left-eye view; the depth effect is explained in text or using other diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display assembly 1-108) for displaying status information of the computer system to the user of the computer system (when the computer system is not worn) and / or to others near the computer system. This status information may optionally be generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, which may optionally be generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor assemblies 1-356 and / or sensor assemblies 1-356) for detecting information about the physical environment of the device. Figure 1I One or more sensors), which can be used (optionally with one or more illuminators, such as Figure 1IThe system combines the illuminators described herein to generate digital pass-through images, capture visual media (e.g., photographs and / or videos) corresponding to the physical environment, or determine the pose (e.g., positioning and / or orientation) of physical objects and / or surfaces in the physical environment, enabling the placement of virtual objects based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor assemblies 1-356 and / or...) for detecting hand positioning and / or movement. Figure 1I One or more sensors), which can be used (optionally with one or more illuminators, such as Figure 1I The illuminators 6-124 described herein (in combination) determine when one or more air gestures are performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I (Eye-tracking and gaze-tracking sensors in the system), these sensors can be used (optionally combined with one or more lights, such as...) Figure 10The light (11.3.2-110) in the image determines attention or gaze localization and / or gaze movement, which can optionally be used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine user facial expressions and / or hand movements for generating an avatar or representation of the user, such as an anthropomorphic avatar or representation for real-time communication sessions, wherein the avatar has facial expressions, hand movements, and / or body movements detected by the user based on or similar to the device. Gaze and / or attention information may optionally be combined with hand tracking information to determine user interaction with one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132 and / or dials or buttons 1-328), knobs (e.g., first buttons 1-128, buttons 11.1.1-114 and / or dials or buttons 1-328), digital crowns (e.g., pressable and twistable or rotatable first buttons 1-128, buttons 11.1.1-114 and / or dials or buttons 1-328), touchpads, touchscreens, keyboards, mice and / or other input devices. One or more buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132, and / or dials or buttons 1-328) may optionally be used to perform system operations, such as recentering content in the user-visible 3D environment of the device, displaying the main user interface for launching an application, initiating a real-time communication session, or initiating the display of a virtual 3D background. A knob or digital crown (e.g., pressable and twistable or rotatable first buttons 1-128, buttons 11.1.1-114, and / or dials or buttons 1-328) may optionally be rotatable to adjust parameters of the visual content, such as the level of immersion of the virtual 3D environment (e.g., the extent to which the virtual content occupies the user's viewport in the 3D environment) or other parameters associated with the 3D environment and the virtual content displayed via optical modules (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b).

[0053] Figure 1BExamples of head-mounted display (HMD) devices 1-100 configured to be worn by a user and provide virtual and altered / mixed reality (VR / AR) experiences are illustrated in front, top, and perspective views. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strip assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 secured at either end to the electronic strip assembly 1-104. The electronic strip assembly 1-104 and the strap 1-106 may be part of a retention assembly configured to wrap around the user's head to hold the display unit 1-102 against the user's face.

[0054] In at least one example, the strap assembly 1-106 may include a first strap 1-116 configured to wrap around the back of the user's head and a second strap 1-117 configured to extend above the top of the user's head. As shown, the second strap may extend between the first electronic strip 1-105a and the second electronic strip 1-105b of the electronic strip assembly 1-104. The strip assembly 1-104 and the strap assembly 1-106 may be part of a fixing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.

[0055] In at least one example, the fixing mechanism includes a first electronic strip 1-105a, which includes a first proximal end 1-134 coupled to a display unit 1-102 (e.g., a housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite to the first proximal end 1-134. The fixing mechanism may also include a second electronic strip 1-105b, which includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite to the second proximal end 1-138. The fixing mechanism may also include a first strip 1-116 and a second strip 1-117, the first strip including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second strip extending between the first electronic strip 1-105a and the second electronic strip 1-105b. Strips 1-105a to b and strip 1-116 may be coupled via a connecting mechanism or assembly 1-114. In at least one example, the second strip 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between a first proximal end 1-134 and a first distal end 1-136, and a second end 1-148 coupled to the second electronic strip 1-105b between a second proximal end 1-138 and a second distal end 1-140.

[0056] In at least one example, the first and second electronic strips 1-105a-b comprise plastic, metal, or other structural materials forming a substantially rigid strip shape. In at least one example, the first strip 1-116 and the second strip 1-117 are formed of an elastic, flexible material including woven textiles, rubber, etc. The first strip 1-116 and the second strip 1-117 may be flexible enough to conform to the shape of the user's head when wearing the HMD 1-100.

[0057] In at least one example, one or more of the first and second electronic stripes 1-105a to b may define an inner strip volume and include one or more electronic components disposed within the inner strip volume. In one example, such as Figure 1B As shown, the first electronic strip 1-105a may include electronic components 1-112. In one example, electronic components 1-112 may include a speaker. In another example, electronic components 1-112 may include computing components, such as a processor.

[0058] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is located in... Figure 1B The section marked 1-152 with dashed lines is because the front cover assembly 1-108 is configured to obscure the first opening 1-152 from the field of view during HMD assembly. The housing 1-150 may also define a rearward second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover disposed in or across the front opening 1-152 to obscure the front opening 1-152, and a display screen (shown in other figures). In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 in general, has a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 can be bent as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, wherein the display unit 1-102 is pressed.

[0059] In at least one example, the housing 1-150 may define a first hole 1-126 between a first opening 1-152 and a second opening 1-154, and a second hole 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first hole 1-128, and a second button 1-132 disposed in the second hole 1-130. The first button 1-128 and the second button 1-132 are pressable through their respective holes 1-126 and 1-130. In at least one example, the first button 1-126 and / or the second button 1-130 may be a rotary dial and a pressable button. In at least one example, the first button 1-128 is a pressable and rotary dial button, and the second button 1-132 is a pressable button.

[0060] Figure 1C A rear perspective view of HMD 1-100 is illustrated. HMD 1-100 may include a light seal 1-110 extending rearwardly around the periphery of housing 1-150 of display unit 1-108, as shown. The light seal 1-110 may be configured to extend from housing 1-150 to the user's face, surrounding the user's eyes, to block external light from being visible. In one example, HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b disposed at or within a rearwardly facing second opening 1-154 defined by housing 1-150 and / or disposed within the internal volume of housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a to b may include a corresponding display screen 1-122a, 1-122b configured to project light toward the user's eyes in a rearward direction through the second opening 1-154.

[0061] In at least one example, reference Figure 1B and Figure 1C Both, the display assembly 1-108 can be a front-facing display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a to b can be configured to project light in a second rearward direction opposite to the first direction. As described above, the light seal 1-110 can be configured to block light from outside the HMD 1-100 from reaching the user's eyes, including a light seal composed of a display screen configured to project light in a second rearward direction opposite to the first direction. Figure 1BThe front perspective view shows the light projected by the front display screen of the display assembly 1-108. In at least one example, the HMD 1-100 may also include a curtain 1-124 that blocks the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a to b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.

[0062] Figure 1B and Figure 1C Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1D to 1F This is in any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1D to 1F Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1B and Figure 1C Examples of devices, features, components, and parts are shown.

[0063] Figure 1D An exploded view of an example HMD 1-200 including its various parts or components, separated according to the modularity and selective coupling of these components. For example, HMD 1-200 may include a strip 1-216 selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first and second strips 1-205a to 1-205b are removably coupled to a display unit 1-202.

[0064] Furthermore, HMD 1-200 may include a light-sealing member 1-210 configured to be removably coupled to display unit 1-202. HMD 1-200 may also include a lens 1-218, which may be removably coupled to display unit 1-202, for example, on a first display assembly including a display screen and a second display assembly. Lens 1-218 may include a custom prescription lens configured for vision correction. As noted, in Figure 1DThe exploded view shows that each component described above can be removably coupled, attached, reattached, and replaced to update the component, or replaced for different users. For example, belts such as belt 1-216, light seals such as light seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a to b can be replaced according to the user, so that these parts are customized to fit and correspond to a single user of HMD 1-200.

[0065] Figure 1D Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1B , Figure 1C and Figures 1E to 1F This is in any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1B , Figure 1C and Figures 1E to 1F Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1D Examples of devices, features, components, and parts are shown.

[0066] Figure 1E An exploded view illustrating an example of a display unit 1-306 of an HMD is shown. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.

[0067] In at least one example, the display unit 1-306 may further include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a to b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a to b has at least one motor, such that the motor is capable of translating the display screens 1-322a to b to match the interpupillary distance of the user's eyes.

[0068] In at least one example, display unit 1-306 may include a dial or button 1-328 that is pressable relative to frame 1-350 and accessible to a user outside frame 1-350. Button 1-328 may be electrically connected to motor assembly 1-362 via a controller, such that button 1-328 can be operated by a user to cause the motor of motor assembly 1-362 to adjust the positioning of display screens 1-322a to b.

[0069] Figure 1E Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1B to 1D and Figure 1F This is in any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1B to 1D and Figure 1F Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1E Examples of devices, features, components, and parts are shown.

[0070] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positioning of the first display sub-assemblies 1-420a and 1-420b of the rear display assembly 1-421, including a first and second corresponding display screen for interpupillary adjustment, as described above.

[0071] References in this article Figures 1B to 1E The following figures, which are referenced in this disclosure, will be used to describe the subject in more detail. Figure 1F The exploded view shows the various parts, systems, and assemblies. Figure 1F The display unit 1-406 shown can be connected with Figures 1B to 1E The shown fixture assembly and integration includes electronic strips, belts, and other components (including light seals, connecting assemblies, etc.).

[0072] Figure 1F Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1B to 1EThis is in any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1B to 1E Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1F Examples of devices, features, components, and parts are shown.

[0073] Figure 1G An example is the front cover assembly 3-100 of the HMD device described herein (e.g., Figure 1G An exploded perspective view of the front cover assembly 3-1) of the HMD 3-100 shown or any other HMD device shown and described herein. Figure 1B The front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or "awning"), an adhesive layer 3-106, a display assembly 3-108 including a biconvex lens panel or array 3-110, and a structural decorative element 3-112. The adhesive layer 3-106 secures the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the decorative element 3-112. The decorative element 3-112 secures the various components of the front cover assembly 3-100 to the frame or base of the HMD device.

[0074] In at least one example, such as Figure 1G As shown, the transparent cover 3-102, the shield 3-104, and the display assembly 3-108 including a biconvex lens array 3-110 can be bent to adapt to the curvature of a user's face. The transparent cover 3-102 and the shield 3-104 can be bent in two or three dimensions, for example, vertically in and out of the Z-plane along the Z direction, and horizontally in and out of the Z-plane along the X direction. In at least one example, the display assembly 3-108 may include the biconvex lens array 3-110 and a display panel with pixels configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 can be bent in at least one direction (e.g., the horizontal direction) to adapt to the curvature of a user's face from one side (e.g., the left) to the other (e.g., the right). In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in the following figures, but may include the biconvex lens array 3-110 and the display layer) may be similarly or concentrically curved in the horizontal direction to accommodate the curvature of the user's face.

[0075] In at least one example, the cover 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the cover 3-104 may include one or more opaque portions, such as opaque ink-printed portions or other opaque film portions on the back of the cover 3-104. When the HMD device is worn, the rear surface may be the surface of the cover 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the cover 3-104 opposite the rear surface. In at least one example, one or more opaque portions of the cover 3-104 may include peripheral portions that visually conceal any components surrounding the outer periphery of the display screen of the display assembly 3-108. In this way, the opaque portions of the cover conceal any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the cover 3-104, including electronic components, structural components, etc.

[0076] In at least one example, the housing 3-104 may define one or more transparent aperture portions 3-120 through which sensors can transmit and receive signals. In one example, portion 3-120 is an aperture through which sensors can extend or transmit and receive signals. In one example, portion 3-120 is a transparent portion, or a portion more transparent than the surrounding translucent or opaque portion of the housing, through which sensors can transmit and receive signals through the housing and via transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.

[0077] Figure 1G Any of the features, components, and / or parts shown herein (including their arrangement and configuration) may be included individually or in any combination of any other example of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination of any other example of the devices, features, components, and parts described herein. Figure 1G Examples of devices, features, components, and parts are shown.

[0078] Figure 1H An exploded view of an example HMD device 6-100 is shown. The HMD device 6-100 may include a sensor array or system 6-102 comprising one or more sensors, cameras, projectors, etc., mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 on which one or more sensors of the sensor system 6-102 may be fixed / secured.

[0079] Figure 1I A portion of an HMD device 6-100, including a front transparent cover 6-104 and a sensor system 6-102, is illustrated. The sensor system 6-102 may include multiple different sensors, transmitters, and receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is illustrated on the front of the sensor system 6-102 to illustrate the relative positioning of the various sensors and transmitters and the orientation of each sensor / transmitter in system 6-102. As referenced herein, "side," "side," "lateral," "horizontal," and other similar terms refer to... Figure 1J The orientation or direction indicated by the X-axis. Terms such as "vertical," "upward," "downward," and similar terms refer to the orientation or direction indicated by... Figure 1J The orientation or direction indicated by the Z-axis. Terms such as "frontward," "rearward," "forward," "backward," and similar terms refer to the orientation or direction indicated by the Z-axis. Figure 1J The orientation or direction indicated by the Y-axis shown.

[0080] In at least one example, a transparent cover 6-104 may define the front outer surface of an HMD device 6-100, and a sensor system 6-102, including various sensors and their components, may be positioned behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through it, including both light detected by the sensor system 6-102 and light emitted therefrom.

[0081] As described elsewhere herein, the HMD device 6-100 may include one or more controllers, which include processors for electrically coupling various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as displays. Furthermore, as will be shown in more detail below with reference to other accompanying drawings, various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to the HMD device 6-100. Figure 1I Various structural frame components, brackets, etc., not shown. For clarity, Figure 1I The components of sensor system 6-102 are shown, which are not attached to or electrically coupled to other components.

[0082] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. These instructions may include, or cause the processor to execute, one or more algorithms for self-correcting the angle and position of the various cameras described herein as the camera's initial positioning, angle, or orientation is affected by collisions or deformations due to accidental drop events or other events over time.

[0083] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. System 6-102 may include two scene cameras 6-102, respectively positioned on either side of the nose bridge or arched structure of the HMD device 6-100, such that each of the two cameras 6-106 approximately corresponds to the positioning of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and, when the HMD device 6-100 is used, provide images and content for MR video pass-through to a display screen facing the user's eyes. The scene cameras 6-106 may also be used for environment and object reconstruction.

[0084] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 that is generally pointing forward in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction as well as user hand and body tracking. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 centrally located along the width of the HMD device 6-100 (e.g., along the X-axis). For example, the second depth sensor 6-110 may be located above the central bridge of the nose or on an adapter structure above the nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction as well as hand and body tracking. In at least one example, the second depth sensor may include a LiDAR sensor.

[0085] In at least one example, the sensor system 6-102 may include a depth projector 6-112, which is typically forward-facing to project electromagnetic waves (e.g., in the form of a predetermined spot pattern) into or within the field of view of the user and / or scene camera 6-106, or into or beyond the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a spot pattern, which are reflected from an object and back into the aforementioned depth sensors, including depth sensors 6-108 and 6-110. In at least one example, the depth projector 6-112 may be used for environment and object reconstruction, as well as hand and body tracking.

[0086] In at least one example, the sensor system 6-102 may include a downward-facing camera 6-114, whose field of view is generally directed downwards relative to the HMD device 6-100 on the Z-axis. In at least one example, the downward-facing camera 6-114 may be positioned as shown on the left and right sides of the HMD device 6-100 and used for hand and body tracking, head-mounted device tracking, and facial avatar detection and creation for displaying a user avatar on the front display screen of the HMD device 6-100 as described elsewhere herein. For example, the downward-facing camera 6-114 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.

[0087] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be positioned as shown on the left and right sides of the HMD device 6-100 and used for hand and body tracking, head-mounted device tracking, and facial avatar detection and creation for displaying a user avatar on the front display screen of the HMD device 6-100 as described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. Used for hand and body tracking, head-mounted device tracking, and facial avatar creation. In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right views along the X-axis or in a direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, head-mounted device tracking, and facial avatar detection and reconstruction.

[0088] In at least one example, the sensor system 6-102 may include multiple eye-tracking and gaze-tracking sensors for determining identity, status, and the user's gaze direction during and / or prior to use. In at least one example, the eye / gaze-tracking sensor may include a nose-eye camera 6-120 positioned on either side of the user's nose and adjacent to the user's nose when wearing the HMD device 6-100. The eye / gaze sensor may also include a bottom eye camera 6-122 positioned below the respective user's eye for capturing images of the eye for use in facial avatar detection and creation, gaze tracking, and iris identification functions.

[0089] In at least one example, sensor system 6-102 may include an infrared illuminator 6-124 that is pointed outward from HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of sensor system 6-102. In at least one example, sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, flicker sensor 6-126 may detect the refresh rate of the overhead light to avoid display flicker. In one example, infrared illuminator 6-124 may include a light-emitting diode and may be specifically designed for low-light environments to illuminate a user's hands and other objects in low light for detection by the infrared sensors of sensor system 6-102.

[0090] In at least one example, multiple sensors (including scene camera 6-106, downward camera 6-114, chin camera 6-116, side camera 6-118, depth projector 6-112, and depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for sizing, thereby improving the hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, as described above and Figure 1I The downward-facing camera 6-114, the jaw camera 6-116, and the side camera 6-118 shown can be wide-angle cameras capable of operating in both the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, and 6-118 can operate solely in black-and-white light detection to simplify image processing and achieve sensitivity.

[0091] Figure 1I Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1J to 1L This is in any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1J to 1LAny of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1I Examples of devices, features, components, and parts are shown.

[0092] Figure 1J A lower perspective view of an example HMD 6-200 including a cover or shield 6-204 fixed to a frame 6-230 is shown. In at least one example, a sensor 6-203 of a sensor system 6-202 may be disposed around the periphery of the HMD 6-200 such that the sensor 6-203 is disposed outwardly around the periphery of the display area or zone 6-232 so as not to obstruct the view of the displayed light. In at least one example, the sensor may be disposed behind the shield 6-204 and aligned with a transparent portion of the shield, thereby allowing light to pass back and forth through the shield 6-204 by the sensor and the projector. In at least one example, an opaque ink or other opaque material or film / layer may be disposed on the shield 6-204 around the display area 6-232 to conceal the components of the HMD 6-200 outside the display area 6-232 rather than through a transparent portion defined by the opaque portion through which the sensor and the projector transmit and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass through the display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area surrounding the periphery of the display and the shield 6-204.

[0093] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent areas 6-209 through which the sensor 6-203 of the sensor system 6-202 transmits and receives signals. In the illustrated examples, the sensor 6-203 of the sensor system 6-202, which transmits and receives signals through the shield 6-204, or more specifically through the transparent area 6-209 defined by the opaque portion 6-207 of the shield 6-204, may include... Figure 1I The examples illustrate those same or similar sensors, such as depth sensors 6-108 and 6-110, depth projector 6-112, first scene camera and second scene camera 6-106, first downward camera and second downward camera 6-114, first side camera and second side camera 6-118, and first infrared illuminator and second infrared illuminator 6-124. These sensors also... Figure 1K and Figure 1L The example is shown. Other sensors, sensor types, number of sensors, and their relative positioning can be included in one or more other examples of the HMD.

[0094] Figure 1J Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1I and Figures 1K to 1L This is in any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1I and Figures 1K to 1L Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1J Examples of devices, features, components, and parts are shown.

[0095] Figure 1K A front view of a portion of an example of an HMD device 6-300, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330, is shown. Figure 1K The examples shown do not include a front cover or shield to illustrate brackets 6-336 and 6-338. For example, Figure 1J The shield 6-204 shown includes an opaque portion 6-207 that visually covers / blocks the view of anything outside the display / display area 6-334 (e.g., radially / peripherally outside the display / display area), including the sensor 6-303 and the bracket 6-338.

[0096] In at least one example, various sensors of sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, scene camera 6-306 includes strict tolerances for angles relative to each other. For example, the tolerance for the mounting angle between two scene cameras 6-306 may be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such strict tolerances, in one example, scene camera 6-306 may be mounted to bracket 6-338 instead of a housing. The bracket may include a cantilever on which scene camera 6-306 and other sensors of sensor system 6-302 may be mounted to maintain their positioning and orientation in the event of a drop event caused by a user that results in any deformation of other brackets 6-226, housing 6-330, and / or housing.

[0097] Figure 1K Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1I to 1J and Figure 1L This is in any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1I to 1J and Figure 1LAny of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1K Examples of devices, features, components, and parts are shown.

[0098] Figure 1L A bottom view illustrating an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is shown. The sensor system 6-402 is compatible with the above and other parts of this document (including references). Figures 1I to 1K Other sensor systems described are similar. In at least one example, the jaw camera 6-416 may be oriented downwards to capture images of the user's lower facial features. In one example, the jaw camera 6-416 may be directly coupled to a frame or housing 6-430 or one or more internal brackets that are directly coupled to the frame or housing 6-430 shown. The frame or housing 6-430 may include one or more holes / openings 6-415 through which the jaw camera 6-416 transmits and receives signals.

[0099] Figure 1L Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1I to 1K This is in any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1I to 1K Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1L Examples of devices, features, components, and parts are shown.

[0100] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is illustrated. This IPD adjustment system includes a first optical module and a second optical module 11.1.1-104a-104a-104a-105a, slidably engaged / coupled to corresponding guide rods 11.1.1-108a ... In at least one example, buttons 11.1.1-114 can be electrically communicated with the first motor and the second motors 11.1.1-110a to b via a processor or other circuit components to activate the first motor and the second motors 11.1.1-110a to b and respectively cause the first optical module and the second optical modules 11.1.1-104a to b to change their positions relative to each other.

[0101] In at least one example, the first and second optical modules 11.1.1-104a to b may include corresponding display screens configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, the user can manipulate (e.g., press and / or rotate) buttons 11.1.1-114 to activate positional adjustment of the optical modules 11.1.1-104a to b to match the interpupillary distance of the user's eyes. The optical modules 11.1.1-104a to b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD, such that the optical modules 11.1.1-104a to b can be adjusted to match the IPD.

[0102] In one example, a user can manipulate buttons 11.1.1-114 to cause automatic positional adjustment of the first and second optical modules 11.1.1-104a to b. In another example, a user can manipulate buttons 11.1.1-114 to cause manual adjustment, moving the optical modules 11.1.1-104a to b further or closer (e.g., when the user rotates buttons 11.1.1-114 in one way or another) until the user visually matches their own IPD. In one example, manual adjustment is communicated electronically via one or more circuits, and the power for moving the optical modules 11.1.1-104a to b via motors 11.1.1-110a to b is supplied by a power source. In another example, the adjustment and movement of the optical modules 11.1.1-104a to b via the manipulation buttons 11.1.1-114 are mechanically actuated via the movement buttons 11.1.1-114.

[0103] Figure 1M Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination in any other example of the devices, features, components, and parts shown and described herein. Similarly, any of the features, components, and / or parts shown and described with reference to any other shown and described, and in any combination herein, may be included individually or in any combination in Figure 1M Examples of devices, features, components, and parts are shown.

[0104] Figure 1N A front perspective view of a portion of HMD 11.1.2-100 is shown, including an outer structural frame 11.1.2-102 defining first and second holes 11.1.2-106a, 11.1.2-106b, and an inner or intermediate structural frame 11.1.2-104. Holes 11.1.2-106a to b are located in... Figure 1N The holes 11.1.2-106a to b are shown in dashed lines because viewing the HMD 11.1.2-100 may be obstructed by one or more other components coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as illustrated. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second holes 11.1.2-106a to b.

[0105] Mounting brackets 11.1.2-108 may include intermediate or central portions 11.1.2-109 coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portions 11.1.2-109 may not be the geometric center or middle of the brackets 11.1.2-108. Instead, the intermediate / central portions 11.1.2-109 may be positioned between a first cantilever extension arm and a second cantilever extension arm extending away from the intermediate portions 11.1.2-109. In at least one example, mounting bracket 108 includes first cantilever arms 11.1.2-112 and second cantilever arms 11.1.2-114 extending away from the intermediate portions 11.1.2-109 of the mounting brackets 11.1.2-108 coupled to the inner frame 11.1.2-104.

[0106] like Figure 1N As shown, the outer frame 11.1.2-102 may define a curved geometry on its lower side to adapt to the user's nose when the user wears the HMD 11.1.2-100. This curved geometry may be referred to as the bridge of the nose 11.1.2-111 and is centrally located on the lower side of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the inner frame 11.1.2-104 between holes 11.1.2-106a and b, such that the cantilever 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the central portion 11.1.2-109 to complement the nose bridge geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to adapt to the user's nose, as described above. The geometry of the bridge of the nose 11.1.2-111 adapts to the nose, as it provides a curvature that conforms to the shape of the user's nose, offering a comfortable fit from above, above, and around.

[0107] The first cantilever 11.1.2-112 may extend in a first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108, and the second cantilever 11.1.2-114 may extend in a second direction opposite to the first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as “cantilever” or “cantilever” arms because each arm 11.1.2-112, 11.1.2-114 includes free distal ends 11.1.2-116, 11.1.2-118, respectively, which are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, arms 11.1.2-112 and 11.1.2-114 extend from the middle section 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102 and 11.1.2-104 are not attached.

[0108] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include multiple sensors 11.1.2-110a-f. Each of the multiple sensors 11.1.2-110a-f may include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a-f may be used for object recognition in three-dimensional space, making it important to maintain the precise relative positioning of two or more of the multiple sensors 11.1.2-110a-f. The cantilever nature of the mounting bracket 11.1.2-108 protects the sensors 11.1.2-110a-f from damage and displacement in the event of an accidental drop by the user. Because the sensors 11.1.2-110a-f cantilevered on the arms 11.1.2-112 and 11.1.2-114 of the mounting bracket 11.1.2-108, the stress and deformation of the internal frame and / or the external frames 11.1.2-104 and 11.1.2-102 are not transmitted to the cantilever arms 11.1.2-112 and 11.1.2-114, and therefore do not affect the relative position of the sensors 11.1.2-110a-f coupled to / mounted to the mounting bracket 11.1.2-108.

[0109] Figure 1NAny of the features, components, and / or parts shown herein (including their arrangement and configuration) may be included individually or in any combination in any other example of the device, feature, or component described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination. Figure 1N Examples of devices, features, components, and parts are shown.

[0110] Figure 10 An example of optical modules 11.3.2-100 for use in electronic devices, such as HMDs, including the HDM devices described herein, is illustrated. As shown in one or more other examples described herein, optical modules 11.3.2-100 may be one of two optical modules within an HMD, wherein each optical module is aligned to project light toward a user's eye. In this way, a first optical module may project light toward a user's first eye via a display screen, and a second optical module of the same device may project light toward a user's second eye via another display screen.

[0111] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a tube or optical module tube. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward the user's eyes when the HMD to which the display module 11.3.2-100 belongs is worn during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.

[0112] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eye during use. In at least one example, the optical module 11.3.2-100 may also include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the camera 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light-emitting diodes (LEDs) or other lights configured to project light toward the user's eyes when the HMD is worn. The individual lights 11.3.2-110 in the light strips 11.3.2-108 may be spaced apart around the light strips 11.3.2-108, and are therefore uniformly or non-uniformly spaced around the display 11.3.2-104 at various locations on the light strips 11.3.2-108 and around the display 11.3.2-104.

[0113] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light onto the user's eyes through the viewing opening 11.3.2-101. In one example, a camera 11.3.2-106 is configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.

[0114] As mentioned above, Figure 10 Each of the components and features of the optical modules 11.3.2-100 shown can be replicated in another (e.g., a second) optical module set up with the HMD to interact with the user’s other eye (e.g., projecting light and capturing images).

[0115] Figure 10 Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1P Any other example of the device, feature, component, and part shown or otherwise described herein. Similarly, refer to... Figure 1P Any of the features, components, and / or parts shown, described, or otherwise referred to herein (including their arrangement and configuration) may be included individually or in any combination. Figure 10 Examples of devices, features, components, and parts are shown.

[0116] Figure 1P A cross-sectional view of an example optical module 11.3.2-200 is shown, which includes a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or channel 11.3.2-212 and a second aperture or channel 11.3.2-214. Channels 11.3.2-212 and 11.3.2-214 can be configured to slidably engage corresponding tracks or guides of an HMD device to allow the optical module 11.3.2-200 to be adjusted and positioned relative to the user's eye to match the user's interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage the guide rod to secure the optical module 11.3.2-200 in the appropriate position within the HMD.

[0117] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user's eye when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eye. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, lenses 11.3.2-216 are positioned above light strips 11.3.2-208 and one or more eye-tracking cameras 11.3.2-206, such that cameras 11.3.2-206 are configured to capture an image of a user's eye through lenses 11.3.2-216, and light strips 11.3.2-208 include lamps configured to project light onto the user's eye through lenses 11.3.2-216 during use.

[0118] Figure 1P Any of the features, components, and / or parts shown herein (including their arrangement and configuration) may be included individually or in any combination of any of the other examples of devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination of any of the other examples of devices, features, components, and parts described herein. Figure 1P Examples of devices, features, components, and parts are shown.

[0119] Figure 2 This is a block diagram of an example controller 110 according to some implementation schemes. Although some specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the implementation schemes disclosed herein. Therefore, as a non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0120] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0121] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 220 may optionally include one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores programs, modules, and data structures, or subsets thereof, including optional operating system 230 and XR experience module 240.

[0122] Operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., single XR experiences for one or more users, or multiple XR experiences for corresponding groups of one or more users). To this end, in various embodiments, XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0123] In some implementations, the data acquisition unit 241 is configured to acquire data from... Figure 1A The data acquisition unit 241 includes at least the display generation component 120, and optionally acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0124] In some implementations, the tracking unit 242 is configured to map scene 105, and the tracking at least shows the generating component 120 relative to... Figure 1A The tracking unit 242 tracks the location / position of scene 105, and optionally tracks the position of one or more of input devices 125, output devices 155, sensors 190, and / or peripheral devices 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the location / position of one or more portions of the user's hand, and / or the position of one or more portions of the user's hand relative to the user's hand. Figure 1A The motion of scene 105 relative to the display generation component 120 and / or relative to a coordinate system (defined relative to the user's hand). The following refers to the motion relative to... Figure 4 The hand tracking unit 244 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the user's gaze (or more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to XR content displayed via display generation component 120. The following description is relative to... Figure 5 The eye-tracking unit 243 is described in more detail.

[0125] In some implementations, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120, and optionally by one or more of output device 155 and / or peripheral device 195. To this end, in various implementations, coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0126] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 248 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0127] Although the data acquisition unit 241, the tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 are shown residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 may reside in a separate computing device.

[0128] also, Figure 2 This is used more as a functional description of various features that can exist in a particular specific implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, individually shown items can be combined, and some items can be separated. For example, Figure 2 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.

[0129] Figure 3This is a block diagram illustrating an example of generating component 120 according to some embodiments. Although some specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZigBee, and / or similar interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0130] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include inertial measurement units (IMUs), accelerometers, gyroscopes, thermometers, one or more physiological sensors (e.g., blood pressure monitors, heart rate monitors, blood oxygen sensors, blood glucose sensors, etc.), one or more microphones, one or more speakers, haptic engines, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).

[0131] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, display generation component 120 (e.g., HMD) includes a single XR display. In another example, display generation component 120 includes XR displays for each of the user's eyes. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting either MR or VR content.

[0132] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to the scene the user would see in the absence of a display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0133] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 may optionally include one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optional operating system 330 and XR rendering module 340.

[0134] Operating system 330 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, XR rendering module 340 includes a data acquisition unit 342, an XR rendering unit 344, an XR mapping generation unit 346, and a data transmission unit 348.

[0135] In some implementations, the data acquisition unit 342 is configured to acquire data from at least... Figure 1A The controller 110 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0136] In some implementations, the XR rendering unit 344 is configured to render XR content via one or more XR displays 312. To this end, in various implementations, the XR rendering unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0137] In some implementations, the XR mapping generation unit 346 is configured to generate XR maps based on media content data (e.g., 3D maps of mixed reality scenes or maps in which computer-generated objects can be placed to generate extended reality physical environments). To this end, in various implementations, the XR mapping generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0138] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 348 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0139] Although the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data transmission unit 348 are shown residing in a single device (e.g., Figure 1A The display generation component 120 is used, but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data transmission unit 348 may be located in a separate computing device.

[0140] also, Figure 3 This serves more as a functional description of various features that may exist in a particular specific implementation, and differs from the structural schematic diagram of the implementation described herein. As those skilled in the art will recognize, individually shown items can be combined, and some items can be separated. For example, Figure 3 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.

[0141] Figure 4 This is a schematic illustration of an example embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) by hand tracking unit 244 ( Figure 2 To control and track the location / position of one or more parts of the user's hand, and / or the location of one or more parts of the user's hand relative to the user's hand. Figure 1AThe scenario 105 refers to movement relative to a portion of the user's surrounding physical environment, relative to display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined relative to the user's hand. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0142] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures at least three-dimensional scene information including the human user's hand 406. The image sensor 404 captures images of the hand at sufficient resolution to distinguish the fingers and their corresponding positions. The image sensor 404 typically captures images of other parts of the user's body, or possibly all parts of the body, and may have scaling capabilities or be a dedicated sensor with increased magnification to capture images of the hand at the desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of scene 105, or serves as the image sensor for capturing the physical environment of scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in a way that uses the field of view of the image sensor 404 or a portion thereof to define an interaction space in which hand movements captured by the image sensor are considered input to the controller 110.

[0143] In some implementations, image sensor 404 outputs a sequence of frames containing 3D image data (and, in addition, possibly color image data) to controller 110, which extracts high-level information from the image data. This high-level information is typically provided via an application interface (API) running on the controller, which in turn drives display generation component 120. For example, a user can interact with software running on controller 110 by moving his hand 406 and changing his hand pose.

[0144] In some embodiments, image sensor 404 projects a speckle pattern onto a scene containing hand 406 and captures an image of the projected pattern. In some embodiments, controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) via triangulation based on the lateral offset of the specks in the pattern. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from image sensor 404. In this disclosure, it is assumed that image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, image sensor 404 (e.g., a hand-tracking device) may use other 3D mapping methods, such as stereo imaging or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.

[0145] In some implementations, hand tracking device 140 captures and processes time-series depth maps containing the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in image sensor 404 and / or controller 110 processes the 3D map data to extract image block descriptors of the hand from these depth maps. The software may match these descriptors with image block descriptors stored in database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D position of the user's hand joints and fingertips.

[0146] The software can also analyze the trajectories of the hand and / or fingers across multiple frames in a sequence to identify gestures. The pose estimation function described herein can be alternated with motion tracking, such that patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to find pose changes occurring in the remaining frames. Pose, motion, and gesture information is provided to an application running on controller 110 via the aforementioned API. The program can, for example, move and modify the image presented on display generation component 120 in response to pose and / or gesture information, or perform other functions.

[0147] In some implementations, gestures include air gestures. An air gesture is a gesture detected without the user touching an input element that is part of the device (e.g., computer system 101, one or more input devices 125 and / or hand tracking device 140) (or independent of an input element that is part of the device) and based on the detected movement of a part of the user's body (e.g., head, one or two arms, one or two hands, one or more fingers and / or one or two legs) through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., including a tapping gesture in which the hand moves a predetermined amount and / or speed in a predetermined pose, or a shaking gesture including a predetermined speed or amount of rotation of a part of the user's body)).

[0148] In some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures for interacting with an XR environment (e.g., a virtual or mixed reality environment) performed by the movement of a user's fingers relative to other fingers or portions of the user's hand. In some embodiments, air gestures are detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and are based on the detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or portion of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).

[0149] In some implementations where the input gesture is an air gesture (e.g., where the input device provides information to the computer system about which user interface element is the target of the user input in the absence of physical contact, such as contact with a user interface element displayed on a touchscreen, or contact with a mouse or touchpad to move the cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Therefore, in implementations involving air gestures, for example, the input gesture is combined with (e.g., simultaneously) movement of the user's fingers and / or hand to detect attention (e.g., gaze) toward a user interface element to perform pinch and / or tap input, as described below.

[0150] In some implementations, input gestures directed to a user interface object are performed, either directly or indirectly, by referencing the user interface object. For example, user input is performed directly on the user interface object based on the user's hand performing an input gesture at a location corresponding to the user interface object's position in the three-dimensional environment (e.g., determined based on the user's current viewpoint). In some implementations, when user attention to the user interface object (e.g., gazing) is detected, input gestures are performed indirectly on the user interface object based on the user's hand not being positioned at a location corresponding to the user interface object's position in the three-dimensional environment while the user is performing the input gesture. For example, for direct input gestures, the user can guide their input to the user interface object by initiating a gesture at or near a location corresponding to the user interface object's display position (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm measured from the outer edge or center of the option). For indirect input gestures, the user can guide their input to the user interface object by focusing on it (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location not corresponding to the user interface object's display position).

[0151] In some implementations, the input gestures (e.g., air gestures) used in the various examples and implementations described herein include pinch input and tap input for interacting with a virtual or mixed reality environment. For example, the pinch input and tap input described below are performed as air gestures.

[0152] In some implementations, pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other, i.e., optionally followed by an immediate (e.g., within 0 to 1 second) interruption of contact. A long pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before an interruption of contact is detected. For example, a long pinch gesture includes the user holding a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some implementations, a double pinch gesture as an air gesture includes two (e.g., more) pinch inputs (e.g., performed by the same hand) that are detected consecutively with each other immediately (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts the contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or within 2 seconds).

[0153] In some embodiments, pinch and drag gestures as air gestures include pinch gestures (e.g., pinching gestures or long pinch gestures) performed in conjunction with (e.g., following) drag input that changes the user's hand position from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the drag). In some embodiments, the user holds the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opening two or more of their fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together to touch each other and uses the drag gesture to move the same hand to the second position in the air). In some implementations, pinch input is performed by the user's first hand, and drag input is performed by the user's second hand (e.g., while the user continues pinch input with the user's first hand, the user's second hand moves in the air from a first position to a second position). In some implementations, input gestures as air gestures include inputs performed using both of the user's hands (e.g., pinch and / or tap inputs). For example, input gestures include two (e.g., more) pinch inputs performed in combination with each other (e.g., simultaneously or within a predefined time period). For example, a first pinch gesture (e.g., pinch input, long pinch input, or pinch and drag input) is performed using the user's first hand, and a second pinch input is performed using the other hand (e.g., the second hand in the user's two hands).

[0154] In some implementations, a tap input performed as an air gesture (e.g., pointing at a user interface element) includes movement of a user's finger toward the user interface element, movement of a user's hand toward the user interface element (optionally, the user's finger extends toward the user interface element), downward movement of a user's finger (e.g., mimicking a mouse click or a tap on a touchscreen), or other predefined movements of the user's hand. In some implementations, the tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tap gesture movement, which is the finger or hand moving away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by the end of the movement. In some implementations, the end of the movement is detected based on changes in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of finger or hand movement, and / or a reversal of the acceleration direction of finger or hand movement).

[0155] In some implementations, the user's attention is determined to be directed to a portion of the 3D environment based on the detection of a gaze directed to that portion of the 3D environment (optionally, no other conditions are required). In some implementations, the user's attention is determined to be directed to that portion of the 3D environment based on the detection of a gaze directed to that portion of the 3D environment using one or more additional conditions, such as requiring the gaze to be directed to that portion of the 3D environment for at least a threshold duration (e.g., dwell time) and / or requiring the gaze to be directed to that portion of the 3D environment when the user's viewpoint is within a distance threshold from that portion of the 3D environment, so that the device determines that the user's attention is directed to that portion of the 3D environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the 3D environment to which the gaze is directed (e.g., until the one or more additional conditions are met).

[0156] In some implementations, the detection of the readiness configuration of a user or a portion of a user is performed by a computer system. The detection of the hand's readiness configuration is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the readiness of the hand is determined based on whether it has a predetermined hand shape (e.g., a pre-pinch shape with the thumb and one or more fingers extended and spaced apart in preparation for a pinch or grasping gesture, or a pre-tap with one or more fingers extended and the back of the hand facing the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., moving towards an area in front of the user above the user's waist and below the user's head, or moving away from the user's body or legs). In some implementations, the readiness state is used to determine whether an interactive element of the user interface responds to attentional (e.g., gaze) input.

[0157] In scenarios where input is described by reference to air gestures, it should be understood that hardware input devices attached to or held by one or both of the user's hands can be used to detect such gestures. Optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the spatial positioning of the hardware input device, and the positioning and / or movement of the hardware input device can be used in place of the positioning and / or movement of the one or two hands corresponding to the air gesture. Similarly, in scenarios where input is described by reference to air pose, it should be understood that hardware input devices attached to or held by one or both of the user's hands can be used to detect such poses. User input can be detected using controls contained in hardware input devices, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or two hand or finger covers that can detect the positioning or changes in positioning of parts of the hand and / or fingers relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, wherein user input using controls contained in the hardware input device replaces hand and / or finger gestures such as air taps or air pinches in corresponding air gestures. For example, selection input described as performed using air taps or air pinches can alternatively be detected using button presses, taps on touch-sensitive surfaces, presses on pressure-sensitive surfaces, or other hardware inputs. As another example, movement input described as performed using air pinches and drags can alternatively be detected based on interaction with hardware input controls (such as button press and hold, touch on touch-sensitive surfaces, presses on pressure-sensitive surfaces, or other hardware inputs following movement of the hardware input device (e.g., together with the hand associated with the hardware input device) through space). Similarly, two-handed input, which includes the movement of hands relative to each other, can be performed using an air gesture and a hardware input device not in which the air gesture is being performed, two hardware input devices held in different hands, or two air gestures performed by different hands using air gestures and / or inputs detected by one or more of the aforementioned hardware input devices.

[0158] In some embodiments, the software may be downloaded to controller 110 electronically, for example, via a network, or alternatively, it may be provided on a tangible, non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, database 408 is also stored in memory associated with controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware, such as custom or semi-custom integrated circuits or programmable digital signal processors (DSPs). Although in Figure 4The controller 110 is shown, but for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be performed by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or by other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or with any other suitable computerized device (such as a game console or media player). The sensing function of the image sensor 404 may also be integrated into a computer or other computerized device controlled by the sensor output.

[0159] Figure 4 It also includes a schematic diagram of a depth map 410 captured by image sensor 404 according to some embodiments. As described above, the depth map comprises a matrix of pixels with corresponding depth values. Pixel 412 corresponding to hand 406 has been segmented from the background and wrist in the map. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from image sensor 404), where gray shadows become darker as depth increases. Controller 110 processes these depth values ​​to identify and segment components of the image that have human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and frame-to-frame motion from the depth map sequence.

[0160] Figure 4 The controller 110 also schematically illustrates, according to some embodiments, the hand skeleton 414 ultimately extracted from the depth map 410 of the hand 406. Figure 4 In this configuration, the hand skeleton 414 is superimposed on the hand background 416, which has already been segmented from the original depth map. In some embodiments, key feature points of the hand, and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, fingertips, the center of the palm, the end of the hand connecting to the wrist, etc.), are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points across multiple image frames to determine, according to some embodiments, the gesture performed by the hand or the current state of the hand.

[0161] Figure 5 An eye-tracking device 130 is illustrated. Figure 1A Example implementation of ). In some implementations, the eye-tracking device 130 consists of an eye-tracking unit 243 ( Figure 2The eye-tracking device 130 is controlled to track the positioning and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as a head-mounted device, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye-tracking device 130 may optionally be a separate device from the handheld device or XR room. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 may optionally be used in conjunction with a display generation component that is also head-mounted or not head-mounted. In some embodiments, the eye-tracking device 130 is not a head-mounted device and may optionally be used in conjunction with a head-mounted display generation component. In some embodiments, the eye-tracking device 130 is not a head-mounted device and may optionally be part of a non-head-mounted display generation component.

[0162] In some embodiments, the display generation component 120 uses display mechanisms (e.g., a left near-eye display panel and a right near-eye display panel) to display frames including left and right images in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or semi-transparent display on which virtual objects are displayed, allowing the user to view the physical environment directly through the transparent or semi-transparent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects may, for example, be projected onto a physical surface or as holograms, allowing an individual to observe virtual objects superimposed on the physical environment using the system. In this case, separate display panels and image frames for the left and right eyes may not be necessary.

[0163] like Figure 5As shown, in some embodiments, eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera) and an illumination source (e.g., an array or ring of IR or NIR light sources, such as LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be pointed at the user's eye to receive IR or NIR light reflected directly from the eye, or alternatively, it may be pointed at "hot" mirrors located between the user's eye and the display panel, which reflect the IR or NIR light from the eye back to the eye-tracking camera while allowing visible light to pass through. Eye-tracking device 130 may optionally capture images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps), analyze these images to generate gaze tracking information, and transmit the gaze tracking information to controller 110. In some embodiments, the user's two eyes are tracked separately using corresponding eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked using corresponding eye-tracking cameras and illumination sources.

[0164] In some implementations, a device-specific calibration procedure is used to calibrate the eye-tracking device 130 to determine parameters for the eye-tracking device in a specific operating environment 100, such as the 3D geometry and parameters of the LEDs, camera, thermal mirror (if present), eye lenses, and display. The device-specific calibration procedure can be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration procedure can be automated or manual. According to some implementations, a user-specific calibration procedure may include estimations of eye parameters for a specific user, such as pupil position, foveal position, optical axis, visual axis, interocular distance, etc. According to some implementations, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, a flash-assisted method can be used to process the images captured by the eye-tracking camera to determine the current visual axis and the user's gaze point relative to the display.

[0165] like Figure 5As shown, the eye-tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system. The gaze tracking system includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye-tracking camera 540 may be pointed toward a mirror 550 located between the user's eye 592 and a display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, projector, etc.). These mirrors reflect the IR or NIR light from the eye 592 while allowing visible light to pass through. Figure 5 (as shown in the top portion), or alternatively, it can be pointed towards the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion), Figure 5 (As shown in the bottom part).

[0166] In some implementations, controller 110 renders AR or VR frames 562 (e.g., left and right frames for the left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye-tracking camera 540 for various purposes, such as processing frame 562 for display. Controller 110 may optionally estimate the user's gaze point on display 510 based on the gaze tracking input 542 obtained from eye-tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated based on gaze tracking input 542 may optionally be used to determine the direction the user is currently looking.

[0167] The following describes several possible use cases for the user's current gaze direction and is not intended to be limiting. As an example use case, controller 110 can render virtual content differently based on the determined user gaze direction. For example, controller 110 can generate virtual content at a higher resolution in the concave region determined according to the user's current gaze direction than in the peripheral region. Alternatively, the controller can position or move virtual content in the view based at least partially on the user's current gaze direction. Also, the controller can display specific virtual content in the view based at least partially on the user's current gaze direction. As another example use case in AR applications, controller 110 can guide an external camera used to capture the physical environment of an XR experience to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently looking at on display 510. As another example use case, eye lens 520 can be a focusable lens, and the controller uses gaze tracking information to adjust the focus of eye lens 520 so that the virtual object the user is currently looking at has appropriate convergence / divergence to match the convergence of the user's eyes 592. The controller 110 can use gaze tracking information to guide the eye lens 520 to adjust its focus so that the nearby object that the user is looking at appears at the correct distance.

[0168] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lens 520), an eye-tracking camera (e.g., eye-tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each lens in the head-mounted device, such as... Figure 5 As shown. In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.

[0169] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. It should be noted that the positions and angles of the eye-tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at different wavelengths (e.g., 940 nm) may be used on each side of the user's face.

[0170] like Figure 5 The gaze tracking system implementations illustrated herein can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.

[0171] Figure 6 Examples of flash-assisted gaze tracking pipelines according to some embodiments are illustrated. In some embodiments, the gaze tracking pipeline uses a flash-assisted gaze tracking system (e.g., such as...) Figure 1A and Figure 5 The eye-tracking device 130 shown is used to implement this. The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and flash in the current frame. When not in tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in tracking state.

[0172] like Figure 6 As shown, the gaze-tracking camera captures left and right images of the user's left and right eyes. The captured images are then fed into a gaze-tracking pipeline for processing to begin at 610. As indicated by the arrow returning to element 600, the gaze-tracking system can continue capturing images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be fed into the pipeline for processing. However, in some embodiments or under certain conditions, not all captured frames are processed by the pipeline.

[0173] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, the image is analyzed to detect the user's pupil and flash, as indicated at 620. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.

[0174] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flashes in part based on previous information from the previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and flashes detected in the current frame. The processing result at element 640 is checked to verify that the tracking or detection result is credible. For example, the result may be checked to determine whether a sufficient number of pupils and flashes used for gaze estimation were successfully tracked or detected in the current frame. At 650, if the result is not credible, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is credible, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.

[0175] Figure 6 This is intended as an example of an eye-tracking technology that can be used in a particular specific implementation. As will be recognized by those skilled in the art, in a computer system 101 for providing an XR experience to a user, other eye-tracking technologies that are currently available or will be developed in the future may be used to replace or in combination with the flash-assisted eye-tracking technology described herein, depending on the various implementations.

[0176] In some implementations, a portion of the captured real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.

[0177] Therefore, this description describes some embodiments of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and virtual objects. For example, the three-dimensional environment may optionally include a representation of a table existing in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and display of a computer system or passively displayed via a transparent or semi-transparent display of a computer system). As previously described, the three-dimensional environment may optionally be a mixed reality system, wherein the three-dimensional environment is based on a physical environment captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system may optionally be able to selectively display portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system may optionally be able to display virtual objects in the three-dimensional environment to appear as if the virtual objects exist in the real world (e.g., the physical environment) by placing virtual objects in the three-dimensional environment at corresponding locations in the real world that have corresponding positions in the three-dimensional environment. For example, the computer system may optionally display a vase such that the vase appears as if a real vase were placed on top of a table in the physical environment. In some implementations, a corresponding location in the three-dimensional environment has a corresponding location in the physical environment. Therefore, when a computer system is described as displaying a virtual object at a corresponding location relative to a physical object (e.g., such as at or near a user's hand or at or near a physical table), the computer system displays the virtual object at a specific location in the three-dimensional environment such that it appears as if the virtual object were at or near a physical object in the physical environment (e.g., the virtual object is displayed in the three-dimensional environment at a location in the physical environment that would be displayed if the virtual object were a real object at that specific location).

[0178] In some implementations, real-world objects that exist in the physical environment and are displayed in a 3D environment (e.g., and / or visible via a display generation component) can interact with virtual objects that exist only in the 3D environment. For example, the 3D environment may include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment, and the vase is a virtual object.

[0179] In a three-dimensional environment (e.g., a real environment, a virtual environment, or a hybrid environment including both real and virtual objects), an object is sometimes referred to as having depth or simulated depth, or as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension other than height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to a user's position or viewpoint, in which case the depth dimension varies based on the user's position and / or the position and angle of the user's viewpoint. In some embodiments where depth is defined relative to the user's location relative to a surface of the environment (e.g., the surface of the environment's floor or ground), objects further away from the user along lines extending parallel to the surface are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from the user's position and parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user's position is at the center of a cylinder extending from the user's head toward the user's feet). In some embodiments where depth is defined relative to the user's viewpoint (e.g., a direction relative to a point in space that determines which part of the environment is visible via a head-mounted device or other display), objects further away from the user's viewpoint along a line extending parallel to the user's viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from a line extending from and parallel to the user's viewpoint (e.g., defining depth in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere extending outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application displaying application and / or system content), where the user interface container has a height and / or width, and depth is a dimension orthogonal to the height and / or width of the user interface container. In some implementations, when a depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or initially displayed (e.g., such that the container's depth dimension extends outward away from the user or the user's viewpoint), the container's height and / or width are typically orthogonal or substantially orthogonal to a straight line extending from the user's location (e.g., the user's viewpoint or the user's position) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some implementations, when a depth is defined relative to a user interface container, the object's depth relative to the user interface container refers to the object's positioning along the depth dimension of the user interface container. In some implementations, multiple different containers may have different depth dimensions (e.g., different depth dimensions extending away from the user or the user's viewpoint in different directions and / or from different starting points).In some implementations, when depth is defined relative to a user interface container, the orientation of the depth dimension remains constant relative to the user interface container as the position of the user interface container changes, or as the user and / or the user's viewpoint changes (e.g., when multiple different viewers are viewing the same container in a 3D environment, such as during a collaborative session and / or when multiple participants are in a real-time communication session with shared virtual content including the container). In some implementations, for curved containers (e.g., containers including areas with curved surfaces or curved contents), the depth dimension may optionally extend into the surface of the curved container. In some cases, z-interval (e.g., the distance between two objects in the depth dimension), z-height (e.g., the distance of one object from another in the depth dimension), z-position (e.g., the position of an object in the depth dimension), z-depth (e.g., the position of an object in the depth dimension), or simulated z-dimensionality (e.g., depth used as a dimension of an object, a dimension of the environment, an orientation in space, and / or an orientation in simulated space) are used to refer to the concept of depth as described above.

[0180] In some implementations, a user may optionally be able to interact with virtual objects in a 3D environment using one or both hands as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the computer system may optionally capture one or both of the user's hands and display a representation of the user's hands in the 3D environment (e.g., in a manner similar to displaying real-world objects in the 3D environment described above). Alternatively, in some implementations, the user's hands may be visible via the display generation component, through the ability to see the physical environment through the user interface, due to the transparency / semi-transparency of a portion of the user interface being displayed by the display generation component, or due to the projection of the user interface onto a transparent / semi-transparent surface or onto the user's eyes or into the user's field of view. Thus, in some implementations, the user's hands are displayed at corresponding locations in the 3D environment and are treated as if they were objects in the 3D environment that could interact with virtual objects in the 3D environment as if these virtual objects were physical objects in the physical environment. In some implementations, the computer system may update the display of the user's hand representation in the 3D environment in conjunction with the movement of the user's hands in the physical environment.

[0181] In some embodiments described below, the computer system may optionally determine the “effective” distance between a physical object in the physical world and a virtual object in a three-dimensional environment, for example, to determine whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grasping, holding, or within a threshold distance of a virtual object). For example, a hand directly interacting with a virtual object may optionally include one or more of the following: a finger pressing a virtual button, a user’s hand grasping a virtual vase, a user’s hand clasped together to pinch / hold the application’s user interface, and two fingers performing any other type of interaction described herein. For example, the computer system may optionally determine the distance between a user’s hand and a virtual object when determining whether and / or how a user is interacting with a virtual object. In some embodiments, the computer system determines the distance between a user’s hand and a virtual object by determining the distance between the position of a hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, if a user's one or both hands are located at a specific location in the physical world, the computer system may optionally capture the one or both hands and display them at a specific corresponding location in a three-dimensional environment (e.g., the location where the hand would be displayed in the three-dimensional environment if it were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment may be compared with the location of a virtual object of interest in the three-dimensional environment to determine the distance between the user's one or both hands and the virtual object. In some embodiments, the computer system may optionally determine the distance between a physical object and a virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining the distance between a user's one or both hands and a virtual object, the computer system may optionally determine the corresponding location of the virtual object in the physical world (e.g., the location where the virtual object would be located in the physical world if it were a physical object rather than a virtual object), and then determine the distance between the corresponding physical location and the user's one or both hands. In some embodiments, the same technique may optionally be used to determine the distance between any physical object and any virtual object. Therefore, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system may optionally perform any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or map the position of the virtual object to the physical environment.

[0182] In some implementations, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what the physical stylus held by the user is pointing at. For example, if the user's gaze is directed at a specific location in the physical environment, the computer system may optionally determine a corresponding location in the three-dimensional environment (e.g., a virtual location of the gaze), and if a virtual object is located at that corresponding virtual location, the computer system may optionally determine that the user's gaze is directed at that virtual object. Similarly, the computer system may optionally be able to determine the direction in which the stylus is pointing in the physical environment based on the orientation of the physical stylus. In some implementations, based on this determination, the computer system determines a corresponding virtual location in the three-dimensional environment corresponding to the location pointed at by the stylus in the physical environment, and optionally determines that the stylus is pointing at the corresponding virtual location in the three-dimensional environment.

[0183] Similarly, the embodiments described herein may refer to the location of a user (e.g., a user of a computer system) in a three-dimensional environment and / or the location of the computer system in a three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system serves as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to a corresponding location in the three-dimensional environment. For example, the location of the computer system would be its location in the physical environment (and its corresponding location in the three-dimensional environment) such that, if the user stands at that location facing the corresponding portion of the physical environment visible via the display generation component, the user will see from that location objects in the physical environment that are positioned, oriented, and / or sized (e.g., in an absolute sense and / or relative to each other) in the same way as objects displayed or visible in the three-dimensional environment by or via the display generation component of the computer system. Similarly, if the virtual objects displayed in a 3D environment are physical objects in the physical environment (e.g., physical objects placed in the physical environment at the same location as these virtual objects in the 3D environment, and physical objects in the physical environment having the same size and orientation as in the 3D environment), then the position of the computer system and / or the user is the position from which the user will see these virtual objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the 3D environment by the display generation components of the computer system.

[0184] In this disclosure, various input methods are described in relation to interaction with a computer system. When an example is provided using one input device or method, and another example is provided using another input device or method, it should be understood that each example is compatible with and optionally utilizes the input device or method described relative to the other example. Similarly, various output methods are described in relation to interaction with a computer system. When an example is provided using one output device or method, and another example is provided using another output device or method, it should be understood that each example is compatible with and optionally utilizes the output device or method described relative to the other example. Similarly, various methods are described in relation to interaction with a virtual or mixed reality environment via a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with and optionally utilizes the methods described relative to the other example. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each example embodiment.

[0185] User interface and related processes Now turn attention to implementations of user interfaces (“UIs”) and associated processes that can be implemented on computer systems (such as portable multifunction devices or head-mounted devices) having display generation components, one or more input devices, and (optionally) one or more cameras.

[0186] Figures 7A to 7H Examples are shown of how a computer system, according to some implementation schemes, generates virtual lighting effects when rendering content items.

[0187] Figure 7A An example is illustrated where a computer system (e.g., an electronic device) 101 displays a three-dimensional environment 702 from the user's viewpoint (e.g., facing the rear wall of the physical environment in which the computer system 101 is located) via a display generation component (e.g., display generation component 120 of FIG. 1). In some embodiments, the computer system 101 includes a display generation component (e.g., a touchscreen) and multiple image sensors (e.g., ...). Figure 3Image sensor 314). The image sensor may optionally include one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of the user or a portion of the user (e.g., one or both of the user's hands) when the user interacts with the computer system 101. In some embodiments, the computer system 101 is kept in a physical environment by the user 716. In some embodiments, the user interface illustrated and described below may also be implemented on a head-mounted display including display generating components for displaying the user interface or a three-dimensional environment to the user, and sensors for detecting the physical environment and / or movement of the user's hands (e.g., external sensors facing outward from the user) and / or sensors for detecting the user's attention (e.g., gaze) (e.g., internal sensors facing inward toward the user's face).

[0188] In some embodiments, computer system 101 displays a user interface for a content application (e.g., streaming, delivery, playback, browsing, library, sharing, etc.) in a three-dimensional environment 702. In some embodiments, the content application includes a mini-player user interface (also referred to as a simplified user interface) and an extended user interface. In some embodiments, the mini-player user interface includes playback control elements that, in response to user input directed to these playback control elements, cause computer system 101 to modify the playback of content items played via the content application, and illustrations (e.g., album art) associated with the currently played content item. In some embodiments, the extended user interface includes a larger number of user interface elements than the mini-player user interface (e.g., containers such as windows, dials, or back panels; selectable options, content, etc.). In some embodiments, the extended user interface includes navigation elements, content browsing elements, and playback elements. In some embodiments, the mini-player user interface includes virtual lighting effects presented in the three-dimensional environment in an area outside the content application user interface, which are not included in the extended user interface elements. References below. Figures 7A to 7H The mini player user interface is described in more detail with reference to Methods 800, 1000 and 1200 below. Figures 7A to 7H It also includes a top view of the three-dimensional environment 702, including the user 716, the computer system 101, and other objects in the three-dimensional environment 702 (e.g., the mini player user interface 704, the table 712, and the sofa 710).

[0189] exist Figure 7AIn this embodiment, computer system 101 presents a three-dimensional environment 702 that includes representations of virtual objects and real objects. For example, virtual objects include a mini-player user interface 704 for a content application. In some embodiments, the mini-player user interface 704 includes an image (e.g., an album art) associated with a content item currently being played via the content application. As another example, representations of real objects include a representation 706 of the floor and a representation 708 of the walls in the physical environment of computer system 101. In some embodiments, representations of real objects are displayed via display generation component 120 (e.g., virtual pass-through or video pass-through), or as a view of real objects through a transparent portion of display generation component 120 (e.g., real pass-through or optical pass-through). In some embodiments, the physical environment of computer system 101 also includes a table and a sofa, and therefore, computer system 101 displays a representation of table 712 and a representation of sofa 710.

[0190] In some implementations, computer system 101 displays titles and artist instructions 718a for content items, as well as multiple user interface elements 718b-718h overlaid on images included in the mini player user interface 704 for modifying playback of content items. Figure 7B User interface elements 718g and 718h are shown. In some embodiments, in response to detecting input to one of the user interface elements 718b-718h, computer system 101 modifies the playback of the content item currently being played via the content application. In some embodiments, in response to detecting a selection of user interface element 718b, computer system 101 jumps back in the content item playback queue to restart the currently playing content item or play the previous item in the content item playback queue. In some embodiments, in response to detecting a selection of user interface element 718c, computer system 101 plays the content item and updates user interface element 718c to the user interface element that, when selected, causes computer system 101 to pause the playback of the content item. Figure 7A As shown, hand 703 in hand state A optionally indicates selection of user interface element 718c. In some embodiments, in response to detecting selection of user interface element 718d, computer system 101 stops playback of the currently playing content item and initiates playback of the next content item in the content item playback queue. In some embodiments, in response to detecting selection of user interface element 718e, computer system 101 stops displaying the mini player user interface 704 and displays Reference Method 1000 and Figure 11A The extended user interface is described in more detail in Figure 11O. In some implementations, in response to detecting a selection of user interface element 718f, the computer system 101 displays time-synchronized lyrics for the currently playing content item, such as... Figure 7D As illustrated. In some implementations, in response to detecting a selection of user interface element 718g, computer system 101 renders and / or updates virtual lighting effects associated with the content item currently being played on computer system 101, as shown in the reference. Figures 7B to 7H Described in more detail. In some implementations, in response to detecting a selection of user interface element 718h, computer system 101 presents another user interface element for adjusting the playback volume of the audio content of the content item and / or presents a menu for modifying the audio output options for the playback of the audio content.

[0191] In some embodiments, computer system 101 detects selection of one of user interface elements 718b-h by detecting indirect selection input, direct selection input, air gesture selection input, or input device selection input. In some embodiments, detecting selection input includes first detecting a ready state corresponding to the type of selection input being detected (e.g., detecting an indirect ready state before detecting indirect selection input, and a direct ready state before detecting direct selection input). In some embodiments, detecting indirect selection input includes detecting a user's gaze pointing at the corresponding user interface element via input device 314, while simultaneously detecting a selection gesture made by the user's hand, such as an air pinch gesture where the user touches their thumb with another finger of their hand. In some embodiments, detecting direct selection input includes detecting a selection gesture made by the user's hand via input device 314, such as a pinch gesture within a predefined threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, or 30 cm) of the location of the corresponding user interface element, or a pressing gesture where the user's hand "presses" onto the location of the corresponding user interface element while forming a hand shape. In some embodiments, detecting air gesture input includes detecting a user's gaze pointing at the corresponding user interface element while detecting a pressing gesture at the location of an air gesture user interface element displayed in the three-dimensional environment 702 via display generation component 120. In some embodiments, detecting input device selection includes detecting manipulation of a mechanical input device (e.g., stylus, mouse, keyboard, touchpad, etc.) in a predefined manner corresponding to the selection of the user interface element when a cursor controlled by the input device is associated with the location of the corresponding user interface element and / or when the user's gaze is pointing at the corresponding user interface element.

[0192] like Figure 7A As shown, computer system 101 detects the selection of selectable option 718c, which, when selected, causes computer system 101 to play the corresponding content item, such as... Figure 7BAs shown. Selecting optional option 718c also causes computer system 101 to update user interface element 718c to a user interface element that, when selected, causes computer system 101 to pause playback of the content item. While selecting optional option 718g displays virtual lighting effects associated with the content item, selecting optional option 718c optionally displays virtual lighting effects associated with the content item and the playback of the corresponding content item. For example... Figure 7B As shown, the virtual lighting effect includes simulated light 720 emitted from the mini player user interface 704 (by... Figure 7B (Represented by dashed lines in the text). The simulated light 720 includes one or more characteristics as described in reference method 800. For example, the simulated light 720 may optionally be two-dimensional and emanate from the back of the mini-player user interface 704. In some embodiments, the simulated light 720 extends 1 cm to 10 cm away from the mini-player user interface 704. In some embodiments, as described in method 800, the simulated light 720 is synchronized with the beat, volume, bass intensity, and / or motion in the video of the content item. In some embodiments, the distance the simulated light 720 extends from the mini-player user interface 704 is determined by the aforementioned factors. For example, during a beat drop or more intense portion of the playback of the corresponding content item, the distance the simulated light 720 extends from the mini-player user interface 704 may be greater than the distance during a quieter, less intense portion of the playback of the corresponding content item. In some embodiments, and as described in method 800, the simulated light 720 is displayed in a color determined by a mood score associated with the content item and further described with reference to method 800.

[0193] Figure 7B This includes a hand 703 in hand state B, which corresponds to the hand shape, pose, position, etc., associated with a ready state or input. The hand 703 is used to select the mini-player user interface 704 using selection input, as discussed above, and then to move the mini-player user interface 704 from a first position to a second position using drag input, such as... Figures 7B to 7C As shown. The update location for the mini player user interface 704 is in... Figure 7CAs shown in the diagram. For example, hand 703 can perform air drag and air release movements and / or air pinch, air drag and air release movements on mini player user interface 704 to move mini player user interface 704 from a first position in a three-dimensional environment to a second position. In some embodiments, the second position may optionally be in front of, behind, to the right of, or to the left of the first position. In some embodiments, moving mini player user interface 704 does not affect the playback of corresponding content items or visual lighting effects (e.g., simulated light 720) and the display of lyrics 722, as discussed in further detail below. Alternatively, in response to the movement of mini player user interface 704, lyrics 722 and / or visual lighting effects may optionally be stopped from displaying.

[0194] In some implementations, while the simulated light 720 is displayed and the content is playing, the hand 703 selects an optional option 718f, such as... Figure 7C As shown. In some implementations, computer system 101 detects a selection of selectable option 718f, which, when selected, causes computer system 101 to display time-synchronized lyrics associated with the content item, such as... Figure 7D As shown. For example, computer system 101 detects indirect selection of option 718f, including detecting a selection gesture made by the user's hand 703 (e.g., Figure 7A The hand state shown (A) and / or the indirect selection of option 718f are detected when the user's gaze is directed at option 718f.

[0195] In some implementations, in response to detecting a selection of lyrics option 718f, computer system 101 updates the 3D environment 702 to include time-synchronized lyrics 722 associated with the content item, such as... Figure 7D As shown. In some implementations and referenced Figure 7D The time-synchronized lyrics 722 are lyrics associated with a content item currently being played via a content application on computer system 101. Computer system 101 may optionally present a portion of the lyrics 722 corresponding to the currently playing portion of the content item on computer system 101 via the content application, and update that portion of the lyrics 722 as the content item continues to play. In some embodiments, the lyrics 722 include a line of lyrics corresponding to the currently playing portion of the content item, one or more lines of lyrics corresponding to a portion of the content item preceding the currently playing portion, and / or one or more lines of lyrics corresponding to a portion of the content item that will play after the currently playing portion. Figure 7DAs shown, lyrics 722 are displayed in the three-dimensional environment 702 outside the boundary of the mini player user interface 704 and / or at a location different from the mini player user interface. In some embodiments, lyrics 722 are displayed adjacent to the mini player user interface 704 (e.g., within 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, 20 cm, 30 cm, or 50 cm, or 1 m, 2 m, 3 m, or 5 m of the mini player user interface). In some embodiments, lyrics 722 are displayed to the left or right of the mini player user interface 704. In some embodiments, lyrics 722 are displayed at the same z-depth (e.g., distance) from the user's viewpoint as the mini player user interface 704 is from the user. In some embodiments, lyrics 722 are in front of or behind the mini player user interface 704 (e.g., closer to or further away from the user's viewpoint than the mini player user interface 704), as described with reference to method 800. In some implementations, lyrics 722 are initially displayed in a three-dimensional environment 702 relative to the mini-player user interface 704 in a predetermined spatial arrangement (e.g., positioning and / or orientation).

[0196] Figure 7D1 Example: User 716's hand 703 is in hand state A. Hand 703 is selecting selectable option 718c to pause playback of a content item. In response to the selection of option 718c, computer system 101 updates selectable option 718c to display a "Play" button, as shown. Figure 7D2 As shown, selecting the play button will resume playback of the content item. In some implementations and as... Figure 7D2 As shown, selecting option 718c to pause playback of the corresponding content item also includes stopping the display of simulated lighting effects. Stopping the display of virtual lighting effects also includes reducing the visual salience of the virtual lighting effects. For example, the virtual lighting effects can optionally become dimmer and more transparent.

[0197] In some implementations, the hand 703 selects an optional option 718g that updates and / or alters the display of virtual lighting effects. For example, selecting the optional option 718g can change the virtual lighting effects from... Figure 7A and Figure 7B The simulated light ray 720 shown is changed to a simulated three-dimensional particle effect 724, as shown. Figure 7E As shown. Figure 7E The illustrated 3D particle effect 724 has a visual appearance similar to fireworks or circular projectiles and / or particles emitted from the sides and / or rear of the mini player user interface 704 toward the viewpoint of the user 716, as discussed in reference method 800 and as... Figure 7EAs shown. In some embodiments, the 3D particle effect 724 emits virtual / simulated light into the 3D environment 702. For example, the 3D particle effect 724 may optionally illuminate / emit light onto objects in the 3D environment (such as table 712 and / or sofa 710). In some embodiments, similar to the simulated light 720, the 3D particle effect 724 includes colors and animations corresponding to the respective content item. For example, mood scores and / or images associated with the respective content item, further discussed in method 800, affect the color of the 3D particle effect 724. Additionally, the beat, volume, bass intensity, and / or motion in the video of the content item may optionally affect the display of the 3D particle effect 724. For example, the 3D particle effect 724 may optionally be larger during bass drops, louder, and / or more intense portions of the playback of the respective content item compared to less intense portions.

[0198] Figure 7F and Figure 7G An example is illustrated of how computer system 101 updates the 3D environment 702 and 3D particle effects 724 in response to the detection of movement of computer system 101, which causes computer system 101 to update the user's viewpoint in the 3D environment 702 and the field of view of computer system 101. For example, user 716 moves to different locations around the physical environment. Therefore, in some embodiments, computer system 101 updates the viewpoint of mini-player user interface 704 in the 3D environment 702 (e.g., the angle from which mini-player user interface 704 is displayed) in response to the user's updated viewpoint. For example, mini-player user interface 704 is shown at different angles to account for the updated viewpoint. However, mini-player user interface 704 may optionally remain in the same position in the 3D environment and will not update its position in response to movement of computer system 101. In some embodiments, updating the field of view of computer system 101 and the user's viewpoint also causes computer system 101 to display representations of table 712 and sofa 710 in the physical environment of computer system 101 from different perspectives corresponding to the user's updated viewpoint.

[0199] In some implementation schemes, and as such Figure 7F As shown, the 3D particle effect 724 follows the viewpoint of user 716. In some embodiments, the 3D particle effect 724 changes direction so that it moves toward the updated viewpoint of user 716, as described in reference method 800. For example, computer system 101 changes the direction and / or speed at which the particle effect 724 moves from the mini-player user interface 704 toward the new position of user 716's viewpoint. Figure 7F It shows that the viewpoint of user 716 moves to the left, and in response to the viewpoint movement, the 3D particle effect 724 is updated to continue pointing and / or moving toward the viewpoint of user 716.

[0200] Alternatively, and in some implementations, the three-dimensional particle effect 724 is displayed independently of the user's 716 viewpoint. For example... Figure 7G As shown, the 3D particle effect 724 is window-centric and unaffected by changes in the user's 716 viewpoint (e.g., Figure 7G The viewpoint of user 716 can be optionally connected to Figure 7F (The viewpoint is the same for user 716). Responding to the user's viewpoint from... Figure 7E Change to Figure 7G One or more properties of the 3D particle effect 724 (e.g., color, brightness, saturation, size and / or orientation) remain unchanged.

[0201] In some implementations, virtual lighting effects (e.g., simulated light 720 and / or 3D particle effects 724) include virtual light spillover 726 displayed on surfaces in a 3D environment 702. For example, virtual light spillover is displayed on surfaces such as a table 712, a sofa 710, and a wall. Figure 7H As shown. In some embodiments, the virtual light spill simulates light emitted from the mini-player user interface 704, including colors and / or virtual lighting effects corresponding to the image associated with the currently playing content item. In some embodiments, the virtual light spill is animated in a manner corresponding to the virtual lighting effects, such as the beat of the audio content of the content item currently being played via the content application (e.g., blinking, changing intensity, and / or color). In some embodiments, both the simulated light 720 and the 3D particle effect 724 include virtual light spill. In some embodiments, if the virtual lighting effect is far from a surface threshold distance (e.g., 1 meter, 2 meters, 3 meters, 5 meters, or 10 meters), the virtual light spill 726 is not displayed on that surface. Figure 7H As shown, in some embodiments, computer system 101 displays virtual light spillover 726 on a representation of a real surface in 3D environment 702. In some embodiments, computer system 101 also displays virtual light spillover 726 on virtual objects (e.g., user interfaces of other applications, user representations, etc.) in 3D environment 702. In some embodiments, displaying virtual lighting effects includes darkening and / or blurring portions of 3D environment 702 that do not include the mini-player user interface 704 and / or virtual lighting effects, such as... Figure 7H As shown. In some implementations, and as... Figure 7HAs shown, in response to a dimly lit environment, a virtual light overflow 726 is displayed on a surface in the three-dimensional environment 702. In some embodiments, the virtual light overflow 726 appears on the surface of the three-dimensional environment 702 when the dimness of the environment reaches a threshold (e.g., below a threshold brightness, such as 1 lumen, 10 lumen, 25 lumen, 100 lumen, 500 lumen, or 1000 lumen). In some embodiments, the virtual light overflow 726 stops displaying when playback of a content item is stopped and / or paused. For example, input (such as hand 703 in hand state A) pauses playback of a content item by selecting option 718c, and as a result of the input, the virtual light overflow 726 stops displaying.

[0202] about Figures 7A to 7H Additional or alternative details of the illustrated implementation schemes are provided below for reference. Figure 8 The method described in 800 is as follows.

[0203] Figure 8 A flowchart illustrating an exemplary method 800 of how a computer system, according to some embodiments, generates virtual lighting effects when presenting content items is described. In some embodiments, method 800 is performed at a computer system (e.g., computer system 101 in FIG. 1, such as a tablet computer, smartphone, wearable computer, or head-mounted device), which includes display generation components (e.g., FIG. 1, ...). Figure 3 and Figure 4 The display generating component 120 (e.g., a head-up display, a monitor, a touchscreen, and / or a projector) and one or more cameras (e.g., one or more cameras pointing forward from or downward from the user's head toward the user's hand, such as color sensors, infrared sensors, and other depth-sensing cameras). In some embodiments, method 800 is performed by one or more processors of a computer system (such as one or more processing units 202 of computer system 101, e.g., ...) stored in a non-transitory computer-readable storage medium and processed by one or more processors of a computer system (e.g., one or more processing units 202 of computer system 101, ...). Figure 1A The controller 110 in the middle manages the instructions executed. Some operations in method 800 may be combined, and / or the order of some operations may be changed.

[0204] In some embodiments, method 800 is performed at a computer system communicating with a display generating component and one or more input devices. These include, for example, mobile devices (e.g., tablets, smartphones, media players, or wearable devices) or computers or other electronic devices. In some embodiments, the display generating component is a display integrated with an electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally integrated or external) used to project a user interface or make the user interface visible to one or more users. In some embodiments, the one or more input devices include the ability to receive user input (e.g., capture user input, detect user input, etc.) and send information associated with that user input to the computer system. Examples of input devices include touchscreens, mice (e.g., external), touchpads (optionally integrated or external), remote control devices (e.g., external), another mobile device (e.g., separate from the computer system), handheld devices (e.g., external), controllers (e.g., external), cameras, depth sensors, eye-tracking devices, and / or motion sensors (e.g., hand-tracking devices, hand motion sensors), etc. In some implementations, the computer system communicates with a hand-tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreens, touchpads)). In some implementations, the hand-tracking device is a wearable device, such as a smart glove. In some implementations, the hand-tracking device is a handheld input device, such as a remote control or stylus.

[0205] In some implementations, the computer system (e.g., Figure 7A The computer system 101 displays (802a) the user interface of the application in a three-dimensional environment via a display generation component, such as Figure 7B The mini-player user interface 704 (e.g., the user interface may optionally be the content player user interface of a content playback application (such as a music player, video player, and / or podcast player application)). In some embodiments, the user interface is the user interface of a mini-player of the content playback application. The content playback application may optionally be associated with an extended user interface and a mini user interface (mini-player). In some embodiments, the mini-player is configured to display a simplified user interface of the content playback application. In some embodiments, the extended user interface may optionally not be displayed when the mini-player is displayed. Additionally, the mini-player may optionally not be displayed when the extended user interface is displayed. In some embodiments, the three-dimensional environment is an extended reality (XR) environment, such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment, and a virtual content container is displayed within the three-dimensional environment.

[0206] In some implementations, the user interface is associated with the playback of corresponding content (such as...). Figure 7B The playback of the content items discussed herein is associated with (for example, the user interface may be a user interface where the computer system detects input from the user of the computer system for controlling the playback of the corresponding content (such as initiating, pausing, and / or skipping the playback of the corresponding content). In some embodiments, the user interface includes one or more selectable controls that can be selected to play or pause the corresponding content, skip forward or backward the corresponding content, or display lyrics simultaneously with a user interface object for the corresponding content. The user interface has one or more of the characteristics of a user interface associated with the playback of the corresponding content described in reference method 1200.

[0207] In some implementations, the relevant content is not currently being played (e.g., in the first instance, the relevant content is paused or the computer system has not yet received input to play the relevant content), and the user interface (such as...) Figure 7A The mini player user interface 704 shown is not displayed in a three-dimensional environment with corresponding simulated lighting effects (e.g., the user interface is displayed separately in the three-dimensional environment without virtual lighting effects, as will be described later).

[0208] In some implementations, when the application's user interface is displayed in a 3D environment and the relevant content is not currently playing and the application's user interface is not displayed with the corresponding simulated lighting effects, the computer system receives a first input corresponding to a request to initiate playback of the relevant content via one or more input devices, such as from... Figure 7A The first input (802b) is input from hand 703 in hand state A. In some embodiments, and discussed in more detail below, the first input includes user interaction with a selectable option included in and / or displayed simultaneously with the user interface object, such as an air pinch gesture of the user's hand from a computer system, including a pinch pointing to a selectable option (e.g., when the user's attention is on the selectable option, e.g., the user's thumb and forefinger are pressed together and touched), a drag (e.g., movement of the hand while the user's hand is in a pinched hand shape), and / or a release (e.g., the user's hand releases the pinch to remove the thumb and forefinger). In other embodiments, the first input includes user interaction as described above with a mouse, touchpad, and / or touchscreen.

[0209] In some implementations, in response to receiving the first input (802c), the computer system initiates (802d) playback of the corresponding content in the three-dimensional environment, such as Figure 7BAs shown. In some implementations, playback of the corresponding content includes playback of audio (e.g., music), one or more videos, one or more podcasts, and / or one or more audiobooks.

[0210] In some implementations, the computer system simulates lighting effects (such as...) in a three-dimensional environment. Figure 7B Simulated light 720 and / or Figure 7E The illustrated 3D particle effect 724) displays a user interface (802e) where one or more characteristics of the corresponding simulated lighting effect are based on playback of the corresponding content. In some embodiments, the characteristics of the corresponding simulated lighting effect include the color of the lighting effect, the movement of the lighting effect, the animation of the lighting effect, the size and / or amplitude of the lighting effect, and / or the brightness of the lighting effect. In some embodiments, displaying the user interface with the corresponding simulated lighting effect also includes displaying the above characteristics of the lighting effect in sync with the beat, music volume, bass intensity, and / or motion in audio playback and / or video playback (e.g., movement and / or color and / or brightness changes). For example, the simulated lighting effect may optionally be shown around the outline or boundary of the user interface. In some embodiments, the characteristics of the corresponding simulated lighting effect are determined by the artist of the corresponding content. For example, the artist may optionally select the hue, saturation, and / or brightness corresponding to the corresponding content. Displaying the user interface with simulated lighting effects based on playback of the corresponding content reduces the resources required to display the simulated lighting effect when the corresponding content is not playing, and reduces the need for manual input to manually enable and / or disable the simulated lighting effect.

[0211] In some implementations, when the user interface is displayed at a first location in the three-dimensional environment (e.g., and when a corresponding simulated lighting effect is displayed in the three-dimensional environment via a display generation component), the computer system receives a second input corresponding to a request to move the user interface in the three-dimensional environment via one or more input devices, such as... Figure 7B The illustration shows input made on the mini player user interface 704 via hand 706b. For example, the second input includes user interaction with a selectable option included in and / or displayed simultaneously with the user interface object, and / or interaction with the user interface itself, such as air pinch gestures of the user's hand from the computer system, including pointing to a selectable option (e.g., when the user's attention is directed to a selectable option) and / or pinching the user interface (e.g., the user's thumb and forefinger touching together), dragging (e.g., movement of the hand while the user's hand is in a pinched shape), and / or releasing (e.g., the user's hand releasing the pinch to remove the thumb and forefinger). In some embodiments, the second input includes input from a mouse, touchpad, and / or touchscreen.

[0212] In some implementations, in response to receiving a second input, the computer system moves the user interface from a first position to a second position in the three-dimensional environment, such as the mini player user interface 704. Figure 7B and Figure 7C The movement between these positions is illustrated. For example, moving the user interface from a first position to a second position also includes moving the corresponding simulated lighting effects in the 3D environment from the first position to the second position. In some embodiments, the user interface is a free-floating entity, where the movement of the user interface does not affect other objects in the 3D environment. Updating the position of the user interface in response to input corresponding to a request for the user interface of a mobile application allows for efficient access to the user interface (thus reducing the resources required to display the user interface) and reduces the likelihood of erroneous input to the user interface.

[0213] In some implementations, such as in Figure 7A If the mini player user interface 704 includes an image corresponding to the corresponding content, then the application's user interface includes the image corresponding to that content. For example, the image includes an album cover of a music album (e.g., the corresponding content is a song included in the music album, and the user interface includes an image of that album). In some embodiments, the color of the corresponding simulated lighting effect is based on the color of the image. Displaying a user interface with an image corresponding to the corresponding content provides feedback about the corresponding content being played, thereby reducing the amount of input required to retrieve additional information related to the corresponding content.

[0214] In some implementations, when displaying the application's user interface without showing a representation of lyrics corresponding to the corresponding content in the three-dimensional environment, the computer system receives input via one or more input devices corresponding to a request to present lyrics corresponding to the relevant content, such as... Figure 7BThe input to option 718f is via hand 703. In some embodiments, the input includes selection of selectable options displayed on the user interface. In some embodiments, selectable options are not displayed unless and until the pose of the corresponding part of the user's hand meets one or more criteria. In some embodiments, the pose of the user's hand meets one or more criteria when the user's hand is within the field of view of the hand tracking device communicating with the computer system. In some embodiments, the pose of the user's hand meets one or more criteria when the user's hand is within a predetermined area of ​​the three-dimensional environment (such as being raised relative to the rest of the user's body (e.g., a raised threshold amount)). In some implementations, when a user's hand is in a pose corresponding to a ready state of the computer system, the user's hand pose satisfies one or more criteria. This ready state corresponds to the initiation of input provided by the user's hand, such as pointing to a hand shape (e.g., one or more fingers extended and one or more fingers curled into the palm) or pre-pinch a hand shape (e.g., the thumb being within a predetermined threshold distance (e.g., 0.1 cm, 0.2 cm, 0.3 cm, 0.5 cm, 1 cm, 2 cm, 3 cm, etc.) of another finger of the hand without touching that finger's hand shape). In some implementations, when the hand tracking device does not detect the hand, the hand pose does not satisfy one or more criteria.

[0215] In some implementations, in response to receiving input corresponding to a request to present lyrics, the computer system simultaneously displays the application's user interface and a representation of the lyrics corresponding to the relevant content in a three-dimensional environment, wherein the lyrics (e.g., Figure 7C The lyrics 722 are displayed in a three-dimensional environment at a location with a predefined spatial arrangement relative to the application's user interface, such as lyrics 722 in... Figure 7CThe lyrics representation is shown near the user interface 704 of the mini player. In some embodiments, the lyrics representation is positioned outside the boundaries of the application's user interface. The lyrics representation may optionally be positioned to the left, right, top, or bottom of the application's user interface. In some embodiments, the lyrics representation is displayed at the same or different distance from the user's viewpoint than the user interface. In some embodiments, in response to input, the lyrics representation is initially displayed relative to the user interface in a predetermined spatial arrangement (e.g., above, below, to the left, or to the right of the user interface, optionally at the same or different distance from the user's viewpoint than the user interface). In some embodiments, the spatial relationship between the user interface and the lyrics representation changes in response to user input pointing to the lyrics representation and / or the user interface. In some embodiments, the lyrics representation includes a time-synchronized representation of the lyrics corresponding to the currently playing portion of the corresponding content. In some embodiments, as the corresponding content continues to play, the lyrics representation is updated to include a representation of the lyrics corresponding to the currently playing portion of the corresponding content. In some embodiments, the lyrics representation has one or more characteristics of the lyrics described with reference to method 1200. Compared to a user interface that displays lyrics in a predefined spatial relationship corresponding to the content, this provides an efficient way to view lyrics and ensures that the lyrics are visible without requiring additional input from the user, thus improving user-device interaction.

[0216] In some implementations, the corresponding simulated lighting effect includes one or more simulated rays emitted from the application's user interface, such as... Figures 7B to 7D The simulated light 720 is shown. In some embodiments, the simulated light is two-dimensional, rather than volumetric. In some embodiments, the simulated light is volumetric. In some embodiments, the generation of the corresponding simulated lighting effect is synchronized with the beat or transition of the corresponding content. In some embodiments, the corresponding simulated light projects virtual light onto the user interface and / or other parts of the three-dimensional environment outside the user interface. In some embodiments, the virtual projected light has various properties corresponding to the characteristics of the simulated light. For example, the virtual projected light has the same color characteristics as the simulated light. Displaying simulated light emanating from the application's user interface provides feedback about the source of the content being played, thereby reducing the possibility of erroneous interaction with the device.

[0217] In some implementations, such as in Figure 7DIn cases where the simulated light 720 is a color (e.g., red, blue, or magenta), displaying one or more simulated lights also includes displaying the color of one or more simulated lights determined by an emotion score. In some embodiments, the emotion score is predefined in a database accessible to the computer system. Emotions may optionally include calm, excitement, positive, or negative. In some embodiments, emotions are combinations of the above. In some embodiments, each emotion determined by the emotion score is associated with a different color. For example, a calm, positive emotion is associated with blue, while an aggressive emotion is associated with red. Alternatively, and in some embodiments, the color of the simulated light depends on the image corresponding to the relevant content. For example, an image with a magenta background results in a magenta color for the simulated light, and an image with a red background results in a red color for the simulated light. In some embodiments, the color of the simulated light is enhanced if the image is not prominent enough. In some embodiments, sufficient prominence is based on color saturation and brightness. For example, magenta may optionally become red because magenta has low brightness and saturation. In some embodiments, a color assignment value is given to the image based on color prominence and compared to a predetermined threshold to determine whether to use the image color or the emotion color. For example, a gray or white image color will not meet a predetermined threshold. For instance, a gray image color has a value of 12, and the threshold is 50. In some embodiments, the predetermined threshold is 25, 50, 100, or 1000. In some embodiments, environmental characteristics determine the brightness and / or final color of the simulated light. In some embodiments, environmental characteristics include the brightness and color of the three-dimensional environment. For example, in a dim environment, the simulated light may optionally be dimmer than in a bright environment. Displaying simulated light in a color determined by a mood score allows the computer system to better align simulated lighting effects with the content, thereby reducing the need for additional input from the user for changing and / or correcting lighting effects.

[0218] In some implementation methods, such as in emotion scores and Figure 7A In the context of the content items discussed, a mood score is associated with the corresponding content. For example, the mood score may optionally be predefined in a database for each corresponding content item, and may optionally be different for different content items. In some implementations, the owner or creator of the corresponding content defines or updates the mood score. Associating mood scores with corresponding content enables an efficient way to determine the color of corresponding simulated lighting effects without user input, and also provides clear feedback on the different content items currently playing on the device, thereby reducing the resources required to determine the color.

[0219] In some implementations, when the user interface is displayed in a three-dimensional environment with corresponding simulated lighting effects and while playback of the corresponding content is in progress, the computer system receives input corresponding to a request to stop the playback of the corresponding content via one or more input devices, such as... Figure 7D1 The input to option 718c is made via hand 703. In some embodiments, the input includes selection of selectable options displayed on the user interface, as described above. In some embodiments, selectable options are not displayed unless and until the pose of the corresponding part of the user meets one or more of the criteria described above. A request to stop playback of the corresponding content may optionally include a request to pause playback of the corresponding content or a request to close the user interface and stop playback and / or display of the corresponding content.

[0220] In some implementations, in response to receiving input corresponding to a request to stop playback of the relevant content, the computer system stops playback of the relevant content and reduces the visual salience of the corresponding simulated lighting effect displayed along with the application's user interface, such as... Figure 7D2 The mini player user interface 704 is no longer displayed as simulated light 720. In some embodiments, visual salience includes brightness, size, color saturation, and / or transparency. In some embodiments, reducing visual salience includes dimming the simulated lighting effect. Dimming the simulated lighting effect may optionally include making the simulated lighting effect transparent. In some embodiments, reducing visual salience includes making the simulated lighting effect smaller in a three-dimensional environment. For example, the simulated lighting effect extends 3 cm away from the user interface. Reducing visual salience may optionally include shrinking the simulated lighting effect to extend 0-1 cm away from the user interface. In some embodiments, reducing visual salience includes changing the color value of the simulated lighting effect below a threshold, as described above. For example, a sufficiently prominent color (e.g., magenta) becomes a less prominent color (e.g., red) to reduce visual salience. In some embodiments, reducing visual salience includes reducing the saturation of the simulated lighting effect and / or increasing the transparency of the simulated lighting effect. In some embodiments, reducing the visual salience of a corresponding lighting effect includes dissipating and / or increasing the transparency of the corresponding lighting effect (e.g., not showing the corresponding lighting effect). By reducing the visual salience of the corresponding simulated lighting effects due to the request to stop playback of the corresponding content, visual interference to the user is reduced before the user provides another input, and clutter in the 3D environment is reduced, thereby reducing interaction errors with the computer system.

[0221] In some implementations, displaying a user interface with corresponding simulated lighting effects in a three-dimensional environment includes displaying the corresponding simulated lighting effects as including simulated three-dimensional particle effects, such as... Figure 7EThe simulated 3D particle effect 724. In some embodiments, the simulated 3D particle effect looks like fireworks or circular projectiles. In some embodiments, the simulated 3D particle effect is displayed on and / or emitted from the sides of the user interface. In some embodiments, the simulated 3D particle effect is displayed as bursting out and / or emanating from the center of the user interface. Displaying simulated 3D particle effects allows the computer system to provide better feedback on the source of content playback, thereby reducing the possibility of erroneous interaction with the computer system.

[0222] In some implementations, the computer system displays simulated 3D particle effects, including animations showing the simulated 3D particle effects moving from the user interface's position in the 3D environment toward the user's viewpoint in the computer system, such as... Figure 7E A 3D particle effect 724 surrounds the user 716. For example, simulated 3D particles are flying around and / or toward the user's viewpoint. In some embodiments, the simulated 3D particle effect encompasses the user's viewpoint, such that the effect forms and / or exists within a sphere surrounding the user's viewpoint. In some embodiments, the computer system detects a new position of the user's viewpoint in the 3D environment and / or movement relative to the user interface, and changes the animation of the particle effect in response to this movement. For example, the computer system changes the direction and / or velocity of the particles moving from the user interface toward the new position of the user's viewpoint. In some embodiments, the 3D particle effect includes simulated particles that, as part of the animation of these 3D particle effects, change their distance from the user's viewpoint (e.g., becoming closer to the user's viewpoint as part of the animation of these 3D particle effects), while the distance of the user interface relative to the user's viewpoint remains unchanged (optionally, there is no user input for changing this distance). Guiding the simulated 3D particle effect toward the position of the user's viewpoint in the computer system allows the computer system to improve the directional feedback of the source of the content playback.

[0223] In some implementations, such as in 3D particle effects 724, Figures 7E to 7G In cases where coloring is determined by the emotion score, the emotion score is used to determine the colors of simulated 3D particle effects. The emotion score has been described above. Displaying simulated light in colors determined by the emotion score allows the computer system to better align simulated lighting effects with the content, thereby reducing the need for additional input from the user for changing and / or correcting lighting effects.

[0224] In some implementations, the computer system displays corresponding simulated lighting effects independently of the user's viewpoint position in the three-dimensional environment, such as through three-dimensional particle effects 724 even when... Figure 7GThe example illustrates how the simulated lighting effect continues to fly forward even when the user's viewpoint has changed. For instance, the simulated lighting effect is window-centric, and the user's viewpoint is independent of the object or simulated lighting effect's location. For example, the brightness, saturation, size, and / or direction of the simulated lighting effect are unaffected by the viewpoint's location. Displaying the simulated lighting effect independently of the user's viewpoint in the 3D environment increases the consistency of feedback presentation in the 3D environment, thereby reducing errors and improving user-device interaction.

[0225] In some implementations, when the user interface is displayed in a three-dimensional environment with corresponding simulated lighting effects, and when the user's viewpoint on the computer system has a first spatial arrangement (e.g., positioning and / or orientation) relative to the user interface, such as Figure 7E As shown, the computer system receives the user's viewpoint via one or more input devices for movement in a three-dimensional environment relative to the user interface in a second spatial arrangement different from the first spatial arrangement (such as through the user 716 in...). Figure 7E and Figure 7F The input corresponds to the movement between (as shown). In some embodiments, the user's viewpoint moves according to the user's movement and / or the movement of the computer system and / or display generation components in the physical environment of the computer system and / or display generation components (e.g., the direction and / or magnitude of the change in viewpoint corresponds to the direction and / or magnitude of the user's movement in the physical environment). In some embodiments, when the user's viewpoint is at a first position, the user has a first spatial arrangement relative to the viewpoint of the user interface.

[0226] In some implementations, in response to receiving input corresponding to a movement of the user's viewpoint in the computer system, the computer system updates the display of the 3D environment to reflect the user's moving viewpoint (e.g., a moving viewpoint at a second location having a second spatial arrangement relative to the user interface in the 3D environment), and continues to display the application's user interface with corresponding simulated lighting effects, such as through... Figure 7FThe three-dimensional particle effect 724 flies toward the updated viewpoint of the user 716. In some embodiments, as the user's viewpoint moves around the user interface, the computer system displays different views of the corresponding lighting effect. For example, in a first spatial arrangement, the first viewpoint is on a first side of the user interface (e.g., in front). For example, in a second spatial arrangement, the moving viewpoint is on a second, different side of the user interface (e.g., opposite or behind the first viewpoint). In some embodiments, the orientation difference between the first viewpoint and the moving viewpoint relative to the user interface is 1 degree, 5 degrees, 10 degrees, 30 degrees, 45 degrees, 90 degrees, 180 degrees, 270 degrees, or 359 degrees. In some embodiments, when the computer system displays the corresponding simulated lighting effect from different perspectives as the user's viewpoint changes, one or more of the characteristics of the corresponding simulated lighting effect remain the same in the three-dimensional environment. For example, the color of the lighting effect, the movement of the lighting effect, the animation of the lighting effect, the size and / or amplitude of the lighting effect, and / or the brightness of the lighting effect remain unchanged. Allowing users to view different spatial arrangements in a 3D environment ensures consistent presentation of feedback within the 3D environment, thereby reducing usage errors and improving user-device interaction.

[0227] Figures 9A to 9E Examples are shown of how a computer system, according to some implementation schemes, generates animated 3D objects when rendering content items.

[0228] Figure 9A An example is illustrated where a computer system (e.g., an electronic device) 101 displays a three-dimensional environment 902 from the viewpoint of a user 916 of the computer system 101 (e.g., facing the rear wall of the physical environment in which the computer system 101 is located) via a display generation component (e.g., display generation component 120 of FIG. 1). In some embodiments, the computer system 101 includes a display generation component (e.g., a touchscreen) and multiple image sensors (e.g., ...). Figure 3 Image sensor 314. The image sensor may optionally include one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of the user or a portion of the user (e.g., one or both of the user's hands) when the user interacts with the computer system 101. In some embodiments, the computer system 101 is held in a physical environment by the user 916. In some embodiments, the user interface illustrated and described below may also be implemented on a head-mounted display including display generation components that display the user interface or three-dimensional environment to the user, and sensors (e.g., external sensors facing outwards from the user) for detecting the physical environment and / or movement of the user's hands, and sensors (e.g., internal sensors facing inwards towards the user's face) for detecting the user's attention (e.g., gaze).

[0229] In some embodiments, computer system 101 displays a user interface for a content application (e.g., streaming, delivery, playback, browsing, library, sharing, etc.) within a three-dimensional environment 902. In some embodiments, the content application includes a mini-player user interface (also referred to as a simplified user interface) and an extended user interface. In some embodiments, the mini-player user interface includes playback control elements that, in response to user input directed to these playback control elements, cause computer system 101 to modify the playback of content items played via the content application, and illustrations (e.g., album art) associated with the currently played content item. In some embodiments, the extended user interface includes a larger number of user interface elements than the mini-player user interface (e.g., containers such as windows, dials, or back panels; selectable options, content, etc.). In some embodiments, the extended user interface includes navigation elements, content browsing elements, and playback elements. In some embodiments, the content application includes three-dimensional virtual objects, such as animated virtual objects 922. Animated objects may optionally be displayed alongside the mini-player user interface. Reference Figures 9A to 9E The mini player user interface is further described in more detail with reference to methods 800, 1000, and 1200 below. Animation objects are described in more detail with reference to method 1000 below.

[0230] Figures 9A to 9E It also includes a top view of the three-dimensional environment 902, including the user 916, the computer system 101, and other objects in the three-dimensional environment 902 (e.g., the mini player user interface 904, the animated virtual object 922, the table 912, and the sofa 910).

[0231] exist Figure 9A In this embodiment, computer system 101 presents a three-dimensional environment 902 that includes representations of virtual objects and real objects. For example, virtual objects include a mini-player user interface 904 for a content application. In some embodiments, the mini-player user interface 904 includes an image (e.g., an album art) associated with a content item currently being played via the content application. As another example, representations of real objects include a representation 906 of the floor and a representation 908 of the walls in the physical environment of computer system 101. In some embodiments, representations of real objects are displayed via display generation component 120 (e.g., virtual pass-through or video pass-through), or as a view of real objects through a transparent portion of display generation component 120 (e.g., real pass-through or optical pass-through). In some embodiments, the physical environment of computer system 101 also includes a table and a sofa, and therefore, computer system 101 displays a representation of table 912 and a representation of sofa 910.

[0232] In some implementations, computer system 101 displays titles and artist instructions 918a for content items, as well as multiple user interface elements 918b-918h overlaid on images included in the mini player user interface 904 for modifying playback of content items. Figure 9B User interface elements 918g and 918h are shown. In some embodiments, in response to detecting input to one of the user interface elements 918b-918h, computer system 101 modifies the playback of the content item currently being played via the content application. In some embodiments, in response to detecting a selection of user interface element 918b, computer system 101 jumps back in the content item playback queue to restart the currently playing content item or play the previous item in the content item playback queue. In some embodiments, in response to detecting a selection of user interface element 918c, computer system 101 plays the content item and updates user interface element 918c to the user interface element that, when selected, causes computer system 101 to pause the playback of the content item. Figure 9A As shown, hand 903 in hand state A can optionally indicate selection of user interface element 918c. In some embodiments, in response to detecting selection of user interface element 918d, computer system 101 stops playback of the currently playing content item and initiates playback of the next content item in the content item playback queue. In some embodiments, in response to detecting selection of user interface element 918e, computer system 101 stops displaying the mini player user interface 904 and displays a reference. Figure 11A The extended user interface is described in more detail with reference to Figure 11O and one or more steps of method 1000. In some embodiments, in response to detecting a selection of user interface element 918f, computer system 101 displays time-synchronized lyrics for the currently playing content item, such as... Figure 9E As illustrated. In some embodiments, in response to detecting a selection of user interface element 918g, computer system 101 stops rendering the three-dimensional virtual object associated with the content item currently playing on computer system 101, and updates user interface element 918g to a user interface element that causes computer system 101 to redisplay the virtual object when selected. In some embodiments, in response to detecting a selection of user interface element 918h, computer system 101 presents another user interface element for adjusting the playback volume of the audio content of the content item and / or presents a menu for modifying audio output options for the playback of the audio content.

[0233] In some embodiments, computer system 101 detects selection of one of the user interface elements 918b-h by detecting indirect selection input, direct selection input, air gesture selection input, or input device selection input. In some embodiments, detecting selection input includes first detecting a ready state corresponding to the type of selection input being detected (e.g., detecting an indirect ready state before detecting indirect selection input, or detecting a direct ready state before detecting direct selection input). In some embodiments, detecting indirect selection input includes detecting a user's gaze pointing at the corresponding user interface element via input device 314, while simultaneously detecting a selection gesture made by the user's hand, such as an air pinch gesture where the user touches their thumb with another finger of their hand. In some embodiments, detecting direct selection input includes detecting a selection gesture made by the user's hand via input device 314, such as a pinch gesture within a predefined threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, or 30 cm) of the location of the corresponding user interface element, or a pressing gesture where the user's hand "presses" onto the location of the corresponding user interface element while forming a hand shape. In some embodiments, detecting air gesture input includes detecting a user's gaze pointing at the corresponding user interface element while detecting a pressing gesture at the location of an air gesture user interface element displayed in the three-dimensional environment 902 via display generation component 120. In some embodiments, detecting input device selection includes detecting manipulation of a mechanical input device (e.g., stylus, mouse, keyboard, touchpad, etc.) in a predefined manner corresponding to the selection of the user interface element when a cursor controlled by the input device is associated with the location of the corresponding user interface element and / or when the user's gaze is pointing at the corresponding user interface element.

[0234] like Figure 9A As shown, computer system 101 detects the selection of selectable option 918c, which, when selected, causes computer system 101 to play the corresponding content item, such as... Figure 9B As shown. In some implementations, selecting optional option 918c also causes computer system 101 to play an animation associated with a virtual object, which is presented together with the mini-player user interface 904. Figure 9BThe animation associated with the virtual object, shown in further detail, may optionally include movement associated with the corresponding content item. For example, the virtual object may optionally dance / move in sync with the beat of music, as discussed further with reference to method 1000. Selecting the selectable option 918c may also optionally cause the computer system 101 to update the user interface element 918c to a user interface element that, when selected, causes the computer system 101 to pause playback of the content item. In some embodiments, selecting the selectable option 918c to pause playback of the corresponding content item also includes pausing the animation associated with the virtual object. In some embodiments, selecting the selectable option 918c to pause playback of the corresponding content also includes blurring and / or increasing the transparency of the animated virtual object 922. In some embodiments, when playback of the corresponding content is paused, the animated virtual object 922 is stopped from being displayed.

[0235] Figure 9B This illustrates how computer system 101 responds to... Figure 9A The example shows the mini-player user interface 904 updating and displaying a virtual object upon detecting input of selection option 918c from hand 903. In some embodiments, the computer system 101 updates the mini-player user interface 904 in response to detecting a user's ready state as described above (e.g., hand 903 in hand state A). Updating the mini-player user interface 904 includes updating the selectable option 918c to a user interface element that pauses playback of content items when selected by the computer system 101.

[0236] In some implementations, in response to Figure 9A When the selection of selectable option 918c is detected, the virtual object appears and becomes animated (e.g., Figure 9B The animated virtual object 922 is shown. In some embodiments, and as described in reference method 1000, the animation of the animated virtual object 922 is derived from movement in a video. In some embodiments, characteristics associated with the animated virtual object 922 may include the appearance of a user (e.g., user 916 and / or a user in the video). For example, the animated virtual object 922 may optionally have the same gender, skin color, hair color, hairstyle, clothing color, clothing style, movement, and facial features as user 916. In some embodiments, user 916 may be in the video on which the animated virtual object 922 is based.

[0237] In some embodiments, the animated virtual object 922 is presented outside the boundary of the mini-player user interface 904. In some embodiments, the animated virtual object 922 is displayed adjacent to the mini-player user interface 904 (e.g., within 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, 20 cm, 30 cm, or 50 cm, or 1 m, 2 m, 3 m, or 5 m of the mini-player user interface 904). In some embodiments, the animated virtual object 922 is displayed to the left or right of the mini-player user interface 904. In some embodiments, the animated virtual object 922 is displayed at the same z-depth (e.g., distance) from the user's viewpoint as the mini-player user interface 904. In some embodiments, the animated virtual object 922 is displayed in front of or behind the mini-player user interface 904 (e.g., closer to or further away from the user's viewpoint than the mini-player user interface 904). In some embodiments, when the animated virtual object 922 is first displayed, it is displayed relative to the mini-player user interface 904 in a predetermined spatial arrangement (e.g., positioning and / or orientation).

[0238] In some implementations, computer system 101 detects that user 916's hand 903 performs an air pinch gesture (e.g., hand state B) pointing at animated virtual object 922, such as Figure 9B As shown. In response to detecting movement of hand 903 in hand state B when hand 903 is in a predefined shape (e.g., a pinched hand shape or an index finger pointing shape), computer system 101 moves animated virtual object 922 to different positions in three-dimensional environment 902, such as... Figure 9C As shown. In some embodiments, the computer system 101 moves the animated virtual object 922 to the location where the hand 903 releases the air pinch gesture or releases the touchscreen. For example, when the hand 903 releases the air pinch hand shape and / or lifts off the touchscreen ( Figure 9C As shown), computer system 101 detects the end of the selection input from hand 903. In some embodiments, the magnitude and / or direction of the movement of hand 903 determines the updated position of animated virtual object 922 and / or the magnitude and / or direction of the movement of animated virtual object.

[0239] In some implementations, the animated virtual object 922 continues to be animated as it is moved, such as Figure 9C As shown. Alternatively, and in some embodiments, the animated virtual object 922 responds to being moved and / or stops being animated while being moved. In some embodiments, the animation fades and / or blurs in response to the animated virtual object 922 moving its position. In some embodiments, the animated virtual object 922 fades from its old position and at the updated position (e.g., Figure 9CThe animation virtual object 922 fades in at the location of the hand 903. Fade-in or fade-out may optionally include increasing or decreasing the transparency, respectively. In some embodiments, the non-animated virtual object is moved in the same manner as described above and in method 1000. In some embodiments, the animated virtual object 922 moves independently of the mini-player user interface 904 and other virtual objects in the 3D environment. For example, when the animated virtual object 922 moves due to input (e.g., using the hand 903), the mini-player user interface 904 remains in the same position. In some embodiments, moving the animated virtual object 922 also includes moving the mini-player user interface 904. For example, input to move the animated virtual object 922 may optionally cause the mini-player user interface 904 to move by the same amount and / or in the same direction. Similarly, input to move the mini-player user interface 904 (from...) Figures 9B to 9C The hand in hand state B (903) can optionally cause the animated virtual object 922 to move with the same amount and direction.

[0240] Figure 9D An example is illustrated of how a computer system 101 updates a three-dimensional environment 902, including a mini-player user interface 904 and animated virtual objects 922, in response to the detection of movement of the computer system 101. This movement causes the computer system 101 to update the user's viewpoint in the three-dimensional environment 902 and the field of view of the computer system 101. In some embodiments, the computer system 101 updates the viewpoint of the mini-player user interface 904 in the three-dimensional environment 902 in response to the user's updated viewpoint. For example, Figure 9D A mini-player user interface 904 is shown at different angles to account for the updated viewpoint. However, the mini-player user interface 904 remains in the same position within the 3D environment 902 regardless of the updated viewpoint. Similarly, animated virtual objects 922 are optionally displayed at different angles to account for the updated viewpoint. Animated virtual objects 922 may optionally remain in the same position within the 3D environment 902 regardless of the updated viewpoint. In some embodiments, the animated virtual object 922 continues to be animated as the user's viewpoint in the 3D environment 902 and the field of view of the computer system 101 are updated. In some embodiments, changing the user's viewpoint in the 3D environment 902 results in the animation being updated to be displayed from the updated perspective associated with the user's new viewpoint.

[0241] Figure 9EAn example is illustrated of time-synchronized lyrics 924 associated with a content item presented alongside a mini-player user interface 904 and an animated virtual object 922. In some embodiments, the time-synchronized lyrics 924 is presented in response to computer system 101 detecting input corresponding to a selection of selectable option 918f. For example, computer system 101 detects indirect selection of option 918f, including detecting user gaze (e.g., using the user's eyes) when a hand (e.g., hand 903) makes a selection gesture (e.g., hand state A or hand state B). In some embodiments, the animated virtual object 922 continues to be animated in response to the presentation of the time-synchronized lyrics 924. Alternatively, the animated virtual object 922 may optionally transform into a non-animated virtual object in response to the presentation of the time-synchronized lyrics 924.

[0242] In some implementations, time-synchronized lyrics 924 are lyrics associated with a content item currently being played via a content application on computer system 101. Computer system 101 may optionally present a portion of lyrics 924 corresponding to the currently playing portion of the content item on computer system 101 via the content application, and update that portion of lyrics 924 as the content item continues to play. In some implementations, lyrics 924 include a line of lyrics corresponding to the currently playing portion of the content item, one or more lines of lyrics corresponding to a portion of the content item preceding the currently playing portion, and / or one or more lines of lyrics corresponding to a portion of the content item that will play after the currently playing portion. Figure 9EAs shown, lyrics 924 are displayed outside the boundary of the mini player user interface 904 in the three-dimensional environment 902 and / or at a location different from the mini player user interface. In some embodiments, lyrics 924 are displayed adjacent to the mini player user interface 904 (e.g., within 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, 20 cm, 30 cm, or 50 cm, or 1 m, 2 m, 3 m, or 5 m of the mini player user interface). In some embodiments, lyrics 924 are displayed to the left or right of the mini player user interface 904 and / or the animated virtual object 922. In some embodiments, lyrics 924 are displayed at the same z-depth (e.g., distance) from the user's viewpoint as the mini player user interface 904 and / or the animated virtual object 922. In some embodiments, from the user's viewpoint 916, the animated virtual object 924 is displayed in front of or behind the mini player user interface 904 and / or the animated virtual object 922 (e.g., closer to or further away from the user's viewpoint than the mini player user interface 904). In some implementations, lyrics 924 are initially displayed in a three-dimensional environment 902 relative to the mini-player user interface 904 and / or animated virtual objects 922 in a predetermined spatial arrangement (e.g., positioning and / or orientation).

[0243] about Figures 9A to 9E Additional or alternative details of the illustrated implementation schemes are provided below for reference. Figure 10 The method described in the description of method 1000.

[0244] Figure 10 A flowchart illustrating an exemplary method 1000 of how a computer system, according to some embodiments, generates animated 3D objects when presenting content items is described. In some embodiments, method 1000 is performed at a computer system (e.g., computer system 101 in FIG. 1, such as a tablet computer, smartphone, wearable computer, or head-mounted device), which includes display generation components (e.g., FIG. 1, ...). Figure 3 and Figure 4 The display generating component 120 (e.g., a head-up display, a monitor, a touchscreen, and / or a projector) and one or more cameras (e.g., one or more cameras pointing forward from or downward from the user's head toward the user's hand, such as color sensors, infrared sensors, and other depth-sensing cameras). In some embodiments, method 1000 is performed by one or more processors of a computer system (such as one or more processing units 202 of computer system 101, e.g., ...) stored in a non-transitory computer-readable storage medium and processed by one or more processors of a computer system (e.g., one or more processing units 202 of computer system 101, ...). Figure 1A The controller 110 in the middle manages the instructions executed. Some operations in method 1000 may be combined, and / or the order of some operations may be changed.

[0245] In some embodiments, method 1000 is performed at a computer system communicating with a display generation component and one or more input devices. In some embodiments, the computer system has one or more of the characteristics of the computer system in method 800. In some embodiments, the display generation component has one or more of the characteristics of the display generation component in method 800. In some embodiments, the one or more input devices have one or more of the characteristics of the one or more input devices in method 800.

[0246] In some implementations, a computer system (e.g., computer system 101) displays (1002a) the user interface of an application in a three-dimensional environment via a display generation component, such as... Figure 9A The mini-player user interface 904 shown (e.g., the user interface, application, and / or 3D environment optionally having one or more of the characteristics of the user interface, application, and / or 3D environment described in reference method 800) is associated with playback of corresponding content (e.g., as described in reference method 800), the corresponding content is not currently playing (e.g., as described in reference method 800), and the user interface is not displayed in the 3D environment as having corresponding animated objects (e.g., 2D or 3D objects), such as... Figure 9A As shown (e.g., the user interface is displayed separately in a 3D environment). In some embodiments, the corresponding animated object is an avatar. In some embodiments, the corresponding animated object depicts a human, animal (e.g., a dog, cat, or bird), and / or plant (e.g., a flower or tree). Alternatively, and in some embodiments, the corresponding animated object is displayed along with the user interface, but is not animated, when the corresponding content is not currently playing. In other embodiments, the corresponding animated object is not displayed at all when the corresponding content is not currently playing. In some embodiments, the corresponding animated object is displayed outside the user interface. For example, optionally, the corresponding animated object is displayed at a location in the 3D environment different from the user interface.

[0247] In some implementations, when the user interface of an application is displayed in a three-dimensional environment, and when the relevant content is not currently playing and the application's user interface is not displayed as having the relevant animated object, the computer system receives a first input (1002b) via one or more input devices corresponding to a request to initiate playback of the relevant content, such as using... Figure 9A The hand 903 shown is in hand state A. In some embodiments, the first input has one or more characteristics of the input described in reference method 800.

[0248] In some implementations, in response to receiving the first input (1002c), the computer system initiates (1002d) playback of the corresponding content in the three-dimensional environment, such as... Figure 9B As shown. In some implementations, playback of the corresponding content includes playback of audio (e.g., music), one or more videos, one or more podcasts, and / or one or more audiobooks.

[0249] In some implementations, the computer system displays (1002e) corresponding animated objects in a three-dimensional environment (e.g., Figure 9B The animated virtual object 922 shown is displayed in a three-dimensional environment relative to the application's user interface with a corresponding spatial arrangement (e.g., positioning and / or orientation), and one or more properties of the corresponding animated object are based on the playback of the corresponding content, such as... Figure 9B As shown, the animated virtual object 922 is located to the left and behind the mini-player user interface 904. In some embodiments, the corresponding animated object is positioned to the left, right, top, bottom, front, or back of the user interface. In some embodiments, the corresponding animated object is closer to the user's viewpoint or further away from the user's viewpoint than the user interface. In some embodiments, the corresponding animated object is animated, moves, and / or changes as the corresponding content is played back. For example, the corresponding animated object dances / moves in sync with the beat of the music, the intensity of the bass, the volume of the music, and / or the movement of the video. Additionally or alternatively, in some embodiments, the corresponding animated object changes its expression, color, and / or mood based on the playback of the corresponding content. In some embodiments, and in response to receiving input to stop or pause the corresponding content, the corresponding animated object is not animated. In some embodiments, the corresponding animated object is still displayed when the corresponding content is paused or stopped. Displaying the user interface as an animated object with playback based on the corresponding content reduces the resources required to display simulated lighting effects when the corresponding content is not playing, and also reduces the need for manual input to enable and / or disable simulated lighting effects.

[0250] In some implementations, before displaying the corresponding animated object, such as when the animated virtual object 922 is imported from a second computer system, the computer system imports the corresponding animated object from the second computer system into the computer system. (In some implementations, importing the corresponding animated object includes importing a USD file containing the corresponding animated object. In some implementations, the second computer system records video and converts it into an animated object. Importing the corresponding animated object reduces the resources required to generate the corresponding animated object on the computer system.)

[0251] In some embodiments, the computer system displays a corresponding animated object, including an animation of the corresponding animated object, wherein, if such animations are used in animated virtual object 922, the animation is also pre-recorded. In some embodiments, recording the animation includes using sensors communicating with the computer system, one or more characteristics associated with the corresponding animated object. The one or more characteristics associated with the corresponding animated object may optionally include those characteristics as described above. In some embodiments, the corresponding animated object is derived from a video of a person. In some embodiments, the characteristics associated with the corresponding animated object include the appearance of the person, such as their gender, skin color, hairstyle, clothing style, movement, and facial features. In some embodiments, the video of the person is a video recorded from a user. The computer system may optionally translate the user's movement in the video into the movement of the avatar. Using sensors to record one or more characteristics associated with the corresponding animated object improves usability, thereby improving user device interaction.

[0252] In some implementations, when the user interface (e.g., mini player user interface 904) is displayed in a three-dimensional environment as having corresponding animated objects, and when the user's viewpoint on the computer system has a first spatial arrangement relative to the corresponding animated objects, by... Figure 9C The spatial arrangement of objects in the environment is shown (e.g., when the user's viewpoint has a first spatial arrangement relative to the corresponding animated object, the computer system displays a first portion of the corresponding animated object in the 3D environment). The computer system receives, via one or more input devices, the movement of the user's viewpoint relative to the corresponding animated object in the 3D environment to have a second spatial arrangement different from the first spatial arrangement (such as moving the position by the user 916 and thus from...). Figures 9C to 9D The input corresponding to the moving viewpoint is described in further detail in method 800.

[0253] In some implementations, in response to receiving input corresponding to the movement of a user's viewpoint in the computer system, the computer system updates the display of the corresponding animated objects in the 3D environment to reflect the user's moving viewpoint, such as through... Figure 9D The updated display of the animated virtual object 922 is shown (e.g., the user's moving viewpoint has a second spatial arrangement relative to the corresponding animated object). In some embodiments, a second portion of the corresponding animated object, different from the first portion, is displayed in a three-dimensional environment from the user's moving viewpoint, such as displaying the corresponding animated object at different angles.

[0254] In some implementations, the computer system continues to display the application's user interface as having corresponding animated objects, such as through... Figure 9DThe updated display of the mini-player user interface is shown. In some embodiments, the computer system displays different views of the corresponding animated object and / or user interface as the user's viewpoint moves around the corresponding animated object and / or user interface. For example, in a first spatial arrangement, the first viewpoint is on a first side of the user interface (e.g., in front). For example, in a second spatial arrangement, the moving viewpoint is on a second, different side of the user interface (e.g., opposite to or behind the first viewpoint). In some embodiments, the orientation difference between the first viewpoint and the moving viewpoint relative to the user interface is 1 degree, 5 degrees, 10 degrees, 30 degrees, 45 degrees, 90 degrees, 180 degrees, 270 degrees, or 359 degrees. In some embodiments, when the computer system displays the corresponding animated object from different perspectives as the user's viewpoint changes, one or more of the characteristics of the corresponding animated object remain unchanged in the three-dimensional environment. For example, color, mood, and / or animation remain unchanged. Allowing users to view corresponding animated objects from different spatial arrangements in the three-dimensional environment achieves consistent presentation of feedback in the three-dimensional environment, thereby reducing usage errors and improving user-device interaction.

[0255] In some implementations, when the application's user interface and corresponding animated objects are displayed at a first location in the 3D environment, the computer system detects input corresponding to a request to move the corresponding animated objects in the 3D environment via one or more input devices, such as through... Figure 9B The input shown is made by hand 903 in hand state B. A request to move the corresponding animated object may optionally include input selecting selectable options displayed in the user interface, as described with reference to method 800. In some embodiments, the magnitude and direction of the input (e.g., pinch, drag, and / or release) directed at the corresponding animated object determine the updated position of the animated object. For example, tapping / touching and then dragging the animated object (in the case of a touch-sensitive surface) indicates a request to move the corresponding animated object. For example, pinching in the air (e.g., using the user's thumb and forefinger) and dragging the corresponding animated object while the hand is in the pinched hand shape (in the case of a wearable device) indicates a request to move the corresponding animated object to the position to which the hand wants to drag it.

[0256] In some implementations, in response to detecting input corresponding to a request to move a corresponding animated object in a three-dimensional environment, the computer system displays the corresponding animated object at a second location in the three-dimensional environment via a display generation component (e.g., including moving the corresponding animated object to a second location in the three-dimensional object according to the input), and continues to display the application's user interface at a first location in the three-dimensional environment, such as via a virtual animated object 922. Figure 9B First position to Figure 9CThe second position movement is illustrated. In some embodiments, the computer system updates the position of the corresponding animated object by initiating the display of the corresponding animated object at the updated position using a fade-in animation effect. In some embodiments, the corresponding animated object fades out from its old position using a fade-out animation effect (e.g., increased opacity) and then fades in at the updated position (e.g., decreased opacity). In some embodiments, the corresponding animated object and the user interface are "world-locked" and remain at their respective positions in the 3D environment unless and until input is received to update the positions of the corresponding animated object and the user interface. Maintaining the position of the displayed user interface in the 3D environment while the position of the corresponding animated object changes provides an efficient way for users to view other user interfaces and content in the 3D environment without moving the application's user interface, thereby improving user device interaction.

[0257] Figures 11A to 11N Examples are shown of how a computer system can display a simplified user interface instead of an extended user interface in response to different inputs.

[0258] Figure 11A An example is illustrated where a computer system (e.g., an electronic device) 101 displays a three-dimensional environment 1102 from the user's viewpoint (e.g., facing the rear wall of the physical environment in which the computer system 101 is located) via a display generation component (e.g., display generation component 120 of FIG. 1). In some embodiments, the computer system 101 includes a display generation component (e.g., a touchscreen) and multiple image sensors (e.g., ...). Figure 3 Image sensor 314). The image sensor may optionally include one or more of the following: a visible light camera; an infrared camera; a depth sensor; or any other sensor that the computer system 101 can use to capture one or more images of the user or a portion of the user (e.g., one or both of the user's hands) when the user interacts with the computer system 101. In some embodiments, the user interface illustrated and described below may also be implemented on a head-mounted display including display generation components for displaying the user interface or a three-dimensional environment to the user, and sensors for detecting the physical environment and / or movement of the user's hands (e.g., external sensors facing outward from the user) and / or sensors for detecting the user's attention (e.g., gaze) (e.g., internal sensors facing inward toward the user's face).

[0259] In some embodiments, computer system 101 displays a user interface for a content application (e.g., streaming, delivery, playback, browsing, library, sharing, etc.) in a three-dimensional environment 1102. In some embodiments, the content application includes a mini-player user interface (also referred to as a simplified user interface) and an extended user interface. In some embodiments, the mini-player user interface includes playback control elements that, in response to user input directed to these playback control elements, cause computer system 101 to modify the playback of content items played via the content application, and illustrations (e.g., album art) associated with the currently played content item. In some embodiments, the extended user interface includes a larger number of user interface elements than the mini-player user interface (e.g., containers such as windows, dials, or back panels; selectable options, content, etc.). In some embodiments, the extended user interface includes navigation elements, content browsing elements, and playback elements. References are provided below. Figure 11A The mini player user interface and extended user interface are described in more detail with reference to Figure 11O and further to Method 1200 below.

[0260] exist Figure 11A In this embodiment, computer system 101 presents a three-dimensional environment 1102 that includes representations of virtual objects and real objects. For example, virtual objects include an extended user interface 1104 for content applications. The extended user interface 1104 optionally includes navigation elements 1124 and content browsing elements 1126. In some embodiments, virtual objects also include playback control elements 1128 displayed below the extended user interface 1104. As another example, representations of real objects include a representation 1106 of the floor and a representation 1108 of the walls in the physical environment of computer system 101. In some embodiments, representations of real objects are displayed via display generation component 120 (e.g., virtual pass-through, active pass-through, or video pass-through), or as a view of real objects through a transparent portion of display generation component 120 (e.g., real pass-through or passive pass-through). In some embodiments, the physical environment of computer system 101 also includes a table and a sofa, and therefore, computer system 101 displays digital representations of table 1112 and sofa 1110.

[0261] Figure 11A Navigation element 1124 of extended user interface 1104 is shown. Navigation element 1124 includes multiple selectable options 1118a-e, which, when selected, navigate the computer system 101 to different user interfaces of the content application within content browsing element 1126 of extended user interface 1104. Figure 11AIn the current implementation, the "Listen Now" option 1118a is selected, therefore the computer system 101 presents a user interface for browsing recommended content items based on the user's content consumption history. In some embodiments, in response to detecting a selection of the "Browse" option 1118b, the computer system 101 presents a content browsing user interface in content browsing element 1126, which includes user interface elements for browsing content items based on genres, artists, playback charts, etc., for all users of the content delivery service associated with the content application. In some embodiments, in response to detecting a selection of the "Radio" option 1118c, the computer system 101 presents a radio user interface in content browsing element 1126, which includes information about playback of Internet-based radio programs and stations available via the content delivery service associated with the content application, and selectable options for initiating such playback. In some embodiments, in response to detecting a selection of the "Library" option 1118d, the computer system 101 presents a user interface in content browsing element 1126, which includes a representation of content items in a content library associated with the user's user account of the computer system 101. In some implementations, in response to detecting a selection of the “search” option 1118e, the computer system 101 displays a search user interface in the content browsing element 1126, which includes user interface elements for providing search terms to be searched in the content delivery service associated with the content application.

[0262] In some implementations, the extended user interface of the content application includes a content browsing element 1126 that displays the aforementioned "Listen Now" user interface. Figure 11A An example is illustrated by a content browsing element 1126 comprising multiple representations, such as representations 1116a and 1116b. In some embodiments, representations 1116a and 1116b are arranged alphabetically (e.g., by artist, content item title, or title of a collection of content items (e.g., album, playlist)), and in response to detecting a selection of a corresponding portion of the letter scroll bar 1120, the computer system 101 scrolls representations 1116a and 1116b to the representation of the content item corresponding to the letter of the corresponding portion of the letter scroll bar 1120.

[0263] Figure 11AThe system also includes a playback control element 1128, which includes an image 1130 (e.g., an album art) corresponding to a content item currently being played via a content application on computer system 101, an indication 1132 of the content item's title, and a plurality of user interface elements 1134a-1134i that, in response to detecting input directed to one of the user interface elements 1134a-1134i, cause computer system 101 to modify the playback of the content item currently being played via the content application. In some embodiments, image 1130 is a selectable option that, when selected, causes computer system 101 to stop displaying an extended user interface of the content application and display a mini-player user interface of the content application, which is described in further detail below and with reference to methods 800 and 1000. In some embodiments, playback control element 1128 includes a backflip option 1134a that, when selected, causes computer system 101 to resume the currently playing content item and / or play the previous content item in the content application's playback queue. In some embodiments, playback control element 1128 includes a pause option 1134b, which, when selected, causes the computer system to pause playback of a content item and updates option 1134b to a play option, which, when selected, causes the computer system 101 to resume playback of the content item (e.g., from a paused playback location). In some embodiments, playback control element 1128 includes a jump-forward option 1134c, which, when selected, causes the computer system 101 to play the next content item in the content application's playback queue. In some embodiments, playback control element 1124 includes a favorites option 1134d, which, when selected, causes the computer system to mark the content item as a favorite of the user account associated with the computer system 101. In some embodiments, favorited content items include updating the characteristics of the content item such that the content item appears in a playlist (e.g., a favorites playlist) and / or a content library (e.g., which can be navigated to by selecting option 1118d). In some embodiments, playback control element 1124 includes option 1134e, which, when selected, causes computer system 101 to display time-synchronized lyrics for the content item. Time-synchronized lyrics may optionally be generated and / or displayed according to one or more steps of method 800, method 1000, and / or method 1200. In some embodiments, playback control element 1124 includes option 1134f, which, when selected, causes computer system 101 to present another user interface element (e.g., a slider) for adjusting the playback volume of the audio content of the content item and / or to present a menu for modifying audio output options for the playback of the audio content. In some embodiments, playback control element 1124 includes option 1134g, which displays one or more audio output settings to configure the output of the audio portion of the content item (e.g., selecting an output device).In some embodiments, playback control element 1124 includes a scribble bar 1134h that indicates the playback position in the content item currently being played via the content application, and in response to input directed to the scribble bar 1134h, causes computer system 101 to update the playback position based on that input rather than on the continued playback of the content item. In some embodiments, and as such... Figure 11A As shown, the playback location of the content item is displayed next to the scrubbing bar 1134h.

[0264] Figure 11AThe illustration includes a user's hand 1103 in hand state A, which corresponds to a hand shape, pose, position, etc., associated with a ready state or input. In some embodiments, the computer system 101 is capable of detecting indirect ready states, direct ready states, air gesture ready states, and / or input device ready states. In some embodiments, detecting an indirect ready state includes detecting a user's hand 1103 in a ready state pose (such as a pre-pinch gesture where the thumb is within a threshold distance (e.g., 0.5 cm, 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, or 30 cm) of a corresponding interactive user interface element (e.g., via one or more input devices in input devices 314) when the hand 1103 is within a predefined threshold distance (e.g., 0.5 cm, 1 cm, 2 cm, 3 cm, 4 cm, or 5 cm) of another finger of the hand but not touching that other finger, or a pointing hand shape where one or more fingers are extended and one or more fingers are curled toward the palm). In some embodiments, detecting an indirect ready state includes detecting the user's hand 1103 in a ready state pose (such as a pre-pinch hand shape) when a user's gaze is detected pointing at a corresponding interactive user interface element (e.g., via one or more input devices in input devices 314). In some embodiments, detecting an air gesture ready state includes detecting the hand 1103 in a ready state pose (such as a pointing hand shape within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, or 30 cm) of an input element displayed via display generation component 120 when a user's gaze is detected pointing at a corresponding interactive user interface element (e.g., via one or more input devices in input devices 314). In some implementations, detecting an input device ready state includes, optionally, detecting a predefined portion of the user (e.g., the user's hand 1103) that is adjacent to a mechanical input device (e.g., a stylus, touchpad, mouse, keyboard, etc.) communicating with the computer system 101 but not providing input to that mechanical input device when a cursor controlled by the input device corresponds to a corresponding interactive user interface element, or optionally, when a user's gaze toward the corresponding interactive user interface element is detected (e.g., via one or more input devices in input device 314).

[0265] In some implementations, computer system 101 detects selection of a corresponding user interface element by detecting indirect selection input, direct selection input, air gesture selection input, or input device selection input. In some implementations, detecting selection input includes first detecting a ready state corresponding to the type of selection input being detected (e.g., detecting an indirect ready state before detecting indirect selection input, and a direct ready state before detecting direct selection input). In some implementations, detecting indirect selection input includes detecting a user's gaze pointing at the corresponding user interface element via input device 314, while simultaneously detecting a selection gesture made by the user's hand, such as a pinch gesture where the user touches their thumb with another finger of their hand. In some implementations, detecting direct selection input includes detecting a selection air gesture made by the user's hand via input device 314, such as a pinch gesture within a predefined threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, or 30 cm) of the location of the corresponding user interface element, or a pressing air gesture where the user's hand "presses" the location of the corresponding user interface element while in a hand-pointing shape. In some embodiments, detecting air gesture input includes detecting a press gesture at the location of an air gesture user interface element displayed in the three-dimensional environment 1102 via display generation component 120, while simultaneously detecting a user's gaze pointing at the corresponding user interface element. In some embodiments, detecting input device selection includes detecting manipulation of a mechanical input device (e.g., stylus, mouse, keyboard, touchpad, etc.) in a predefined manner corresponding to the selection of the user interface element when a cursor controlled by the input device is associated with the location of the corresponding user interface element and / or when the user's gaze is directed at the corresponding user interface element. Figure 11A An example is shown of selecting image 1130 using hand 1103 in hand state A.

[0266] Figure 11B An example is illustrated where computer system 101 updates 3D environment 1102 in response to selection of image 1130 to stop displaying extended user interface 1104 and display mini-player user interface 1136. In some embodiments, mini-player user interface 1136 appears in the same location within 3D environment 1102 where extended user interface 1104 is located. Alternatively, in some embodiments, mini-player user interface 1136 appears in a different location within 3D environment 1102 than where extended user interface 1104 is located. In some embodiments, in response to selection of image 1130, extended user interface 1104 fades out to stop displaying and mini-player user interface 1136 fades in. In some embodiments, mini-player user interface 1136 is movable according to method 800.

[0267] Figure 11B The mini player user interface 1136 includes titles for content items and instructions for artists 1138a, as well as multiple user interface elements 1138b-1138h overlaid on the image for modifying playback of content items, such as... Figures 7A to 7H And as described in methods 800 and 1000. In some embodiments, in response to detecting a selection of user interface element 1138b, computer system 101 stops displaying the mini player user interface 1136 and redisplays the extended user interface 1104. In some embodiments, in response to detecting a selection of user interface element 1138c, computer system 101 displays time-synchronized lyrics for a content item according to one or more steps of method 1200. In some embodiments, the mini player user interface 1136 includes a scrubbing bar 1138h that indicates the playback position of computer system 101 in the content item currently being played via the content application, and in response to input to the scrubbing bar 1138d, causes computer system 101 to update the playback position based on that input rather than based on the continued playback of the content item. In some embodiments, the playback position of the content item is displayed next to the scrubbing bar 1138d. For example, the remaining time of the content item is displayed. In some embodiments, the mini player user interface 1136 includes, for example, the remaining time of the content item. Figure 7A The backward jump option 1138e is described in further detail below. In some embodiments, the mini player user interface 1136 includes, for example, Figure 7A The pause option 1138f is described in further detail below. In some implementations, the mini player user interface 1136 includes, for example, the pause option 1138f. Figure 7A The forward jump option 1138g is described in further detail below. In some embodiments, the mini player user interface 1136 includes a slider 1138h that, in response to input manipulating the slider 1138h, causes the computer system 101 to modify the playback volume of content items on the computer system 101.

[0268] Figure 11B A hand 1103 in hand state B is also illustrated at a corner of the mini player user interface 1136. In some embodiments, hand 1103 makes the selection gesture described above (pointing to the corner of the mini player user interface) toward the corner of the mini player user interface 1136. In some embodiments, making the selection gesture includes making a corresponding hand shape (e.g., hand state B), such as making a pinch hand shape as part of performing a pinch gesture.

[0269] exist Figure 11CIn this implementation, after selecting a corner of the mini-player user interface 1136, the user moves their hand 1103 while maintaining the corresponding hand shape (e.g., hand state B). In response to detecting movement of the hand 1103 while in hand state B, the computer system 101 zooms in on the mini-player user interface 1136 based on the movement of the hand 1103 (or based on movement of input, such as from the user's gaze). For example, if the hand 1103 moves diagonally upwards and to the right, the computer system 101 zooms in on the mini-player user interface 1136 by the same amount and direction as the hand movement. In some embodiments, if the hand 1103 moves downwards and to the left, the computer system zooms out on the mini-player user interface 1136 by the same amount and direction. In some embodiments, multiple user interface elements 1138a-1138h also scale up or down based on the movement of the hand 1103. In some embodiments, other virtual objects remain the same size unless the computer system 101 detects input pointing to an object for modifying the object's size.

[0270] Figure 11D An example is illustrated of a user interface 1140 for a second application other than the content application, displayed alongside the mini-player user interface 1136 of the content application. In some embodiments, after the display of the extended user interface 1104 is stopped, the computer system 101 detects input from the user for displaying the user interface 1140. In some embodiments, the user interface 1140 is a user interface for a web browsing application, a file application, a document editing application, a media viewing application, and / or other applications present on the computer system 101. In some embodiments, in response to input from the user (e.g., air pinch or gaze), the user interface 1140 is positioned in a location within the three-dimensional environment 1102. In some embodiments, the user interface 1140 is presented outside the boundary of the mini-player user interface element 1136. In some embodiments, the user interface 1140 is displayed adjacent to the mini-player user interface 1136 (e.g., within 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, 20 cm, 30 cm, or 50 cm, or 1 m, 2 m, 3 m, or 5 m of the mini-player user interface). In some embodiments, the user interface 1140 is displayed to the left or right of the mini player user interface 1136. In some embodiments, the user interface 1140 is displayed at the same z-depth (e.g., distance) from the user's viewpoint as the mini player user interface 1136. In some embodiments, the user interface 1140 is displayed in front of or behind the mini player user interface 1136. In some embodiments, the mini player user interface 1136 moves independently of the user interface 1140. In some embodiments, the mini player user interface 1136 moves in conjunction with the user interface 1140.

[0271] Figure 11D It is also illustrated that hand 1103 in hand state A selects user interface element 1138b. Computer system 101 detects hand 1103 in a selection gesture (or other input in a selection gesture, such as a user's gaze), as described above. In some embodiments, the user uses indirect selection (e.g., using the user's gaze) to select user interface element 1138b. As a result of the selection of user interface element 1138b, the computer system redisplays... Figure 11E Extended user interface 1104. In some implementations, in the first position (e.g., Figure 11A The extended user interface 1104 is displayed at the location (in the context). In response to the selection of element 1138b, the mini-player user interface 1136 and user interface 1140 cease to be displayed by the computer system 101. In some embodiments, if user interface 1140 is not located at the first position where the extended user interface 1104 was previously displayed and is now redisplayed, user interface 1140 continues to be displayed when the computer system 101 redisplays the extended user interface 1104. In some embodiments, applications other than content applications are displayed simultaneously with the mini-player user interface 1136. However, in some embodiments, applications other than content applications are not displayed simultaneously with the mini-player user interface 1136. In some embodiments, the mini-player user interface 1136 and / or user interface 1140 cease to be displayed by fading out from the 3D environment 1102. Fading out may optionally include blurring the mini-player user interface 1136 and / or user interface 1140 and increasing the semi-transparency. In some embodiments, the selection of element 1138b does not affect the playback of content items or the position of other objects in the 3D environment 1102.

[0272] Figure 11E An example is illustrated using hand 1103d in hand state A to select option 1134e. In some embodiments, hand 1103d makes a selection gesture, as described above. In some embodiments, computer system 101 detects selection input directed toward option 1134e, such as a user's gaze. In response to detecting input pointing toward option 1134e, time-synchronized lyrics 1142 are displayed, such as... Figure 11FAs shown. Optionally, time-synchronized lyrics 1142 may be displayed instead of extended user interface 1104, and extended user interface 1104 may cease display. In some embodiments, playback control element 1128 continues to be displayed. For example, time-synchronized lyrics 1142 may be displayed at the location where navigation element 1124 and content browsing element 1126 were previously displayed. Optionally, time-synchronized lyrics 1142 may be displayed at the location of one or more steps of reference method 1200. In other examples, time-synchronized lyrics 1142 may be displayed at a location different from the location where navigation element 1124 and content browsing element 1126 were displayed. In some embodiments, selection of option 1134e does not affect playback of the content item. For example, the content item continues to play during and after selection of option 1134e.

[0273] In some implementations, time-synchronized lyrics 1142 are lyrics associated with a content item currently being played via a content application on computer system 101. Computer system 101 may optionally present a portion of lyrics 1142 corresponding to the currently playing portion of the content item on computer system 101 via the content application, and update that portion of lyrics 1142 as the content item continues to play. In some implementations, lyrics 1142 include a line of lyrics corresponding to the currently playing portion of the content item, one or more lines of lyrics corresponding to a portion of the content item preceding the currently playing portion, and / or one or more lines of lyrics corresponding to a portion of the content item that will play after the currently playing portion.

[0274] Figure 11E The example also illustrates input using hand 1103 (or indirect input, such as a user's gaze) to move an object (e.g., playback control element 1128 and / or lyrics 1142). For example, computer system 101 detects direct selection of playback control element 1128 using hand 1103 in hand state C and / or direct selection of lyrics 1142 using hand 1103 in hand state C, such as... Figure 11F As shown. In some embodiments, hand 1103 performs indirect selection of playback control element 1128 or lyrics 1142. Direct and indirect selection have been described in further detail above. In some embodiments, hand 1103 moves (e.g., to the right) playback control element 1128 or lyrics 1142. In some embodiments, computer system 101 independently uses hand 1103 as described herein to respond to each input.

[0275] like Figure 11G As shown, due to Figure 11F The computer system 101 detected the input, and the playback control element 1128 and lyrics 1142 were placed in new positions within the three-dimensional environment 1102. It should be understood that, although... Figure 11F An example of selecting both playback control element 1128 and lyrics 1142 is given, but in some implementations, the input is detected at different times rather than simultaneously. Furthermore, although... Figure 11G The illustration shows the movement of both playback control element 1128 and lyrics 1142 due to two inputs. However, in some embodiments, a movement input to playback control element 1128 causes both playback control element 1128 and lyrics 1142 to move by the same amount and direction. Similarly, in some embodiments, a movement input to lyrics 1142 causes both playback control element 1128 and lyrics 1142 to move by the same amount and direction. In other words, playback control element 1128 and lyrics 1142 may optionally move together. In some embodiments, while a movement input to playback control element 1128 causes playback control element 1128 and lyrics 1142 to move together, a movement input to lyrics 1142 causes lyrics 1142 to move independently of playback control element 1128 (e.g., not moving playback control element 1128). The movement of playback control elements 1128 and lyrics 1142 may optionally not affect the playback of content items.

[0276] Figure 11H An example is illustrated where hand 1103g selects element 1138c. Hand 1103g is making a selection gesture (e.g., hand state A). In some embodiments, input pointing to element 1138c causes computer system 101 to update the 3D environment. Figure 11I As shown, due to Figure 11H In response to the selection of element 1138c, computer system 101 updates the 3D environment 1102 to include time-synchronized lyrics 1144. Time-synchronized lyrics 1144 have one or more of the attributes of lyrics 1142 described above. Lyrics 1144 are displayed outside the boundaries of the mini-player user interface 1136. In some embodiments, lyrics 1144 are displayed adjacent to the mini-player user interface 1136 (e.g., within 1 cm, 2 cm, 3 cm, 5 cm, 10 cm, 15 cm, 20 cm, 30 cm, or 50 cm, or 1 m, 2 m, 3 m, or 5 m of the mini-player user interface). In some embodiments, lyrics 1144 are displayed to the left or right of the mini-player user interface 1136. In some embodiments, lyrics 1144 are displayed at the same z-depth (e.g., distance) from the user's viewpoint as the mini-player user interface 1136.

[0277] Figure 11I The example also illustrates how hand 1103, in hand state C, performs selection and movement gestures on lyrics 1144 and / or the mini-player user interface 1136. For instance, moving hand 1103 to the left causes lyrics 1144 to move to the left. Moving hand 1103 to the right causes the mini-player user interface 1136 to move to the right, as... Figure 11J As shown. In some embodiments, direct or indirect input to lyrics 1144 and / or mini-player user interface 1136 achieves the results described below. In some embodiments, movement input to mini-player user interface 1136 causes mini-player user interface 1136 to move and keeps lyrics 1144 in the same position. Similarly, in some embodiments, movement input to lyrics 1144 causes lyrics 1144 to move and keeps mini-player user interface 1136 in the same position. In other words, mini-player user interface 1136 and lyrics 1144 may optionally be independently controllable and / or movable. It should be noted that although multiple inputs are shown simultaneously, in some embodiments, each input may be performed individually. For example, mini-player user interface 1136 and lyrics 1144 may optionally be moved independently with two different hands at different times (and in different directions).

[0278] In some implementations, lyrics 1144 are adjustable in size. Figure 11J An example is shown of hand 1103 at the corner of lyrics 1144. In some embodiments, and as... Figure 11J and Figure 11K As shown, hand 1103 performs a selection gesture (e.g., a pinched hand shape as part of a pinch gesture) and a drag gesture (e.g., hand state B) at the corner of lyrics 1144. Because hand 1103 selects the lower left corner and drags the input away from the center of lyrics 1144, the size of lyrics 1144 increases, as... Figure 11K As shown. In some embodiments, the input is indirect. In some embodiments, the input reduces the size of the lyrics 1144. For example, input by hand 1103 selecting the lower left corner and dragging towards the center of the lyrics will result in a smaller representation of the lyrics 1144. In some embodiments, the mini-player user interface 1136 does not change size in response to input from hand 1103. Therefore, the mini-player user interface 1136 and the lyrics 1144 may optionally be independently scalable.

[0279] In some implementations, the lyrics 1144 are scrollable. In some implementations, the computer system 101 displays a scroll bar 1146 next to the lyrics 1144. For example, in Figure 11LIn the lyrics 1144, scroll bar 1146 is located to the right of the lyrics 1144. In some embodiments, scroll bar 1146 may be to the left, top, or bottom of the lyrics 1144. In some embodiments, scroll bar 1146 may be within 0.1 cm, 1 cm, 2 cm, or 10 cm of the text of the lyrics 1144. In some embod...

Claims

1. A method, the method comprising: At the computer system that communicates with the display generation components and one or more input devices: An extended user interface of the application is displayed at a first location in the three-dimensional environment via the display generation component, wherein: The application is controlling the playback of the first content item on the computer system. The extended user interface includes a first selectable user interface object, which can be selected to initiate playback of a second content item, different from the first content item, on the computer system. The extended user interface and the second selectable user interface object are displayed simultaneously, while the second selectable user interface object and the extended user interface are displayed separately at a second location in the three-dimensional environment; When the extended user interface of the application is displayed at the first location in the three-dimensional environment and the second selectable user interface object is displayed at the second location in the three-dimensional environment, a first input corresponding to the selection of the second selectable user interface object is received via the one or more input devices; and In response to receiving the first input: A simplified user interface for controlling the playback of the first content item on the computer system is displayed in the three-dimensional environment, wherein the simplified user interface is displayed at a third location in the three-dimensional environment different from the first and second locations; and Stop displaying the extended user interface and the second selectable user interface object in the three-dimensional environment.

2. The method according to claim 1, further comprising: The display generation component simultaneously displays, in the three-dimensional environment, a playback control user interface of the application, separate from the extended user interface, which is used to control the playback of the first content item.

3. The method of claim 2, wherein the playback control user interface includes an image corresponding to the first content item and a third selectable user interface object for modifying the playback of the first content item.

4. The method of claim 3, wherein the second selectable user interface object includes the image corresponding to the first content item.

5. The method according to claim 1, further comprising: When displaying the simplified user interface of the application, a second input is received via the one or more input devices, the second input corresponding to a request to display a second user interface in the three-dimensional environment that corresponds to a second application different from the application; as well as In response to receiving the second input corresponding to displaying the second user interface: The simplified user interface of the application and the second user interface corresponding to the second application are displayed simultaneously in the three-dimensional environment.

6. The method according to claim 5, further comprising: When displaying the second user interface corresponding to the second application, a third input is received via the one or more input devices, the third input corresponding to a request to display a third user interface corresponding to a third application that is different from the application and the second application; And in response to receiving the third input corresponding to displaying the third user interface: Stop the display of the second user interface and display the third user interface.

7. The method of claim 5, wherein the simplified user interface includes a third selectable user interface object, the third selectable user interface object being selectable to redisplay the extended user interface, and the method further includes: When displaying the third selectable user interface object that is displayed simultaneously with the simplified user interface, a fourth input corresponding to the selection of the third selectable user interface object is received via the one or more input devices; as well as In response to receiving the fourth input corresponding to the selection of the third selectable user interface object: Display the extended user interface and stop displaying the second user interface corresponding to the second application.

8. The method according to claim 1, further comprising: When the simplified user interface is displayed, a second input corresponding to a request to adjust the size of the simplified user interface is received via the one or more input devices; as well as In response to receiving the second input corresponding to the request to adjust the size of the simplified user interface, the size of the simplified user interface in the three-dimensional environment is updated according to the second input corresponding to the request to adjust the size of the simplified user interface.

9. The method of claim 2, wherein the simplified user interface includes a third selectable user interface object that can be selected to display a representation of the lyrics of the first content item, and the playback control user interface includes a fourth selectable user interface object that can be selected to display a representation of the lyrics of the first content item.

10. The method according to claim 9, further comprising: When displaying the simplified user interface including the third selectable user interface object, a second input corresponding to the selection of the third selectable user interface object is received via the one or more input devices; as well as In response to receiving the second input, the representation of the lyrics of the first content item is displayed simultaneously in the three-dimensional environment and the simplified user interface, wherein the representation of the lyrics is displayed at a fourth position in the three-dimensional environment, and the simplified user interface is displayed at a fifth position in the three-dimensional environment, different from the fourth position.

11. The method according to claim 9, further comprising: When the representation of the lyrics is displayed at a fourth location in the three-dimensional environment and the simplified user interface is located at a fifth location in the three-dimensional environment, a second input corresponding to a request to move the simplified user interface is received via the one or more input devices; as well as In response to receiving the second input: The simplified user interface is moved to a sixth position in the three-dimensional environment based on the second input; as well as The representation of the lyrics of the first content item is maintained at the fourth position in the three-dimensional environment.

12. The method according to claim 9, further comprising: When the representation of the lyrics is displayed at a fourth location in the three-dimensional environment and the simplified user interface is located at a fifth location in the three-dimensional environment, a second input corresponding to a request to move the representation of the lyrics of the first content item is received via the one or more input devices; as well as In response to receiving the second input corresponding to the representation of the lyrics that moved the first content item: Move the representation of the lyrics of the first content item to the sixth position; and The simplified user interface continues to be displayed at the fifth position.

13. The method according to claim 9, further comprising: When the representation of the lyrics is displayed at a first size in the three-dimensional environment, a second input corresponding to a request to modify the size of the representation of the lyrics of the first content item is received via the one or more input devices; as well as In response to receiving the second input: Based on the second input, the representation of the lyrics of the first content item is modified to have a second size in the three-dimensional environment that is different from the first size; as well as Maintain the dimensions of the simplified user interface within the three-dimensional environment.

14. The method according to claim 1, further comprising: When the extended user interface is displayed and a third selectable user interface object that can be selected to display the lyrics of the first content item, wherein the extended user interface is displayed at the first position in the three-dimensional environment, a second input corresponding to the selection of the third selectable user interface object is received via the one or more input devices; And in response to receiving the second input: Stop displaying the extended user interface at the first location in the three-dimensional environment; as well as The representation of the lyrics of the first content item is displayed at the first location in the three-dimensional environment.

15. The method of claim 14, wherein the representation of displaying the lyrics of the first content item at the first location comprises displaying the representation of the lyrics of the first content item at a size different from the size of the extended user interface displayed when the second input is detected.

16. The method of claim 14, wherein when the second input is detected, the extended user interface and the playback control user interface including the second selectable user interface object are simultaneously displayed, the method further comprising: In response to receiving the second input, the representation of the lyrics of the first content item is displayed simultaneously with the playback control user interface; When the representation of the lyrics of the first content item is displayed simultaneously with the playback control user interface, a third input corresponding to a request to move the playback control user interface in the three-dimensional environment is received via the one or more input devices; as well as In response to receiving the third input: The playback control user interface is moved in the three-dimensional environment according to the third input; as well as The representation of the lyrics of the first content item being moved in the three-dimensional environment according to the third input.

17. A computer system communicating with a display generation component and one or more input devices, the computer system comprising: One or more processors; Memory; and One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods according to claims 1 to 16.

18. A non-transitory computer-readable storage medium storing one or more programs, said one or more programs comprising instructions that, when executed by one or more processors of a computer system in communication with a display generation component and one or more input devices, cause the computer system to perform any one of the methods according to claims 1 to 16.