Devices, methods, and graphical user interfaces for content applications
The system addresses inefficiencies in augmented and virtual reality interactions by utilizing touch-sensitive displays, eye-tracking, and hand-tracking, along with virtual lighting effects, to streamline user input and conserve power.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2025-12-26
- Publication Date
- 2026-05-26
AI Technical Summary
Existing methods for interacting with augmented and virtual reality environments are cumbersome, inefficient, and require excessive user input, leading to increased cognitive burden and energy consumption.
The system employs improved methods and interfaces that reduce the number and types of user inputs by using touch-sensitive displays, eye-tracking, hand-tracking, and voice input, along with virtual lighting effects and three-dimensional content navigation, to enhance user interaction and reduce power consumption.
This approach provides a more efficient and intuitive user interface, reducing the need for multiple inputs and minimizing power consumption, thereby enhancing the user experience and extending battery life.
Smart Images

Figure 2026086400000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 261,564, filed on September 23, 2021, the content of which is incorporated herein by reference in its entirety for all purposes.
[0002] This generally relates to a computer system having one or more input devices that present a graphical user interface, including, but not limited to, a display generation component and a display generation component that includes a user interface for presenting and browsing content through which a graphical user interface is presented.
Background Art
[0003] The development of computer systems for augmented reality has advanced significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch - sensitive surfaces, and touch - screen displays for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include virtual objects that include digital images, videos, text, icons, and control elements such as buttons and other graphics.
Summary of the Invention
[0004] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and restrictive. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in augmented reality environments, and systems where manipulating virtual objects is complex and error-prone impair the user's cognitive burden and detract from the virtual / augmented reality experience. In addition, these methods are unnecessarily time-consuming, thereby wasting energy. This latter consideration is particularly important in battery-powered devices.
[0005] Therefore, there is a need for computer systems with improved methods and interfaces to provide users with computer-generated experiences that make interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces can optionally complement or replace conventional methods of providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or types of user input by helping the user understand the connection between the inputs provided and the device response to those inputs, thereby generating a more efficient human-machine interface.
[0006] The above-mentioned defects and other problems relating to the user interface for a computer system having a display generation component and one or more input devices are mitigated or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, tablet computer, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generation component, the output devices include one or more tactile output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing multiple functions. In some embodiments, the user interacts with the GUI (and / or computer system) through stylus and / or finger touch and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the user's body as captured by a camera and other motion sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gameplay, making phone calls, video conferencing, sending emails, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing those functions optionally include temporary computer-readable storage media and / or non-temporary computer-readable storage media, or other computer program products configured to be executed by one or more processors.
[0007] There is a need for electronic devices having improved methods and interfaces for navigating user interfaces. Such methods and interfaces can complement or replace conventional methods for interacting with graphical user interfaces. Such methods and interfaces reduce the number, extent, and / or types of user input, resulting in a more efficient human-machine interface.
[0008] In some embodiments, the electronic device generates virtual lighting effects while presenting content items. In some embodiments, the electronic device enhances navigation to individual playback positions of content items. In some embodiments, the electronic device displays media content in a three-dimensional environment. In some embodiments, the electronic device presents media content in different presentation modes.
[0009] It should be noted that the various embodiments described herein can be combined with any other embodiments described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, in particular, in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention. [Brief explanation of the drawing]
[0010] To better understand the various embodiments described, the following “Modes for Carrying Out the Invention” should be referenced in conjunction with the following drawings, and similar reference numbers throughout the following drawings refer to the corresponding parts.
[0011] [Figure 1] This block diagram shows the operating environment of a computer system for providing an XR experience, according to several embodiments.
[0012] [Figure 2] Block diagram showing a controller for a computer system configured to manage and adjust the user's XR experience, according to several embodiments.
[0013] [Figure 3] This block diagram shows display generation components of a computer system configured to provide users with visual components of an XR experience, according to several embodiments.
[0014] [Figure 4] This is a block diagram showing a hand tracking unit for a computer system configured to capture user gesture input, according to several embodiments.
[0015] [Figure 5]A block diagram showing an eye-tracking unit of a computer system configured to capture a user's gaze input according to some embodiments.
[0016] [Figure 6A] A flowchart showing a grint-assisted gaze tracking pipeline according to some embodiments.
[0017] [Figure 6B] An exemplary environment of an electronic device for providing an XR experience according to some embodiments is shown.
[0018] [Figure 7A] An example of a method for generating a virtual lighting effect while an electronic device presents a content item according to some embodiments is shown. [Figure 7B] An example of a method for generating a virtual lighting effect while an electronic device presents a content item according to some embodiments is shown. [Figure 7C] An example of a method for generating a virtual lighting effect while an electronic device presents a content item according to some embodiments is shown. [Figure 7D] An example of a method for generating a virtual lighting effect while an electronic device presents a content item according to some embodiments is shown. [Figure 7E] An example of a method for generating a virtual lighting effect while an electronic device presents a content item according to some embodiments is shown.
[0019] [Figure 8A] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8B] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8C]A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8D] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8E] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8F] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8G] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8H] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8I] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8J] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8K] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8L] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8M] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8N] A flowchart showing a method for generating a virtual lighting effect while presenting a content item according to some embodiments. [Figure 8O] This flowchart shows a method for generating virtual lighting effects while presenting content items, according to several embodiments.
[0020] [Figure 9A] This document describes exemplary methods for displaying media content in a three-dimensional environment, according to several embodiments. [Figure 9B] This document describes exemplary methods for displaying media content in a three-dimensional environment, according to several embodiments. [Figure 9C] This document describes exemplary methods for displaying media content in a three-dimensional environment, according to several embodiments. [Figure 9D] This document describes exemplary methods for displaying media content in a three-dimensional environment, according to several embodiments. [Figure 9E] This document describes exemplary methods for displaying media content in a three-dimensional environment, according to several embodiments.
[0021] [Figure 10A] This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment. [Figure 10B] This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment. [Figure 10C] This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment. [Figure 10D] This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment. [Figure 10E] This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment. [Figure 10F] This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment. [Figure 10G]This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment. [Figure 10H] This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment. [Figure 10I] This flowchart shows several embodiments of methods for displaying media content in a three-dimensional environment.
[0022] [Figure 11A] This document presents examples of how electronic devices can improve navigation to individual playback positions of content items, based on several embodiments. [Figure 11B] This document presents examples of how electronic devices can improve navigation to individual playback positions of content items, based on several embodiments. [Figure 11C] This document presents examples of how electronic devices can improve navigation to individual playback positions of content items, based on several embodiments. [Figure 11D] This document presents examples of how electronic devices can improve navigation to individual playback positions of content items, based on several embodiments. [Figure 11E] This document presents examples of how electronic devices can improve navigation to individual playback positions of content items, based on several embodiments.
[0023] [Figure 12A] This flowchart shows methods for improving navigation to individual playback positions of content items, according to several embodiments. [Figure 12B] This flowchart shows methods for improving navigation to individual playback positions of content items, according to several embodiments. [Figure 12C]This flowchart shows methods for improving navigation to individual playback positions of content items, according to several embodiments.
[0024] [Figure 13A] This disclosure illustrates exemplary methods for presenting media content in immersive and non-immersive presentation modes, according to several embodiments of this disclosure. [Figure 13B] This disclosure illustrates exemplary methods for presenting media content in immersive and non-immersive presentation modes, according to several embodiments of this disclosure. [Figure 13C] This disclosure illustrates exemplary methods for presenting media content in immersive and non-immersive presentation modes, according to several embodiments of this disclosure. [Figure 13D] This disclosure illustrates exemplary methods for presenting media content in immersive and non-immersive presentation modes, according to several embodiments of this disclosure. [Figure 13E] This disclosure illustrates exemplary methods for presenting media content in immersive and non-immersive presentation modes, according to several embodiments of this disclosure.
[0025] [Figure 14A] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14B] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14C] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14D] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14E] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14F]This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14G] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14H] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14I] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Figure 14J] This flowchart shows methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments. [Modes for carrying out the invention]
[0026] This disclosure relates to user interfaces that provide a computer-generated (CGR) experience to a user, in several embodiments.
[0027] The systems, methods, and GUIs described herein provide improved methods for electronic devices to present content corresponding to physical locations indicated within navigation user interface elements.
[0028] In some embodiments, a computer system displays a content application containing content items in a three-dimensional environment. In some embodiments, an electronic device applies virtual lighting effects to the three-dimensional environment while displaying a content application containing content items. In some embodiments, the virtual lighting effects are based on the content items played through the content application (e.g., including the colors contained in the images associated with the content items). Presenting the content application user interface with virtual lighting provides the user with an immersive and less distracting experience while consuming content items, which further reduces power consumption and improves the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently.
[0029] In some embodiments, the computer system presents media content in a three-dimensional environment using different presentation modes, including an extended presentation mode and a picture-in-picture presentation mode. In some embodiments, the computer system updates the position and / or orientation of the media content in the three-dimensional environment as the user's viewpoint in the three-dimensional environment changes. In some embodiments, whether the computer system updates the position and / or orientation of the media content in the three-dimensional environment is based on the presentation mode associated with the media content when the computer system detects a shift in the user's viewpoint in the three-dimensional environment. Changing the pose and / or orientation of the media content as the user's viewpoint in the three-dimensional environment changes provides an efficient method for providing continuous access to the media content regardless of the user's current viewpoint in the three-dimensional environment, which further reduces power consumption and improves the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently.
[0030] In some embodiments, the computer system enhances navigation to individual parts of a content item. In some embodiments, while presenting a content item, the electronic device detects when the user's attention (e.g., gaze) is no longer directed towards the content item. In some embodiments, in response to detecting that the user's attention has been directed towards the content item after being away from it, the electronic device, when selected, presents a selectable option that causes the electronic device to navigate to an individual playback position of the content item associated with the playback position of the content item that was being played when the user's attention was diverted from the content item. Presenting an option to navigate to an individual playback position of a content item provides an efficient way to navigate the content item, which further reduces power consumption, improves the battery life of the electronic device, and reduces errors in use that would otherwise need to be corrected with further user input, by enabling the user to use the electronic device more quickly and efficiently.
[0031] In some embodiments, the computer system presents immersive and non-immersive media content in a three-dimensional environment. In some embodiments, the computer system presents immersive content in immersive and non-immersive presentation modes. In some embodiments, while the computer system is presenting immersive content in non-immersive presentation mode, the computer system displays, if selected, a selectable option to transition the presentation of immersive content from non-immersive presentation mode to immersive presentation mode. Providing a selectable option to transition the presentation of content from non-immersive to immersive presentation mode provides an efficient way to access different presentation modes associated with immersive content, which further reduces power consumption and improves the battery life of the electronic device by enabling the user to use the electronic device more quickly and efficiently.
[0032] Figures 1 to 6 provide a description of exemplary computer systems for providing users with an XR experience (as described below with respect to methods 800, 1000, 1200, and 1400). Figures 7A to 7E show exemplary techniques for generating virtual lighting effects while presenting content items, according to several embodiments. Figures 8A to 8O are flowcharts of methods for generating virtual lighting effects while presenting content items, according to several embodiments. Figures 9A to 9E show exemplary techniques for displaying media content in a three-dimensional environment, according to several embodiments. Figures 10A to 10I are flowcharts of methods for displaying media content in a three-dimensional environment, according to several embodiments. Figures 11A to 11E show exemplary techniques for improving navigation to individual playback positions of content items, according to several embodiments. Figures 12A to 12C are flowcharts of methods for improving navigation to individual playback positions of content items, according to several embodiments. Figures 13A to 13E show exemplary techniques for presenting media content in immersive and non-immersive presentation modes, according to several embodiments of the present disclosure. Figures 14A to 14J are flowcharts illustrating methods for presenting media content in immersive and non-immersive presentation modes according to several embodiments.
[0033] The processes described below enhance the usability of the device and streamline the user-device interface by providing users with improved visual feedback, reducing the number of inputs required to perform operations, offering additional control options without cluttering the user interface with additional controls displayed, performing operations without requiring further user input when a set of conditions is met, improving privacy and / or security, and / or other technologies. These technologies also reduce power consumption and improve the device's battery life by enabling users to use the device more quickly and efficiently.
[0034] Furthermore, in any method described herein that is conditional on one or more conditions being met in one or more steps, it should be understood that the method described can be repeated in multiple iterations such that all the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires that a first step be performed if a condition is met, and a second step be performed if the condition is not met, a person skilled in the art will understand that the steps described in the claim are repeated in a specific order until the conditions are met and then not met. Thus, a method described in one or more steps that depends on one or more conditions being met can be rewritten as a method that is repeated until each of the conditions described in the method is met. However, this is not required for a claim of a system or computer-readable medium that includes instructions for performing a conditional operation based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency has been met without explicitly repeating the steps of the method until all the conditions that the steps of the method are conditional on are met. Those skilled in the art will also understand that, as with a method having conditional steps, a system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.
[0035] In some embodiments, as shown in Figure 1, the XR experience is provided to the user via an operating environment 100 which includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or remote server), display generation components 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the display generation component 120 (for example, within a head-mounted device or handheld device).
[0036] When describing an XR experience, various terms are used to refer individually to several related but distinct environments that the user perceives and / or interacts with (for example, using inputs detected by the computer system 101, which causes the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101 that generates the XR experience). The following is a subset of these terms.
[0037] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the help of electronic systems. Examples of physical environments, such as a physical park, include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses of sight, touch, hearing, taste, and smell.
[0038] Augmented Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through an electronic system. In XR, a subset of a person's bodily movements or their representations are tracked, and accordingly, one or more properties of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics. For example, an XR system may detect a person's head rotation and, accordingly, adjust the graphic content and sound field presented to the person in a similar manner to how such views and sounds would change in a physical environment. In some circumstances (e.g., for reasons of accessibility), adjustments to the properties(s) of virtual objects(s) in the XR environment may be made in response to representations of bodily movements (e.g., voice commands). A person may perceive and / or interact with XR objects using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may perceive and / or interact with audio objects that create a 3D or spatially expansive audio environment, providing the perception of a point source in 3D space. In another example, audio objects may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may perceive and / or interact with audio objects only.
[0039] Examples of XR include virtual reality and mixed reality.
[0040] Virtual reality: A virtual reality (VR) environment refers to a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of their presence within the computer-generated environment and / or through a simulation of a subset of their physical movement within the computer-generated environment.
[0041] Mixed Reality: A mixed reality (MR) environment is a simulated environment designed to incorporate sensory input or its representation from a physical environment, in addition to including computer-generated sensory input (e.g., virtual objects), in contrast to a virtual reality (VR) environment designed to rely entirely on computer-generated sensory input. On a virtual continuum, a mixed reality environment is any place between, but not including, the complete physical environment at one end and the virtual reality environment at the other end. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Also, some electronic systems for presenting an MR environment may track location and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, a system may account for movement so that a virtual tree appears stationary relative to the physical ground.
[0042] Examples of mixed reality include augmented reality and augmented virtual reality.
[0043] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on or onto a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display that allows a person to directly view the physical environment. The system may also be configured to present virtual objects on the transparent or translucent display, thereby allowing a person to use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system composites the image or video with the virtual objects and presents the composite on the opaque display. A person uses this system to perceive the virtual objects superimposed on the physical environment by indirectly viewing the physical environment through the image or video of the physical environment. As used herein, a video of the physical environment shown on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into or onto the physical environment, so that a person can use the system to perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to an imitation environment in which the representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may transform one or more sensor images to plane a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be transformed by graphically modifying (e.g., enlarging) a portion of it, so that the modified portion is a non-photorealistic altered version of the original captured image. As yet another example, the representation of the physical environment may be transformed by graphically removing or obscuring a portion of it.
[0044] Augmented Virtual: An Augmented Virtual (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. These sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and virtual buildings, while people with faces are realistically reproduced from images of real people. Another example is that a virtual object might adopt the shape or color of a physical article captured by one or more imaging sensors. A further example is that a virtual object might adopt shadows that correspond to the position of the sun in the physical environment.
[0045] Viewpoint-locked virtual objects: A virtual object is viewpoint-locked when the computer system displays the virtual object in the same location and / or position within the user's view, even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked in the forward direction of the user's head (e.g., the user's viewpoint is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's viewpoint remains fixed even if the user's gaze moves, without moving the user's head. In embodiments where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object displayed in the upper-left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) will continue to be displayed in the upper-left corner of the user's viewpoint even if the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position in which a viewpoint-locked virtual object is displayed from the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so that the virtual object is also referred to as a "head-locked virtual object."
[0046] Environment-Locked Virtual Objects: A virtual object is environment-locked (or "world-locked") when a computer system displays it at a location and / or position in the user's viewpoint that is based on (e.g., selected by reference to and / or fixed to) a location and / or object in a three-dimensional environment (e.g., a physical or virtual environment). As the user's viewpoint shifts, the location and / or object in the environment relative to the user's viewpoint changes, and as a result, the environment-locked virtual object will appear at a different location and / or position in the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear centered in the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes left-leaning in the user's viewpoint (e.g., the tree's position in the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear left-leaning in the user's viewpoint. In other words, the location and / or position in which an environment-locked virtual object is displayed in the user's viewpoint depends on the location and / or object's position and / or orientation in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a fixed location in the physical environment and / or a coordinate system fixed to an object) to determine the position in which the environment-locked virtual object is displayed from the user's viewpoint. The environment-locked virtual object can be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object) or to a moving part of the environment (e.g., a vehicle, animal, person, or a representation of a part of the user's body that moves independently of the user's viewpoint, such as the user's hands, wrists, arms, or feet), so that the virtual object moves as the viewpoint or the part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.
[0047] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed tracking behavior, reducing or delaying its movement in response to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed tracking behavior, the computer system detects movement of the reference point that the virtual object is following (e.g., a part of the environment, a viewpoint, or a point fixed to the viewpoint, such as a point between 5 and 300 cm from the viewpoint) and intentionally delays the movement of the virtual object. For example, when the reference point (e.g., a part of the environment or the viewpoint) moves at a first velocity, the virtual object is moved by the device so as to remain locked to the reference point, but at a second velocity slower than the first velocity (e.g., the virtual object begins to catch up to the reference point until the reference point stops or slows down). In some embodiments, when a virtual object exhibits delayed tracking behavior, the device ignores small movements of the reference point (e.g., ignoring movements of the reference point that are below a threshold movement amount, such as a movement of 0 to 5 degrees or a movement of 0 to 50 cm). For example, when the reference point (e.g., the part of the environment or viewpoint from which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a different viewpoint or part of the environment from which the virtual object is locked), and when the reference point (e.g., the part of the environment or viewpoint from which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a different viewpoint or part of the environment from which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a "delayed tracking" threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object that maintains a substantially fixed position with respect to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or forward / behind the position of the reference point).
[0048] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to receive an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eye. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, the controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to Figure 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some embodiments, the controller 110 is communicably coupled to a display generation component 120 (e.g., an HMD, display, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within a housing (e.g., a physical housing) of one or more of the display generation components 120 (e.g., a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0049] In some embodiments, the display generation component 120 is configured to provide the user with an XR experience (e.g., at least the visual components of the XR experience). In some embodiments, the display generation component 120 includes a preferred combination of software, firmware, and / or hardware. The display generation component 120 is described in more detail below with reference to Figure 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.
[0050] According to some embodiments, the display generation component 120 provides the user with an XR experience while the user is virtually and / or physically present in the scene 105.
[0051] In some embodiments, the display generation component is mounted on a part of the user's body (e.g., their head or hand). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally placed in a housing mounted on the user's head. In some embodiments, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content when the user is not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interaction with XR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD where the interaction occurs in the space in front of the HMD and the XR content response is displayed through the HMD. Similarly, a user interface showing interaction with CRG content triggered based on the movement of a handheld or tripod-mounted device relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may be implemented similarly to an HMD where the movement is triggered by the movement of the HMD relative to the physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).
[0052] While relevant features of the operating environment 100 are shown in Figure 1, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the exemplary embodiments disclosed herein.
[0053] Figure 2 is a block diagram of an example of the controller 110 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0054] In some embodiments, one or more communication buses 204 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0055] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some embodiments, memory 220, or the non-temporary computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and XR experience module 240.
[0056] The operating system 230 handles various basic system services and includes instructions for performing hardware-dependent tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for each group of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a data transmission unit 248.
[0057] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1, and optionally from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0058] In some embodiments, the tracking unit 242 is configured to map scene 105 and track the position / location of at least the display generation component 120 relative to scene 105 in Figure 1, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for this purpose, as well as heuristics and metadata for this purpose. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand relative to scene 105 in Figure 1, relative to the display generation component 120, and / or relative to a coordinate system defined for the user's hand. The hand tracking unit 244 is described in more detail below with respect to Figure 4. In some embodiments, the eye-tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hands)) or to XR content displayed via the display generation component 120. The eye-tracking unit 243 is described in more detail below with reference to Figure 5.
[0059] In some embodiments, the adjustment unit 246 is configured to manage and adjust the XR experience presented to the user by the display generation component 120 and optionally by one or more of the output devices 155 and / or peripheral devices 195. For this purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0060] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 248 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0061] While the data acquisition unit 241, tracking unit 242 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 are shown as residing on a single device (e.g., a controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 may be located in separate computing devices.
[0062] Furthermore, Figure 2 is intended to illustrate the function of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules separately shown in Figure 2 can be realized in a single module, and the various functions of a single functional block can be realized by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0063] Figure 3 is a block diagram of an example of a display generation component 120 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. For that purpose, in some non-limiting examples, the display generation component 120 (e.g., HMD) may include one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, Bluetooth, ZiGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional in-facing and / or out-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0064] In some embodiments, one or more communication buses 304 include circuits for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).
[0065] In some embodiments, one or more XR displays 312 are configured to provide the user with an XR experience. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, one or more XR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, a display generation component 120 (e.g., HMD) includes a single XR display. In another example, a display generation component 120 (e.g., HMD) includes an XR display for each of the user's eyes. In some embodiments, one or more XR displays 312 can present MR or VR content. In some embodiments, one or more XR displays 312 can present MR or VR content.
[0066] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally, at least a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would view if a display generation component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.
[0067] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some embodiments, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and XR presentation module 340.
[0068] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to the user via one or more XR displays 312. For this purpose, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.
[0069] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1. For this purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0070] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. For this purpose, in various embodiments, the XR presentation unit 344 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0071] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (for example, a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate augmented reality) based on media content data. For this purpose, in various embodiments, the XR map generation unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0072] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral devices 195. For this purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0073] While the data acquisition unit 342, XR presentation unit 344, XR map generation unit 346, and data transmission unit 348 are shown as existing on a single device (e.g., the display generation component 120 in Figure 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, XR presentation unit 344, XR map generation unit 346, and data transmission unit 348 may be located in separate computing devices.
[0074] Furthermore, Figure 3 is intended to illustrate the functionality of various features that may be present in a particular implementation, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 3 can be realized within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0075] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (Figure 1) is controlled by the hand tracking unit 244 (Figure 2) to track the location / position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to the scene 105 of Figure 1 (e.g., relative to a part of the physical environment surrounding the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to the user's hand) in a defined coordinate system. In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0076] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures a hand image with sufficient resolution to allow for the distinction of fingers and their respective positions. The image sensor 404 can typically capture images of other parts of the user's body, or images of the entire body, and may have either a zoom function or a dedicated sensor with high magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures a 2D color video image of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.
[0077] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), which drives the display generation components 120 accordingly. For example, a user can interact with the software running on the controller 110 by moving their hand 406 to change the orientation of their hand.
[0078] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spot in the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a given reference plane at a specific distance from the image sensor 404. In this disclosure, it is assumed that the image sensor 404 defines a set of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to a z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on one or more cameras or other types of sensors.
[0079] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or the processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand within these depth maps. Based on a previous learning process, the software matches these descriptors against patch descriptors stored in the database 408 to estimate the hand pose in each frame. The pose typically includes the 3D location of the user's wrist and fingertips.
[0080] The software can also analyze the trajectory of the hand and / or fingers across multiple frames in a sequence to identify gestures. The pose estimation function described herein may be interleaved with the motion tracking function, so that patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to detect changes in pose that occur over the remaining frames. Pose, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify the image presented on the display generation component 120, or perform other functions, depending on the pose and / or gesture information.
[0081] In some embodiments, the gestures include air gestures. Air gestures are gestures detected without (or independently of) the user touching an input element that is part of a device (e.g., a computer system 101, one or more input devices 125, and / or a hand tracking device 140), and are based on detected movements of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including the user's body movement relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), the user's body movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the user's other hand, and / or the movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movements of a part of the user's body (e.g., a tap gesture including the movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body).
[0082] In some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures, as in some embodiments, performed by moving one or more of the user's fingers relative to other fingers or parts of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, an air gesture is a gesture detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device), and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture involving movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture involving rotation of a part of the user's body by a predetermined speed or amount).
[0083] In some embodiments where the input gesture is an air gesture (i.e., without physical contact with an input device that provides the computer system with information about which user interface element is the target of user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor over a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of user input (e.g., in the case of direct input, as described below). Thus, in implementations involving air gestures, the input gesture is the detected attention (e.g., gaze) to the user interface element in combination (e.g., simultaneously) with the movement of the user's fingers (one or more) and / or hand to perform pinch and / or tap input, as described in more detail below.
[0084] In some embodiments, input gestures directed towards a user interface object are performed directly or indirectly by reference to the user interface object. For example, user input is performed directly towards the user interface object in response to the user performing an input gesture with their hand at a position corresponding to the user interface object's position in a three-dimensional environment (e.g., determined based on the user's current viewpoint). In some embodiments, the input gesture is performed indirectly towards the user interface object according to the user performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in a three-dimensional environment, while detecting the user's attention (e.g., gaze) to the user interface object. For example, in the case of a direct input gesture, the user can direct their input towards the user interface object by initiating the gesture at or near a position corresponding to the user interface object's display position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm from the optional outer edge or optional central portion). In the case of indirect input gestures, the user can direct their input towards the user interface object by paying attention to the user interface object (for example, by gazing at the user interface object), and while paying attention to the options, the user initiates the input gesture (for example, at any position detectable by the computer system) (for example, at a position that does not correspond to the display position of the user interface object).
[0085] In some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch and tap inputs for interacting with virtual or mixed reality environments, as in some embodiments. For example, the pinch and tap inputs described later are performed as air gestures.
[0086] In some embodiments, a pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture involves moving two or more fingers of a hand to touch each other, i.e., including an optional interruption (e.g., within 0 to 1 second) immediately after the touch. A long pinch gesture that is an air gesture involves moving two or more fingers of a hand to touch each other for at least a threshold time amount (e.g., at least 1 second) before detecting an interruption of contact between them. For example, a long pinch gesture includes the user holding a pinch gesture (e.g., if two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected directly and consecutively (e.g., within a predetermined period of time) to each other. For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks contact between two or more fingers), and then performs a second pinch input within a predetermined period (e.g., within 1 second or 2 seconds) after releasing the first pinch input.
[0087] In some embodiments, an air gesture, a pinch-and-drag gesture, includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in relation to (e.g., after) a drag input that changes the user's hand position from a first position (e.g., a drag initiation position) to a second position (e.g., a resistance termination position). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers) to terminate the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together and touches them to each other, and then moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by the user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from the first position to the second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of the user's hands. For example, an input gesture includes two (e.g., or more) pinch inputs performed in relation to each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using the user's first hand, and a second pinch input performed using the other hand (e.g., a second hand of the user's hands) in relation to performing a pinch input using the first hand. In some embodiments, movement between the user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).
[0088] In some embodiments, a tap input performed as an air gesture (e.g., directed towards a user interface element) includes the movement of one or more of the user's fingers toward the user interface element, the movement of the user's hand toward the user interface element with the user's fingers (one or more) optionally extended toward the user interface element, a downward movement of the user's fingers (e.g., mimicking a mouse click or a tap on a touchscreen), or other default movements of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand that performs the tap gesture movement away from the user's viewpoint and / or toward the object that is the target of the tap input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand that performs the tap gesture (e.g., away from the user's viewpoint and / or the end of the movement toward the object that is the target of the tap input, a reversal of the direction of the finger or hand movement, and / or a reversal of the direction of acceleration of the finger or hand movement).
[0089] In some embodiments, the user's attention is determined to be directed towards a part of the three-dimensional environment based on the detection of a gaze directed towards that part of the three-dimensional environment (optionally, without requiring any other conditions). In some embodiments, for the device to determine that the user's attention is directed towards a part of the three-dimensional environment, the device determines that the user's attention is directed towards a part of the three-dimensional environment based on the detection of a gaze directed towards a part of the three-dimensional environment, with one or more additional conditions such as the gaze being directed towards the part of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the part of the three-dimensional environment, and / or the gaze being directed towards a part of the three-dimensional environment. If one of the additional conditions is not met, the device determines that the user's attention is not directed towards the part of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).
[0090] In some embodiments, the detection of a ready state configuration of the user or a part of the user is detected by the computer system. The detection of a ready state configuration of the hand is used by the computer system as an indication that the user is likely to be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap shape where one or more fingers are extended and the palm is facing away from the user), whether the hand is in a predetermined position relative to the user's line of sight (e.g., below the user's head, above the user's waist, or extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular way (e.g., moved towards the area in front of the user above the user's waist, below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of the user interface is responsive to attention (e.g., gaze) input.
[0091] In some embodiments, the software may be downloaded electronically to the controller 110, for example, over a network, or instead, it may be provided on a tangible non-temporary medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the computer's described functions may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although the controller 110 is shown in Figure 4, for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor (e.g., a hand-tracking device 402), or in other ways. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device), or by any other suitable computerized device such as a game console or media player. The sensing function of the image sensor 404 can also be integrated into a computer or other computerized device controlled by the sensor output.
[0092] Figure 4 further includes schematic diagrams of depth maps 410 captured by image sensor 404 according to several embodiments. The depth map includes a matrix of pixels, each having a depth value, as described above. Pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The brightness of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from image sensor 404, with the gradation becoming richer as the depth increases. Controller 110 processes these depth values to identify and segment image components (i.e., adjacent pixel groups) that have the characteristics of a human hand. These characteristics may include, for example, the overall size, shape, and frame-to-frame movement of the depth map sequence.
[0093] Figure 4 also schematically shows the hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to several embodiments. In Figure 4, the skeleton 414 is superimposed on the hand background 416, which has been segmented from the original depth map. In some embodiments, the hand (e.g., finger joints, fingertips, center of the palm, end of the hand connected to the wrist), and optionally major feature points on the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these major feature points across multiple image frames are used by the controller 110 to determine, according to several embodiments, a hand gesture performed by the hand or the current state of the hand.
[0094] Figure 5 shows an exemplary embodiment of the eye-tracking device 130 (Figure 1). In some embodiments, the eye-tracking device 130 is controlled by an eye-tracking unit 243 (Figure 2) to track the position and movement of the user's gaze toward the scene 105 or toward the XR content displayed via the display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device positioned in a wearable frame, the head-mounted device includes both a component for generating XR content for user viewing and a component for tracking the user's gaze toward the XR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or XR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used with a display generation component that is mounted on the head or a display generation component that is not mounted on the head. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally part of a non-head-mounted display generation component.
[0095] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames containing left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include, or be coupled to, one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects on a transparent or translucent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as holograms, so that the individual can use the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0096] As shown in Figure 5, in some embodiments, the eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) camera or a near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be directed toward the user's eye to receive reflected IR or NIR light from the light source directly from the eye, or alternatively, it may be directed toward a "hot" mirror positioned between the user's eye and a display panel that reflects IR or NIR light from the eye to the eye-tracking camera while allowing visual light to pass through. The eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a separate eye-tracking camera and light source.
[0097] In some embodiments, the eye-tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye-tracking device for a specific operating environment 100, e.g., the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at the factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automated calibration process or a manual calibration process. The user-specific calibration process may include estimating the eye parameters of a particular user, e.g., pupil location, central visual location, optical axis, visual axis, interpupillary distance. According to some embodiments, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, the images captured by the eye-tracking camera can be processed using a glint-assisted method to determine the user's current visual axis and viewpoint relative to the display.
[0098] As shown in Figure 5, the eye-tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and an eye-tracking system which includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes(s) 592. The eye-tracking camera 540 is positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display or projector of a handheld device) and may be directed towards a mirror 550 that transmits visible light while reflecting IR or NIR light from the eye(s) 592 (e.g., as shown at the top of Figure 5), or may be directed towards the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of Figure 5).
[0099] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames of left and right display panels) and provides the frames 562 to the display 510. For various purposes, for example, when processing the frames 562 for display, the controller 110 uses gaze tracking input 542 from the eye-tracking camera 540. The controller 110 optionally uses a glint-assisted method or other appropriate method to estimate the user's viewpoint on the display 510 based on the gaze tracking input 542 obtained from the eye-tracking camera 540. The viewpoint estimated from the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0100] The following describes, but is not intended to be limiting, several possible use cases of the user's current gaze direction. As an exemplary use case, the controller 110 may render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content within the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may capture the physical environment of the XR experience and orient an external camera to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently viewing on the display 510. In another exemplary use case, the eyepiece 520 may be a focusing lens, and the controller uses eye-tracking information to adjust the focus of the eyepiece 520 so that the virtual object currently being viewed by the user has appropriate binocular coordination to match the convergence of the user's eye 592. The controller 110 can use the eye-tracking information to orient and adjust the focus of the eyepiece 520 so that the nearby object being viewed by the user appears at the correct distance.
[0101] In some embodiments, the eye-tracking device is part of a head-mounted device mounted on a wearable housing, which includes a display (e.g., display 510), two eyepieces (e.g., one or more eyepieces 520), an eye-tracking camera (e.g., one or more eye-tracking cameras 540), and a light source (e.g., a light source 530 (e.g., an IR LED or NIR LED)). The light source emits light (e.g., IR light or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in Figure 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.
[0102] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, thus not introducing noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given as examples and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is positioned on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0103] Embodiments of eye-tracking systems, such as those shown in Figure 5, can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0104] Figure 6A shows a glint-assisted eye-tracking pipeline according to several embodiments. In some embodiments, the eye-tracking pipeline is implemented by a glint-assisted eye-tracking system (e.g., an eye-tracking device 130 as shown in Figures 1 and 5). The glint-assisted eye-tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the glint-assisted eye-tracking system tracks the pupil contour and glint in the current frame by using prior information from previous frames when analyzing the current frame. When not in tracking state, the glint-assisted eye-tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in tracking state for the next frame.
[0105] As shown in Figure 6A, the eye-tracking camera can capture left and right images of the user's left and right eyes. The captured images are then fed into the eye-tracking pipeline for processing, which begins at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0106] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as shown in 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If they are not successfully detected, the method returns to element 610 and processes the next image of the user's eyes.
[0107] At 640, if the process proceeds from element 610, the current frame is analyzed to track the pupil and glint based in part on previous information from the previous frame. At 640, if the process proceeds from element 630, the tracking state is initialized based on the detected pupil and glint in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results may be checked to determine whether a sufficient number of glints for pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are unreliable, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze.
[0108] Figure 6A is intended to serve as an example of an eye-tracking technology that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking technologies that currently exist or may be developed in the future may be used in computer system 101 to provide users with XR experiences in various embodiments, either in place of or in combination with the Glint-assisted eye-tracking technology described herein.
[0109] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with an XR experience, for example, a mixed reality environment in which one or more virtual objects are superimposed on a representation of the real-world environment 602.
[0110] Figure 6B shows an exemplary environment for an electronic device 101 to provide an XR experience, according to several embodiments. In Figure 6B, the real-world environment 602 includes the electronic device 101, a user 608, and real-world objects (e.g., a table 604). As shown in Figure 6B, the electronic device 101 is optionally mounted on a tripod or otherwise fixed to the real-world environment 602 such that one or more of the user 608's hands are free (e.g., the user 608 is not optionally holding the device 101 with one or more hands). As described above, the device 101 optionally has one or more groups of sensors located on different sides of the device 101. For example, the device 101 optionally includes sensor groups 612-1 and 612-2 located on the "rear" and "front" sides of the device 101, respectively (e.g., information can be captured from each side of the device 101). As used herein, the front side of device 101 is the side facing user 608, and the rear side of device 101 is the side facing away from user 608.
[0111] In some embodiments, the sensor group 612-2 includes an eye-tracking unit (e.g., the eye-tracking unit 245 described above with reference to Figure 2) which includes one or more sensors for tracking the user's eyes and / or gaze, and the eye-tracking unit can "look" at user 608 and track user 608's eyes (one or more) in the manner described above. In some embodiments, the eye-tracking unit of device 101 can capture the movement, orientation, and / or gaze of user 608's eyes and process the movement, orientation, and / or gaze as input.
[0112] In some embodiments, the sensor group 612-1 includes a hand tracking unit (e.g., the hand tracking unit 243 described above with reference to Figure 2) that can track one or more hands of user 608 held on the "rear" side of device 101, as shown in Figure 6B. In some embodiments, a hand tracking unit is optionally included in sensor group 612-2 so that user 608 can additionally or alternatively hold one or more hands on the "front" side of device 101 while device 101 tracks the position of one or more hands. As described above, the hand tracking unit of device 101 can capture the movement, position, and / or gestures of one or more hands of user 608 and process the movement, position, and / or gestures as input.
[0113] In some embodiments, the sensor group 612-1 optionally includes one or more sensors (e.g., the image sensor 404 described above with reference to Figure 4) configured to capture images of the real-world environment 602, including the table 604. As described above, the device 101 can capture images of a portion (e.g., part or all) of the real-world environment 602 and present the captured portion of the real-world environment 602 to the user via one or more display generating components of the device 101 (e.g., the display of the device 101 optionally located on the user-facing side of the device 101, opposite to the side of the device 101 facing the captured portion of the real-world environment 602).
[0114] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with an XR experience, for example, a mixed reality environment in which one or more virtual objects are superimposed on a representation of the real-world environment 602.
[0115] Accordingly, this description describes several embodiments of three-dimensional environments (e.g., XR environments) that include representations of real-world objects and virtual objects. For example, a three-dimensional environment optionally includes a representation of a table existing in a physical environment, which is captured and displayed within the three-dimensional environment (e.g., actively via a computer system's camera and display, or passively via a computer system's transparent or translucent display). As previously stated, a three-dimensional environment optionally is a mixed reality system based on a physical environment in which the three-dimensional environment is captured by one or more sensors of a device and displayed via a display generation component. As a mixed reality system, a computer system can optionally selectively display parts and / or objects of the physical environment such that each part and / or object of the physical environment appears to exist in the three-dimensional environment displayed by the electronic device. Similarly, a computer system can optionally display virtual objects in a three-dimensional environment such that the virtual objects appear to exist in the real world (e.g., the physical environment) by placing virtual objects in each location within the three-dimensional environment that has a corresponding location in the real world. For example, the computer system may optionally display a vase so that it appears as if a real vase were placed on a table in a physical environment. In some embodiments, each location in the three-dimensional environment has a corresponding location in the physical environment. Therefore, when the computer system is described as displaying a virtual object in a separate location relative to a physical object (e.g., the user's hand or its vicinity, or a physical table or its vicinity), the computer system displays the virtual object in a specific location in the three-dimensional environment so that it appears as if the virtual object is in or near the physical object in the physical world (for example, the virtual object is displayed in a location in the three-dimensional environment that corresponds to the location in the physical environment where the virtual object would be displayed if it were a real object at that particular location).
[0116] In some embodiments, real-world objects existing in a physical environment displayed within a three-dimensional environment (e.g., real-world objects that can be seen through and / or display-generating components) can interact with virtual objects that exist only within the three-dimensional environment. For example, the three-dimensional environment may include a table and a vase placed on the table, where the table is a view (or representation) of a physical table in the physical environment, and the vase is a virtual object.
[0117] Similarly, just as virtual objects are real objects in a physical environment, the user can optionally interact with virtual objects in a three-dimensional environment using one or more hands. For example, as described above, one or more sensors in the computer system can optionally capture one or more of the user's hands and display a representation of the user's hands in a three-dimensional environment (in a similar manner to, for example, displaying real-world objects in a three-dimensional environment as described above), or, in some embodiments, the user's hands can be seen through the display-generating components, due to the ability to see the physical environment through the user interface, due to the transparency / transparency of some of the display-generating components displaying the user interface, or the projection of the user interface onto a transparent / translucent surface, or the projection of the user interface onto the user's eyes or field of view. Thus, in some embodiments, the user's hands are displayed at separate locations in the three-dimensional environment and are treated as if they were objects in a three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were actual physical objects in the physical environment. In some embodiments, the computer system can update the display of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.
[0118] In some of the embodiments described below, for example, to determine whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grasping, or holding a virtual object, or whether it is within a threshold distance from the virtual object), the computer system may optionally determine the "effective" distance between the physical object in the physical world and the virtual object in the three-dimensional environment. For example, a hand directly interacting with a virtual object may optionally include one or more of the fingers of a hand pressing a virtual button, a user's hand grasping a virtual vase, two fingers of a user's hand pinching / holding an application's user interface together, and other types of interactions described herein. For example, when determining whether a user is interacting with a virtual object and / or how a user is interacting with a virtual object, the computer system may optionally determine the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of the virtual object of interest in the three-dimensional environment. For example, one or more of the user's hands are positioned in a specific location in the physical world, which the computer system optionally captures and displays at a specific corresponding location in a three-dimensional environment (e.g., the position in the three-dimensional environment where the hands are displayed, if the hands are virtual hands rather than physical hands). The position of the hands in the three-dimensional environment is optionally compared to the position of a target virtual object in the three-dimensional environment to determine the distance between the user's one or more hands and the virtual object. In some embodiments, the computer system optionally determines the distance between a physical object and a virtual object by comparing the position in the physical world (as opposed to comparing the position in the three-dimensional environment).For example, when determining the distance between one or more of a user's hands and a virtual object, the computer system optionally determines the corresponding location of the virtual object in the physical world (e.g., the position in the physical world where the virtual object is located as if it were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and one or more of the user's hands. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Thus, when determining whether a physical object is in contact with a virtual object, or whether a physical object is within a threshold distance of a virtual object, as described herein, the computer system optionally performs one of the techniques described above to map the location of the physical object to a three-dimensional environment and / or to map the location of the virtual object to a physical environment.
[0119] In some embodiments, the same or similar techniques are used to determine where and what the user's gaze is directed, and / or where and what the physical stylus held by the user is directed. For example, if the user's gaze is directed to a particular position in the physical environment, the computer system optionally determines the corresponding position in the three-dimensional environment (e.g., the virtual position of the gaze), and if a virtual object is located at that corresponding virtual position, the computer system optionally determines that the user's gaze is directed to that virtual object. Similarly, the computer system optionally determines, based on the orientation of the physical stylus, where in the physical environment the stylus is pointing. In some embodiments, based on this determination, the computer system optionally determines the corresponding virtual position in the three-dimensional environment corresponding to the location in the physical environment that the stylus is pointing to, and optionally determines that the stylus is pointing to the corresponding virtual position in the three-dimensional environment.
[0120] Similarly, embodiments described herein may refer to the location of a user (e.g., a user of a computer system) and / or the location of a computer system in a three-dimensional environment. In some embodiments, the user of a computer system is holding, wearing, or otherwise positioned near the computer system. Thus, in some embodiments, the location of the computer system is used as a proxy for the user's location. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to individual locations in the three-dimensional environment. For example, if a user stands at a location facing an individual part of the physical environment displayed by a display-generating component, the location of the computer system is the location in the physical environment (and its corresponding location in the three-dimensional environment) where the user will see objects in the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) as the objects are displayed by the display-generating component of the computer system in the three-dimensional environment. Similarly, if a virtual object displayed in a three-dimensional environment is a physical object in a physical environment (for example, the virtual object is located in the same physical environment as it is in the three-dimensional environment, and has the same size and orientation as it does in the three-dimensional environment), then the computer system and / or user's location is the position from which the user views the virtual object in the physical environment in the same position, orientation, and / or size (for example, absolutely, and / or relative to each other, and in relation to real-world objects) as it was displayed by the computer system's display generation components in the three-dimensional environment.
[0121] This disclosure describes various input methods for interaction with computer systems. Where one example is provided using one input device or method, and another example is provided using a different input device or method, each example may be compatible with the input device or method described in the other example, and their use should be considered optional. Similarly, various output methods for interaction with computer systems are described. Where one example is provided using one output device or method, and another example is provided using a different output device or method, each example may be compatible with the output device or method described in the other example, and their use should be considered optional. Similarly, various methods for interaction with virtual or mixed reality environments via computer systems are described. Where one example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, each example may be compatible with the method described in the other example, and their use should be considered optional. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.
[0122] User interface and related processes Here, we focus on embodiments of a user interface ("UI") and related processes that may be performed in a computer system such as a portable multifunction device or head-mounted device, which comprises display generation components, one or more input devices, and (optionally) one or more cameras.
[0123] Figures 7A to 7E illustrate examples of methods, according to several embodiments, for generating virtual lighting effects while an electronic device presents content items.
[0124] Figure 7A shows an electronic device 101 that displays a three-dimensional environment 702 via a display generation component 120. In some embodiments, it should be understood that the electronic device 101 utilizes one or more techniques described with reference to Figures 7A to 7E in a two-dimensional environment without departing from the scope of this disclosure. As described above with reference to Figures 1 to 6, the electronic device 101 optionally includes a display generation component 120 (e.g., a touchscreen) and a plurality of image sensors 314. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the display generation component 120 is a touchscreen that can detect the user's hand gestures and movements. In some embodiments, the user interface described below can also be implemented in a head-mounted display that includes a display generating component for displaying the user interface to the user, and sensors for detecting the physical environment and / or the movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward towards the user's face).
[0125] In Figure 7A, the electronic device 101 displays a content item 704 in a three-dimensional environment 702. The three-dimensional environment 702 further includes representations of real objects in the physical environment of the electronic device 101, including a representation of a table 706a, a representation of a sofa 706b, a representation of a wall 708a, and a representation of a ceiling 708b. The electronic device 101 detects, via one or more sensors 314, the user's gaze 713a directed towards the content item 704, which is video content currently being played on the electronic device 101. In some embodiments, since the user's gaze 713a is directed towards the content item 704 while the content item 704 is being played, the electronic device 101 displays the three-dimensional environment 702 using one or more virtual lighting effects.
[0126] In some embodiments, the virtual lighting effect generated by the electronic device 101 includes visually highlighting the content item by blurring and / or darkening portions of the three-dimensional environment 702 that do not contain the content item 704, and by displaying a virtual light spill emanating from the content item 704. The virtual light spill includes virtual lighting 710a displayed on a wall representation 708a, virtual lighting 710b displayed on a ceiling representation 708b, virtual lighting 710c displayed on a table representation 706a, and virtual lighting 710d displayed on a sofa representation 706b. In some embodiments, the virtual lighting 710a-d is based on the video content of the content item 704. For example, the color, intensity, etc. of the virtual lights 710a-d are optionally based on the color, intensity, etc. of the video content of content item 704 in order to simulate that the virtual lights 710a-d are reflections of the video content of content item 704 and / or that the virtual lights 710a-d are radiating from content item 704 on various surfaces in the three-dimensional environment 702. In some embodiments, the sizing of the virtual lights 710a, 710b, 710c, and 710d is based on the distance from the individual surfaces of the virtual lighting effect to content item 704, and the positions of the virtual lights 710a, 710b, 710c, and 710d are based on the position of content item 704 in the three-dimensional environment 702.
[0127] In Figure 7A, the user's hand 703a is in a position that prevents the electronic device 101 from displaying one or more selectable options for controlling the playback of content item 704, as will be explained in more detail below with reference to Figures 7B to 7E. Therefore, the electronic device 101 stops displaying the selectable options for controlling the playback of content item 704 in Figure 7A.
[0128] Figure 7B shows the display of several selectable options 712a–L for controlling the playback of content item 704, and user input corresponding to a request to resize content item 704 in a three-dimensional environment 702. In some embodiments, the electronic device 101 displays the selectable options 712a–L in response to detecting the user's hand 703b in a ready-state pose while the user's line of sight 713a is directed at content item 704. In some embodiments, detecting the hand 703b in a ready-state pose includes detecting the hand 703b in a pre-pinch hand shape where the thumb is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, or 5 centimeters) of another finger of the hand 703b but not touching another finger of the hand 703b, or detecting the hand 703b in a pointing hand shape where one or more fingers are extended and one or more fingers are curled toward the palm.
[0129] In Figure 7B, several selectable options 712c-L for controlling the playback of content item 704 are displayed on a separate user interface element 711 from content item 704. In some embodiments, the user interface element 711 is angled toward the user's viewpoint at a different angle than the angle at which content item 704 is displayed relative to the user's viewpoint, as will be described in more detail below with reference to Method 1400. For example, the user interface element 711 is displayed at a lower height than content item 704 in the three-dimensional environment 702, so the electronic device 101 orients the user interface element 711 upwards relative to the angle of content item 704. As shown in Figure 7B, a portion of the user interface element 711 is visually or spatially overlaid on content item 704, and the user interface element 711 is optionally displayed closer to the user's viewpoint in the three-dimensional environment 702 than content item 704.
[0130] Next, selectable options 712c to L included in the user interface element 711 will be described. Option 712c, when selected, causes the electronic device 101 to display content item 704 in immersive content mode according to one or more steps of method 1400. Option 712d, when selected, causes the electronic device 101 to display content item 704 in a picture-in-picture element according to one or more steps of method 1000. Option 712e, when selected, causes the electronic device 101 to present a content playback queue containing one or more content items configured for the electronic device 101 to play next (for example, in a user interface element separate from content item 704 and / or instead of content item 704). Option 712f, when selected, causes the electronic device 101 to adjust the playback position of content item 704 back by a predetermined amount (for example, 5, 10, 15, 30, or 60 seconds). Option 712g, when selected, causes the electronic device 101 to pause playback of content item 704. In some embodiments, while content item 704 is paused, the electronic device 101 ceases displaying content item 704 with greater visual emphasis on the rest of the three-dimensional environment. In some embodiments, while content item 704 is paused, the electronic device 101 ceases displaying virtual lighting effects, including blurring and / or dimming the three-dimensional environment 702 other than content item 704, and / or displaying virtual light spills emanating from content item 704. Option 712h, when selected, causes the electronic device 101 to adjust the playback position of content item 704 forward by a predetermined amount (e.g., 5, 10, 15, 30, or 60 seconds). Option 712i, when selected, causes the electronic device 101 to present subtitle options associated with content item 704. Option 712j, when selected, causes the electronic device 101 to adjust virtual lighting effects such as blurring and / or dimming and / or light spill effects, as will be described in more detail below with reference to Figures 7C to 7E.Option 712k, when selected, causes the electronic device 101 to adjust the playback volume of the audio content contained in content item 704. User interface element 711 further includes a scrub bar 712L that includes an indicator of the current playback position of content item 704, and in response to an input that moves the current playback position indicator, causes the electronic device 101 to adjust the playback position of content item 704 and resume playback of content item 704 from the adjusted playback position.
[0131] In addition to the selectable options 712c to 712L displayed on the user interface element 711, the electronic device 101 further displays, when selected, selectable option 712a which causes the electronic device 101 to stop displaying content item 704 (and optionally, user interface element 711 and options 712a and 712b), and selectable option 712b which causes the electronic device 101 to resize content item 704 in the three-dimensional environment 702 when the electronic device 101 detects input directed towards it. Selectable option 712a for closing content item 704 is displayed as an overlay on content item 704 outside the user interface element 711. Selectable option 712b for resizing content item 704 is displayed outside content item 704 and user interface element 711.
[0132] As shown in Figure 7B, the electronic device 101 receives input provided by hand 703a and gaze 713b directed to selectable option 712b for resizing content item 704. In some embodiments, detecting input includes detecting that the user performs a selectable pose with hand 703a, including a predetermined hand shape such as a pinch hand shape where the thumb of hand 703a touches another finger of hand, or a pointing hand shape where one or more fingers of hand 703a are extended and one or more fingers of hand 703a are bent toward the palm of hand 703a while gaze 713b is directed toward option 712b. In some embodiments, while maintaining the predetermined hand shape, the user moves their hand, and in response to detecting the movement, the electronic device 101 resizes the content item 704 according to the movement of hand 703a (e.g., speed, duration, distance, direction, etc.). As will be explained in more detail below with reference to Figure 7C, when the electronic device 101 resizes the content item 704 according to the input provided by the hand 703a and gaze 713b, it does not resize element 711.
[0133] Figure 7C shows how the electronic device 101 resizes the content item 704 in response to the input shown in Figure 7B without changing the size of the user interface element 711. In response to the input shown in Figure 7B, the electronic device 101 displays the content item 704 in Figure 7C at a smaller size than the size in Figure 7B, while maintaining the size of the user interface element 711. The electronic device 101 also updates one or more properties (e.g., size) of the virtual lighting 710a-710c in accordance with the updated size of the content item 704, for example, by reducing the size of the virtual lighting 710a-710c in the three-dimensional environment 702 to correspond to the reduced size of the content item 704. In Figure 7C, the electronic device 101 optionally continues to display the user interface element 711 and other selectable options in response to detecting the hand 703b which is in a ready state as described above.
[0134] Figure 7C also shows an electronic device 101 that displays a user interface element 714 indicating the functionality of the virtual lighting effect option 712j, depending on whether the user's gaze 713d is directed towards option 712j. In some embodiments, the electronic device 101 displays the user interface element 714 associated with the lighting effect option 712j because the lighting effect option 712j is particularly relevant to displaying content items in a three-dimensional environment 702. As shown in Figure 7C, if the user's gaze 713c is instead directed towards option 712f to adjust the playback position of the content item back by a predetermined amount, the electronic device 101 optionally discontinues displaying the user interface element associated with option 712f because option 712f is less relevant to presenting content items in a three-dimensional environment than to presenting content items in other environments or user interfaces. In some embodiments, a first set of interactive elements within the user interface element 711 are associated with a visual indication similar to the indication 714, while a second set of interactive elements within the user interface element 711 are not associated with a visual indication similar to the indication 714.
[0135] In Figure 7C, the electronic device 101 detects input provided by a hand 703a and gaze 713a directed towards the content item 704. In some embodiments, the input corresponds to a request to update the position of the content item in the three-dimensional environment 702. In some embodiments, the electronic device 101 displays a user interface element other than the content item itself, which, when selected, causes the electronic device 101 to initiate the process of repositioning the content item 704 in the three-dimensional environment 702. In some embodiments, the repositioning user interface element, as well as the resizing user interface element 714, is displayed outside the content item 704 and outside the user interface element 711. In some embodiments, the repositioning user interface element is a horizontal bar or line aligned along the bottom of the user interface element 711. In some embodiments, the repositioning user interface element is either contained within the user interface element 711 or overlaid on the content item 704 without being contained within the user interface element 711.
[0136] In some embodiments, detecting input corresponding to a request to rearrange a content item 704 within a three-dimensional environment 702 includes detecting that the user has made a default hand shape with their hand 703a, such as the pinch hand shape or pointing hand shape described above, while detecting a line of sight 713a directed towards the content item 704. In some embodiments, while the user is making a default hand shape, the electronic device 101 detects the movement of the hand 703a and, accordingly, moves the content item 704 and user interface elements 711 according to the movement of the hand 703a (e.g., distance, duration, speed, direction, etc.), as shown in Figure 7D.
[0137] Figure 7D shows an electronic device 101 that displays a content item 704 and a user interface element 711 in an updated location within a three-dimensional environment 702, according to the input shown in Figure 7C. According to the input shown in Figure 7C, the electronic device 101 displays the content item 704 and the user interface element 711 in Figure 7D in a location within the three-dimensional environment 702 that is closer to the user's viewpoint from which the three-dimensional environment 702 is displayed, than the location of the content item 704 and the user interface element 711 in Figure 7C. By displaying the content item 704 closer to the user's viewpoint in Figure 7D, the electronic device 101 displays the content item 704 in a larger angular size in Figure 7D than in Figure 7C (for example, occupying more space within the field of view of the user and / or display generation component 120), however, in some embodiments, the size of the content item 704 in the three-dimensional environment 702 is the same in Figure 7C and Figure 7D (for example, the size of the content item 704 in the three-dimensional environment 702 does not change from Figure 7C to Figure 7D). As shown in Figures 7C and 7D, the electronic device 101 does not update the angular size of the user interface element 711, even if the user interface element 711 is closer to the user's viewpoint in Figure 7D than in Figure 7C. In some embodiments, discontinuing the angular size update of the user interface element 711 includes updating the size of the user interface element 711 in the three-dimensional environment 702. For example, if the user interface element 711 is closer to the user's viewpoint in Figure 7D than in Figure 7C, but is displayed at the same angular size in both Figures 7C and 7D, the electronic device 101 reduces the size of the user interface element 711 in the environment 702 in Figure 7D compared to the size of the user interface element 711 in the environment 702 in Figure 7C. As shown in Figure 7D, in accordance with the movement of the content item 704, the electronic device 101 updates the positions of the virtual lights 710a, 710b, and 710d on the rest of the three-dimensional environment 702.
[0138] In Figure 7D, the electronic device 101 detects input directed at the line of sight 713e and the hand 703a to the lighting effect option 712j. In some embodiments, detecting input includes detecting the line of sight 713e directed at option 712j while detecting that the hand 703a forms the pinch hand shape or the pointing hand shape described above. In some embodiments, in response to detecting a pinch hand shape or pointing hand shape less than a predetermined time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds), the electronic device 101 switches the lighting effect (e.g., blurring and / or darkening lighting effect, and / or light spill lighting effect) on or off. In some embodiments, in response to detecting a pinch hand shape or pointing hand shape exceeding the time threshold, the electronic device 101 presents a slider element 716 that, when interacted with, allows the user to adjust the level (e.g., intensity) of the lighting effect. In some embodiments, both blur and / or dimming lighting effects and light spill lighting effects are updated according to inputs directed to lighting effect options 712j and / or slider 716. In some embodiments, the three-dimensional environment 702 includes separate elements for adjusting each lighting effect. As shown in Figure 7D, the electronic device 101 detects the user's line of sight 713e directed to slider 716 and also detects the movement of the hand 703a while the hand 703a is pinching or making a pointing hand shape. In response to the inputs shown in Figure 7E, the electronic device 101 reduces the intensity of the blur and / or dimming lighting effects and light spill lighting effects, as shown in Figure 7E.
[0139] Figure 7E shows the three-dimensional environment 702 with reduced visual emphasis on the content item 704 relative to the rest of the three-dimensional environment 702, for example, with reduced intensity of blurring and / or darkening lighting effects and light spill effects. For example, in Figure 7E, the amount of darkening and / or blurring applied to areas of the three-dimensional environment 702 other than the content item 704 is reduced, and the size and / or intensity of virtual illuminations 710a, 710b, and 710d are reduced. In some embodiments, the electronic device 101 reduces the intensity of virtual illumination effects in response to the input shown in Figure 7D. In some embodiments, the electronic device 101 reduces the intensity of virtual illumination effects (or stops displaying them) in response to detecting a user's line of sight 713f directed away from the content item 704. In some embodiments, the electronic device reduces the intensity of virtual illumination effects (or stops displaying them) in response to detecting that the content item has been paused. Figure 7E also shows an electronic device 101 that stops displaying the selectable options 712a to 712L and the user interface element 711 in response to no longer detecting the user's hand in a ready state.
[0140] Additional or alternative details relating to the embodiments shown in Figures 7A to 7E are provided below in the description of Method 800, which is described with reference to Figures 8A to 8O.
[0141] Figures 8A to 8O are flowcharts illustrating methods for generating virtual lighting effects while presenting content items, according to several embodiments. In some embodiments, Method 800 is performed in a computer system (e.g., computer system 101 in Figure 1) which includes a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing downwards from the user's hands (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, Method 800 is stored in a non-temporary computer-readable storage medium and controlled by instructions executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., control unit 110 in Figure 1A). Some operations of Method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0142] In some embodiments, such as Figure 7A, the method 800 is performed in an electronic device (e.g., 101) that communicates with a display generation component (e.g., 120) and one or more input devices (e.g., 314) (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, or television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users. In some embodiments, the one or more input devices include an electronic device or component that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include touchscreens, mice (e.g., external), trackpads (optionally integrated or external), touchpads (optionally integrated or external), remote control devices (e.g., external), another mobile device (e.g., separate from the electronic device), handheld devices (e.g., external), controllers (e.g., external), cameras, depth sensors, eye-tracking devices, and / or motion sensors (e.g., hand-tracking devices, hand motion sensors). In some embodiments, the electronic device communicates with the hand-tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreen, trackpad)). In some embodiments, the hand-tracking device is a wearable device such as a smart glove. In some embodiments, the hand-tracking device is a handheld input device such as a remote control or stylus.
[0143] In some embodiments, such as Figure 7B, while presenting a content item (e.g., 704) within a three-dimensional environment (e.g., 702), an electronic device (e.g., 101) displays a user interface (e.g., 711) associated with the content item via a display generation component (e.g., 120) (802a), the user interface (e.g., 711) includes one or more user interface elements (e.g., 712f) for modifying the playback of the content item and separate user interface elements (e.g., 712j) for modifying virtual lighting effects affecting the appearance of the three-dimensional environment (e.g., 702). In some embodiments, the content item is video content, and the one or more user interface elements for modifying the playback of the content item include play / pause, skip ahead, skip back, subtitles, and audio options. In some embodiments, the one or more user interface elements include options for displaying the content item in a picture-in-picture user interface element according to one or more steps of Method 1000, or for displaying the content item in an immersive (e.g., fullscreen) mode according to one or more steps of Method 1400.
[0144] In some embodiments, a content item is a video content item such as a movie, an episode in a series of episodic content, or a video clip, being played / displayed within a three-dimensional environment, or a content item is an audio content item such as music, a podcast, or an audiobook being played within a three-dimensional environment. In some embodiments, the three-dimensional environment includes virtual objects such as application windows, operating system elements, representations of other users, and / or representations of content items and physical objects in the physical environment of the electronic device. In some embodiments, representations of physical objects are displayed within the three-dimensional environment via a display-generating component (e.g., virtual or video passthrough). In some embodiments, a representation of a physical object is a view of the physical object in the physical environment of the electronic device seen through the transparency of the display-generating component (e.g., true passthrough or real passthrough). In some embodiments, the electronic device displays the three-dimensional environment from the user's viewpoint at a location in the three-dimensional environment corresponding to the physical location of the electronic device in the physical environment of the electronic device. In some embodiments, the three-dimensional environment is generated, displayed, or otherwise made visible by a device (for example, a computer-generated reality (XR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment).
[0145] In some embodiments, such as those shown in Figure 7D, while displaying a user interface (e.g., 711) associated with a content item (e.g., 704), an electronic device (e.g., 101) receives user input (802b) directed to individual user interface elements (e.g., 716) via one or more input devices, the user input responding to requests to modify virtual lighting effects. In some embodiments, the input responds to requests to change the amount of virtual lighting effects to a different (e.g., non-zero) amount. In some embodiments, the input responds to requests to display a three-dimensional environment without virtual lighting effects. In some embodiments, the virtual lighting effects are content item-based, such as light spill effects on representations of virtual and / or real objects in a three-dimensional environment having color, patterns, and / or motion based on the content item's image and / or video content. In some embodiments, the virtual lighting effects are independent of the content item, such as the degree of dimming and / or blurring applied to areas of the three-dimensional environment other than the user interface, including one or more user interface elements for modifying the content item or its reproduction (described later).
[0146] In some embodiments, such as Figure 7E, in response to receiving user input (802c), the electronic device (e.g., 101) continues to present content items (e.g., 704) within a three-dimensional environment (e.g., 702) (802d).
[0147] In some embodiments, such as Figure 7E, in response to receiving user input (802c), an electronic device (e.g., 101) applies a virtual lighting effect to a three-dimensional environment (e.g., 702) (802e). For example, in response to a request to display a three-dimensional environment without using a dimming virtual lighting effect, the electronic device displays the three-dimensional environment such that all areas of the three-dimensional environment, including areas where content is presented and areas where no content is presented, have the same relative dimming level. As another example, in response to a request to display a three-dimensional environment with an increased amount of light spill virtual lighting effect, the electronic device increases the size and / or brightness of the light spill on objects in the three-dimensional environment (e.g., from content items). In some embodiments, the virtual light spill changes over time in accordance with changes to content items (e.g., visual content contained in content items).
[0148] Modifying the amount of virtual lighting effects used to display a three-dimensional environment provides an efficient way to switch between an immersive experience and a less distracting virtual environment, thereby reducing the cognitive burden on the user both when engaging with content items and when engaging with other content or applications within the three-dimensional environment.
[0149] In some embodiments, before receiving user input directed to a separate user interface (e.g., 711 in Figure 7C), an electronic device (e.g., 101) displays a three-dimensional environment (e.g., 702) without using virtual lighting effects via a display generation component (e.g., 120) (804a). In some embodiments, virtual lighting effects include one or more of the following: displaying a virtual light spill effect emanating from a content item in areas of the three-dimensional environment other than the content item; dimming areas of the three-dimensional environment other than the content item relative to the content item; and / or blurring areas of the three-dimensional environment other than the content item relative to the content item. In some embodiments, displaying a three-dimensional environment without virtual lighting effects includes discontinuing the display of a virtual light spill effect emanating from the content item; displaying areas of the three-dimensional environment other than the content item at the same blur level as the content item; and / or displaying areas of the three-dimensional environment other than the content item at the same dimming level as the content item.
[0150] In some embodiments, upon receiving user input, an electronic device (e.g., 101) displays a three-dimensional environment (e.g., 702) using virtual lighting effects, such as those shown in Figure 7A, via a display generation component (e.g., 120) (804b). In some embodiments, the virtual lighting effects include one or more of the following: displaying a virtual light spill effect emanating from a content item in areas of the three-dimensional environment other than the content item; dimming areas of the three-dimensional environment other than the content item relative to the content item; and / or blurring areas of the three-dimensional environment other than the content item relative to the content item. In some embodiments, upon detecting a first input directed towards an individual user interface element, the electronic device switches the display of the three-dimensional environment with or without virtual lighting effects. For example, the first input is a selection of an individual user interface element (e.g., a primary selection such as "click"). In some embodiments, the first input includes detecting a default part of the user in a default shape (e.g., the user's hand in a pinching hand shape) for a period of time less than a specific threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds) while the user's gaze is directed towards the specific user interface element. In some embodiments, in response to a second input directed towards the specific user interface element (e.g., a secondary selection, an input similar to a "long click"), the electronic device updates the specific user interface element to a user interface element for adjusting the amount of lighting effect applied to the three-dimensional environment, as described below. For example, detecting the second input includes detecting that the default part of the user (e.g., the user's hand) performs a default gesture, such as a pinch hand gesture, which involves touching the thumb to another finger for a threshold time amount (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds) before the thumb is lifted from the finger.
[0151] Turning on virtual lighting effects in response to input directed at individual user interface elements provides an efficient way to switch virtual lighting effects on and off, thereby reducing the cognitive burden on the user when switching between an immersive experience and a three-dimensional environment where other user interface elements are clearly displayed.
[0152] In some embodiments, before receiving user input directed to a separate user interface (e.g., 711) such as Figure 7iD, the electronic device (e.g., 101) displays a three-dimensional environment (e.g., 702) using a first amount of virtual illumination effect via a display generation component (e.g., 120) (806a). In some embodiments, displaying a three-dimensional environment using a first amount of virtual illumination effect includes one or more of the following: displaying a virtual light spill effect emanating from a content item in areas of the three-dimensional environment other than the content item, with a first size, intensity, sharpness, etc.; dimming areas of the three-dimensional environment other than the content item by a first amount relative to the content item; and / or blurring areas of the three-dimensional environment other than the content item by a first amount relative to the content item. In some embodiments, the first amount of virtual illumination effect is zero (e.g., the electronic device displays the three-dimensional environment without using virtual illumination effect).
[0153] In some embodiments, in response to receiving user input, an electronic device (e.g., 101) displays a three-dimensional environment (e.g., 702) using a second amount of virtual illumination effect, different from a first amount, via a display generation component (e.g., 120), as shown in Figure 7E (806b). In some embodiments, the electronic device modifies multiple virtual illumination effects (e.g., light spill, blur, dimming) by the same amount in response to the input. For example, in response to an input to increase the virtual illumination effect by a specific amount, the electronic device increases the size, intensity, sharpness, etc. of the light spill by a specific amount, increases the blur by a specific amount, and increases the dimming by a specific amount. As another example, in response to an input to decrease the virtual illumination effect by a specific amount, the electronic device decreases the size, intensity, sharpness, etc. of the light spill by a specific amount, decreases the blur by a specific amount, and decreases the dimming by a specific amount. In some embodiments, the electronic device adjusts the amount of virtual lighting effect according to a determination that input directed to an individual user interface element satisfies one or more criteria, such as a criterion that is met when the electronic device detects a default part of the user (e.g., the user's hand) in a default hand shape (e.g., a pinching hand shape) for a threshold period (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds). In some embodiments, in response to the detection of a default pose beyond the threshold period, the electronic device presents an interactive slider that controls the amount of virtual lighting effect that displays the three-dimensional environment. For example, continued input directed to an individual user interface element, including movement of a default part of the user, causes the electronic device to adjust the amount of virtual lighting effect and the position of the slider's indicator.
[0154] Adjusting the amount of virtual lighting effects applied to the three-dimensional environment in response to input directed at individual user interface elements provides an efficient way to choose between the immersion of content items and the clarity of elements in the three-dimensional environment other than content items, thereby reducing the cognitive burden on the user when interacting with various elements in the three-dimensional environment.
[0155] In some embodiments, such as Figure 7C, a region of the three-dimensional environment (e.g., 702) that does not contain a content item (e.g., 704) is displayed at a first brightness level (808a) before receiving user input. In some embodiments, the content item is displayed at a brightness level higher than the first brightness level (e.g., dimming visual effects are active). In some embodiments, the content item is displayed at the same brightness level as the region of the three-dimensional environment that does not contain a content item (e.g., dimming visual effects are not active). In some embodiments, the three-dimensional environment includes representations of virtual and / or real objects.
[0156] In some embodiments, such as Figure 7E, displaying a three-dimensional environment (e.g., 702) using virtual lighting effects in response to user input includes displaying areas of the three-dimensional environment (e.g., 702) that do not contain content items (e.g., 704) at a second luminance level different from (e.g., lower, higher) a first luminance level (808b). In some embodiments, the electronic device modifies the luminance level of areas that do not contain content items without modifying the luminance level of the content items. In some embodiments, the electronic device turns the dimming visual effect on or off in response to input. In some embodiments, the electronic device adjusts the amount of the dimming visual effect in response to input. In some embodiments, the entire three-dimensional environment consists of areas that do not contain content items and content items (e.g., the luminance of the entire three-dimensional environment excluding the content items is adjusted). In some embodiments, the three-dimensional environment includes a third area that does not contain content items and is not affected by the adjustment of the amount of luminance. In some embodiments, the three-dimensional environment includes one or more virtual objects and / or regions outside the content item, and / or representations of real objects and / or regions outside the content item that are dimmed when the three-dimensional environment is displayed using virtual lighting effects.
[0157] Adjusting the brightness levels of areas in a three-dimensional environment that do not contain content items in response to input directed at individual user interface elements provides an efficient way to strike a trade-off between an immersive experience using content items and the ability to see and interact with elements in the three-dimensional environment other than the content items.
[0158] In some embodiments, such as Figure 7A, displaying a three-dimensional environment (e.g., 702) using virtual lighting effects in response to user input includes displaying individual virtual lighting effects (e.g., 710a) (e.g., virtual light spills) emanating from a content item (e.g., 704) on one or more objects (e.g., 708a) within the three-dimensional environment (e.g., 704) (810). In some embodiments, the individual virtual lighting effects change over time based on the content. For example, the individual virtual lighting effects include one or more colors, intensities, patterns, animations, etc., currently contained in the video and / or image content of the content item. In some embodiments, the individual virtual lighting effects are virtual light spills that simulate the reflection of light emanating from the image and / or video content of the content item on one or more (e.g., real, virtual) surfaces and / or objects in the three-dimensional environment outside the content item. In some embodiments, one or more objects include virtual objects. In some embodiments, one or more objects include representations of real objects in the physical environment of electronic devices and / or display generating components. In some embodiments, the representation of a real object includes one or more true or real passthroughs and / or video or virtual passthroughs, as described in more detail above. In some embodiments, virtual light spills are more strongly displayed (e.g., displayed with higher brightness, sharpness, and / or size) in locations within the three-dimensional environment closer to the content item than in locations within the three-dimensional environment further away from the content item.
[0159] In some embodiments, presenting individual virtual lighting effects emanating from a content item onto one or more objects in a three-dimensional environment provides an immersive experience for the content item, thereby reducing user distraction and cognitive burden while consuming the content item.
[0160] In some embodiments, such as Figure 7A, before receiving user input directed to a separate user interface (e.g., 711), the electronic device (e.g., 101) displays the three-dimensional environment (e.g., 702) using a first amount of virtual lighting effect (e.g., 812a), which includes displaying a region of the three-dimensional environment (e.g., 702) that does not contain content items at a first brightness level via a display generation component (e.g., 120), and displaying a first amount of individual virtual lighting effects (e.g., 710a) (e.g., virtual light spills) emitted from content items (e.g., 704) on one or more objects (e.g., 708a) in the three-dimensional environment. In some embodiments, the virtual lighting effect includes dimming the region of the three-dimensional environment other than the content item relative to the content item, and displaying the virtual light spills emitted from the content item on other objects in the three-dimensional environment. In some embodiments, the first amount of virtual lighting effect is zero (e.g., the electronic device displays the three-dimensional environment without using virtual lighting effects).
[0161] In some embodiments, such as Figure 7E, in response to receiving user input, an electronic device (e.g., 101) displays the three-dimensional environment (e.g., 702) using a second amount of virtual lighting effect, which includes displaying a region of the three-dimensional environment (e.g., 702) that does not contain a content item (e.g., 704) at a second brightness level via a display generation component (e.g., 120), and displaying a second amount of individual virtual lighting effect (e.g., 710a) (e.g., virtual light spill) emanating from the content item (e.g., 704) on one or more objects in the three-dimensional environment (e.g., 702) (812b). In some embodiments, individual user interface elements control both the level of dimming of the region of the three-dimensional environment other than the content item relative to the content item, and the level of virtual light spill (e.g., brightness, size, translucency, etc.) emanating from the content item on other objects in the three-dimensional environment. In some embodiments, in response to a first input directed to an individual user interface element, the electronic device switches the dimming and light spill virtual lighting effect on or off. In some embodiments, in response to a second input directed to an individual user interface element, the electronic device adjusts the level and level(s) (e.g., brightness, size, translucency, etc.) of the dimming of virtual light spills emanating from a content item on other objects outside the content item in a three-dimensional environment.
[0162] Adjusting both the amount and dimming of individual virtual lighting effects emitted from content items in response to inputs directed at individual user interface elements provides an efficient method for adjusting multiple characteristics of virtual lighting in a three-dimensional environment with a single input, thereby reducing the amount of time and the number of inputs required to make adjustments.
[0163] In some embodiments, such as Figure 7A, displaying a three-dimensional environment (e.g., 702) using virtual lighting effects in response to user input includes (815b) (814a) displaying the three-dimensional environment (e.g., 702) using a first amount of virtual lighting effect via a display generating component (e.g., 120) in accordance with detecting that the user's attention (e.g., gaze 713a) of an electronic device (e.g., 101) is directed towards a first area (e.g., including a content item 704) of the three-dimensional environment (e.g., 702) via one or more input devices (e.g., 314) (e.g., detecting that the user's gaze is directed towards the first area via an eye-tracking device). In some embodiments, in response to detecting that the user's attention (e.g., gaze) is directed towards a content item, the electronic device increases the amount of virtual lighting effect on which the three-dimensional environment is displayed.
[0164] In some embodiments, such as Figure 7E, displaying a three-dimensional environment (e.g., 702) using virtual lighting effects in response to user input includes (814c) (814a) displaying the three-dimensional environment using a second amount of virtual lighting effect different from a first amount via a display generating component (e.g., 120) in accordance with detecting via one or more input devices (e.g., 314) that the user's attention (e.g., gaze 713f) is directed to a second area of the three-dimensional environment (e.g., 702) that is different from a first area (e.g., via an eye-tracking device that the user's gaze is directed to a second area that does not contain content items). In some embodiments, in response to detecting that the user's attention (e.g., gaze) is not directed to content items, the electronic device reduces the amount of virtual lighting effect used to display the three-dimensional environment. In some embodiments, in response to detecting that the user's attention (e.g., gaze) is not directed to content items, the electronic device updates the three-dimensional environment to display it without using virtual lighting effects.
[0165] Adjusting the amount of virtual lighting effects depending on the area of the three-dimensional environment the user is focusing their attention on provides an efficient way to automatically adjust the level of immersion in content items based on the user's attention, which reduces the user's cognitive burden and the number of inputs when the user changes the area of the three-dimensional environment the user is focusing their attention on.
[0166] In some embodiments, while displaying a three-dimensional environment without using virtual lighting effects, an electronic device (e.g., 101) receives a first input via one or more input devices (e.g., 314) (816a) directed to a separate user interface element (e.g., 712j in Figure 7C), which includes detecting a default part of the user of the electronic device (e.g., hand 703b) in a default pose for less than a predetermined time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds) via one or more input devices (e.g., 314). In some embodiments, the default pose of the default part of the user is that the user's hand makes a pinch shape where the thumb touches another finger of the hand. In some embodiments, the default pose of a user's default portion is that the user's hand makes a pointing hand shape with one or more fingers extended and one or more fingers bent toward the palm while the user's hand is within a threshold distance (e.g., 1, 2, 3, 5, 10, 15, 30, or 50 centimeters) of a location in a three-dimensional environment corresponding to an individual user interface element. In some embodiments, detecting a first input further includes detecting, via one or more input devices, that the user's attention is directed toward an individual user interface element. In some embodiments, detecting a first input further includes detecting, via eye-tracking devices of one or more input devices, that the user's gaze is directed toward an individual user interface element.
[0167] In some embodiments, upon receiving a first input, the electronic device (e.g., 101) displays a three-dimensional environment (e.g., 702) using a virtual lighting effect, such as that shown in Figure 7A, via a display generation component (e.g., 120) (816b). In some embodiments, upon detecting a first input while displaying the three-dimensional environment without using the virtual lighting effect, the electronic device switches on the virtual lighting effect.
[0168] In some embodiments, while displaying a three-dimensional environment using virtual lighting effects, an electronic device (e.g., 101) receives a second input directed to a specific user interface element (e.g., 712j in Figure 7C) via one or more input devices (e.g., 314) (816c), which includes detecting a default part of the user (e.g., hand 703b) of the electronic device (e.g., 101) in a default pose for less than a predetermined time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds). In some embodiments, detecting the second input further includes detecting, via one or more input devices, that the user's attention is directed to the specific user interface element. In some embodiments, detecting the second input further includes detecting, via eye-tracking devices of one or more input devices, that the user's gaze is directed to the specific user interface element.
[0169] In some embodiments, upon receiving a second input, the electronic device (e.g., 101) displays a three-dimensional environment (e.g., 314) without using virtual lighting effects via a display generation component (e.g., 120) (816d). In some embodiments, upon detecting a second input while displaying the three-dimensional environment with virtual lighting effects, the electronic device switches off the virtual lighting effects.
[0170] Turning virtual lighting effects on or off in response to detecting that a user's default pause in a default part of the environment is below a threshold time amount provides an efficient way to switch between an immersive experience using content items and viewing other parts of the three-dimensional environment with less distraction, thereby reducing the cognitive burden on the user when interacting with the three-dimensional environment.
[0171] In some embodiments, such as Figure 7D, while displaying a three-dimensional environment (e.g., 702) using a first quantity of virtual lighting effects, the electronic device (e.g., 101) receives input directed to a separate user interface element (e.g., 712j) via one or more input devices (e.g., 314) (818a), which includes detecting the movement of a default part of the user (e.g., hand 703a) of the electronic device (e.g., 101) via one or more input devices (e.g., 314) while the default part of the user (e.g., hand 703a) is in a default pose. In some embodiments, detecting input further includes detecting that the user's attention (e.g., gaze) is directed to the separate user interface element via one or more input devices (e.g., eye-tracking devices). In some embodiments, the default pose is the user's hand in the pinch-hand shape described above. In some embodiments, the default pose is the user's hand in the pointing-hand shape described above while the user's hand is within a threshold distance (e.g., 5, 10, 15, 30, or 50 centimeters) of the location of the separate user interface element.
[0172] In some embodiments, in response to input directed to a separate user interface element (e.g., 712j in Figure 7D), an electronic device (e.g., 101) displays a three-dimensional environment (e.g., 702) via a display generation component (e.g., 120) using a second amount of virtual lighting effect, as shown in Figure 7E (818b), where the second amount is based on the movement (e.g., velocity, duration, distance, etc.) of a default part of the user (e.g., the hand 703a in Figure 7D) while the default part of the user is in a default pose. In some embodiments, in response to detecting movement of the default part of the user in a first direction (e.g., downward, left), the electronic device decreases the amount of virtual lighting effect. In some embodiments, in response to detecting movement of the default part of the user in a second direction (e.g., upward, right), the electronic device increases the amount of virtual lighting effect. In some embodiments, in response to movement having a first magnitude (e.g., velocity, duration, distance), the electronic device changes the amount of virtual lighting effect by a first amount corresponding to the first magnitude. In some embodiments, in response to movement having a second magnitude (e.g., speed, duration, distance), the electronic device changes the amount of virtual illumination effect by a second amount corresponding to the first magnitude. For example, in response to downward movement having a relatively small magnitude, the electronic device decreases the amount of virtual illumination effect by a relatively small amount. In another example, in response to upward movement having a relatively large magnitude, the electronic device increases the amount of virtual illumination effect by a relatively large amount.
[0173] Adjusting the amount of virtual lighting effects based on the movement of a default portion of the user's input directed towards individual user interface elements provides an efficient way for users to make trade-offs between an immersive experience with content items and the clarity of the rest of the three-dimensional environment, thereby reducing the cognitive burden on the user when interacting with elements within the three-dimensional environment.
[0174] In some embodiments, while a content item is being played (820a), an electronic device (e.g., 101) displays a three-dimensional environment (e.g., 702) using a virtual lighting effect, such as that shown in Figure 7A (820b). In some embodiments, the virtual lighting effect includes one or more of blurring and / or dimming virtual light spills emanating from areas of the three-dimensional environment that do not contain the content item and / or from the content item as described above.
[0175] In some embodiments, while a content item is being played (820a), an electronic device (e.g., 101) receives user input (e.g., selection of option 712g in Figure 7B) via one or more input devices in response to a request to pause the content item (e.g., 704) (820c). For example, the input in response to a request to pause the content item is a selection of individual user interface elements from one or more user interface elements displayed in the user interface associated with the content item to modify the playback of the content item.
[0176] In some embodiments, upon receiving user input corresponding to a request to pause a content item (820d), the electronic device (e.g., 101) pauses the content item (e.g., 704 in Figure 7A) (820e). In some embodiments, the electronic device continues to display the paused content item (e.g., displays a frame of the video content at the playback position where the video content was paused).
[0177] In some embodiments, upon receiving user input corresponding to a request to pause a content item (e.g., 704 in Figure 7A) (820d), the electronic device (e.g., 101) displays the three-dimensional environment (e.g., 702) via a display generation component (e.g., 120) without virtual lighting effects (or with lighting effects of reduced magnitude) (820f). In some embodiments, while the content item is paused and the electronic device is displaying the three-dimensional environment without virtual lighting effects, upon receiving input to resume playback of the content item, the electronic device resumes playback of the content item and displays the three-dimensional environment with virtual lighting effects. In some embodiments, upon receiving input to pause the content item, the electronic device reduces the amount of virtual lighting effect displayed on the three-dimensional environment without ceasing the display of virtual lighting effects. Ceasing the display of virtual lighting effects in response to input to pause the content item provides an efficient method to improve the readability of elements in the three-dimensional environment other than the content item while the content item is paused, thereby reducing the number of inputs required to switch between engaging with the content item and engaging with other elements in the three-dimensional environment.
[0178] In some embodiments, such as Figure 7C, an electronic device (e.g., 101) receives individual user inputs (822a) via one or more input devices (e.g., 314) directed to a second individual interface element of one or more user interface elements (e.g., 712f) for modifying the playback of a content item (e.g., 704).
[0179] In some embodiments, upon receiving individual user input (822b), the electronic device (e.g., 101) switches the playback or pause state of the content item (e.g., 704) according to the determination that a second individual user interface element, when selected, is a user interface element (e.g., 712g in Figure 7B) that causes the electronic device (e.g., 101) to switch the playback or pause state of the content item (e.g., 704) (822c). In some embodiments, the electronic device selects a user interface element in response to detecting that the user's attention (e.g., gaze) is directed towards a user interface element via one or more input devices (e.g., eye-tracking devices), while detecting that a default part of the user (e.g., hand) is in a default pose via one or more input devices (e.g., eye-tracking devices). For example, detecting a default part of the user in a default pose includes detecting the user's hand in a pinching hand shape, as described above. In some embodiments, the electronic device pauses the content in response to detecting input directed to a user interface element that causes the electronic device to switch between playing and pausing while the content item is being played. In some embodiments, the electronic device resumes playing the content in response to detecting input directed to a user interface element that causes the electronic device to switch between playing and pausing while the content item is paused.
[0180] In some embodiments, upon receiving individual user input (822b), the electronic device (e.g., 101) updates the playback position of the content item (e.g., 704) according to the individual user input (822d), based on the determination that a second individual user interface element is a user interface element (e.g., 712f, 712h in Figure 7B) that, when selected, causes the electronic device (e.g., 101) to update the playback position of the content item (e.g., 704) according to the individual user input (822d). In some embodiments, the selection of an individual user interface element causes the electronic device to change the playback position of the content item at a rate different from the rate at which the playback position of the content item changes while the content item is being played. In some embodiments, an individual user interface element for modifying virtual lighting effects is displayed in the user interface associated with the content item, which includes the first and second user interface elements. In some embodiments, the user interface further includes selectable options for accessing audio and subtitle settings for content items, switching picture-in-picture elements according to one or more steps of method 1000, switching immersive content modes according to one or more steps of method 1400, and browsing a content item playback queue for electronic devices.
[0181] Displaying separate user interface elements within a user interface that includes a user interface element for switching between playback and pause states, and a user interface element for updating the playback position of a content item, provides an efficient method for facilitating the modification of content item playback and the modification of the three-dimensional environment, thereby reducing the cognitive burden on the user when interacting with content items and the three-dimensional environment.
[0182] In some embodiments, an electronic device (e.g., 101) receives individual user inputs via one or more input devices (e.g., 314) directed to a second individual interface element (e.g., 712k in Figure 7B) of one or more user interface elements for modifying the playback of content items (824a).
[0183] In some embodiments, upon receiving individual user input, if a second individual user interface element is selected, the electronic device (e.g., 101) modifies the volume of the audio content of a content item (e.g., 704) according to the determination that it is a user interface element (e.g., 712k in Figure 7B) that causes the electronic device (e.g., 101) to modify the volume of the audio content according to the individual input (824b). In some embodiments, upon receiving a first input directed to a user interface element for modifying the volume of the audio content of a content item, the electronic device updates the user interface element to include a slider user interface element that causes the electronic device to change the volume of the audio content according to the updated position of the indicator when the position of the slider indicator changes. In some embodiments, the user interface element for modifying the volume of the audio content is displayed in the user interface associated with the content item, together with an individual user interface element for modifying virtual lighting effects.
[0184] Displaying separate user interface elements within a user interface that have a user interface element for adjusting the volume of audio content provides an efficient method for facilitating the adjustment of content item playback and the adjustment of the three-dimensional environment, thereby reducing the cognitive burden on the user when interacting with content items and the three-dimensional environment.
[0185] In some embodiments, such as Figure 7B, a user interface (e.g., 711) associated with a content item (e.g., 704) is a separate user interface from the content item (e.g., 704) and is displayed between the content item (e.g., 704) and the user's viewpoint of an electronic device (e.g., 101) in a three-dimensional environment (e.g., 702) via a display generation component (e.g., 120) (826). In some embodiments, the content item and the user interface are displayed in separate windows in the three-dimensional environment. In some embodiments, the user interface associated with the content item is partially overlaid on the content item. In some embodiments, the user interface associated with the content item is not overlaid on the content item.
[0186] Displaying a user interface associated with a content item between the user's viewpoint and the content item in a three-dimensional environment, separate from the content item itself, provides an efficient way to facilitate user interaction with the user interface and reduces the cognitive load and time required to modify the playback of the content item through interaction with the user interface.
[0187] In some embodiments, a content item (e.g., 704 in Figure 7B) is displayed via a display generation component (e.g., 120) at a first angle relative to the user's viewpoint in a three-dimensional environment (828a). In some embodiments, the first angle includes a lateral angle in the three-dimensional environment (e.g., tilting to the left or right of the user). In some embodiments, the first angle includes a vertical angle in the three-dimensional environment (e.g., tilting upward or downward from the user's viewpoint).
[0188] In some embodiments, a user interface (e.g., 711 in Figure 7B) associated with a content item (e.g., 704) is displayed via a display generation component (e.g., 120) at a second angle different from a first angle with respect to the user's viewpoint in a three-dimensional environment (e.g., 702) (828b). In some embodiments, the second angle includes a lateral angle in the three-dimensional environment (e.g., tilting to the left or right of the user). In some embodiments, the second angle includes a vertical angle in the three-dimensional environment (e.g., tilting upward or downward from the user's viewpoint). For example, a user interface associated with a content item is displayed at a location in the three-dimensional environment below the content item at an angle higher relative to the user's viewpoint than the angle at which the content item is displayed to the user.
[0189] Displaying content items and their associated user interfaces from different angles relative to the user's viewpoint in a three-dimensional environment provides an efficient way to present content items and their associated user interfaces in a recognizable manner when they are in different positions within the three-dimensional environment, thereby reducing the time and effort required to interact with content items and their associated user interfaces.
[0190] In some embodiments, such as Figure 7B, the electronic device (e.g., 101) displays one or more user interface elements (e.g., 712f, 712g) for modifying the playback of a content item (e.g., 704) in response to detecting a default part of the user of the electronic device (e.g., hand 703b) in a pose that satisfies one or more criteria (830) via one or more input devices (e.g., 314) (e.g., a hand tracking device). In some embodiments, detecting a pose of a default part of the user that satisfies one or more criteria includes detecting the movement of the user's hand from a position close to the user's body to an elevated location (e.g., within a default area of a three-dimensional environment). In some embodiments, detecting a pose of a default part of the user that satisfies one or more criteria includes detecting the user's hand in a default hand shape, such as the pointing hand shape described above, or a pre-pinch hand shape where the thumb of the hand is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 centimeters) but not touching another finger of the hand. In some embodiments, one or more criteria are met when the electronic device detects, via one or more input devices (e.g., eye-tracking devices), that the user's attention (e.g., gaze) is directed towards a content item. In some embodiments, while the electronic device does not detect a default portion of the user in a pose that satisfies one or more criteria, the electronic device discontinues displaying one or more user interface elements for correcting the playback of the content item. In some embodiments, the electronic device continues to present (and play) the content item.
[0191] Displaying one or more user interface elements to correct the playback of a content item in response to detecting a default position of the user in a pose that meets one or more criteria provides an efficient way to selectively facilitate interaction with one or more user interface elements, thereby reducing the time and input required to correct the playback of a content item.
[0192] In some embodiments, while displaying one or more user interface elements (e.g., 712f, 712g in Figure 7B) for modifying the playback of a content item (e.g., 704) (e.g., in response to detecting a default part of the user in a pose that satisfies one or more criteria), an electronic device (e.g., 101) detects a default part of the user (e.g., hand 703a) in a pose that does not satisfy one or more criteria, as shown in Figure 7A, via one or more input devices (e.g., 314) (e.g., hand tracking device) (832a). In some embodiments, one or more input devices (e.g., hand tracking device, eye tracking device) detect that one or more of the criteria described above are not met. In some embodiments, one or more input devices (e.g., hand tracking device) do not detect a default part of the user (e.g., hand) (e.g., because the default part of the user is outside the range of one or more input devices (e.g., hand tracking device)). In some embodiments, an electronic device detects a default part of the user in a pose (e.g., shape and / or position) that does not satisfy one or more criteria. For example, an electronic device can detect when a user drops their hand on their knee or beside them.
[0193] In some embodiments, upon detecting a default portion of the user's pose that does not meet one or more criteria (e.g., 703a), as shown in Figure 7A, the electronic device (e.g., 101) reduces the visual emphasis displayed by the electronic device (e.g., 101) of one or more user interface elements (e.g., 712f in Figure 7B) for correcting the playback of the content item via a display generating component (e.g., 120) (832b). In some embodiments, the electronic device reduces the opacity of one or more user interface elements to correct the playback of the content item. In some embodiments, the electronic device stops displaying one or more user interface elements for correcting the playback of the content item. In some embodiments, the electronic device continues to present (and play) the content item.
[0194] In response to detecting a default position in a user that does not meet one or more criteria, discontinuing the display of one or more user interface elements to correct the playback of a content item provides an efficient way to reduce distraction while the user is consuming a content item without indicating an intention to interact with one or more interactive elements, thereby reducing the cognitive burden on the user while consuming the content item.
[0195] In some embodiments, such as Figure 7B, while displaying a content item (e.g., 704) at a first size and a user interface associated with the content item (e.g., 711) at a second size via a display generation component (e.g., 120), the electronic device (e.g., 101) receives inputs via one or more input devices (e.g., 314) corresponding to a request to resize the content item (e.g., 704) (834a). In some embodiments, the input corresponding to the request to resize the content item is a request to change the virtual size of the content item in the three-dimensional environment, either by changing or without changing the position of the content item in the three-dimensional environment. In some embodiments, the input corresponding to the request to resize the content item is a request to change the angular size (e.g., a portion of the display generation component occupied by the content item) either by changing or without changing the position and / or virtual size of the content item in the three-dimensional environment.
[0196] In some embodiments, such as Figure 7C, upon receiving input corresponding to a request to resize a content item (e.g., 704) (834b), the electronic device (e.g., 101) displays the content item (e.g., 704) at a third size different from a first size, in accordance with the input corresponding to the request to resize the content item (e.g., 704) via a display generation component (e.g., 120) (834c). In some embodiments, the electronic device resizes the content item by an amount corresponding to the magnitude of the movement of the user's default part (e.g., speed, duration, distance, etc.) in a direction corresponding to the direction of movement of the user's default part (e.g., hand).
[0197] In some embodiments, such as Figure 7C, upon receiving input corresponding to a request to resize a content item (834b), the electronic device (e.g., 101) displays the user interface (e.g., 711) associated with the content item (e.g., 704) at a second size (834d) via a display generation component (e.g., 120). In some embodiments, the (e.g., angular) size of the user interface associated with the content item remains constant even when the content item is resized. In some embodiments, if the input is a request to resize a content item without changing the position of the content item, the electronic device maintains the angular size and virtual size of the user interface associated with the content item, and also maintains the position of the user interface in the three-dimensional environment. In some embodiments, if the input is a request to change the size and position of a content item in the three-dimensional environment, the electronic device maintains the angular size of the user interface and updates the virtual size of the user interface according to the updated position of the user interface to maintain the angular size of the user interface in the three-dimensional environment.
[0198] Maintaining the size of the user interface associated with a content item when updating the size of the content item provides an efficient way to maintain the readability of the user interface in a three-dimensional environment, which reduces the cognitive burden on the user when interacting with the user interface.
[0199] In some embodiments, such as Figure 7C, a display generation component (e.g., 120) displays a content item (e.g., 704) at a first (e.g., angular) size and a first distance from the user's viewpoint in a three-dimensional environment (e.g., 702), and while displaying a user interface (e.g., 711) associated with the content item (e.g., 704) at a second (e.g., angular) size and a second distance from the user's viewpoint in the three-dimensional environment (e.g., 702), the electronic device (e.g., 101) receives input via one or more input devices (e.g., 314) in response to a request to reposition the content item (e.g., 704) (e.g., and user interface) in the three-dimensional environment (e.g., 702) (836a). In some embodiments, the electronic device repositions the content item and user interface in accordance with the movement of a default part of the user (e.g., hand) while the user is providing input. For example, an electronic device moves content items and user interfaces in a direction and by an amount corresponding to the direction and amount (e.g., speed, distance, duration, etc.) of movement of a predefined part of the user (e.g., hand) while providing input. In some embodiments, detecting input includes detecting the user's gaze directed towards a content item or a user interface element for repositioning a content item, while detecting the user's hand in a predefined hand shape such as a pinch-hand shape or a pointing-hand shape.
[0200] In some embodiments, such as Figure 7D, upon receiving an input corresponding to a request to reposition a content item (e.g., 704) (836b), the electronic device (e.g., 101) displays the content item (e.g., 704) via a display generating component (e.g., 120) at a third distance different from a first distance from the user's viewpoint in a three-dimensional environment (e.g., 702) with a third (e.g., angular) size, in accordance with the input corresponding to the request to reposition the content item (e.g., 704) (836c). In some embodiments, the electronic device maintains the virtual size of the content item in response to a request to move the content item in the three-dimensional environment, thereby updating the angular size of the content item (e.g., the portion of the display generating component occupied by the content item) in accordance with the change in distance between the content item and the user's viewpoint. For example, in response to a request to move the content item further away from the user, the electronic device displays the content item with a smaller angular size, and in response to a request to move the content item closer to the user, the electronic device displays the content item with a larger angular size. In some embodiments, the electronic device updates the virtual size of the content item at a rate smaller than the change in the content item's angular size (for example, to maintain the display of the content item within the maximum and minimum sizes if updating the angular size of the content item without changing the virtual size of the content item would cause the angular size of the content item to fall outside the maximum or minimum angular size).
[0201] In some embodiments, such as Figure 7D, upon receiving input corresponding to a request to reposition a content item (e.g., 704) (836b), the electronic device (e.g., 101) displays the user interface (e.g., 711) associated with the content item (e.g., 704) at a second (e.g., angular) size and a fourth distance from the user's viewpoint, via a display generation component (e.g., 120), in accordance with the input corresponding to the request to reposition the content item (e.g., 704) (836d). In some embodiments, the electronic device updates the virtual size of the user interface in accordance with updating the distance between the user interface and the user's viewpoint in the three-dimensional environment, while maintaining the angular size of the user interface.
[0202] Maintaining the size of the user interface associated with content items provides an efficient way to maintain the readability of the user interface in a three-dimensional environment, which reduces the cognitive burden on the user when interacting with the user interface.
[0203] In some embodiments, such as Figure 7B, a content item (e.g., 704) is separate from the user interface (e.g., 711) associated with the content item (e.g., 704) in a three-dimensional environment (e.g., 702) (838a). In some embodiments, the content item and the user interface are displayed in separate containers (e.g., platters, user interface elements, windows, etc.) located in separate locations within the three-dimensional environment.
[0204] In some embodiments, such as Figure 7B, an electronic device (e.g., 101) displays (838b) one or more second user interface elements (e.g., 712a) for modifying the playback of a content item (e.g., 704) via a display generation component (e.g., 120), and one or more second user interface elements (e.g., 712a) are displayed as an overlay on the content item (e.g., 704) in a three-dimensional environment (e.g., 702). In some embodiments, one or more second user interface elements are displayed within a container containing the content item. In some embodiments, one or more second user interface elements are visually (e.g., from the user's perspective) or spatially overlaid on the content item in the three-dimensional environment. For example, the electronic device displays a user interface element via a display generation component that, when selected, causes the electronic device to stop displaying the content item overlaid on the content item in the content item's container and to display one or more other user interface elements (e.g., for modifying the playback of the content item) in a user interface in a separate container from the content item.
[0205] Displaying one or more second user interface elements overlaid on content items within a three-dimensional environment provides an efficient way for users to interact with one or more second user interface elements while browsing content items, thereby reducing the cognitive burden on the user.
[0206] In some embodiments, such as Figure 7C, an electronic device (e.g., 101) detects, via one or more input devices (e.g., 314) (e.g., eye-tracking devices), that the user's attention (e.g., gaze 713c, 713d) directed to individual user interface elements (e.g., 712f, 712j) of one or more user interface elements satisfies one or more first criteria (840a). In some embodiments, one or more first criteria include criteria that are satisfied when the user's attention (e.g., gaze) is directed to the individual user interface element for at least a threshold period (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 1, 2, or 3 seconds). In some embodiments, one or more criteria are satisfied at the moment the user's attention (e.g., gaze) is directed to the individual user interface element.
[0207] In some embodiments, such as Figure 7C, the electronic device (e.g., 101) displays a visual indication (e.g., 714) that identifies the function of the individual user interface element (e.g., 712j) via a display generation component (e.g., 120) in response to detecting that the user's attention (e.g., gaze 713d) directed towards an individual user interface element (e.g., 712j) satisfies one or more first criteria (840b), and in accordance with the determination that the individual user interface element (e.g., 712j) satisfies one or more second criteria (840c). In some embodiments, the individual user interface element satisfies one or more second criteria when the individual user interface element is included in a predetermined subset of one or more user interface elements included in the user interface associated with a content item. For example, one or more of the immersive content options by one or more steps of Method 1400, the picture-in-picture options by one or more steps of Method 1000, and individual user interface elements for modifying virtual lighting effects are included in the predetermined subset. In some embodiments, individual user interface elements satisfy one or more second criteria if the individual user interface element is associated with a function related to the presentation of content items in a three-dimensional environment (for example, it may not be commonly associated with the presentation of content items in another environment). In some embodiments, a visual indication that identifies the function of an individual user interface element includes text that describes the function of the individual user interface element. For example, a visual indication that identifies an individual optional function for modifying virtual lighting effects includes text that says “lighting effects,” etc. In some embodiments, the electronic device stops displaying the visual indication in response to detecting that the user's attention has been directed away from the individual user interface element. In some embodiments, the electronic device displays the visual indication via one or more input devices in response to detecting that the user is ready while their gaze is directed towards the individual user interface element.
[0208] In some embodiments, such as Figure 7C, the electronic device (e.g., 101) detects that the user's attention (e.g., gaze 713c) directed towards an individual user interface element (e.g., 712f) satisfies one or more first criteria (840b), and in accordance with the determination that the individual user interface element (e.g., 712f) does not satisfy one or more second criteria, the electronic device (e.g., 101) discontinues displaying the visual indication that identifies the function of the individual user interface element (e.g., 712f) (840d). In some embodiments, an individual user interface element does not satisfy one or more second criteria when it is not included in a predetermined subset of one or more user interface elements that are included in the user interface associated with a content item. For example, one or more of the following are not included in a predetermined subset of one or more user interface elements: the playback queue option, the option to skip back or head to the playback position of the content item, the option to play / pause the content item, the option to browse the subtitle options for the content item, and the option to browse the audio options for the content item. In some embodiments, individual user interface elements do not satisfy one or more of the second criteria when the individual user interface elements are associated with a function related to the presentation of content items in general (for example, they may not be specifically associated with the presentation of content items in a three-dimensional environment).
[0209] Displaying visual indicators that identify the functionality of one or more individual user interface elements that meet the second criterion provides an efficient way to show the user the actions that will be performed in response to further input directed at those individual user interface elements, thereby reducing the number of user errors and the inputs required to correct them.
[0210] In some embodiments, such as Figure 7B, an electronic device (e.g., 101) displays a separate user interface element (e.g., 712b) displayed separately from a content item (e.g., 711) and a user interface (e.g., 711) associated with the content item (e.g., 704) via a display generation component (e.g., 120) (842a). In some embodiments, the content item and user interface are displayed in separate containers (e.g., platters, user interface elements, windows, etc.) in separate locations in a three-dimensional environment, and the separate user interface element is displayed outside these containers in a different location in the three-dimensional environment than the location of the content item and user interface. In some embodiments, the electronic device displays the separate user interface element simultaneously with one or more selectable elements for modifying the playback of the content item, in response to detecting a default part of the user in a pose that satisfies one or more criteria, while optionally detecting the user's gaze directed towards the content item. For example, while detecting a user's gaze optionally directed at a content item, if the device detects a default part of the user in a pose that meets one or more criteria, it displays individual user interface elements and multiple selectable elements for modifying the playback of the content item.
[0211] In some embodiments, such as Figure 7B, while displaying an individual user interface element (e.g., 712b), an electronic device (e.g., 101) receives input directed to the individual user interface element (e.g., 712b) via one or more input devices (e.g., 314) (842b).
[0212] In some embodiments, such as Figure 7B, upon detecting input directed to an individual user interface element (e.g., 712b), the electronic device (e.g., 101) initiates a process of resizing a content item (e.g., 704) in a three-dimensional environment (e.g., 702) according to the input directed to the individual user interface element (e.g., 712b) (842c). In some embodiments, upon detecting selection of an individual user interface element, the electronic device initiates a process of resizing the individual user interface element according to the movement of the user's default part (e.g., hand) after the selection of the individual user interface element. In some embodiments, upon detecting movement of the user's default part (e.g., hand) after selection of an individual user interface element, the electronic device resizes the content item according to the movement of the user's default part (e.g., hand) without resizing the individual user interface element containing one or more selectable elements for modifying the playback of the content item, as described above.
[0213] Displaying separate user interface elements for resizing content items, distinct from the content items and their associated user interfaces, provides an efficient way to view the size of content items while providing input for resizing them. This improves the ergonomics of resizing input and reduces the time and number of inputs required to resize content items to the desired size.
[0214] Figures 9A to 9E illustrate exemplary methods for displaying media content in a three-dimensional environment.
[0215] Figure 9A shows a three-dimensional environment 904 displayed by the display generation component 120 of the electronic device 101, and an overhead view 920 of the three-dimensional environment 904. As described above with reference to Figures 1 to 6, the electronic device 101 optionally includes a display generation component (e.g., a touchscreen 120) and a plurality of image sensors (e.g., the image sensor 314 in Figure 3). The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the user interface shown below may also be implemented on a head-mounted display that includes a display generation component that displays the user interface to the user, and sensors for detecting the physical environment, the user's hand movements (e.g., external sensors facing outward from the user), and / or the user's line of sight (e.g., internal sensors facing inward towards the user's face).
[0216] In Figure 9A, the electronic device 101 displays a three-dimensional environment 904, which includes a user interface 906 for application 1, a user interface 910 for application 2, a user interface 912 for application 3, and a table representation 918, which is a physical table in the physical environment of the device 101. In some embodiments, applications 1-3 are optionally media applications, game applications, social applications, navigation applications, streaming applications, etc., as will be described in more detail below. In some embodiments, the representation 918 and user interfaces 906, 910, and 912 are displayed by the electronic device 101 because these objects are within the field of view of the user 922 of the three-dimensional environment 904 from their current viewpoint. For example, as shown in the overhead view 920 of Figure 9A, the user 922's current viewpoint of the three-dimensional environment 904 corresponds to the central position of the three-dimensional environment 904 and is oriented towards the top / rear of the three-dimensional environment 904. For the sake of facilitating the explanation in the remainder of this disclosure, the position / pose of user 922 within the three-dimensional environment 904 will be referred to herein as the current viewpoint of user 922 within the three-dimensional environment 904, or more simply, the viewpoint of user 922 as shown in the overhead view 920.
[0217] Therefore, as shown in Figure 9A, the electronic device 101 displays the representation 918 and user interfaces 906, 910, and 912 via the display generation component 120 because these objects are within the field of view of the user 922 in the three-dimensional environment 904 from their current viewpoint (as shown in the overhead view 920). Conversely, the electronic device 101 does not display the representation 924 of the sofa, the representation 930 of the corner table, the representation 932 of the coffee table, and user interfaces 926 and 928 via the display generation component 120 because these objects are not within the field of view of the user 922 in the three-dimensional environment 904 from their current viewpoint (as shown in the overhead view 920).
[0218] In some embodiments, the user 922's viewpoint in the three-dimensional environment 904 corresponds to the user 922's physical location in the physical environment 902 (e.g., the operating environment 100) of the electronic device 101. For example, the user 922's viewpoint is optionally the viewpoint shown in the overhead view 920, as the user 922 is currently oriented toward the back wall in the physical environment 902 and is located at the center of the physical environment 902 while holding (e.g., or wearing) the electronic device 101 if the device 101 is a head-mounted device.
[0219] As shown in Figure 9A, the electronic device 101 is currently playing TV program A in the user interface 906. In some embodiments, TV program A is played in the user interface 906 in response to the electronic device detecting a request to start playing TV program A. In some embodiments, as will be described in more detail below, the electronic device 101 can present TV program A in different presentation modes, including picture-in-picture presentation mode and / or extended presentation mode (e.g., different from picture-in-picture presentation mode). In the example in Figure 9A, the electronic device 101 is currently presenting TV program A in extended presentation mode in the user interface 906. It should be understood that while TV program A is presented in extended presentation mode, the user interface 906 may also display other locations in the three-dimensional environment 904 that are within the user's field of view from the user's current viewpoint in the three-dimensional environment 904.
[0220] In Figure 9B, while the electronic device 101 is playing TV program A in extended presentation mode, the electronic device 101 detects that the user 922's viewpoint in the three-dimensional environment 904 has moved from the viewpoint shown in Figure 9A to the viewpoint shown in Figure 9B. In some embodiments, the user's viewpoint in the three-dimensional environment 904 has moved to the viewpoint shown in Figure 9B because the user 922 has moved to a corresponding pose and / or position in the physical environment 902. As shown in Figure 9B, in response to the electronic device 101 detecting the shift in the user 922's viewpoint in the three-dimensional environment 904, the electronic device 101 displays the three-dimensional environment 904 from the user's new viewpoint in the three-dimensional environment 904, as shown in the overhead view 920 of Figure 9B.
[0221] Specifically, as a result of user 922's viewpoint shift from the viewpoint shown in Figure 9A to the viewpoint shown in Figure 9B, the user interface 912 is no longer within the user's field of view of the three-dimensional environment 904 (as shown in the overhead view 920 of Figure 9B), and therefore the electronic device 101 no longer presents the user interface 912 of application 3 as previously shown in Figure 9A. In addition, as a result of user 922's viewpoint shift, user 922's viewpoint in the three-dimensional environment 904 has shifted to the right from user 922's viewpoint shown in Figure 9A, and therefore the electronic device 101 displays the table representation 918, user interface 910, and user interface 906 in locations within the three-dimensional environment that are further to the left of the user's field of view compared to Figure 9A.
[0222] In some embodiments, when media content is presented in extended presentation mode, the location of the user interface presenting the media content does not change within the three-dimensional environment 904 even when the user's viewpoint in the three-dimensional environment 904 moves. For example, when the viewpoint of user 922 in the three-dimensional environment 904 moved from the viewpoint shown in Figure 9A to the viewpoint shown in Figure 9B, the location of user interface 906 within the three-dimensional environment 904 did not change (as shown in the overhead view 920 in Figures 9A and 9B) because user interface 906 was presenting TV program A in extended presentation mode.
[0223] In Figure 9B, the electronic device 101 also presents a playback control user interface 908. In some embodiments, the electronic device 101 displays the playback control user interface 908 in response to the electronic device 101 detecting that user 922's hand 916 is in a “pointing” pose (e.g., one or more fingers of hand 1331 are extended and one or more fingers of hand 916 are curled toward the palm of hand 916) or a “pre-pinch” pose (e.g., the thumb of hand 916 is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, or 5 centimeters) of another finger of hand 916, but is not touching). In some embodiments, the user interface elements 908a to 908j displayed in the playback control user interface 908 are similar to the selectable user interface options 712c to L described in the series of Figures 7.
[0224] In addition, in Figure 9B, while the electronic device 101 is playing TV program A in extended presentation mode, the electronic device 101 detects a request to start playing TV program A in picture-in-picture presentation mode. In some embodiments, the electronic device 101 detects a request to start playing TV program A in picture-in-picture presentation mode because the user's hand 916 was in a “pointing” or “pinching” pose (for example, the thumb and index finger of hand 916 converge on each other within a threshold distance (e.g., 0.2, 0.5, 1, 1.5, 2, or 2.5 centimeters)) while the user's gaze 914 was directed at the user interface element 908b.
[0225] In Figure 9C, upon detecting a request to start the presentation of TV program A in the picture-in-picture presentation mode of Figure 9B, the electronic device 101 stops playback of TV program A in the user interface 906 and starts the presentation of TV program A in the picture-in-picture user interface 934. In some embodiments, when the playback of TV program A transitions from the media user interface 906 to the picture-in-picture user interface 934, the electronic device 101 displays an animation in which the picture-in-picture user interface fades in and the playback of TV program A fades out in the user interface 906.
[0226] As shown in the overhead view 920, the picture-in-picture user interface 934 is displayed at a location within the three-dimensional environment 904, in a position within the three-dimensional environment 904 that is in front of and to the right of the user's current viewpoint in the three-dimensional environment 904. In some embodiments, the electronic device 101 displays the picture-in-picture user interface 934 at the location shown in the overhead view 920 because the location within the three-dimensional environment 904 is within a threshold distance (e.g., 0.1, 0.2, 0.5, 1, 1.5, or 3 feet) from the user's current viewpoint in the three-dimensional environment 904 and / or occupies a predetermined portion (e.g., lower right, lower left, upper right, upper left) of the field of view from the user 922's viewpoint in the three-dimensional environment 904. Thus, in some embodiments, the user interface 934 is displayed at a location within the three-dimensional environment 904 based on the user 922's viewpoint, while the user interface 906 is displayed at a location within the three-dimensional environment 904 that is not based on the user 922's viewpoint.
[0227] In some embodiments, while the electronic device 101 is presenting media content in picture-in-picture presentation mode, the electronic device optionally displays one or more representations of media items that can be selected for playback. For example, in Figure 9C, in response to the electronic device 101 receiving a request in Figure 9B to transition the presentation of TV program A from extended presentation mode to picture-in-picture presentation mode, the electronic device 101 updates the user interface 906 (the user interface that previously presented TV program A in extended presentation mode) to include multiple representations 940-958 of individual media content. The multiple representations 940-958 are optionally selectable such that when the electronic device 101 detects a selection of one of the representations 940-958, the media item corresponding to the selected representation begins playback in the user interface 906 (optionally, without interrupting playback of TV program A in the media user interface 934) and / or in the picture-in-picture user interface 934.
[0228] In some embodiments, multiple representations 940-958 of individual media content are displayed in one or more groups (e.g., columns) within the user interface 906. For example, in Figure 9C, representations 940-946 are displayed in a first column of the user interface 906 because the corresponding media items were selected for display based on the user's content consumption history. Similarly, representations 948-958 are displayed in a second column of the user interface 906 because the corresponding media items correspond to popular / currently trending content items (e.g., more users have recently viewed the media content corresponding to representations 948-958 over the past hour, day, week, month, etc.).
[0229] In some embodiments, the electronic device 101 updates the type / category of media content displayed in the user interface 906. For example, in Figure 9C, the electronic device 101 displays the user interface 936 (which is also optionally displayed in response to the electronic device 101 receiving a request to begin displaying TV program A in picture-in-picture display mode, as previously mentioned). The user interface element 936 includes, when selected, selectable option 936a which causes the electronic device 101 to display representations of currently popular and / or recommended media content based on the user 922's content consumption history (as shown in user interface 906 in Figure 9C); selectable option 936b which, when selected, causes the electronic device 101 to display one or more representations of media content corresponding to one or more TV programs within user interface 906; selectable option 936c which, when selected, causes the electronic device 101 to display one or more representations of media content corresponding to one or more movies within user interface 906; selectable option 936d which, when selected, causes the electronic device 101 to display one or more representations of media content corresponding to one or more (e.g., live) sports games within user interface 906; and selectable option 936e which, when selected, causes the electronic device 101 to present a user interface for searching for specific media content within user interface 906.
[0230] In Figure 9D, the electronic device 101 detects that the user 922's viewpoint has moved from the viewpoint shown in Figure 9C to the viewpoint shown in Figure 9D. In some embodiments, the user 922's viewpoint in the three-dimensional environment 904 has moved to the viewpoint shown in Figure 9D because the user 922 has moved to a corresponding pose and / or location within the physical environment 902. As shown in Figure 9D, in response to detecting that the user 922's viewpoint in the three-dimensional environment 904 has moved to the viewpoint shown in Figure 9D, the electronic device 101 displays the three-dimensional environment 904 from the user's new viewpoint in the three-dimensional environment 904. Specifically, the display generation component 120 of the device 101 is now displaying the user interfaces 926 and 928 and representations 924 and 932 because these elements are now within the field of view from the user's viewpoint shown in Figure 9D.
[0231] In some embodiments, when the viewpoint of the user 922 in the three-dimensional environment 904 changes, the electronic device 101 updates the location of the picture-in-picture user interface 934 based on the new viewpoint of the user 922 in the three-dimensional environment 904. For example, as shown in the overhead views 920 of Figures 9C and 9D, when the electronic device 101 detects that the viewpoint of the user in the three-dimensional environment 904 has moved from the viewpoint shown in Figure 9C to the viewpoint shown in Figure 9D, the electronic device 101 moves the location of the picture-in-picture user interface 934 from the location shown in Figure 9C to the location shown in Figure 9D. In some embodiments, the electronic device 101 moves the picture-in-picture user interface 934 from its location in the three-dimensional environment 904 shown in the overhead view 920 of Figure 9C because, based on the user's current viewpoint in the three-dimensional environment 904, the location of the user interface 934 shown in the overhead view 920 is no longer within a threshold distance (e.g., 0.1, 0.2, 0.5, 1, 1.5, or 3 feet) of the user's new viewpoint in the three-dimensional environment 904, and / or its location is no longer in a predetermined (e.g., lower right) portion of the user's field of view from the user's new viewpoint in the three-dimensional environment 904. In some embodiments, the picture-in-picture user interface 934 is optionally displayed at its location in the three-dimensional environment 904 shown in the overhead view 920 for similar reasons as described with reference to Figure 9C. In addition, as shown in the overhead view 920, the location of the user interface 906 in the three-dimensional environment 904 does not change for similar reasons as previously described with reference to Figure 9B.
[0232] In some embodiments, while the electronic device 101 is presenting content in picture-in-picture presentation mode, playback controls are displayed overlaid or integrated onto the picture-in-picture user interface. For example, since user interface 934 is currently presenting TV program A in picture-in-picture presentation mode, the electronic device 101 displays user interface elements 936-948 overlaid on user interface 934 (as described with reference to Figure 9B, when playback controls are presented in a separate user interface while TV program A is presented in extended presentation mode). In some embodiments, as shown in Figure 9D, user interface elements 936-948 are displayed on the media user interface 934 when the electronic device detects that the hand 916 is in a “pre-pinch” pose and optionally when the user 922’s gaze is directed towards user interface 934. If the electronic device 101 does not detect that the hand 916 is in a “pre-pinch” pose, user interface elements 936-948 are optionally not displayed.
[0233] Next, the functions related to user interface elements 936 to 948 will be described. User interface element 936 is optionally selectable, and when selected, it causes the electronic device 101 to stop displaying TV program A in picture-in-picture display mode and start displaying TV program A in extended display mode. User interface element 938 is optionally selectable, and when selected, it causes the electronic device 101 to stop playing TV program A (and optionally stop displaying user interface 934). User interface element 940 is optionally selectable, and when selected, it causes the electronic device 101 to fast-forward TV program A by a predetermined amount (for example, 10, 15, 20, 30, 40, or 60 seconds). User interface element 942 is optionally selectable, and when selected, it causes the electronic device 101 to rewind TV program A by a predetermined amount (for example, 10, 15, 20, 30, 40, or 60 seconds). The user interface element 944 is optional and, when selected, causes the electronic device 101 to pause playback of TV program A (for example, if TV program A is currently playing) or to start playback of TV program A (for example, if TV program A is currently paused). Finally, the media user interface 934 includes a scrub bar 946 which includes an indicator 948 that shows the current playback position of TV program A. Further details of the scrub bar 908j and the operation associated with the scrub bar 908j will be described with reference to Method 1400 and Figures 13A to 13E.
[0234] In addition, as shown in Figure 9D, while user interface elements 936-946 are displayed and TV program A is being played in picture-in-picture presentation mode, the electronic device receives a request to present TV program A in extended presentation mode (indicated by the selection of user interface 936). In some embodiments, the input for selecting user interface element 936 is the same as the input for selecting user interface element 918b in Figure 9B. In some embodiments, upon receiving a request to start presenting TV program A in extended presentation mode, the electronic device 101 stops presenting TV program A in media user interface 934 and starts displaying TV program A in extended presentation mode in user interface 906 as shown in Figure 9C (and optionally stops displaying representations 940-958 in user interface 906 and user interface 934 as shown in Figure 9C).
[0235] In some embodiments, when the electronic device 101 receives a request to transition the presentation mode of TV program A from picture-in-picture presentation mode to extended presentation mode, if the user interface 906 (for example, a user interface that presents TV program A in extended presentation mode) is not within the user's field of view from the user's current viewpoint in the three-dimensional environment 904, the electronic device 101 updates the location of the user interface 906 so that it is within the user's field of view from the user's current viewpoint in the three-dimensional environment 904, as shown in Figure 9E. Conversely, if the electronic device 101 receives a request to transition TV program A from being presented in picture-in-picture mode to extended presentation mode while the user interface 906 is currently in a location within the user's field of view in the three-dimensional environment 904 (as shown in Figure 9C), the electronic device 101 optionally does not update the location of the user interface 906 in the three-dimensional environment 904.
[0236] In some embodiments, when playback of a media item in extended presentation mode has finished, the electronic device 101 displays one or more representations of one or more suggested media items to be viewed next. For example, in Figure 9E, the electronic device 101 detects that playback of TV program A in the user interface 906 has finished, or that a specific position in playback has been reached (e.g., 0.25, 0.5, 1, 2, 3, or 5 minutes from the end of playback). Accordingly, the electronic device 101 displays the user interface 909, which includes representations 946-950 of the respective media items. In some embodiments, the media items corresponding to representations 946-950 are selected for display in the user interface 909 based on the user 922's content consumption history. In some embodiments, representations 946-950 are selectable, and when selected, the electronic device 101 plays the corresponding media item in the user interface 906. For example, when the electronic device 101 detects that user 916's hand is in a "pointing" or "pinching" position (as described above) while user 922's gaze 914 is directed towards the representation 950, the electronic device 101 optionally starts playback of media item C in the user interface 906.
[0237] Additionally or alternatively, media items corresponding to representations 946-950 are selectively selected for playback based on the user's gaze 914 (without detecting input from the user's hand). For example, as shown in Figure 9E, the user's gaze 914 is currently directed towards representation 946 of item A. In some embodiments, when the electronic device detects that the user's gaze 914 is directed towards representation 914, the electronic device 101 begins playback of item A in the user interface 906. Alternatively, in some embodiments, the electronic device 101 begins playback of media item A only when the user's gaze 914 has been directed towards representation 946 for at least a threshold time (e.g., 15, 30, 60, 90, or 200 seconds). For example, in Figure 9E, the electronic device 101 has not begun playback of media item A in the user interface 906 because the user's gaze 914 has not been directed towards representation 946 for the aforementioned threshold time.
[0238] In some embodiments, the electronic device 101 displays an indication 915 showing the amount of time remaining before the user's gaze 914 initiates playback of a media item on the electronic device 101. For example, in Figure 9E, the electronic device 101 displays a circular visual indication 915 within the representation 946. In some embodiments, when the user's gaze 914 remains directed toward the representation 946, the electronic device 101 updates the visual indication 915 (e.g., in real time) to occupy a space ranging from 0 degrees (e.g., when the user's gaze 914 is not directed toward the representation 914) to 360 degrees (e.g., when the user's gaze 914 has been directed toward the representation 914 for the threshold time amount described above). For example, in Figure 9E, the visual indication 915 occupies a 180-degree angular distance, indicating that the user's gaze 914 has been directed toward the representation 946 for half of the threshold time amount described above.
[0239] Additional or alternative details relating to the embodiments illustrated in Figures 9A to 9E are provided below in the description of Method 1000, which is described with reference to Figures 10A to 10I.
[0240] Figures 10A to 10I are flowcharts illustrating methods for displaying media content in a three-dimensional environment according to several embodiments. In some embodiments, Method 1000 is performed in a computer system (e.g., computer system 101 in Figure 1) which includes a display generation component (e.g., display generation component 120 in Figures 1, 3, and 4) (e.g., a head-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing downwards from the user's hands (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, Method 1000 is stored in a non-temporary computer-readable storage medium and controlled by instructions executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., a control unit 110 in Figure 1A). Some operations of Method 1000 are optionally combined, and / or the order of some operations is optionally changed.
[0241] In some embodiments, Method 1000 is performed in an electronic device that communicates with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, or television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users. In some embodiments, the one or more input devices include an electronic device or component that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include touchscreens, mice (e.g., external), trackpads (optionally integrated or external), touchpads (optionally integrated or external), remote control devices (e.g., external), another mobile device (e.g., separate from the electronic device), handheld devices (e.g., external), controllers (e.g., external), cameras, depth sensors, eye-tracking devices, and / or motion sensors (e.g., hand-tracking devices, hand motion sensors). In some embodiments, the electronic device communicates with the hand-tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreen, trackpad)). In some embodiments, the hand-tracking device is a wearable device such as a smart glove. In some embodiments, the hand-tracking device is a handheld input device such as a remote control or stylus.
[0242] In some embodiments, an electronic device (e.g., device 101 in Figures 9A to 9E) presents content (e.g., media) via a display generation component and displays a three-dimensional environment (e.g., the three-dimensional environment is a computer-generated reality (XR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment) located at a first distinct location in the three-dimensional environment (1002a). For example, in Figure 9A, the electronic device 101 displays a three-dimensional environment 904 that includes "TV program A" presented to the user interface 906. In some embodiments, while the first media user interface is displayed in the three-dimensional environment, the first media user interface presents movies, TV programs, music videos, and / or other types of video or audio content. In some embodiments, the first media user interface is positioned at a first individual location in the three-dimensional environment because the first individual location is the default launch location for the application associated with the first media user interface (for example, the first media user interface is displayed at the first individual location in the three-dimensional environment in response to the launch of the individual application). In some embodiments, the first media user interface is positioned at a first individual location in the three-dimensional environment because the user of the electronic device has moved the first media user interface to the first individual location. In some embodiments, if the three-dimensional environment is displayed while the user's viewpoint is a first viewpoint (for example, while the electronic device is oriented to a first area in the physical environment), the three-dimensional environment optionally includes representations of some or all of the objects located at the first location in the physical environment, and / or representations of virtual objects (for example, objects that are not in the physical environment but are displayed so that the user's viewpoint corresponds to the first viewpoint).In other words, different viewpoints of users of an electronic device optionally cause different virtual and / or physical object representations to exist within the user's field of view while displaying a three-dimensional environment from individual viewpoints. For example, a first location in the physical environment may include one or more physical objects such as a chair, sofa, and table, and the three-dimensional environment may include representations of one or more of those chairs, sofas, and tables. Similarly, a first media user interface optionally occupies a location in the three-dimensional environment such that the first media user interface is within the user's field of view from the first viewpoint when the user's viewpoint corresponds to the first viewpoint.
[0243] In some embodiments, the electronic device displays the three-dimensional environment using a first media user interface at a first distinct location in the three-dimensional environment that has a pose (e.g., position and / or orientation) within a distinct range of poses relative to the user's first viewpoint (for example, in some embodiments, the distinct range of poses relative to the user's first viewpoint includes all (or a subset thereof) of poses in the three-dimensional environment that are within the user's field of view of the three-dimensional environment from the user's first viewpoint). In some embodiments, the position of the first media user interface in the three-dimensional environment is not within the distinct range of poses relative to the first viewpoint if that position is not within the user's field of view from the first viewpoint. (1002b) (for example, the electronic device detects that the user of the electronic device has begun to look at a different location in the physical environment (e.g., the orientation of the user's field of view into the three-dimensional environment has changed)). For example, in Figure 9B, the electronic device 101 detects that the user's viewpoint of the three-dimensional environment 904 has changed from the viewpoint shown in the overhead view 920 of Figure 9A to the viewpoint shown in the overhead view 920 of Figure 9B. In some embodiments, portions of the three-dimensional environment that were within the user's field of view from the first viewpoint are optionally still within the user's field of view from the second viewpoint (e.g., at least a portion of the three-dimensional environment displayed via the display generation component while the three-dimensional environment was presented from the first viewpoint is displayed via the display generation component while the three-dimensional environment is presented from the second viewpoint). Alternatively, in some embodiments, while the three-dimensional environment is presented from the user's second viewpoint, areas / portions of the three-dimensional environment that were within the user's field of view from the first viewpoint are optionally no longer within the user's field of view from the second viewpoint. In some embodiments, the user's viewpoint changes as the user moves within the physical environment (e.g., walks, runs, etc.) and / or looks towards different areas within the physical environment (e.g., while remaining stationary).
[0244] In some embodiments, upon detecting a shift in the user's viewpoint from a first viewpoint to a second viewpoint, the electronic device displays the three-dimensional environment from the second viewpoint via a display generation component (1002c) (for example, electronically updating the display of the three-dimensional environment to correspond to the user's new viewpoint—the second viewpoint). In some embodiments, the display of the three-dimensional environment from the second viewpoint is similar to the display of the three-dimensional environment from the user's first viewpoint, as described above.
[0245] In some embodiments, according to the determination that the content is presented in a first presentation mode (for example, in some embodiments, the content is presented in the first presentation mode when the content is not presented in a picture-in-picture (PiP) UI; in some embodiments, the content is presented in the first presentation mode when the content is presented in a default presentation mode (for example, playing natively in a video player application associated with the first media user interface); in some embodiments, the content is presented in the first presentation mode when the content is presented in a video aspect ratio higher than 9:16 or 16:9, and the electronic device 101 maintains the first media user interface in a first separate location within the three-dimensional environment (1002d), and the first media user interface is no longer within a separate range of poses relative to the user's second viewpoint. For example, the viewpoint of user 922 in the three-dimensional environment 904 is between Figures 9A and 9B. Even if there was a change in the three-dimensional environment 904, the location of the user interface 906 in the three-dimensional environment 904 did not change (as shown in the overhead view 920 in Figures 9A and 9B). For example, if the content in the first media user interface is presented in a non-PiP presentation mode (e.g., presented natively within a video player application), the movement of the user's viewpoint does not change the location of the first media user interface in the three-dimensional environment. Therefore, in some embodiments, when the three-dimensional environment is displayed from the user's second viewpoint, the first media user interface is optionally no longer located within the user's field of view. In some embodiments, when the three-dimensional environment is presented from the user's second viewpoint, the location (e.g., position) of the first media user interface in the three-dimensional environment is not within the user's field of view from the user's second viewpoint of the three-dimensional environment, and therefore the first media user interface is no longer within the individual range of poses relative to the user's second viewpoint.
[0246] In some embodiments, upon determination that the content is presented in a second presentation mode different from a first presentation mode (for example, if the content is presented in picture-in-picture (PiP) format), the electronic device displays the first media user interface in a second separate location different from a first separate location in the three-dimensional environment (1002e), and by displaying the first media user interface in the second separate location, the first media user interface is displayed in a pose that falls within a separate range of poses relative to the user's second viewpoint, such as the location of user interface 934 which moves from the location shown in Figure 9C to the location shown in Figure 9D as a result of the user's viewpoint shift (for example, in some embodiments, the range of poses relative to the user's second viewpoint includes all (or a subset thereof) of poses in the three-dimensional environment that are within the user's field of view from the user's second viewpoint). For example, when the content within the first media user interface is presented in picture-in-picture (PiP) presentation mode, the location of the first media user interface in the three-dimensional environment changes as the user's viewpoint in the three-dimensional environment changes, so that the first media user interface is always displayed in the currently displayed portion of the three-dimensional environment (e.g., corresponding to the user's current viewpoint). In some embodiments, when the first media user interface is presented in a second presentation mode, the first media user interface is displayed in a location in the three-dimensional environment such that it appears within a threshold distance (0.5, 1, 2, 4, or 6 feet) of the user of the electronic device (e.g., a given portion) (or within a threshold distance (0.5, 1, 2, 4, or 6 feet) of individual body parts of the user (e.g., right or left hip, right or left shoulder)). For example, when the first media user interface is presented in the second presentation mode, the first media user interface is displayed in the lower right portion of the user's field of view in the three-dimensional environment, regardless of the viewpoint location and / or orientation.In some embodiments, if the first media user interface is not presented in a second presentation mode (for example, if it is presented in a first presentation mode), the first media user interface is optionally displayed in a location within a three-dimensional environment that is not within the threshold distance (0.5, 1, 2, 4, or 6 feet) of the user of the electronic device (or within the threshold distance (0.5, 1, 2, 4, or 6 feet) of individual body parts of the user (e.g., right or left hip, right or left shoulder)).
[0247] Changing the location of the first media user interface when the user's viewpoint in the three-dimensional environment changes provides an efficient way to provide continuous access to a specific user interface in the three-dimensional environment, regardless of the user's current viewpoint in the three-dimensional environment, when the content of the first media user interface is presented in a second presentation mode, thereby reducing the cognitive burden on the user both when engaging with the first media user interface and when engaging with other content or applications in the three-dimensional environment.
[0248] In some embodiments, the second individual location is based on the second viewpoint (1004a) (for example, the location of the first media user interface is no longer based on the location of the user's first viewpoint, but rather on the location of the user's second viewpoint). In some embodiments, the second individual location is a location within a threshold distance (e.g., 0.1, 0.2, 0.5, 1, 1.5, or 3 feet) from the second viewpoint or within a three-dimensional environment located there. In some embodiments, the second individual location is a location within the user's field of view from the second viewpoint. In some embodiments, the second individual location within a three-dimensional environment corresponds to a location within a threshold distance of a particular body part of the user (e.g., a part) or within 0.1, 0.2, 0.3, 1, 2, or 3 feet of the user's buttocks, hand, head, foot, or knee in the physical environment. In some embodiments, the second individual location corresponds to the lower right portion (or lower left portion or upper right portion) of the user's field of view from the user's second viewpoint. In some embodiments, displaying the first media user interface at a second separate location means that the user's viewpoint movement after moving to the second viewpoint satisfies one or more criteria (for example, in some embodiments, the user's viewpoint in the three-dimensional environment does not change by more than a predetermined amount following the user's viewpoint moving to the second viewpoint in the three-dimensional environment (for example, the user's viewpoint does not move by more than a threshold amount (e.g., less than 1cm, 2cm, 5cm, 10cm, 50cm, 100cm, 300cm, or 1000cm)). (1004b) (1004a) the first media user interface is displayed at a second separate location in accordance with the determination that one or more criteria are met if the second viewpoint is corresponding for a threshold time period (e.g., 0.5, 1, 3, 7, 10, 20, or 30 seconds). For example, in Figure 9D, if the viewpoint of user 922 in the three-dimensional environment 904 meets the above criteria, the user interface 934 is displayed at a location in the three-dimensional environment 904 that is within the field of view of user 922.For example, if the user's viewpoint in the three-dimensional environment corresponds to the second viewpoint for at least a threshold time amount (e.g., 0.5, 1, 3, 7, 10, 20, or 30 seconds), and / or moves less than a threshold amount (e.g., less than 1 cm, 2 cm, 5 cm, 10 cm, 50 cm, 100 cm, 300 cm, or 1000 cm), the first media user interface will be displayed at the second separate location within the three-dimensional environment.
[0249] In some embodiments, displaying the first media user interface at a second separate location includes discontinuing the display of the first media user interface at the second separate location if the user's viewpoint movement after moving to the second viewpoint does not meet one or more criteria (for example, in some embodiments, one or more criteria is not met if, following the user's viewpoint in the three-dimensional environment moving to the second viewpoint, the user's viewpoint in the three-dimensional environment moves by more than a threshold amount (e.g., the user's viewpoint moves by more than 1 cm, 2 cm, 5 cm, 10 cm, 50 cm, 100 cm, 300 cm, or 1000 cm), and / or if the user has not corresponded to the second viewpoint for at least a threshold amount (e.g., 0.1, 1, 3, 7, 10, 20, or 30 seconds)) (1004c) (1004a). For example, in Figure 9D, if the user 922's viewpoint in the three-dimensional environment 904 does not meet the above criteria, the user interface 934 will not be displayed at a location within the three-dimensional environment 904 that is within the user 922's field of view. For example, if the user's viewpoint in the three-dimensional environment has not corresponded to the second viewpoint (and / or moved beyond a threshold movement amount) for at least a threshold amount (e.g., 0.1, 1, 3, 7, 10, 20, or 30 seconds), the first media user interface will not be displayed at the second individual location until the user's viewpoint in the three-dimensional environment meets one or more criteria. In some embodiments, if one or more criteria are not met (e.g., the user's viewpoint in the three-dimensional environment has not corresponded to the second viewpoint for at least the threshold time amount described above), the first media user interface will continue to be displayed at a location within the three-dimensional environment based on the user's first viewpoint. In some embodiments, when the user's viewpoint shifts to a second viewpoint, the first media user interface in the first individual location fades out and fades back in after one or more criteria are met. In some embodiments, the first media user interface does not change its location in the three-dimensional environment until one or more criteria are met.Therefore, while the electronic device discontinues displaying the first media user interface in the second separate location, the electronic device optionally remains displayed in the first separate location within the three-dimensional environment.
[0250] Displaying the first media user interface following a shift in the user's viewpoint within a three-dimensional environment, or delaying its display, provides an efficient method for displaying the first media user interface to the user's new viewpoint after the user's movement has settled, thereby reducing the cognitive burden on the user both when engaging with the first media user interface and when engaging with other content or applications within the three-dimensional environment.
[0251] In some embodiments, the first media user interface is associated with a separate application (e.g., a video application, a media application, or a streaming application). In some embodiments, the orientation of the first media user interface in a second presentation mode (e.g., whether the first media user interface is displayed in portrait or landscape mode) is defined by the application associated with the first media user interface. In some embodiments, the type of application associated with the first media user interface defines the orientation of the first media user interface in a second presentation mode. In some embodiments, during a second presentation mode, the first media user interface is automatically oriented toward the user's viewpoint (e.g., perpendicular to the user's viewpoint) so that the content within the first media user interface is displayed toward the user's viewpoint in a three-dimensional environment (e.g., angled).
[0252] In some embodiments, while the first media user interface is presenting content in a second presentation mode, the first media user interface includes one or more selectable user interface elements to modify the playback of the content, such as user interface elements 936-948 in user interface 934 in Figure 9D (1006a). In some embodiments, while the first media user interface is displaying one or more user interface elements, the electronic device receives input via one or more input devices corresponding to the selection of individual user interface elements of one or more user interface elements, such as the selection of user interface element 936 in Figure 9D (1006b).
[0253] In some embodiments, upon receiving input, the electronic device modifies content playback according to the selection of individual user interface elements (1006c). For example, upon detection of the selection of user interface element 936 by the electronic device 101, the electronic device 101 transitions the playback of TV program A from picture-in-picture presentation to extended presentation mode, as described in more detail with reference to Figure 9D. For example, when content is presented in picture-in-picture mode (e.g., second presentation mode), user interface elements for modifying content playback are displayed overlaid on the first media user interface. In some embodiments, the user interface elements are integrated within the first media user interface (as opposed to overlaying the first media user interface) so that the content presented within the first media user interface and the user interface elements are at the same Z-depth in a three-dimensional environment. In some embodiments, user interface elements for modifying content playback include user interface elements for playback, pause, fast forward, rewind, subtitle display, and / or audio associated with the content presented in the first media user interface. In some embodiments, the first media user interface also includes user interface elements associated with playing the content in a third (e.g., immersive) presentation mode if the content presented in the first media user interface is immersive content (as described in more detail in Method 1400). In some embodiments, one or more user interface elements are displayed in the first media user interface after the electronic device detects that the user of the electronic device has performed a pinch gesture (e.g., with the thumb and index finger of the user's hand) while the user's gaze is directed towards the first media user interface element.In some embodiments, the user interface element is displayed on the first media user interface when only the user's line of sight is directed towards the first media user interface and / or when the user's line of sight is directed towards the first media user interface while the user's hand is performing the start of a pinch gesture (e.g., when the thumb and index finger of the user's hand are separated by more than a threshold distance (e.g., 0.5, 1, 1.5, 3, or 6 cm) and have not yet converged within the above-mentioned threshold distance of each other, etc.).
[0254] Displaying the user interface element on the first media user interface (e.g., being overlaid or integrated) provides an efficient way to modify the playback of the content and display the user interface elements associated with interacting with such controls, thereby reducing the user's cognitive burden when engaging with the first media user interface and when modifying the playback of the first media user interface.
[0255] In some embodiments, while the first media user interface presents content in the first presentation mode, the three-dimensional environment includes a playback control user interface separate from the first media user interface that includes one or more user interface elements selectable to modify the playback of the content, and the first media user interface does not include one or more user interface elements selectable to modify the playback of the content (1008a). For example, in FIG. 9B, the playback control user interface 908 is displayed separately from the user interface 906 during the extended presentation mode. In some embodiments, the content is presented in the first presentation mode if the content is not presented in the picture-in-picture presentation mode. In some embodiments, the content is presented in the first presentation mode if the content is presented in a display size larger than the display size of the content during the second presentation mode. In some embodiments, the content is presented in the first presentation mode when the content is presented in the default presentation mode (e.g., playing natively in a video player application associated with the first media user interface).
[0256] In some embodiments, while displaying a playback control user interface, the electronic device 101 receives input via one or more input devices corresponding to the selection of individual user interface elements of one or more user interface elements (1008b). For example, in Figure 9B, the electronic device 101 detects the selection of user interface element 908b. In some embodiments, in response to receiving the input, the electronic device modifies the playback of the content according to the selection of individual user interface elements (1008c). For example, in Figure 9C, in response to the electronic device 101 detecting the selection of user interface element 908b in Figure 9B, the electronic device 101 displays TV program A in the picture-in-picture user interface 934. For example, during the presentation of content in a first presentation mode, user interface elements associated with modifying the playback of content presented in the first media user interface are displayed in a playback control user interface separate from the first media user interface (e.g., not integrated with and / or overlaid on the first media user interface). In some embodiments, one or more user interface elements include options to play / pause content, advance content by a predetermined amount (e.g., 15, 30, 60, 90 seconds), and rewind content by a predetermined amount (e.g., 15, 30, 60, 90 seconds) in order to modify the display of subtitles and / or audio associated with content presented in the first media user interface. In some embodiments, the playback control user interface is angled toward the user's viewpoint in a three-dimensional environment in a different way than the first media user interface (e.g., perpendicular to the user's viewpoint). For example, in some embodiments, the playback control user interface is displayed at an upward tilt relative to a fixed reference frame, while the first media user interface is displayed parallel to a fixed reference frame. In some embodiments, both the playback control user interface and the first media user interface are perpendicular to the user's viewpoint.In some embodiments, the playback control user interface includes an option to initiate playback of content in a third presentation mode (e.g., immersive presentation) if the content is immersive content, as will be described in more detail with reference to Method 1400. In some embodiments, the playback control user interface is displayed in the three-dimensional environment after the electronic device detects that the user of the electronic device has performed a pinch gesture (e.g., with the thumb and index finger of the user's hand) while the user's gaze is directed towards a first media user interface element. In some embodiments, the playback control user interface is displayed in the three-dimensional environment when only the user's gaze is directed towards the first media user interface and / or when the user's gaze is directed towards the first media user interface while the user's hand is performing the initiation of a pinch gesture (e.g., when the thumb and index finger of the user's hand are separated by more than a threshold distance (e.g., 0.5, 1, 1.5, 3, 6 cm) and have not yet converged to each other within the aforementioned threshold distances).
[0257] Displaying user interface elements for modifying the playback of content within the first media user interface in a separate user interface during the first presentation mode provides an efficient way to access and interact with such user interface elements during the first presentation mode, thereby reducing the cognitive burden on the user when engaging with the first media user interface and modifying the playback of the first media user interface.
[0258] In some embodiments, the content is presented in the first presentation mode while the content is displayed in the first media user interface (for example, in some embodiments, the content is presented in the first presentation mode if the content is not presented in picture-in-picture presentation mode). In some embodiments, the content is presented in the first presentation mode if the content / first media user interface is presented at a larger display size than the content / first media user interface in the second presentation mode. In some embodiments, the content is presented in the default presentation mode (for example, in the video player application associated with the first media user interface). When the content is being played natively in the context, it is presented in a first presentation mode. While simultaneously displaying a first media user interface that has first individual user interface elements that can be selected to transition the content from the first presentation mode to a second presentation mode (for example, a user interface element that can be selected to transition from the state in which the content is presented in the first presentation mode to a second presentation mode (for example, a picture-in-picture presentation mode)), the electronic device 101 receives a first input via one or more input devices that corresponds to the selection of a first individual user interface element, such as the selection of user interface element 908b in Figure 9B (1010a).
[0259] In some embodiments, in response to the reception of a first input (e.g., the discontinuation of content presentation in the media user interface) (1010b), the electronic device displays a second media user interface (e.g., different from the first media user interface) presenting the content in a second presentation mode via a display generation component (1010c). For example, in response to the electronic device 101 detecting the selection of user interface element 908b in Figure 9B, the electronic device 101 transitions the display of TV program A from user interface 906 to picture-in-picture user interface 934 in Figure 9C. For example, after receiving a first input, the content transitions from playback in the first media user interface of the media application to a second media user interface (e.g., picture-in-picture user interface). In some embodiments, the second media user interface is smaller in size than the first media user interface (e.g., has a smaller width and / or height than the first media user interface in a three-dimensional environment). In some embodiments, the portion of the user's field of view occupied by the second media user interface is smaller than the portion of the user's field of view occupied by the first media user interface.
[0260] In some embodiments, upon receiving a first input (1010b), the electronic device displays in the first media user interface (1010d) one or more selectable representations of one or more content items, including a first selectable representation of the first content item, such as representations 940-958 in Figure 9C, which are selectable to play the first content item in the first or second media user interface. For example, after receiving the first input, the first media user interface (e.g., the one that presented the content before it began playing in the second media user interface) begins displaying the selectable representations of the content item to trigger playback of the corresponding content item (e.g., within the first or second media user interface). In some embodiments, the content items corresponding to one or more selectable representations correspond to content items recommended based on the user's content consumption history. In some embodiments, the content items corresponding to one or more selectable representations correspond to trending, popular, and / or newly released content items.
[0261] Updating the first media user interface to include representations of additional content items when content presented in the first media user interface begins to be displayed in a different user interface provides an efficient way to access additional content items simultaneously with the content being presented in the second presentation mode (without requiring additional input), thereby reducing the cognitive burden on the user when engaging with the first media user interface and when modifying the presentation of content presented in the first media user interface.
[0262] In some embodiments, upon receiving a first input, before presenting the content in a second presentation mode in a second media user interface, the electronic device displays an animation of the content transitioning from a first presentation mode in the first media user interface to a second presentation mode in the second media user interface (1012a). For example, an animation is displayed when the electronic device 101 transitions the presentation of TV program A from the extended presentation mode shown in Figure 9B to the picture-in-picture presentation mode shown in Figure 9C. For example, after receiving a request to change the presentation of content from a state where it is presented in a first presentation mode to a second presentation mode, an animation is displayed indicating that the content is transitioning from the first presentation mode to the second presentation mode. In some embodiments, the animation includes content that is not visually highlighted (e.g., fades out) in the first media user interface and / or content that is visually highlighted (e.g., fades in) in the second media user interface. In some embodiments, the animation includes visually highlighting the second media user interface to indicate that the content is currently being presented in the second media user interface. In some embodiments, the second media user interface continues to be highlighted or visually emphasized until the content is presented to the second media user interface for a threshold amount of time (e.g., 5, 10, 20, 40, 60, 120 seconds) or until the user's attention is directed to the second media user interface (e.g., the user's gaze is directed to the second media user interface). In some embodiments, the animation includes content that shrinks and / or moves in a three-dimensional environment from the location of the first media user interface to the location of the second media user interface to be displayed.
[0263] Displaying animation when content transitions from a first presentation mode to a second presentation mode provides an efficient way to indicate the current presentation mode associated with the content, thereby reducing the cognitive burden on the user when engaging with content presented in the first media user interface.
[0264] In some embodiments, while content is being presented in a second presentation mode in a second media user interface (for example, while content is being presented in a picture-in-picture user interface), the electronic device receives a second input via one or more input devices that corresponds to a request to change the presentation of the content from a second presentation mode to a first presentation mode, such as an input for selecting a user interface element 936 in Figure 9D (1014a). In some embodiments, the request to transition the presentation of the content from a second presentation mode to a first presentation mode is received when a user interface element displayed in or with the second media user interface is selected as described above. In some embodiments, in response to receiving a second input (1014b), the electronic device ceases to display the second media user interface and one or more selectable representations within the second media user interface (1014c). For example, the second user interface (e.g., a picture-in-picture user interface) stops being displayed in the three-dimensional environment when the presentation mode associated with the content switches from a second presentation mode to a first presentation mode. In some embodiments, an electronic device presents content in a first media user interface (1014d), and the content is presented in a first presentation mode while the content is displayed in the first media user interface. For example, if electronic device 101 detects a request to transition the playback of TV program A from picture-in-picture presentation mode to extended presentation mode, electronic device 101 replaces expressions 940-958 in user interface 906 with the playback of TV program A. For example, when the presentation mode associated with the content switches to the first presentation mode, the first media user interface in the three-dimensional environment begins presenting the content.In some embodiments, when the content presentation mode switches to a first presentation mode, the location of the content in the three-dimensional environment changes from a location corresponding to a second media user interface (e.g., a picture-in-picture user interface) to a location in the three-dimensional environment corresponding to the location of the application that is currently facilitating playback of the content. In some embodiments, when the content is presented in the first presentation mode, the content is presented in a first media user interface that is larger in size than the second media user interface (e.g., the content is therefore displayed in a larger size compared to the size presented in the second media user interface). In some embodiments, if a second input is received while the content is in the first playback position, the electronic device begins to present the content from the first playback position to the first media user interface.
[0265] Presenting content in different media user interfaces based on its content presentation mode provides an efficient way to indicate the current presentation mode associated with the content, thereby reducing the cognitive burden on the user when engaging with content presented in a first media user interface.
[0266] In some embodiments, upon receiving a second input and before presenting the content in a first presentation mode in the first media user interface, the electronic device displays an animation of the content transitioning from a second presentation mode in the second media user interface to a first presentation mode in the first media user interface (1016a). For example, in response to a selection of user interface element 936, the animation is displayed when the electronic device 101 is transitioning the presentation of TV program A from the picture-in-picture presentation mode shown in Figure 9D to an extended presentation mode. For example, upon receiving a request to change the content from being presented in a second presentation mode to a first presentation mode, an animation is displayed indicating that the content is transitioning from a second presentation mode to a first presentation mode. In some embodiments, the animation includes content that fades out (e.g., is not visually emphasized) in the second media user interface and / or content that fades in (e.g., is visually emphasized) in the first media user interface. In some embodiments, the animation includes visually highlighting the first media user interface to indicate that content is currently being presented in the first media user interface (rather than the second media user interface). In some embodiments, the first media user interface remains highlighted or visually emphasized until the content is presented within the first media user interface for a threshold amount of time (e.g., 5, 10, 20, 40, 60, 120 seconds) or until the user's attention is directed to the second media user interface (e.g., the user's gaze is directed to the first media user interface). In some embodiments, the animation includes content expanding and / or moving within a three-dimensional environment from the location of the second media user interface to the location of the first media user interface.Displaying animation when content transitions from a second presentation mode to a first presentation mode provides an efficient way to indicate the current presentation mode associated with the content, thereby reducing the cognitive burden on the user when engaging with content presented in the first media user interface.
[0267] In some embodiments, while content is being presented in a first presentation mode in a first media user interface (for example, in some embodiments, the content is presented in the first presentation mode if the content is not presented in picture-in-picture presentation mode; in some embodiments, the content is presented in the first presentation mode if the content is presented at a display size larger than the display size of the content in a second presentation mode; in some embodiments, the content is presented in the first presentation mode when the content is presented in a default presentation mode (for example, when it is being played natively in a video player application associated with the first media user interface), the electronic device receives a first input via one or more input devices that corresponds to a request to display a first user interface of a first application (1018a) (for example, a request to launch a new application in a three-dimensional environment is received). In some embodiments, in response to receiving the first input (1018b), the electronic device displays the first user interface of the first application in a three-dimensional environment (1018c). For example, when an electronic device receives a request to open / launch a first application in a three-dimensional environment, the user interface of the first application is displayed in the three-dimensional environment. In some embodiments, in response to receiving a first input (1018b), the electronic device ceases to present content in a first presentation mode in the first media user interface (1018d). For example, if an application in the three-dimensional environment is launched while content is being presented in the first presentation mode, the content stops being presented in the first presentation mode. In some embodiments, the first media user interface also ceases to be displayed in the three-dimensional environment.In some embodiments, upon receiving a first input (1018b), the electronic device displays a second media user interface presenting content in a three-dimensional environment (1018e), and the content is presented in a second presentation mode while the content is presented in the second media user interface. For example, in Figure 9A, if the electronic device 101 receives a request to launch a new application in the three-dimensional environment 904, the electronic device 101 automatically transitions the presentation of TV program A from extended presentation mode to picture-in-picture presentation mode. In some embodiments, the second media user interface is displayed simultaneously with the first user interface of the first application. For example, if a request to launch a new application is received in the three-dimensional environment while the content is presented in the first presentation mode, the content begins playback in a different user interface (e.g., picture-in-picture user interface) and a different presentation mode (e.g., second presentation mode). In some embodiments, the default launch location of the first application corresponds to the current location of the first media user interface in the three-dimensional environment, so launching the first application in the three-dimensional environment transitions the content from a first presentation mode to a second presentation mode. In some embodiments, the display location of the first user interface occludes (or partially occludes) the first media user interface, so launching the first application in the three-dimensional environment transitions the content from a first presentation mode to a second presentation mode. In some embodiments, the content presented in the second user interface is not occluded by the first user interface of the first application.
[0268] Switching the content presentation mode from a first presentation mode to a second presentation mode when a request to launch a new application in a three-dimensional environment is received provides an efficient way to continue displaying content in the three-dimensional environment when displaying a new user interface in the three-dimensional environment, thereby reducing the cognitive burden on the user when interacting with content presented in the first media user interface.
[0269] In some embodiments, the electronic device detects that content playback has reached a predetermined playback threshold (e.g., playback is complete, or playback of a content item is within a threshold time period from the end (e.g., content playback ends in 0.5, 1, 1.5, 3, 5, 10, or 20 minutes)) (1020a). In some embodiments, in response to detecting that content playback has reached a predetermined playback threshold, the electronic device displays a second user interface in the three-dimensional environment that, when selected, includes one or more representations of recommended content to initiate playback of the corresponding content in the first media user interface (1020b). For example, in Figure 9E, the electronic device 101 displays the user interface 946 in response to detecting that playback of TV program A has been completed. For example, when content playback reaches a predetermined playback threshold, a second media user interface is displayed in the three-dimensional environment that includes one or more representations of selectable content to initiate playback of new content in the first media user interface. In some embodiments, a second user interface is displayed when content is presented in a first presentation mode, and not displayed when content is not presented in a first presentation mode. In some embodiments, content corresponding to one or more representations corresponds to recommended content based on the electronic device user's content consumption history and / or because the user has previously saved / favorited the content. In some embodiments, a playback control user interface is displayed simultaneously with the first media user interface and / or below the first media user interface presenting the content.
[0270] Displaying a second user interface, including a selection of recommended content, when the content being played in the first media user interface reaches a predetermined playback position provides an efficient way to provide access to other content that can be played in the first media user interface, thereby reducing the cognitive burden on the user when engaging with the content presented in the first media user interface.
[0271] In some embodiments, one or more representations of recommended content include a first individual representation of the first recommended content (1022a) (for example, as described above, the representation corresponding to the first content item is displayed when the content reaches a predetermined playback threshold). In some embodiments, while the user's gaze is directed to the first individual representation (1022b), the electronic device 101 begins playback of the first recommended content (for example, in the first media user interface) (1022c) upon determination that the user's gaze has been directed to the first individual representation for a threshold amount of time (for example, without considering any other inputs / gestures performed by the user of the electronic device, such as without detecting input from the user's hand directed to the first individual representation and / or any other element in the three-dimensional environment). For example, if the user's gaze has been directed to the first individual representation for a threshold amount of time (for example, 5, 7, 9, 10, 20, 30, 60 seconds), the content corresponding to the first individual representation (the first recommended content) begins playback in the three-dimensional environment. In some embodiments, while the user's gaze is directed towards a first individual representation (1022b), the electronic device refrains from starting playback of the first recommended content (e.g., in the first media user interface) (1022d) based on the determination that the user's gaze has not been directed towards the first individual representation for a threshold amount of time (1022d). For example, the electronic device 101 starts playing item A if the user's gaze 914 has been directed towards the corresponding representation 946 for the aforementioned threshold amount of time, and does not start playing the item if the user's gaze 914 has not been directed towards the corresponding representation 946 for the aforementioned threshold amount of time. For example, if the user's gaze has not been directed towards the first individual representation for a threshold amount of time (e.g., 5, 7, 9, 10, 20, 30, 60 seconds), the content corresponding to the first individual representation (first recommended content) does not start playing in the three-dimensional environment until the user's gaze has been directed towards the first individual representation for the aforementioned threshold amount of time.Initiating content playback based on the amount of time the user's gaze was directed towards a corresponding representation displayed in a three-dimensional environment provides an efficient way to initiate content playback in a three-dimensional environment without requiring the user to perform additional gestures (e.g., hand gestures), thereby reducing the cognitive burden on the user when engaging with content presented in a second media user interface.
[0272] In some embodiments, while the user's gaze is directed to a first individual representation (for example, while the user's gaze is not directed to the first individual representation beyond a threshold time amount), the electronic device displays a visual indication in relation to the first individual representation (1024a), which is updated when the user's gaze remains directed to the first individual representation to show progress toward reaching a threshold time amount, such as the visual indication 915 in Figure 9E. For example, a visual indication is displayed showing the amount of time remaining before the user's gaze is directed to the first individual representation for at least the aforementioned threshold time amounts (e.g., 5, 7, 9, 10, 20, 30, 60 seconds). In some embodiments, the progress indicator is displayed as an overlay on the first individual presentation when the user's gaze is directed to the first individual representation, and is not displayed when the user's gaze is not directed to the first individual representation. In some embodiments, upon determination that the user's gaze has been directed to a first individual representation for at least a threshold amount of time, the visual indication stops updating and content corresponding to the first individual representation (e.g., first recommended content) begins to play in the three-dimensional environment. In some embodiments, the visual indication is a symbol (e.g., an arrow) that extends in a circular shape when the user's gaze is directed to the first individual representation.
[0273] Providing an indication of when new content will begin playing in a three-dimensional environment, based on the amount of time the user's gaze is directed towards the corresponding representation of the content, provides an efficient way to show when new content will play, thereby reducing the cognitive burden on the user when engaging with content presented within a second media user interface.
[0274] In some embodiments, while presenting content in a second presentation mode, the pose of the first media user interface at a first individual location relative to the first viewpoint is the same as the pose of the first media user interface at a second individual location relative to the user's second viewpoint (1026a). For example, the first media user interface is displayed in a predetermined portion of the user's field of view (e.g., bottom right, top right, bottom left, top left, or bottom center) regardless of the user's current viewpoint in the three-dimensional environment. In some embodiments, the first media user interface is displayed in the same relative position and / or orientation with respect to the user's viewpoint in the three-dimensional environment (e.g., always).
[0275] Displaying the first media user interface in the same pose (e.g., position and / or orientation) relative to the user's viewpoint provides an efficient method for displaying the first media user interface in a uniform manner, regardless of the user's viewpoint in a three-dimensional environment, thereby reducing the cognitive burden on the user when interacting with the first media user interface.
[0276] In some embodiments, the user's first perspective corresponds to a first location within the physical environment of the electronic device, and the user's second perspective corresponds to a second location within the physical environment that is different from the first location (1028a). In some embodiments, the electronic device displays a three-dimensional environment from the user's perspective at a location within a three-dimensional environment corresponding to the physical location of the electronic device within the physical environment of the electronic device. In some embodiments, detecting a movement of the user's perspective includes detecting a movement of at least a portion of the user within the physical environment (e.g., the user's head, torso, or hand). In some embodiments, detecting a movement of the user's perspective includes detecting a movement of the electronic device or a display generation component within the physical environment. In some embodiments, displaying a three-dimensional environment from the user's perspective includes displaying the three-dimensional environment from a perspective view associated with the location of the user's perspective within the three-dimensional environment. In some embodiments, updating the user's perspective causes the electronic device to display a plurality of virtual objects from a perspective view associated with the location of the updated user's perspective. For example, if the electronic device detects a leftward movement in the physical environment, the user's perspective moves left in the three-dimensional environment, and the electronic device updates the positions of the plurality of virtual objects displayed via the display generation component to move right.
[0277] Displaying a three-dimensional environment from a perspective based on the user's physical location provides an efficient way to interact with the three-dimensional environment based on the actual pose and / or location of the user within the physical environment, thereby reducing the cognitive burden on the user when engaging with the three-dimensional environment.
[0278] In some embodiments, while content is presented in a second presentation mode in a first media user interface, and the second media user interface is displayed in a third separate location in a three-dimensional environment (for example, while content in the first media user interface is presented in a picture-in-picture user interface in a three-dimensional environment, and the second media user interface is displaying a content recommendation representation), the electronic device receives a second input via one or more input devices corresponding to a request to change the content presentation from the second presentation mode to the first presentation mode (1030a). In some embodiments, the electronic device receives the input because the user has selected a user interface element associated with changing the content presentation mode from the second presentation mode to the first presentation mode, as described above. In some embodiments, in response to receiving the second input (1030b), the electronic device stops displaying the first media user interface (1032c). In some embodiments, upon receiving a second input (1030b), the electronic device presents the content within the second media user interface to a third distinct location (1032d), based on a determination that the second media user interface is within a second distinct range of pauses relative to the user's second viewpoint. For example, if the electronic device 101 receives a request to transition to playback of TV program A in Figure 9C, the location of user interface 906 does not change because user interface 906 is currently within the field of view of user 922 from the user's current viewpoint in the three-dimensional environment. For example, if a request is received to transition the content from a second presentation mode to a first presentation mode while the second media user interface (e.g., a video player / video application user interface) is within the user's field of view, the content begins playback within the second media user interface without a change in the location of the second media user interface in the three-dimensional environment.In some embodiments, the second distinct range of poses relative to the user's second viewpoint includes all (or a subset thereof) of poses in the three-dimensional environment that are within the user's field of view from the user's second viewpoint. In some embodiments, in accordance with the determination that the second media user interface is not within the second distinct range of poses relative to the user's second viewpoint (1032e) (for example, in some embodiments, if the second media user interface is not in a location in the three-dimensional environment that is within the user's field of view from the second viewpoint, then the second media user interface is not within the second distinct range of poses relative to the user's second viewpoint), the electronic device displays the second media user interface in a fourth distinct location different from a third distinct location in the three-dimensional environment (1032f), and displaying the second media user interface in the fourth distinct location causes the second media user interface to be displayed in a distinct pose that is within the second distinct range of poses relative to the second viewpoint, and the second media user interface includes content. For example, if electronic device 101 detects a request to transition the playback of a TV program from picture-in-picture presentation mode to extended presentation mode, electronic device 101 updates the location of user interface 906 so that it is within the user's field of view from the user's current viewpoint in the three-dimensional environment 904, and presents TV program A at the new location of user interface 906 in the three-dimensional environment 904. For example, if a request to transition content from a second presentation mode to a first presentation mode is received while a second media user interface (e.g., a video player / video application user interface) is not within the user's field of view from the second viewpoint, the location of the media user interface moves to a location within the user's field of view from the user's second viewpoint.Moving the location of the second media user interface within a three-dimensional environment when the second user interface is not within the user's field of view provides an efficient way to move the second media user interface into the user's field of view when a request to present content within the second user interface is received (when the first media user interface is not currently within the user's field of view), thereby reducing the cognitive burden on the user when interacting with the second media user interface.
[0279] Figures 11A to 11E illustrate examples of how electronic devices can improve navigation to individual playback positions of content items, according to several embodiments.
[0280] Figure 11A shows an electronic device 101 that displays a three-dimensional environment 1102 via a display generation component 120. In some embodiments, it should be understood that the electronic device 101 utilizes one or more techniques described with reference to Figures 11A to 11E in a two-dimensional environment without departing from the scope of this disclosure. As described above with reference to Figures 1 to 6, the electronic device 101 optionally includes a display generation component 120 (e.g., a touchscreen) and a plurality of image sensors 314. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and the electronic device 101 can be used to capture one or more images of the user or a part of the user while the user is interacting with the electronic device 101. In some embodiments, the display generation component 120 is a touchscreen that can detect the user's hand gestures and movements. In some embodiments, the user interface described below can also be implemented in a head-mounted display that includes a display generating component for displaying the user interface to the user, and sensors for detecting the physical environment and / or the movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward towards the user's face).
[0281] In Figure 11A, the electronic device 101 presents a content item 1104 in a three-dimensional environment 1102. In some embodiments, the content item 1104 is an item of video content. In addition to the content item 1104, the three-dimensional environment 1102 includes representations of real objects in the physical environment of the electronic device 101, such as a wall representation 1108a, a ceiling representation 1108b, a table representation 1106a, and a sofa representation 1106b. The electronic device 101 displays the content item 1104 with increased visual emphasis over the rest of the three-dimensional environment 1102, such as displaying an area of the three-dimensional environment 1102 that does not contain the content item 1104 and has a greater amount of blur and / or darkening than the content item 1104, and displays virtual lighting effects 1110a-d to simulate light spills emanating from the content item 1104.
[0282] As shown in Figure 11A, the user's gaze 1113a is directed towards content item 1104 while the content item is being played. The playback position 1106 of content item 1104 in Figure 11A advances as playback of the content item continues. In some embodiments, if the user's attention is diverted from content item 1104, the electronic device 101 continues to play the content item. In some embodiments, when the user returns attention to content item 1104 after diverting their attention, the electronic device 101 presents a selectable option, which, if selected, causes the electronic device 101 to update the playback position to the playback position that was playing at the time the user's attention was diverted from content item 1104.
[0283] For example, in Figure 11B, while the playback position 1106 of the content item is the playback position shown in the figure, the user diverts their attention from the content item 1104. In some embodiments, detecting that the user has diverted their attention from the content item 1104 includes detecting the user's line of sight 1113b directed to a location in the three-dimensional environment 1102 other than the content item 1104. In some embodiments, in response to detecting the user's line of sight 1113b directed away from the content item 1104, the electronic device 101 reduces the amount of visual emphasis on the content item 1104 relative to the rest of the three-dimensional environment 1102. In some embodiments, the user's line of sight 1113b must be directed away from the content item 1104 for a threshold period (e.g., 1, 2, 3, 5, 10, 15, 30, or 45 seconds, 1, 2, 3, or 5 minutes) for the electronic device 101 to determine that the user's attention is directed away from the content item 1104. In some embodiments, the electronic device 101 determines that the user's attention is being directed away from the content item 1104 at the moment the user's gaze 1113b is directed away from the content item 1104. In some embodiments, detecting that the user has diverted their attention from the content item 1104 optionally includes detecting that the user closes their eyes, as shown in Legend 1123, for at least a threshold period corresponding to, for example, the user falling asleep (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, 5, 10, 15, 30, or 45 seconds, 1, 2, 3, or 5 minutes). In some embodiments, the electronic device 101 continues to play the content item 1104 after detecting that the user's attention is being directed away from the content item 1104, advancing the playback position 1106 beyond the point shown in Figure 11B, which is the playback position of the content item at the point in time when the electronic device 101 determines that the user's attention is being directed away from the content item 1104.
[0284] Figure 11C shows the electronic device 101 presenting selectable options 1112a and 1112b, when selected, to resume playback of content item 1104 from a playback position 1106 corresponding to the moment the user diverted their attention from content item 1104. In some embodiments, the electronic device 101 presents selectable options 1112a and / or selectable options 1112b in response to detecting the user's attention directed towards content item 1104 and / or the user's hand 1103b in a ready state pose. In some embodiments, detecting the user's attention directed towards content item 1104 includes detecting the user's line of sight 1103d directed towards content item 1104. In some embodiments, in response to detecting the line of sight 1103d directed towards content item 1104, the electronic device 101 increases the visual emphasis of content item 1104 against the rest of the three-dimensional environment 1102. In some embodiments, detecting a hand 1103b in a ready-to-play pose includes detecting a hand 1103b in a pre-pinch hand shape where the thumb is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, or 5 centimeters) of another finger of the hand 1103b but not touching another finger of the hand 1103b, or detecting a hand 1103b in a pointing hand shape where one or more fingers are extended and one or more fingers are curled toward the palm. As shown in Figure 11C, the electronic device 101 also displays options 1114a and 1114b and user interface element 1116, which include additional options for modifying the playback of content item 1104, in response to detecting a gaze line 1103d directed towards content item 1104 and / or hand 1103b in a ready-to-play pose, as described in more detail above with reference to Method 800.
[0285] In some embodiments, the electronic device 101 presents both options 1112a and 1112b. In some embodiments, the electronic device 101 presents either option 1112a or option 1112b, but not both. Option 1112a is displayed outside of the user interface element 1116 overlaid on the content item 1104. Option 1112b is displayed as part of the scrub bar 1111 included in the user interface element 1116. The scrub bar 1111 includes an indication 1113 of the current playback position of the content item 1104 (which optionally continues playback while options 1112a and / or 1112b are displayed). The electronic device 101 displays option 1112b at the location of the scrub bar 1111 corresponding to the playback position where playback of the content item 1104 resumes, depending on the selection of option 1112b. In some embodiments, the playback position at which playback of content item 1104 resumes is the playback position in Figure 11B at which the user's attention is diverted from content item 1104. In some embodiments, the playback position at which playback of content item 1104 resumes is a predetermined time (e.g., 1, 2, 3, 5, 10, 15, or 30 seconds) before or after the playback position in Figure 11B at which the user's attention is diverted from content item 1104.
[0286] As shown in Figure 11C, the user selects option 1112a with gaze 1103c and hand 1103a, for example, via indirect input. In some embodiments, detecting the selection of option 1112a via indirect input includes detecting that hand 1103a has made a pinch gesture in which the thumb of hand 1103a touches another finger of hand while gaze 1103c is directed towards option 1112a. Figure 11C shows a first input state of hand 1103a and gaze 1103d, and a second input state of hand 1103b and gaze 1103c, but it should be understood that in some embodiments these input states are detected at different times. In some embodiments other selection inputs are possible. In response to detecting the selection of option 1112a, as will be described in more detail below with reference to Figure 11E, the electronic device 101 updates the playback position 1106 of the content item to the playback position associated with the time the user diverted their attention from the content item 1104.
[0287] Figure 11D illustrates, for example, the selection of selectable option 1112b via direct input. In some embodiments, detecting the selection of selectable option 1112b via direct input includes detecting the user's hand 1103a within a predetermined threshold distance of option 1112b while the hand 1103a is in a predetermined shape. In some embodiments, the predetermined shape is a pinch hand shape where the thumb touches another finger of the hand. In some embodiments, the predetermined shape is a pointing hand shape where one or more fingers are extended and one or more fingers are bent toward the palm. In some embodiments, the direct input includes detecting the user's line of sight 1103e directed toward option 1112b, and in some embodiments, the direct input does not include detecting the user's line of sight 1103e dire...
Claims
1. It is a method, In an electronic device that communicates with a display generation component and one or more input devices, While presenting a content item in a three-dimensional environment, the display generation component displays a user interface associated with the content item, the user interface comprising one or more user interface elements for modifying the playback of the content item, and separate user interface elements for modifying virtual lighting effects affecting the appearance of the three-dimensional environment. While displaying the user interface associated with the content item, the system receives user input via one or more input devices directed to the individual user interface elements, wherein the user input corresponds to a request to modify the virtual lighting effect. Upon receiving the aforementioned user input, In the aforementioned three-dimensional environment, the content item continues to be presented. A method comprising applying the virtual lighting effect to the three-dimensional environment.
2. Before receiving the user input directed to the individual user interface, the electronic device displays the three-dimensional environment via the display generation component without using the virtual lighting effect. The method according to claim 1, wherein, in response to receiving the user input, the electronic device displays the three-dimensional environment using the virtual lighting effect via the display generation component.
3. Before receiving the user input directed to the individual user interface, the electronic device displays the three-dimensional environment using a first amount of the virtual lighting effect via the display generation component. The method according to claim 1 or 2, wherein, in response to receiving the user input, the electronic device displays the three-dimensional environment using a second amount of the virtual lighting effect, wherein the second amount is different from the first amount, via the display generation component.
4. Before receiving the user input, the area of the three-dimensional environment that does not contain the content item is displayed at a first brightness level. The method according to any one of claims 1 to 3, wherein displaying the three-dimensional environment using the virtual lighting effect in response to the user input includes displaying the area of the three-dimensional environment that does not contain the content item at a second brightness level different from the first brightness level.
5. The method according to any one of claims 1 to 4, wherein displaying the three-dimensional environment using the virtual lighting effect in response to user input includes displaying individual virtual lighting effects emanating from the content item on one or more objects in the three-dimensional environment.
6. Before receiving the user input directed to the individual user interface, the electronic device displays the three-dimensional environment having a first amount of virtual lighting effect, which includes, via the display generation component, displaying a region of the three-dimensional environment that does not contain the content item at a first brightness level, and displaying a first amount of individual virtual lighting effect emanating from the content item on one or more objects in the three-dimensional environment. The method according to any one of claims 1 to 5, wherein, in response to receiving the user input, the electronic device displays the area of the three-dimensional environment that does not contain the content item at a second brightness level via the display generation component, and displays the three-dimensional environment having a second amount of the individual virtual lighting effect emanating from the content item on one or more objects in the three-dimensional environment.
7. Displaying the three-dimensional environment using the virtual lighting effect in response to the user input is, In accordance with the detection that the attention of the user of the electronic device is directed to a first area of the three-dimensional environment via one or more input devices, the display generation component displays the three-dimensional environment using a first amount of the virtual lighting effect, The method according to any one of claims 1 to 6, comprising: detecting, via one or more input devices, that the user's attention is directed to a second area of the three-dimensional environment different from the first area, and then displaying the three-dimensional environment via the display generation component using a second amount of the virtual lighting effect different from the first amount.
8. While the three-dimensional environment is being displayed without using the virtual lighting effect, receiving a first input directed to the individual user interface element via one or more input devices, which includes detecting a default portion of the user of the electronic device that is in a default pause for a period of less than a predetermined time threshold via one or more input devices, In response to receiving the first input, the three-dimensional environment is displayed using the virtual lighting effect via the display generation component, While the three-dimensional environment is being displayed using the virtual lighting effect, a second input directed to the individual user interface element is received via one or more input devices, which includes detecting the default portion of the user of the electronic device in the default pose for a period of less than a predetermined time threshold. The method according to any one of claims 1 to 7, further comprising displaying the three-dimensional environment via the display generation component without using the virtual lighting effect in response to receiving the second input.
9. Receiving input directed to the individual user interface elements via one or more input devices, including detecting the movement of the user's default portion of the electronic device via one or more input devices while the three-dimensional environment is displayed using a first amount of the virtual lighting effect, while the user's default portion is in a default pose, The method according to any one of claims 1 to 8, further comprising, in response to the input directed to the individual user interface elements, displaying the three-dimensional environment via the display generation component using a second amount of virtual lighting effect, the second amount being based on the movement of the default portion of the user while the default portion of the user is in the default pose.
10. While the aforementioned content item is being played, The three-dimensional environment is displayed using the virtual lighting effect via the display generation component. Receiving user input corresponding to a request to pause the content item via one or more input devices, Upon receiving the user input corresponding to the request to pause the content item, Pause the aforementioned content item, The method according to any one of claims 1 to 9, further comprising displaying the three-dimensional environment via the display generation component without using the virtual lighting effect.
11. Receiving individual user inputs via one or more input devices directed to a second individual interface element of one or more user interface elements for modifying the playback of the content item, In response to receiving the individual user inputs, When the second individual user interface element is selected, the system determines that it is a user interface element that causes the electronic device to switch between playing and pausing the content item, and switches the playback or pause state of the content item accordingly. The method according to any one of claims 1 to 10, further comprising updating the playback position of the content item in accordance with the individual user input, based on the determination that the second individual user interface element, when selected, is a user interface element that causes the electronic device to update the playback position of the content item.
12. Receiving individual user inputs via one or more input devices directed to a second individual interface element of one or more user interface elements for modifying the playback of the content item, In response to receiving the individual user inputs, The method according to any one of claims 1 to 11, further comprising: modifying the volume of the audio content according to the individual input, based on the determination that the second individual user interface element, when selected, is a user interface element that causes the electronic device to modify the volume of the audio content of the content item.
13. The method according to any one of claims 1 to 12, wherein the user interface associated with the content item is a separate user interface from the content item and is displayed between the content item and the viewpoint of the user of the electronic device in the three-dimensional environment via the display generation component.
14. The content item is displayed via the display generation component at a first angle with respect to the user's viewpoint in the three-dimensional environment. The method according to claim 13, wherein the user interface associated with the content item is displayed via the display generation component at a second angle different from the first angle with respect to the user's viewpoint in the three-dimensional environment.
15. The method according to any one of claims 1 to 14, wherein the electronic device displays one or more user interface elements for modifying the playback of the content item in response to detecting a default portion of the user of the electronic device in a pose that satisfies one or more criteria via one or more input devices.
16. While displaying one or more user interface elements for correcting the playback of the content item, the system detects, via one or more input devices, the default portion of the user in a pose that does not meet one or more criteria, The method of any one of claim 15, further comprising: detecting the default portion of the user in the pose that does not meet one or more of the criteria, reducing the visual emphasis of one or more user interface elements for modifying the playback of the content item, which the electronic device displays via the display generation component.
17. While the content item is being displayed at a first size and the user interface associated with the content item at a second size via the display generation component, input corresponding to a request to resize the content item is received via one or more input devices. Upon receiving the input corresponding to the request to resize the content item, The content item is displayed in a third size different from the first size, according to the input corresponding to the request to resize the content item via the display generation component. The method according to any one of claims 1 to 16, further comprising displaying the user interface associated with the content item in the second size via the display generation component.
18. While displaying the content item at a first size and first distance from the user's viewpoint in the three-dimensional environment via the display generation component, and displaying the user interface associated with the content item at a second size and second distance from the user's viewpoint in the three-dimensional environment, the system receives input via one or more input devices corresponding to a request to rearrange the content item in the three-dimensional environment. Upon receiving the input corresponding to the request to rearrange the content items, In accordance with the input corresponding to the request to rearrange the content item via the display generation component, the content item is displayed in the three-dimensional environment at a third distance different from the first distance from the user's viewpoint and at a third size. The method according to any one of claims 1 to 17, further comprising displaying the user interface associated with the content item at a second size and a fourth distance from the user's viewpoint, in accordance with the input corresponding to the request to rearrange the content item via the display generation component.
19. The content item is separate from the user interface associated with the content item in the three-dimensional environment, and the method is The method according to any one of claims 1 to 18, further comprising displaying one or more second user interface elements for modifying the playback of the content item via the display generation component, wherein the one or more second user interface elements are displayed as overlays on the content item in the three-dimensional environment.
20. The detection of whether the user's attention directed to individual user interface elements of the one or more user interface elements satisfies one or more first criteria via the one or more input devices, In response to the detection that the user's attention directed towards the individual user interface element satisfies one or more of the first criteria, In accordance with the determination that the individual user interface elements satisfy one or more second criteria, a visual indication identifying the function of the individual user interface elements is displayed via the display generation component. The method according to any one of claims 1 to 19, further comprising: discontinuing the display of the visual indication that identifies the function of the individual user interface element in accordance with the determination that the individual user interface element does not satisfy one or more second criteria.
21. To display individual user interface elements that are displayed separately from the content item and the user interface associated with the content item via the display generation component, While the individual user interface elements are being displayed, input directed to the individual user interface elements is received via one or more input devices. The method according to any one of claims 1 to 20, further comprising detecting the input directed to the individual user interface element and initiating a process for resizing the content item in the three-dimensional environment according to the input directed to the individual user interface element.
22. It is an electronic device, One or more processors, Memory and An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions, and the instructions are While presenting a content item in a three-dimensional environment, a user interface associated with the content item is displayed via a display generation component, the user interface comprising one or more user interface elements for modifying the playback of the content item, and separate user interface elements for modifying virtual lighting effects affecting the appearance of the three-dimensional environment. While displaying the user interface associated with the content item, the system receives user inputs via one or more input devices directed to the individual user interface elements, wherein the user inputs correspond to a request to modify the virtual lighting effect. Upon receiving the aforementioned user input, In the aforementioned three-dimensional environment, the content item continues to be presented. An electronic device for applying the aforementioned virtual lighting effect to the three-dimensional environment.
23. A non-temporary computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of an electronic device, the electronic device... While a content item is being presented in a three-dimensional environment, a user interface associated with the content item is displayed via a display generation component, the user interface comprising one or more user interface elements for modifying the playback of the content item, and separate user interface elements for modifying virtual lighting effects that affect the appearance of the three-dimensional environment. While displaying the user interface associated with the content item, the system receives user inputs directed to the individual user interface elements via one or more input devices, the user inputs corresponding to requests to modify the virtual lighting effect. Upon receiving the aforementioned user input, The content item is continuously displayed in the aforementioned three-dimensional environment. A non-temporary computer-readable storage medium for applying the aforementioned virtual lighting effect to the three-dimensional environment.
24. It is an electronic device, One or more processors, Memory and While a content item is being presented in a three-dimensional environment, means for displaying a user interface associated with the content item via a display generation component, the user interface comprising one or more user interface elements for modifying the playback of the content item, and individual user interface elements for modifying virtual lighting effects affecting the appearance of the three-dimensional environment. While displaying the user interface associated with the content item, means for receiving user input directed to the individual user interface elements via one or more input devices, wherein the user input corresponds to a request to modify the virtual lighting effect, Upon receiving the aforementioned user input, In the aforementioned three-dimensional environment, the content item continues to be presented. An electronic device comprising means for applying the virtual lighting effect to the three-dimensional environment.
25. An information processing device for use in an electronic device, wherein the information processing device is While a content item is being presented in a three-dimensional environment, means for displaying a user interface associated with the content item via a display generation component, the user interface comprising one or more user interface elements for modifying the playback of the content item, and individual user interface elements for modifying virtual lighting effects affecting the appearance of the three-dimensional environment. While displaying the user interface associated with the content item, means for receiving user input directed to the individual user interface elements via one or more input devices, wherein the user input corresponds to a request to modify the virtual lighting effect, Upon receiving the aforementioned user input, In the aforementioned three-dimensional environment, the content item continues to be presented. An information processing apparatus comprising means for applying the virtual lighting effect to the three-dimensional environment.
26. It is an electronic device, One or more processors, Memory and An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method according to any one of claims 1 to 21.
27. A non-temporary computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 21.
28. It is an electronic device, One or more processors, Memory and An electronic device comprising means for carrying out the method according to any one of claims 1 to 21.
29. An information processing device for use in an electronic device, wherein the information processing device is An information processing apparatus comprising means for performing the method described in any one of claims 1 to 21.
30. It is a method, In an electronic device that communicates with a display generation component and one or more input devices, The content is presented via the display generation component, and the three-dimensional environment is displayed, which includes a first media user interface located at a first individual location within the three-dimensional environment. While the three-dimensional environment is being displayed using the first media user interface at a first individual location in the three-dimensional environment having a pose within an individual range of poses relative to the first viewpoint of the user of the electronic device, the movement of the user's viewpoint in the three-dimensional environment from the first viewpoint to a second viewpoint different from the first viewpoint is detected. A method comprising: detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, and displaying the three-dimensional environment from the second viewpoint via the display generation component, wherein the display is In accordance with the determination that the content is presented in a first presentation mode, the first media user interface is maintained in the first individual location within the three-dimensional environment, wherein the first media user interface is no longer within the individual range of poses relative to the user's second viewpoint. A method comprising: displaying the first media user interface in a second separate location in the three-dimensional environment, different from the first separate location, in accordance with the determination that the content is being presented in a second presentation mode different from the first presentation mode, wherein displaying the first media user interface in the second separate location causes the first media user interface to be displayed in a pose within the separate range of poses with respect to the second viewpoint of the user.
31. The second individual location is based on the second viewpoint, and displaying the first media user interface in the second individual location is In accordance with the determination that the user's movement of the viewpoint after moving to the second viewpoint satisfies one or more criteria, the first media user interface is displayed at the second individual location, The method according to claim 30, further comprising: discontinuing the display of the first media user interface at the second individual location in accordance with the determination that the movement of the user's viewpoint after moving to the second viewpoint does not satisfy one or more of the criteria.
32. While the first media user interface is presenting the content in the second presentation mode, the first media user interface includes one or more user interface elements that can be selected to modify the playback of the content, and the method While displaying one or more user interface elements within the first media user interface, the system receives inputs via one or more input devices corresponding to the selection of individual user interface elements of the one or more user interface elements, The method according to claim 30 or 31, further comprising modifying the playback of the content in accordance with the selection of the individual user interface elements in response to receiving the aforementioned input.
33. While the first media user interface is presenting the content in the first presentation mode, the three-dimensional environment includes a separate playback control user interface from the first media user interface, which includes one or more selectable user interface elements for modifying the playback of the content, and the first media user interface does not include the one or more selectable user interface elements for modifying the playback of the content, and the method is While the playback control user interface is displayed, input corresponding to the selection of individual user interface elements of the one or more user interface elements is received via the one or more input devices. The method according to any one of claims 30 to 32, further comprising modifying the playback of the content in accordance with the selection of the individual user interface elements in response to receiving the aforementioned input.
34. The aforementioned method, While the content is displayed in the first media user interface in the first presentation mode, and while the first media user interface is simultaneously displayed together with a first individual user interface element that can be selected to transition the content from the first presentation mode to the second presentation mode, a first input corresponding to the selection of the first individual user interface element is received via one or more input devices. Upon receiving the first input, A second media user interface presenting the content in the second presentation mode is displayed via the display generation component. The method according to any one of claims 30 to 33, further comprising displaying one or more selectable representations of one or more content items in the first media user interface, including a first selectable representation of the first content item which is selectable to cause playback of the first content item in the first or second media user interface.
35. The method according to claim 34, further comprising displaying an animation of the content transitioning from the first presentation mode in the first media user interface to the second presentation mode in the second media user interface, before presenting the content in the second presentation mode in the second media user interface in response to receiving the first input.
36. While the content is being presented in the second presentation mode in the second media user interface, a second input is received via one or more input devices that corresponds to a request to change the presentation of the content from the second presentation mode to the first presentation mode. Upon receiving the second input, The second media user interface and the one or more selectable representations are discontinued from being displayed on the second media user interface. The method according to claim 34 or 35, further comprising presenting the content in the first media user interface, wherein the content is presented in the first presentation mode while the content is displayed in the first media user interface.
37. The method according to claim 36, further comprising, in response to receiving the second input, displaying an animation of the content transitioning from the second presentation mode in the second media user interface to the first presentation mode in the first media user interface before presenting the content in the first presentation mode.
38. While the content is being presented in the first media user interface in the first presentation mode, a first input corresponding to a request to display the first user interface of the first application is received via one or more input devices, Upon receiving the first input, In the aforementioned three-dimensional environment, the first user interface of the first application is displayed. In the first media user interface, the presentation of the content in the first presentation mode is discontinued. The method according to claims 30 to 37, further comprising, in the three-dimensional environment, the content, wherein the content displays a second media user interface presenting the content, which is presented in the second presentation mode, while the content is presented in the second media user interface.
39. The process involves detecting when the playback of the aforementioned content has reached a predetermined playback threshold, The method according to claims 30 to 38, further comprising: detecting that the playback of the content has reached a predetermined playback threshold, in the three-dimensional environment, displaying a second user interface that, when selected, includes one or more representations of recommended content that initiate playback in the first media user interface for the corresponding content.
40. The one or more expressions of the recommended content include the first individual expression of the first recommended content, and the method is While the user's gaze is directed towards the first individual representation, In accordance with the determination that the user's gaze has been directed towards the first individual representation for a threshold amount of time, playback of the first recommended content is started. The method according to claim 39, further comprising, in accordance with the determination that the user's gaze was not directed at the first individual representation for a threshold amount of time, the method for which playback of the first recommended content is discontinued.
41. The aforementioned method, The method according to claim 40, further comprising displaying a visual indication in relation to the first individual representation, while the user's gaze is directed to the first individual representation, the visual indication being updated to indicate progress toward reaching the threshold time amount while the user's gaze remains directed to the first individual representation.
42. The method according to any one of claims 30 to 41, wherein, while the content is being presented in the second presentation mode, the pose of the first media user interface in the first individual location relative to the first viewpoint is the same as the pose of the first media user interface in the second individual location relative to the second viewpoint of the user.
43. The method according to any one of claims 30 to 42, wherein the first viewpoint of the user corresponds to a first location in the physical environment of the electronic device, and the second viewpoint of the user corresponds to a second location in the physical environment that is different from the first location.
44. While the content is presented in the second presentation mode in the first media user interface, and while the second media user interface is displayed in a third separate location in the three-dimensional environment, a second input is received via one or more input devices, corresponding to a request to change the presentation of the content from the second presentation mode to the first presentation mode. Upon receiving the second input, The display of the first media user interface is discontinued. In accordance with the determination that the second media user interface is within a second individual range of the user's second viewpoint, the content is presented within the second media user interface at the third individual location. In accordance with the determination that the second media user interface is not within the second individual range of the pose relative to the user's second viewpoint, The method according to any one of claims 30 to 43, further comprising displaying the second media user interface in a fourth separate location different from the third separate location in the three-dimensional environment, wherein the display of the second media user interface in the fourth separate location is the second media user interface, the second media user interface includes the content, and causes the second media user interface to be displayed in a separate pose within the second separate range of poses relative to the second viewpoint.
45. It is an electronic device, One or more processors, Memory and An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions, and the instructions are Content is presented via a display generation component, and the three-dimensional environment is displayed, which includes a first media user interface located at a first individual location within the three-dimensional environment. While the three-dimensional environment is being displayed using the first media user interface at a first individual location within the three-dimensional environment having a pose within an individual range of poses relative to the first viewpoint of the user of the electronic device, a movement of the user's viewpoint within the three-dimensional environment from the first viewpoint to a second viewpoint different from the first viewpoint is detected. An electronic device that, in response to detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, displays the three-dimensional environment from the second viewpoint via the display generation component, wherein the display is In accordance with the determination that the content is presented in a first presentation mode, the first media user interface is maintained in the first individual location within the three-dimensional environment, wherein the first media user interface is no longer within the individual range of poses relative to the user's second viewpoint. An electronic device that includes, in accordance with the determination that the content is presented in a second presentation mode different from the first presentation mode, displaying the first media user interface in a second separate location different from the first separate location in the three-dimensional environment, wherein displaying the first media user interface in the second separate location causes the first media user interface to be displayed in a pose within the separate range of poses with respect to the second viewpoint of the user.
46. A non-temporary computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of an electronic device, the electronic device... Content is presented via a display generation component, and the three-dimensional environment is displayed, which includes a first media user interface located at a first individual location within the three-dimensional environment. While the three-dimensional environment is being displayed using the first media user interface at a first individual location within the three-dimensional environment having a pose within an individual range of poses relative to the first viewpoint of the user of the electronic device, the movement of the user's viewpoint within the three-dimensional environment from the first viewpoint to a second viewpoint different from the first viewpoint is detected. A non-temporary computer-readable storage medium that, in response to detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, displays the three-dimensional environment from the second viewpoint via the display generation component, wherein the display is In accordance with the determination that the content is presented in a first presentation mode, the first media user interface is maintained in the first individual location within the three-dimensional environment, wherein the first media user interface is no longer within the individual range of poses relative to the user's second viewpoint. A non-temporary computer-readable storage medium, comprising: displaying the first media user interface in a second separate location different from the first separate location in the three-dimensional environment, in accordance with the determination that the content is being presented in a second presentation mode different from the first presentation mode, wherein displaying the first media user interface in the second separate location includes displaying the first media user interface in a pose within the separate range of poses relative to the second viewpoint of the user.
47. It is an electronic device, One or more processors, Memory and A means for displaying a three-dimensional environment, which includes a first media user interface located at a first individual location within the three-dimensional environment, presenting content via a display generation component, Means for detecting a movement of the user's viewpoint within the three-dimensional environment from a first viewpoint to a second viewpoint different from the first viewpoint, while the three-dimensional environment is being displayed using the first media user interface at a first individual location within the three-dimensional environment having a pose within an individual range of poses relative to the first viewpoint of the user of the electronic device; An electronic device comprising means for displaying the three-dimensional environment from the second viewpoint via the display generation component in response to detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, wherein the display is In accordance with the determination that the content is presented in a first presentation mode, the first media user interface is maintained in the first individual location within the three-dimensional environment, wherein the first media user interface is no longer within the individual range of poses relative to the user's second viewpoint. An electronic device that includes, in accordance with the determination that the content is presented in a second presentation mode different from the first presentation mode, displaying the first media user interface in a second separate location different from the first separate location in the three-dimensional environment, wherein displaying the first media user interface in the second separate location includes displaying the first media user interface in a pose within the separate range of poses with respect to the second viewpoint of the user.
48. An information processing device for use in an electronic device, wherein the information processing device is Means for detecting a movement of the user's viewpoint within the three-dimensional environment from a first viewpoint to a second viewpoint different from the first viewpoint, while the three-dimensional environment is being displayed using the first media user interface at a first individual location within the three-dimensional environment having a pose within an individual range of poses relative to the first viewpoint of the user of the electronic device; An information processing apparatus comprising means for displaying the three-dimensional environment from the second viewpoint via the display generation component in response to detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, wherein the display is In accordance with the determination that the content is presented in a first presentation mode, the first media user interface is maintained in the first individual location within the three-dimensional environment, wherein the first media user interface is no longer within the individual range of poses relative to the user's second viewpoint. Information processing device, which includes, in accordance with the determination that the content is presented in a second presentation mode different from the first presentation mode, displaying the first media user interface in a second separate location different from the first separate location in the three-dimensional environment, wherein displaying the first media user interface in the second separate location includes displaying the first media user interface in a pose within the separate range of poses with respect to the second viewpoint of the user.
49. It is an electronic device, One or more processors, Memory and An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method according to any one of claims 30 to 44.
50. A non-temporary computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of an electronic device, the electronic device causes the electronic device to perform the method according to any one of claims 30 to 44.
51. It is an electronic device, One or more processors, Memory and An electronic device comprising means for carrying out the method described in any one of claims 30 to 44.
52. An information processing device for use in an electronic device, wherein the information processing device is An information processing apparatus comprising means for performing the method described in any one of claims 30 to 44.
53. It is a method, In an electronic device that communicates with a display generation component and one or more input devices, While presenting content items that change over time, and while displaying the user interface associated with the content items via the display generation component, While the playback position within the content item is at a first playback position, it is detected that one or more criteria are met, including criteria that are met when the attention of the user of the electronic device is not directed to the user interface associated with the content item via one or more input devices. After detecting that one or more of the above criteria are met, While the playback position of the content item is a second playback position different from the first playback position, it is detected via one or more input devices that the user's attention is directed to the user interface associated with the content item. If selected, after one or more of the above criteria have been met, and in response to detecting that the user's attention is directed to the user interface associated with the content item, the electronic device displays, via the display generation component, selectable options to present the content item from individual playback positions associated with the first playback position. While the selectable options are displayed, an input corresponding to the selection of the selectable options is detected via one or more input devices. A method comprising updating the playback position of the content item to the individual playback position associated with the first playback position in response to the detection of the aforementioned input.
54. The method according to claim 53, wherein the one or more criteria include a criterion that is satisfied when the electronic device detects, via the one or more input devices, that one or more of the user's eyes are closed for a predetermined threshold period of time.
55. The method according to claim 53 or 54, wherein the one or more criteria include criteria that are satisfied when the electronic device detects, via the one or more input devices, that the user's gaze is directed away from the user interface associated with the content item.
56. The method according to any one of claims 53 to 55, wherein the user interface associated with the content item includes a scrub bar corresponding to the playback of the content item, and the selectable options are displayed via the display generation component at locations in the scrub bar corresponding to the individual playback positions associated with the first playback position.
57. Detecting that the user's attention is directed to the user interface associated with the content item after one or more of the above criteria are met includes detecting, via one or more input devices, a default portion of the user in a pose that satisfies one or more second criteria, and the method is The method according to any one of claims 53 to 56, further comprising detecting, after detecting that one or more of the above criteria are met, that the display of the selectable options is discontinued in response to detecting that the default portion of the user is in a position that does not meet one or more of the above second criteria.
58. In response to detecting that the user's attention is directed to the user interface associated with the content item after one or more of the above criteria have been met, The method according to any one of claims 53 to 57, further comprising simultaneously displaying one or more selectable elements for controlling the playback of the content item together with the selectable options via the display generation component.
59. The method according to any one of claims 53 to 58, further comprising continuing to play the content item from the second playback position while the selectable options are displayed and before the selection of the selectable options is detected.
60. The method according to any one of claims 53 to 59, wherein detecting the input corresponding to the selection of the selectable option includes detecting the user's gaze directed towards the selectable option via one or more input devices while the default portion of the user is performing separate gestures.
61. It is an electronic device, One or more processors, Memory and An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions, and the instructions are While presenting content items that change over time, and while displaying the user interface associated with the content items via the display generation component, While the playback position within the content item is at a first playback position, it is detected that one or more criteria are met, including criteria that are met when the attention of the user of the electronic device is not directed to the user interface associated with the content item, via one or more input devices. After detecting that one or more of the above criteria are met, While the playback position of the content item is a second playback position different from the first playback position, it is detected via one or more input devices that the user's attention is directed to the user interface associated with the content item. If selected, after one or more of the above criteria have been met, and in response to detecting that the user's attention is directed to the user interface associated with the content item, the electronic device displays, via the display generation component, selectable options to present the content item from individual playback positions associated with the first playback position. While the selectable options are displayed, an input corresponding to the selection of the selectable options is detected via one or more input devices. An electronic device that, upon detecting the aforementioned input, updates the playback position of the content item to the individual playback position associated with the first playback position.
62. A non-temporary computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of an electronic device, the electronic device... While presenting content items that change over time, and while displaying the user interface associated with the content items via the display generation component, While the playback position within the content item is a first playback position, one or more criteria are detected via one or more input devices, including criteria that are met when the user's attention of the electronic device is not directed towards the user interface associated with the content item. After detecting that one or more of the above criteria are met, While the playback position of the content item is a second playback position different from the first playback position, one or more input devices are used to detect that the user's attention is directed to the user interface associated with the content item. If selected, after one or more of the above criteria have been met, and in response to the detection that the user's attention is directed to the user interface associated with the content item, the electronic device will display, via the display generation component, selectable options for presenting the content item from individual playback positions associated with the first playback position. While the selectable options are displayed, the system detects an input corresponding to the selection of one of the selectable options via one or more input devices. A non-temporary computer-readable storage medium that, upon detection of the aforementioned input, updates the playback position of the content item to the individual playback position associated with the first playback position.
63. It is an electronic device, One or more processors, Memory and While presenting content items that change over time, and while displaying the user interface associated with the content items via the display generation component, While the playback position within the content item is at a first playback position, it is detected that one or more criteria are met, including criteria that are met when the attention of the user of the electronic device is not directed to the user interface associated with the content item, via one or more input devices. After detecting that one or more of the above criteria are met, While the playback position of the content item is a second playback position different from the first playback position, it is detected via one or more input devices that the user's attention is directed to the user interface associated with the content item. If selected, after one or more of the above criteria have been met, and in response to detecting that the user's attention is directed to the user interface associated with the content item, the electronic device displays, via the display generation component, selectable options to present the content item from individual playback positions associated with the first playback position. While the selectable options are displayed, an input corresponding to the selection of the selectable options is detected via one or more input devices. An electronic device comprising means for updating the playback position of the content item to the individual playback position associated with the first playback position in response to the detection of the aforementioned input.
64. An information processing device for use in an electronic device, wherein the information processing device is While presenting content items that change over time, and while displaying the user interface associated with the content items via the display generation component, While the playback position within the content item is at a first playback position, it is detected that one or more criteria are met, including criteria that are met when the attention of the user of the electronic device is not directed to the user interface associated with the content item, via one or more input devices. After detecting that one or more of the above criteria are met, While the playback position of the content item is a second playback position different from the first playback position, it is detected via one or more input devices that the user's attention is directed to the user interface associated with the content item. If selected, after one or more of the above criteria have been met, and in response to detecting that the user's attention is directed to the user interface associated with the content item, the electronic device displays, via the display generation component, selectable options to present the content item from individual playback positions associated with the first playback position. While the selectable options are displayed, an input corresponding to the selection of the selectable options is detected via one or more input devices. An information processing device comprising means for updating the playback position of the content item to the individual playback position associated with the first playback position in response to the detection of the aforementioned input.
65. It is an electronic device, One or more processors, Memory and An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method according to any one of claims 53 to 60.
66. A non-temporary computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of an electronic device, the electronic device causes the electronic device to perform the method according to any one of claims 53 to 60.
67. It is an electronic device, One or more processors, Memory and An electronic device comprising means for carrying out the method described in any one of claims 53 to 60.
68. An information processing device for use in an electronic device, wherein the information processing device is An information processing apparatus comprising means for performing the method described in any one of claims 53 to 60.
69. It is a method, In an electronic device that communicates with a display generation component and one or more input devices, Displaying a three-dimensional environment via the aforementioned display generation component is: Displaying a media user interface object simultaneously including a first content item presented in a first presentation mode, wherein, during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint of the electronic device, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, according to the determination that one or more criteria are met, wherein one or more criteria include the requirement that the first content includes immersive content, and during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, Displaying the three-dimensional environment via the display generation component, which includes, in accordance with the determination that one or more of the above criteria are not met, displaying the media user interface object containing the first content item presented in the first presentation mode, without displaying the first user interface element for transitioning the presentation of the first content item to the second presentation mode; While displaying the three-dimensional environment including the media user interface and the first user interface elements, a first input corresponding to the selection of the first user interface elements is received via one or more input devices. A method comprising displaying the first content item in a second presentation mode in the three-dimensional environment in response to receiving the first input, wherein during the presentation of the first content item in the second presentation mode, the first content item extends from the user's viewpoint of the electronic device to at least one edge of the field of view.
70. The method according to claim 69, wherein during the presentation of the first content item in the second presentation mode, the first content item extends from the user's viewpoint of the electronic device to at least a plurality of respective edges of the field of view.
71. The method according to claim 69 or 70, wherein, during the presentation of the first content item in the second presentation mode, the first content item extends beyond at least one edge of the field of view from the user's viewpoint of the electronic device.
72. The method according to any one of claims 69 to 71, wherein, while the media user interface object is presenting the first content item in the first presentation mode, a playback control user interface comprising one or more user interface elements including the first user interface element, for the purpose of modifying the playback of the first content item in the three-dimensional environment, the playback control user interface further comprises displaying a playback control user interface that is displayed at a first location in the three-dimensional environment based on the location of the media user interface object.
73. During the presentation of the first content item in the first presentation mode, the user's viewpoint corresponds to the first viewpoint, and the method is While the first content item is being presented in the three-dimensional environment from the first viewpoint in the first presentation mode, the movement of the user's viewpoint from the first viewpoint to the second viewpoint is detected. In response to detecting the movement of the user's viewpoint to the second viewpoint, The three-dimensional environment is displayed from the user's second viewpoint. The method according to claim 72, further comprising maintaining the display of the playback control user interface at the first location in the three-dimensional environment.
74. The media user interface object is displayed within the three-dimensional environment, and while the playback control user interface is displayed at the first location, a second input is received via one or more input devices corresponding to a request to move the media user interface object to a different location within the three-dimensional environment. Upon receiving the second input, Move the media user interface object to the different locations in the three-dimensional environment. The method according to claim 72 or 73, further comprising displaying the playback control user interface in a second location in the three-dimensional environment, which is different from the first location in the three-dimensional environment, and which is based on the different location in the three-dimensional environment.
75. While the first content item is displayed in the first presentation mode, a playback control user interface is displayed at a first location in the three-dimensional environment, the first location being a first distance from the user's viewpoint, which includes one or more user interface elements for modifying the playback of the first content item. The method according to any one of claims 69 to 74, further comprising, while presenting the first content item in the second presentation mode, displaying the playback control user interface at a second location in the three-dimensional environment, wherein the second location in the three-dimensional environment is at a second distance from the user's viewpoint in the three-dimensional environment that is closer than the first distance.
76. The method according to claim 75, wherein the second location in the three-dimensional environment is based on the location of the user's viewpoint.
77. The aforementioned method, The user's viewpoint is the first viewpoint, the media user interface object is displayed in the second presentation mode, and the playback control user interface is displayed at the second location in the three-dimensional environment, while the movement of the user's viewpoint from the first viewpoint to a second viewpoint different from the first viewpoint is detected, and in response to the detection of the movement of the user's viewpoint from the first viewpoint to the second viewpoint, The three-dimensional environment is displayed from the second viewpoint of the user of the electronic device via the display generation component. The method according to claim 75 or 76, further comprising displaying the playback control user interface in a third location different from the second location in the three-dimensional environment, wherein the third location is based on the location of the user's second viewpoint.
78. While the first content item is being presented in an individual presentation mode, the system receives a second input via one or more input devices, the second input including the movement of an individual part of the user, wherein the second input corresponds to a request to scrub the first content item. The method according to any one of claims 69 to 77, further comprising: receiving the second input, scrubbing the first content item in accordance with the movement of the individual part of the user.
79. Scrubing the first content item described above is: The method according to claim 78, further comprising, while detecting the second input, displaying content within the media user interface object that corresponds to the current scrubbing position in the first content item, which changes as the individual part of the user moves, according to a determination that one or more second criteria are met.
80. Scrubing the first content item described above is: While detecting the second input, if one or more of the second criteria are not met, and the content is determined to be displayed as immersive content, Pause playback of the first content item in the media user interface object. In the three-dimensional environment, the individual content within the first content item, which is separate from the immersive content, displays a visual indication of the individual content that changes as the individual part of the user moves, without changing the appearance of the immersive content as the individual part of the user moves, and that changes as the individual part of the user moves. The method according to claim 78 or 79, comprising: detecting the end of the second input, ceasing the display of the visual indication of the individual content, and playing the first content item in the media user interface object as immersive content, starting from the individual scrub position where the end of the second input was detected.
81. The method according to claim 79 or 80, wherein the one or more second criteria include a criterion that is satisfied when the second input is received while the first content item occupies less than a threshold portion of the user's field of view, and that is not satisfied when the second input is received while the first content item occupies more than the threshold portion of the user's field of view.
82. The method according to any one of claims 69 to 81, wherein during the presentation of the first content item in the first presentation mode, the first content item is displayed in the three-dimensional environment at a first size, and during the presentation of the first content item in the second presentation mode, the first content item is displayed in the three-dimensional environment at a second size that is larger than the first size.
83. The method according to any one of claims 69 to 82, wherein presenting the first content item in the first presentation mode includes presenting the first content in the three-dimensional environment while the representation of the first portion of the physical environment of the electronic device has a first level of visual emphasis, and presenting the first content item in the second presentation mode includes reducing the visual emphasis of the first portion of the physical environment to a second level of visual emphasis that is lower than the first level of visual emphasis.
84. The method according to claim 69, further comprising displaying a second user interface element in the three-dimensional environment that transitions the presentation of the first content item from the second presentation mode to the first presentation mode while the first content item is being presented in the second presentation mode.
85. The aforementioned method, While the first content item is being presented in the first presentation mode, and while the user's viewpoint corresponds to the first viewpoint, and while the first portion of the first content item, not the second portion of the first content item, is displayed on the media user interface object, the movement of the user's viewpoint from the first viewpoint to a second viewpoint different from the first viewpoint is detected. The method according to any one of claims 69 to 84, further comprising: detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, displaying the first content item within the media user interface object from the user's second viewpoint, wherein displaying the first content item within the media user interface object from the user's second viewpoint includes displaying the second portion of the first content item within the media user interface object.
86. It is an electronic device, One or more processors, Memory and An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions, and the instructions are Displaying a three-dimensional environment through display generation components is possible. Displaying a media user interface object simultaneously including a first content item presented in a first presentation mode, wherein, during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint of the electronic device, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, according to the determination that one or more criteria are met, wherein one or more criteria include the requirement that the first content includes immersive content, and during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, The three-dimensional environment is displayed via the display generation component, which includes displaying the media user interface object containing the first content item presented in the first presentation mode, without displaying the first user interface element that transitions the presentation of the first content item to the second presentation mode, in accordance with the determination that one or more of the above criteria are not met. While displaying the three-dimensional environment including the media user interface and the first user interface elements, a first input corresponding to the selection of the first user interface elements is received via one or more input devices. An electronic device that, in response to receiving the first input, displays the first content item in the three-dimensional environment in the second presentation mode, wherein, during the presentation of the first content item in the second presentation mode, the first content item extends from the user's viewpoint of the electronic device to at least one edge of the field of view.
87. A non-temporary computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of an electronic device, the electronic device... Displaying a three-dimensional environment through display generation components is possible. Displaying a media user interface object simultaneously including a first content item presented in a first presentation mode, wherein, during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint of the electronic device, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, according to the determination that one or more criteria are met, wherein one or more criteria include the requirement that the first content includes immersive content, and during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, The three-dimensional environment is displayed via the display generation component, which includes displaying the media user interface object containing the first content item presented in the first presentation mode, without displaying the first user interface element that transitions the presentation of the first content item to the second presentation mode, in accordance with the determination that one or more of the above criteria are not met. While the three-dimensional environment including the media user interface and the first user interface elements is being displayed, a first input corresponding to the selection of the first user interface elements is received via one or more input devices. A non-temporary computer-readable storage medium that, in response to receiving the first input, displays the first content item in the three-dimensional environment in the second presentation mode, wherein, during the presentation of the first content item in the second presentation mode, the first content item extends from the user's viewpoint of the electronic device to at least one edge of the field of view.
88. It is an electronic device, One or more processors, Memory and Displaying a three-dimensional environment through display generation components is possible. Displaying a media user interface object simultaneously including a first content item presented in a first presentation mode, wherein, during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint of the electronic device, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, according to the determination that one or more criteria are met, wherein one or more criteria include the requirement that the first content includes immersive content, and during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, A means for displaying the three-dimensional environment via the display generation component, which includes displaying the media user interface object containing the first content item presented in the first presentation mode, without displaying the first user interface element that transitions the presentation of the first content item to the second presentation mode, in accordance with the determination that one or more of the above criteria are not met; While displaying the three-dimensional environment including the media user interface and the first user interface elements, means for receiving a first input corresponding to the selection of the first user interface elements via one or more input devices, An electronic device comprising means for displaying the first content item in the three-dimensional environment in the second presentation mode in response to receiving the first input, wherein during the presentation of the first content item in the second presentation mode, the first content item extends from the user's viewpoint of the electronic device to at least one edge of the field of view.
89. An information processing device for use in an electronic device, wherein the information processing device is Displaying a three-dimensional environment through display generation components is possible. Displaying a media user interface object simultaneously including a first content item presented in a first presentation mode, wherein, during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint of the electronic device, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, according to the determination that one or more criteria are met, wherein one or more criteria include the requirement that the first content includes immersive content, and during the presentation of the first content item in the first presentation mode, the first content item occupies a first portion of the field of view from the user's viewpoint, and a second portion of the field of view from the user's viewpoint is occupied by other elements of the three-dimensional environment, and a first user interface element that transitions the presentation of the first content item to a second presentation mode, A means for displaying the three-dimensional environment via the display generation component, which includes displaying the media user interface object containing the first content item presented in the first presentation mode, without displaying the first user interface element that transitions the presentation of the first content item to the second presentation mode, in accordance with the determination that one or more of the above criteria are not met; While displaying the three-dimensional environment including the media user interface and the first user interface elements, means for receiving a first input corresponding to the selection of the first user interface elements via one or more input devices, An information processing device comprising means for displaying the first content item in the three-dimensional environment in the second presentation mode in response to receiving the first input, wherein during the presentation of the first content item in the second presentation mode, the first content item extends from the user's viewpoint of the electronic device to at least one edge of the field of view.
90. It is an electronic device, One or more processors, Memory and An electronic device comprising one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method according to any one of claims 69 to 85.
91. A non-temporary computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of an electronic device, the electronic device causes the electronic device to perform the method according to any one of claims 69 to 85.
92. It is an electronic device, One or more processors, Memory and An electronic device comprising means for carrying out the method described in any one of claims 69 to 85.
93. An information processing device for use in an electronic device, wherein the information processing device is An information processing apparatus comprising means for performing the method described in any one of claims 69 to 85.