DEVICE, METHOD, AND GRAPHICAL USER INTERFACE FOR CONTENT APPLICATIONS - Patent application
The computer system addresses inefficiencies in augmented and virtual reality interactions by using advanced interfaces and virtual lighting effects to streamline user inputs, enhancing efficiency and reducing power consumption.
Patent Information
- Application Number
- JP2024518498
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-23
- Filing Date
- 2022-09-16
- Publication Date
- 2026-01-15
- Estimated Expiration
- 2042-09-16
AI Technical Summary
Existing methods for interacting with augmented and virtual reality environments are cumbersome, inefficient, and require excessive user input, leading to cognitive burden and energy wastage, particularly in battery-operated devices.
A computer system with enhanced user interfaces that utilize touch-sensitive displays, eye-tracking, hand-tracking, and virtual lighting effects to reduce the number and complexity of user inputs, providing improved navigation and presentation modes for content items in three-dimensional environments.
The system enhances user interaction efficiency, reduces power consumption, and improves battery life by minimizing unnecessary inputs and providing intuitive feedback, allowing users to navigate and interact with content more quickly and effectively.
Smart Images

Figure 0007799819000001 
Figure 0007799819000002 
Figure 0007799819000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 261,564, filed September 23, 2021, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0002] It generally relates to a computer system having a display generating component and one or more input devices that present a graphical user interface, including but not limited to an electronic device, through the display generating component that includes a user interface for presenting and browsing content. [Background technology]
[0003] The development of computer systems for augmented reality has progressed significantly in recent years. Exemplary augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with the virtual / augmented reality environment. Exemplary virtual elements include virtual objects, including digital images, video, text, icons, and control elements such as buttons and other graphics. Summary of the Invention
[0004] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which manipulating virtual objects is complex and error-prone create a significant cognitive burden for users and detract from the experience of the virtual / augmented reality environment. In addition, these methods are unnecessarily time-consuming, thereby wasting energy. This latter consideration is particularly important in battery-operated devices.
[0005] Therefore, there is a need for a computer system having improved methods and interfaces for providing users with computer-generated experiences that make interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces can optionally complement or replace conventional methods of providing users with computer-generated reality experiences. Such methods and interfaces reduce the number, extent, and / or type of inputs from the user by helping the user understand the connection between the input provided and the device response to that input, thereby creating a more efficient human-machine interface.
[0006] The above-mentioned deficiencies and other problems associated with user interfaces for computer systems having display generating components and one or more input devices are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generating components, the output devices including one or more tactile output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing a plurality of functions. In some embodiments, a user interacts with the GUI through stylus and / or finger contacts and gestures on the touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body as captured by cameras and other movement sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through the interactions optionally include image editing, drawing, presenting, word processing, spreadsheet creation, game playing, making phone calls, video conferencing, emailing, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note taking, and / or digital video playback, and executable instructions to perform those functions are optionally contained in a transitory and / or non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.
[0007] There is a need for electronic devices with improved methods and interfaces for navigating user interfaces. Such methods and interfaces can complement or replace conventional methods for interacting with graphical user interfaces. Such methods and interfaces reduce the number, extent, and / or type of input from a user, creating a more efficient human-machine interface.
[0008] In some embodiments, the electronic device generates virtual lighting effects while presenting a content item. In some embodiments, the electronic device enhances navigation to individual playback positions of a content item. In some embodiments, the electronic device displays media content in a three-dimensional environment. In some embodiments, the electronic device presents media content in different presentation modes.
[0009] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification, and claims. Furthermore, it should be noted that the language used in this specification has been selected solely for the purposes of readability and explanation, and not to define or limit the subject matter of the present invention. [Brief explanation of the drawings]
[0010] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout:
[0011] [Figure 1] FIG. 1 is a block diagram illustrating an operating environment for a computer system for providing an XR experience, according to some embodiments.
[0012] [Figure 2] FIG. 1 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience, according to some embodiments.
[0013] [Figure 3] FIG. 1 is a block diagram illustrating display generation components of a computer system configured to provide a user with visual components of an XR experience, according to some embodiments.
[0014] [Figure 4] FIG. 1 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input, according to some embodiments.
[0015] [Figure 5]FIG. 1 is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input, according to some embodiments.
[0016] [Figure 6A] 1 is a flowchart illustrating a glint-assisted gaze tracking pipeline, according to some embodiments.
[0017] [Figure 6B] 1 illustrates an exemplary environment of an electronic device for providing an XR experience, according to some embodiments.
[0018] [Figure 7A] 1 illustrates an example of a method for generating virtual lighting effects while an electronic device is presenting a content item, according to some embodiments. [Figure 7B] 1 illustrates an example of a method for generating virtual lighting effects while an electronic device is presenting a content item, according to some embodiments. [Figure 7C] 1 illustrates an example of a method for generating virtual lighting effects while an electronic device is presenting a content item, according to some embodiments. [Figure 7D] 1 illustrates an example of a method for generating virtual lighting effects while an electronic device is presenting a content item, according to some embodiments. [Figure 7E] 1 illustrates an example of a method for generating virtual lighting effects while an electronic device is presenting a content item, according to some embodiments.
[0019] [Figure 8A] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8B] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8C]1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8D] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8E] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8F] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8G] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8H] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8I] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8J] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8K] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8L] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8M] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8N] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments. [Figure 8O] 1 is a flowchart illustrating a method for generating virtual lighting effects during presentation of a content item, according to some embodiments.
[0020] [Figure 9A] 1 illustrates an exemplary method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 9B] 1 illustrates an exemplary method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 9C] 1 illustrates an exemplary method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 9D] 1 illustrates an exemplary method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 9E] 1 illustrates an exemplary method for displaying media content in a three-dimensional environment, according to some embodiments.
[0021] [Figure 10A] 1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 10B] 1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 10C] 1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 10D] 1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 10E] 1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 10F] 1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 10G]1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 10H] 1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. [Figure 10I] 1 is a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments.
[0022] [Figure 11A] 1 illustrates an example of how an electronic device can enhance navigation to individual playback positions of a content item, according to some embodiments. [Figure 11B] 1 illustrates an example of how an electronic device can enhance navigation to individual playback positions of a content item, according to some embodiments. [Figure 11C] 1 illustrates an example of how an electronic device can enhance navigation to individual playback positions of a content item, according to some embodiments. [Figure 11D] 1 illustrates an example of how an electronic device can enhance navigation to individual playback positions of a content item, according to some embodiments. [Figure 11E] 1 illustrates an example of how an electronic device can enhance navigation to individual playback positions of a content item, according to some embodiments.
[0023] [Figure 12A] 1 is a flowchart illustrating a method for enhancing navigation to individual playback positions of a content item, according to some embodiments. [Figure 12B] 1 is a flowchart illustrating a method for enhancing navigation to individual playback positions of a content item, according to some embodiments. [Figure 12C]1 is a flowchart illustrating a method for enhancing navigation to individual playback positions of a content item, according to some embodiments.
[0024] [Figure 13A] 1 illustrates an exemplary method for presenting media content in immersive and non-immersive presentation modes according to some embodiments of the present disclosure. [Figure 13B] 1 illustrates an exemplary method for presenting media content in immersive and non-immersive presentation modes according to some embodiments of the present disclosure. [Figure 13C] 1 illustrates an exemplary method for presenting media content in immersive and non-immersive presentation modes according to some embodiments of the present disclosure. [Figure 13D] 1 illustrates an exemplary method for presenting media content in immersive and non-immersive presentation modes according to some embodiments of the present disclosure. [Figure 13E] 1 illustrates an exemplary method for presenting media content in immersive and non-immersive presentation modes according to some embodiments of the present disclosure.
[0025] [Figure 14A] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14B] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14C] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14D] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14E] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14F]1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14G] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14H] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14I] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. [Figure 14J] 1 is a flowchart illustrating a method for presenting media content in immersive and non-immersive presentation modes, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0026] The present disclosure relates to a user interface that provides a computer-generated (CGR) experience to a user, according to some embodiments.
[0027] The systems, methods, and GUIs described herein provide improved ways for electronic devices to present content that corresponds to a physical location indicated within a navigation user interface element.
[0028] In some embodiments, the computer system displays a content application including a content item in a three-dimensional environment. In some embodiments, the electronic device applies virtual lighting effects to the three-dimensional environment while displaying the content application including the content item. In some embodiments, the virtual lighting effects are based on the content item played through the content application (e.g., including colors included in an image associated with the content item). Presenting the content application user interface with virtual lighting provides a user with an immersive and less distracting experience while consuming the content item, which further reduces power usage and improves the battery life of the electronic device by allowing the user to use the electronic device more quickly and efficiently.
[0029] In some embodiments, the computer system presents media content in a three-dimensional environment in different presentation modes, including an augmented presentation mode and a picture-in-picture presentation mode. In some embodiments, the computer system updates the position and / or orientation of the media content within the three-dimensional environment as a user's viewpoint of the three-dimensional environment changes. In some embodiments, whether the computer system updates the position and / or orientation of the media content in the three-dimensional environment is based on the presentation mode associated with the media content at the time the computer system detects a movement of the user's viewpoint in the three-dimensional environment. Changing the pose and / or orientation of the media content as a user's viewpoint of the three-dimensional environment changes provides an efficient way of providing continuous access to media content regardless of the user's current viewpoint of the three-dimensional environment, which further reduces power usage and improves battery life of the electronic device by allowing a user to use the electronic device more quickly and efficiently.
[0030] In some embodiments, the computer system enhances navigation to individual portions of a content item. In some embodiments, while presenting a content item, the electronic device detects that the user's attention (e.g., gaze) is no longer directed at the content item. In some embodiments, in response to detecting that the user's attention has been directed at the content item after having been directed away from the content item, the electronic device presents a selectable option that, when selected, causes the electronic device to navigate to an individual playback position of the content item associated with the playback position of the content item that was being played when the user averted their attention from the content item. Presenting an option to navigate to an individual playback position of a content item provides an efficient way of navigating a content item, which allows a user to use the electronic device more quickly and efficiently, further reducing power usage, improving the battery life of the electronic device, and reducing errors in usage that must be corrected with further user input.
[0031] In some embodiments, a computer system presents immersive and non-immersive media content in a three-dimensional environment. In some embodiments, the computer system presents the immersive content in an immersive presentation mode and a non-immersive presentation mode. In some embodiments, while the computer system is presenting the immersive content in a non-immersive presentation mode, the computer system displays a selectable option that, when selected, causes the computer system to transition the presentation of the immersive content from the non-immersive presentation mode to an immersive presentation mode. Providing a selectable option to transition the presentation of the content from the non-immersive presentation mode to the immersive presentation mode provides an efficient way to access different presentation modes associated with the immersive content, which further reduces power usage and improves battery life of the electronic device by allowing a user to use the electronic device more quickly and efficiently.
[0032] FIGS. 1-6 provide an illustration of an exemplary computer system for providing an XR experience to a user (as described below with respect to methods 800, 1000, 1200, and 1400). FIGS. 7A-7E illustrate an exemplary technique for generating virtual lighting effects while presenting a content item, according to some embodiments. FIGS. 8A-8O are a flowchart illustrating a method for generating virtual lighting effects while presenting a content item, according to some embodiments. FIGS. 9A-9E illustrate an exemplary technique for displaying media content in a three-dimensional environment, according to some embodiments. FIGS. 10A-10I are a flowchart illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. FIGS. 11A-11E illustrate an exemplary technique for enhancing navigation to individual playback positions of a content item, according to some embodiments. FIGS. 12A-12C are a flowchart illustrating a method for enhancing navigation to individual playback positions of a content item, according to some embodiments. FIGS. 13A-13E illustrate an exemplary technique for presenting media content in immersive and non-immersive presentation modes, according to some embodiments of the present disclosure. 14A-14J are flowcharts illustrating methods for presenting media content in immersive and non-immersive presentation modes according to some embodiments.
[0033] The processes described below enhance the usability of the device and streamline the user-device interface (e.g., by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the device) through various techniques, including providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional controls that are displayed, performing an operation without requiring further user input when a set of conditions is met, improving privacy and / or security, and / or other techniques. These techniques also reduce power usage and improve the device's battery life by allowing the user to use the device more quickly and efficiently.
[0034] Furthermore, for methods described herein in which one or more steps are conditioned on one or more conditions being satisfied, it should be understood that the described method can be repeated in multiple iterations, such that over the course of the iterations, all of the conditions on which the method steps are conditioned are satisfied in different iterations of the method. For example, if a method requires performing a first step if a condition is satisfied and a second step if the condition is not satisfied, one skilled in the art will understand that the steps recited in the claim are repeated in a particular order until the conditions are satisfied and then no longer satisfied. Thus, a method described with one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that is repeated until each condition recited in the method is satisfied. However, this is not required for system or computer-readable medium claims in which the system or computer-readable medium includes instructions for performing a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency is met without explicitly repeating the method steps until all conditions on which the method steps are conditioned are satisfied. Those skilled in the art will also understand that, as with methods having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.
[0035] 1, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a consumer electronics device, a wearable device, etc.). In some embodiments, one or more of input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with display generation component 120 (e.g., within a head-mounted or handheld device).
[0036] When describing an XR experience, various terms are used to individually refer to several related, but distinct, environments that a user senses and / or can interact with (e.g., using inputs detected by computer system 101 that cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to computer system 101 generating the XR experience). The following is a subset of these terms:
[0037] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through their senses, such as sight, touch, hearing, taste, and smell.
[0038] Augmented reality: In contrast, an extended reality (XR) environment refers to a wholly or partially mimicked environment that people sense and / or interact with through electronic systems. In XR, a subset of a person's body movements or representations thereof are tracked, and one or more properties of one or more virtual objects simulated within the XR environment are adjusted accordingly to behave according to at least one law of physics. For example, an XR system may detect a person's head rotation and adjust the graphical content and sound field presented to the person accordingly, in a manner similar to how such views and sounds change in a physical environment. In some circumstances (e.g., for accessibility reasons), adjustments to property(ies) of virtual object(s) in the XR environment may be made in response to representations of body movements (e.g., voice commands). A person may sense and / or interact with an XR object using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may sense and / or interact with audio objects that create a 3D or spatially expansive audio environment that provides the perception of a point sound source in 3D space. In another example, audio objects may enable audio transparency that selectively incorporates ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may sense and / or interact with only audio objects.
[0039] Examples of XR include virtual reality and mixed reality.
[0040] Virtual Reality: A virtual reality (VR) environment refers to an emulated environment designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in the VR environment through a simulation of the person's presence in the computer-generated environment and / or through a simulation of a subset of the person's physical movement within the computer-generated environment.
[0041] Mixed Reality: A mixed reality (MR) environment refers to a mimetic environment designed to incorporate sensory input from or representations of a physical environment in addition to including computer-generated sensory input (e.g., virtual objects), as opposed to a VR environment designed to be based entirely on computer-generated sensory input. On the virtual continuum, a mixed reality environment is anywhere between, but not including, a fully physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Some electronic systems for presenting MR environments may also track location and / or orientation relative to the physical environment to allow virtual objects to interact with real objects (i.e., physical items from the physical environment or representations thereof). For example, the system may account for movement so that a virtual tree appears stationary relative to the physical ground.
[0042] Examples of mixed reality include augmented reality and augmented virtuality.
[0043] Augmented reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person using the system perceives the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with virtual objects and presents the composite on the opaque display. The person uses the system to indirectly view the physical environment through the images or videos of the physical environment and perceive the virtual objects superimposed on the physical environment. As used herein, video of a physical environment shown on an opaque display is referred to as "pass-through video," meaning that the system captures images of the physical environment using one or more image sensors and uses those images in presenting the AR environment on the opaque display. Alternatively, the system may include a projection system that projects virtual objects, e.g., as holograms, into a physical environment or onto a physical surface, such that a person using the system perceives the virtual objects superimposed on the physical environment. Augmented reality environments also refer to mimic environments in which a representation of a physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, a system may distort one or more sensor images to impose a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, a representation of a physical environment may be distorted by graphically modifying (e.g., enlarging) portions thereof, thereby rendering the modified portions a non-photorealistic, altered version of the originally captured image. As a further example, a representation of a physical environment may be distorted by graphically removing or obscuring portions thereof.
[0044] Augmented Virtual: An augmented virtual (AV) environment refers to a mimicking environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, while people with faces are realistically recreated from images taken of physical people. As another example, virtual objects may adopt the shape or color of physical items imaged by one or more imaging sensors. As a further example, virtual objects may adopt shadows that match the position of the sun in the physical environment.
[0045] Perspective-Locked Virtual Object: A virtual object is perspective-locked when the computer system displays the virtual object in the same location and / or position within the user's perspective, even as the user's perspective shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user's perspective is locked to the forward-facing orientation of the user's head (e.g., the user's perspective is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's perspective remains fixed even as the user's line of sight moves without moving the user's head. In embodiments in which the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's perspective is the augmented reality view being presented to the user on the display generating component of the computer system. For example, a perspective-locked virtual object that is displayed in the upper left corner of the user's perspective when the user's perspective is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's perspective when the user's perspective changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position at which a viewpoint-locked virtual object is displayed in a user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."
[0046] Environment-Locked Virtual Object: A virtual object is environment-locked (or "world-locked") when a computer system displays the virtual object at a location and / or position within a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) locations and / or objects within a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint shifts, the locations and / or objects within the environment relative to the user's viewpoint change, resulting in the environment-locked virtual object appearing at a different location and / or position within the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear centered within the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes more left-leaning within the user's viewpoint (e.g., the position of the tree within the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear more left-leaning within the user's viewpoint. In other words, the location and / or position at which the environment-locked virtual object appears within the user's viewpoint depends on the position and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system fixed to a fixed location and / or object in the physical environment) to determine a position at which to display an environment-locked virtual object in the user's viewpoint. The environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object) or can be locked to a moving portion of the environment (e.g., a vehicle, an animal, a person, or a representation of a part of the user's body that moves independent of the user's viewpoint, such as the user's hand, wrist, arm, or leg), so that the virtual object moves as the viewpoint or part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.
[0047] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed-following behavior, which reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed-following behavior, the computer system intentionally delays the movement of the virtual object when it detects movement of the reference point that the virtual object is following (e.g., a part of the environment, the viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 and 300 cm from the viewpoint). For example, when the reference point (e.g., a part of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but at a second speed that is slower than the first speed (e.g., until the reference point stops or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits delayed-following behavior, the device ignores small amounts of movement of the reference point (e.g., ignores movement of the reference point that is less than a threshold amount of movement, such as movement between 0 and 5 degrees or movement between 0 and 50 cm). For example, when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a “delayed following” threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object maintaining a substantially fixed position relative to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or forward / backward relative to the position of the reference point).
[0048] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (e.g., similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display rather than an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser-scanned light source, or any combination of these technologies. The medium may be a light guide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or as physical surfaces.In some embodiments, controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Controller 110 is described in more detail below with reference to FIG. 2. In some embodiments, controller 110 is a computing device that is local or remote to scene 105 (e.g., the physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within the housing (e.g., physical housing) of one or more of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the foregoing.
[0049] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of the XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with reference to FIG. 3. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.
[0050] According to some embodiments, the display generation component 120 provides an XR experience to the user while the user is virtually and / or physically present in the scene 105.
[0051] In some embodiments, the display generating component is worn on a part of the user's body (e.g., on their head, their hand, etc.). Thus, display generating component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, display generating component 120 surrounds the user's field of view. In some embodiments, display generating component 120 is a handheld device (e.g., a smartphone or tablet) configured to present XR content, where the user holds the device with a display pointed toward the user's field of view and a camera pointed toward scene 105. In some embodiments, the handheld device is optionally located within a housing worn on the user's head. In some embodiments, the handheld device is optionally located on a support (e.g., a tripod) in front of the user. In some embodiments, display generating component 120 is an XR chamber, housing, or room configured to present XR content without the user wearing or holding display generating component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with XR content triggered based on interactions occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD in which the interactions occur in the space in front of the HMD and the XR content responses are displayed via the HMD. Similarly, a user interface showing interactions with CRG content triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)) may be implemented similarly to an HMD in which the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)).
[0052] While relevant features of operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that for the sake of brevity, various other features are not shown so as to not obscure more pertinent aspects of the exemplary embodiments disclosed herein.
[0053] 2 is a block diagram of an example controller 110, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more pertinent aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0054] In some embodiments, one or more communication buses 204 include circuitry that interconnects and controls communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0055] Memory 220 includes high-speed random-access memory, such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from the one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240:
[0056] Operating system 230 includes instructions for handling various basic system services and performing hardware-dependent tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for respective groups of one or more users). To that end, in various embodiments, XR experience module 240 includes a data acquisition unit 242, a tracking unit 244, an adjustment unit 246, and a data transmission unit 248.
[0057] 1 , and optionally one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data acquisition unit 241 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0058] In some embodiments, tracking unit 242 is configured to map scene 105 and track the position / location of at least display generating component 120 relative to scene 105 of FIG. 1 , and optionally relative to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, tracking unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, tracking unit 242 includes hand tracking unit 244 and / or eye tracking unit 243. In some embodiments, hand tracking unit 244 is configured to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 , relative to display generating component 120, and / or relative to a coordinate system defined relative to the user's hand. Hand tracking unit 244 is described in more detail below with respect to FIG. 4. In some embodiments, eye tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or the user (e.g., the user's hands)), or relative to XR content displayed via display generation component 120. Eye tracking unit 243 is described in more detail below with respect to FIG. 5.
[0059] In some embodiments, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120 and, optionally, by one or more of output devices 155 and / or peripheral devices 195. To that end, in various embodiments, coordination unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0060] In some embodiments, data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least display generation component 120, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0061] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 may be located within separate computing devices.
[0062] Furthermore, Figure 2 is intended more to illustrate the functionality of various features that may be present in particular embodiments, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 2 can be implemented in a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary depending on implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0063] 3 is a block diagram of an example of a display generation component 120, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that, for the sake of brevity, various other features are not shown so as to not obscure more pertinent aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward-facing and / or outward-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0064] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.
[0065] In some embodiments, the one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emissive element display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), MEMS, and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to a waveguide display, such as a diffractive, reflective, polarized, holographic, etc. For example, the display generation component 120 (e.g., HMD) includes a single XR display. In another example, the display generation component 120 (e.g., HMD) includes an XR display for each eye of the user. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content.
[0066] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as eye-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as hand-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene as the user would view it if the display generating component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, and / or the like.
[0067] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and an XR presentation module 340:
[0068] The operating system 330 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. To that end, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.
[0069] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of Figure 1. To that end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0070] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. To that end, in various embodiments, the XR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0071] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an augmented reality) based on the media content data. To that end, in various embodiments, the XR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0072] In some embodiments, data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least controller 110, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0073] Although the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., the display generation component 120 of FIG. 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 may be located in separate computing devices.
[0074] Furthermore, Figure 3 is intended more to illustrate the functionality of various features that may be present in particular implementations, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 3 can be implemented within a single module, and various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functions and how functions are allocated among them, will vary from implementation to implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0075] 4 is a schematic diagram of an example embodiment of a hand tracking device 140. In some embodiments, hand tracking device 140 (FIG. 1) is controlled by hand tracking unit 244 (FIG. 2) to track the location / position and / or movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 (e.g., relative to a portion of the physical environment surrounding the user, relative to display generating components 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to the user's hand). In some embodiments, hand tracking device 140 is part of display generating components 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, hand tracking device 140 is separate from display generating components 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0076] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images with sufficient resolution to allow for differentiation of the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either zoom capabilities or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor, or a portion thereof, is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.
[0077] In some embodiments, image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to controller 110, which extracts high-level information from the map data. This high-level information is provided, typically via an application program interface (API), to an application running on the controller, which drives display generation component 120 accordingly. For example, a user can interact with software running on controller 110 by moving their hand 406 and changing the posture of their hand.
[0078] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the pattern's spots. This approach is advantageous in that it does not require the user to hold or wear any type of beacon, sensor, or other marker. This provides depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In this disclosure, the image sensor 404 is assumed to define a set of orthogonal x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) can use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurement, based on single or multiple cameras or other types of sensors.
[0079] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps including the user's hand while the user moves the hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or a processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software matches these descriptors with patch descriptors stored in the database 408, based on a previous learning process, to estimate the pose of the hand in each frame. The pose typically includes the 3D locations of the user's wrist joints and fingertips.
[0080] The software can also analyze hand and / or finger trajectories across multiple frames in a sequence to identify gestures. The pose estimation functionality described herein may be interleaved with motion tracking functionality, whereby patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to discover pose changes that occur across the remaining frames. Pose, motion, and gesture information is provided to an application program running on controller 110 via the API described above. This program can, for example, move and modify an image presented on display generation component 120 or perform other functions in response to the pose and / or gesture information.
[0081] In some embodiments, the gesture includes an air gesture, which is detected without (or independent of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to another of the user's hands, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of the user's body part (e.g., a tap gesture involving movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture involving a predetermined speed or amount of rotation of the user's body part).
[0082] In some embodiments, input gestures used in various examples and embodiments described herein include air gestures performed by movement of a user's finger(s) relative to other finger(s) or part(s) of the user's hand to interact with an XR environment (e.g., a virtual or mixed reality environment), according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture that includes movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes rotation of a part of the user's body at a predetermined speed or amount).
[0083] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides a computer system with information about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is detected attention (e.g., gaze) to a user interface element in combination with (e.g., simultaneous with) movement of the user's finger(s) and / or hand to perform pinch and / or tap input, as described in more detail below.
[0084] In some embodiments, an input gesture directed at a user interface object is performed directly or indirectly with reference to the user interface object. For example, user input is performed directly at a user interface object in response to performing an input gesture with the user's hand at a position corresponding to the user interface object's position in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, an input gesture is performed indirectly at a user interface object in response to detecting the user's attention (e.g., gaze) to the user interface object while performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in the three-dimensional environment. For example, in the case of a direct input gesture, a user can direct the user's input at a user interface object by initiating the gesture at or near a position corresponding to the user interface object's displayed position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from an optional outer edge or optional central portion). For indirect input gestures, a user can direct their input to a user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user initiates an input gesture (e.g., at any position detectable by the computer system) (e.g., at a position that does not correspond to the displayed position of the user interface object).
[0085] In some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment, according to some embodiments. For example, pinch inputs and tap inputs, as described below, are performed as air gestures.
[0086] In some embodiments, the pinch input is part of an air gesture, including one or more of a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture includes moving two or more fingers of a hand to contact each other, i.e., optionally with a short break (e.g., within 0-1 second) after contact with each other. A long pinch gesture that is an air gesture includes moving two or more fingers of a hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before detecting a break in contact with each other. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., when two or more fingers are in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) that are detected immediately in succession (e.g., within a predetermined period of time) after each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaking contact between two or more fingers), and performs a second pinch input within a predetermined period of time (e.g., within 1 second or 2 seconds) after releasing the first pinch input.
[0087] In some embodiments, a pinch-and-drag gesture that is an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., followed by) a drag input that changes the position of a user's hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, a user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers apart) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., a user pinches two or more fingers together and moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by a user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from a first position to a second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of a user's hands. For example, the input gesture includes two (e.g., or more) pinch inputs performed in conjunction with each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using a first hand of the user and a second pinch input performed using the other hand (e.g., a second of the user's hands) in conjunction with performing the pinch input using the first hand. In some embodiments, a movement between a user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).
[0088] In some embodiments, a tap input (e.g., directed toward a user interface element) performed as an air gesture includes movement(s) of a user's finger(s) toward the user interface element, movement of a user's hand toward a user interface element, optionally with the user's finger(s) extended toward the user interface element, a downward movement of a user's finger (e.g., mimicking a mouse click action or a tap on a touchscreen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture, moving the finger or hand away from the user's viewpoint and / or toward the object that is the target of the tap input followed by an end of the movement. In some embodiments, an end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's viewpoint and / or toward the object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the direction of acceleration of the movement of the finger or hand).
[0089] In some embodiments, the user's attention is determined to be directed to a portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, the device determines that the user's attention is directed to the portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment with one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, and / or requiring the gaze to be directed to the portion of the three-dimensional environment, and if one of the additional conditions is not met, the device determines that the user's attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).
[0090] In some embodiments, detection of a ready configuration of a user or a portion of a user is detected by a computer system, and detection of a ready configuration of the hands is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs performed with the hands (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand geometry (e.g., a pre-pinch geometry with the thumb and one or more fingers extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap geometry with one or more fingers extended and the palm facing away from the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., above the user's waist, moved toward an area in front of the user below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface is responsive to attentional (e.g., gaze) input.
[0091] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or alternatively may be provided on a tangible, non-transitory medium, such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively, or additionally, some or all of the described functionality of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). While the controller 110 is shown in FIG. 4 as, by way of example, a separate unit from the image sensor 404, some or all of the controller's processing functions may be performed by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor (e.g., the hand tracking device 402), or otherwise associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device) or using any other suitable computerized device, such as a game console or media player. The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device that is controlled by the sensor output.
[0092] FIG. 4 also includes a schematic diagram of a depth map 410 captured by the image sensor 404, according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. A pixel 412 corresponding to the hand 406 is segmented from the background and wrist in this map. The intensity of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from the image sensor 404, with increasing gray levels as depth increases. The controller 110 processes these depth values to identify and segment components of the image (i.e., groups of adjacent pixels) that have characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.
[0093] 4 also schematically illustrates a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to some embodiments. In FIG. 4, the skeleton 414 is superimposed on a hand background 416 that was segmented from the original depth map. In some embodiments, key feature points on the hand (e.g., knuckles, fingertips, center of the palm, end of the hand where it connects to the wrist, etc.), and optionally the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these key feature points over multiple image frames are used by the controller 110 to determine hand gestures performed by the hand or the current state of the hand, according to some embodiments.
[0094] FIG. 5 shows an exemplary embodiment of eye tracking device 130 ( FIG. 1 ). In some embodiments, eye tracking device 130 is controlled by eye tracking unit 243 ( FIG. 2 ) to track the position and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, eye tracking device 130 is integrated with display generation component 120. For example, in some embodiments, if display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed in a wearable frame, the head-mounted device includes both components for generating XR content for viewing by the user and components for tracking the user's gaze relative to the XR content. In some embodiments, eye tracking device 130 is separate from display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, eye tracking device 130 is optionally a device separate from the handheld device or the XR chamber. In some embodiments, eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, head-mounted eye tracking device 130 is optionally used in conjunction with head-mounted or non-head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally used in combination with head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally part of non-head-mounted display generating components.
[0095] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display that allows the user to view the physical environment directly and display virtual objects on the transparent or translucent display. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, allowing an individual using the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0096] As shown in FIG. 5 , in some embodiments, eye tracking device 130 (e.g., gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) camera or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be aimed at the user's eyes to receive reflected IR or NIR light from the light source directly from the eyes, or alternatively, may be aimed at a “hot” mirror positioned between the user's eyes and a display panel that reflects IR or NIR light from the eyes to the eye tracking camera while allowing visual light to pass through. Eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate eye tracking information, and communicates the eye tracking information to controller 110. In some embodiments, the user's eyes are tracked separately by their respective eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by a separate eye-tracking camera and lighting source.
[0097] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the eye tracking device's parameters for the particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automatic or manual calibration process. The user-specific calibration process may include estimation of a particular user's eye parameters, such as pupil location, central visual location, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 have been determined, images captured by the eye tracking camera can be processed using glint-assisted methods to determine the user's current visual axis and viewpoint relative to the display.
[0098] As shown in FIG. 5, eye tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and a gaze tracking system including at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking occurs and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 may be positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, a projector, etc.) and may be directed at a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (e.g., as shown at the top of FIG. 5), or may be directed at the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of FIG. 5).
[0099] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, such as in processing frames 562 for display. Controller 110 optionally estimates the user's viewpoint on display 510 based on gaze tracking input 542 obtained from eye tracking camera 540, using a glint-assisted method or other suitable method. The viewpoint estimated from gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0100] Some possible use cases of the user's current gaze direction are described below, but are not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content with higher resolution in a central visual area determined from the user's current gaze direction than in a peripheral area. As another example, the controller may position or move virtual content within a view based at least in part on the user's current gaze direction. As another example, the controller may display particular virtual content within a view based at least in part on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 can orient an external camera to capture the physical environment of the XR experience and focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface within the environment the user is currently viewing on the display 510. As another exemplary use case, eyepiece 520 may be a focusable lens, and eye-tracking information is used by the controller to adjust the focus of eyepiece 520 so that the virtual object the user is currently looking at has the proper binocular coordination to match the convergence of the user's eyes 592. Controller 110 can utilize the eye-tracking information to orient and focus eyepiece 520 so that close objects the user is looking at appear at the correct distance.
[0101] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye tracking camera (e.g., eye tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) attached to the wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520, as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.
[0102] In some embodiments, the display 510 emits light in the visible light range and not in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. Note that the location and angle of the eye tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0103] Embodiments of an eye tracking system such as that shown in FIG. 5 may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience.
[0104] FIG. 6A illustrates a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the glint-assisted gaze tracking system tracks the pupil contour and glint in the current frame using prior information from the previous frame when analyzing the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in the tracking state to the next frame.
[0105] As shown in FIG. 6A, an eye-tracking camera may capture left and right images of a user's left and right eyes. The captured images are then input into an eye-tracking pipeline for processing beginning at 610. As indicated by the arrow returning to element 600, the eye-tracking system may continue to capture images of the user's eyes at a rate of, for example, 60-120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0106] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. If the tracking status is no at 610, the image is analyzed to detect the user's pupil and glint in the image, as shown at 620. If the pupil and glint are successfully detected at 630, the method proceeds to element 640. If not, the method returns to element 610 to process the next image of the user's eyes.
[0107] At 640, proceeding from element 610, the current frame is analyzed to track pupils and glints based in part on previous information from the previous frame. At 640, proceeding from element 630, a tracking state is initialized based on the detected pupils and glints in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results can be checked to determine whether a sufficient number of glints are successfully tracked or detected in the current frame to perform pupil and gaze estimation. At 650, if the results are not reliable, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and pupil and glint information is passed to element 680 to estimate the user's gaze.
[0108] 6A is intended to serve as an example of eye-tracking technology that may be used in particular implementations. As will be recognized by those skilled in the art, other eye-tracking technologies, now existing or developed in the future, may be used in place of or in combination with the glint-assisted eye-tracking technology described herein in computer system 101 to provide a user with an XR experience according to various embodiments.
[0109] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with an XR experience, e.g., a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.
[0110] 6B shows an exemplary environment for the electronic device 101 for providing an XR experience, according to some embodiments. In FIG. 6B , a real-world environment 602 includes the electronic device 101, a user 608, and a real-world object (e.g., a table 604). As shown in FIG. 6B , the electronic device 101 is optionally tripod-mounted or otherwise secured to the real-world environment 602 so that one or more hands of the user 608 are free (e.g., the user 608 is optionally not holding the device 101 with one or more hands). As described above, the device 101 optionally has one or more groups of sensors located on different sides of the device 101. For example, the device 101 optionally includes a sensor group 612-1 and a sensor group 612-2 located on the “rear” and “front” sides of the device 101, respectively (e.g., capable of capturing information from each side of the device 101). As used herein, the front side of the device 101 is the side that faces the user 608 and the back side of the device 101 is the side that faces away from the user 608 .
[0111] In some embodiments, sensor group 612-2 includes an eye tracking unit (e.g., eye tracking unit 245 described above with reference to FIG. 2) that includes one or more sensors for tracking the eyes and / or gaze of a user, and the eye tracking unit can "watch" user 608 and track the eye(s) of user 608 in the manner described above. In some embodiments, the eye tracking unit of device 101 can capture the movement, orientation, and / or gaze of the eyes of user 608 and process the movement, orientation, and / or gaze as input.
[0112] In some embodiments, sensor group 612-1 includes a hand tracking unit (e.g., hand tracking unit 243 described above with reference to FIG. 2) that can track one or more hands of user 608 held on the “back” side of device 101, as shown in FIG. 6B. In some embodiments, a hand tracking unit is optionally included in sensor group 612-2 so that user 608 can additionally or alternatively hold one or more hands on the “front” side of device 101 while device 101 tracks the position of the one or more hands. As described above, the hand tracking unit of device 101 can capture the movements, positions, and / or gestures of one or more hands of user 608 and process the movements, positions, and / or gestures as input.
[0113] In some embodiments, sensor group 612-1 optionally includes one or more sensors (e.g., image sensor 404 described above with reference to FIG. 4 ) configured to capture images of real-world environment 602, including table 604. As described above, device 101 can capture images of portions (e.g., part or all) of real-world environment 602 and present the captured portions of real-world environment 602 to the user via one or more display generation components of device 101 (e.g., a display of device 101 optionally located on a side of device 101 facing the user, opposite the side of device 101 facing the captured portions of real-world environment 602).
[0114] In some embodiments, the captured portion of the real-world environment 602 is used to provide the user with an XR experience, e.g., a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.
[0115] Accordingly, the description herein describes several embodiments of three-dimensional environments (e.g., XR environments) that include representations of real-world objects and representations of virtual objects. For example, the three-dimensional environment optionally includes a representation of a table present in a physical environment that is captured and displayed within the three-dimensional environment (e.g., actively via a camera and display of the computer system, or passively via a transparent or translucent display of the computer system). As mentioned above, the three-dimensional environment is optionally a mixed reality system based on a physical environment, where the three-dimensional environment is captured by one or more sensors of the device and displayed via a display generation component. As a mixed reality system, the computer system can optionally selectively display portions and / or objects of the physical environment such that each portion and / or object of the physical environment appears to exist within the three-dimensional environment displayed by the electronic device. Similarly, the computer system can optionally display virtual objects within the three-dimensional environment such that the virtual objects appear to exist within the real world (e.g., the physical environment) by placing the virtual objects at respective locations within the three-dimensional environment that have corresponding locations in the real world. For example, the computer system optionally displays a vase so that it appears as if the real vase were placed on a table in the physical environment. In some embodiments, each location in the three-dimensional environment has a corresponding location in the physical environment. Thus, when a computer system is described as displaying a virtual object at a distinct location relative to a physical object (e.g., a location at or near the user's hand, or at or near the physical table, etc.), the computer system displays the virtual object at a particular location in the three-dimensional environment so that it appears as if the virtual object is at or near the physical object in the physical world (e.g., the virtual object is displayed at a location in the three-dimensional environment that corresponds to the location in the physical environment where the virtual object would appear if it were a real object at that particular location).
[0116] In some embodiments, real-world objects that exist in the physical environment (e.g., and / or that are viewable via display generation components) that are displayed in the three-dimensional environment can interact with virtual objects that exist only in the three-dimensional environment. For example, the three-dimensional environment can include a table and a vase placed on the table, where the table is a view (or representation) of the physical table in the physical environment and the vase is a virtual object.
[0117] Similarly, a user can optionally use one or more hands to interact with virtual objects in the three-dimensional environment as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the computer system optionally capture one or more of the user's hands and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above), or in some embodiments, due to the transparency / translucency of the user interface, or the projection of the user interface onto a transparent / translucent surface, or the portion of the display generating components displaying the projection of the user interface to the user's eyes or field of view of the user's eyes, the user's hands are visible through the display generating components by the ability to see the physical environment through the user interface. Thus, in some embodiments, the user's hands are displayed at discrete locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment as if they were actual physical objects in the physical environment. In some embodiments, the computer system can update the display of the representation of the user's hands in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.
[0118] In some of the embodiments described below, for example, for purposes of determining whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grabbing, holding, etc., a virtual object, or whether it is within a threshold distance from the virtual object), the computer system can optionally determine an “effective” distance between the physical object in the physical world and the virtual object in the three-dimensional environment. For example, a hand directly interacting with a virtual object optionally includes one or more of the fingers of a hand pressing a virtual button, a user's hand grasping a virtual vase, two fingers of a user's hand pinching / holding an application's user interface together, and any other types of interactions described herein. For example, when determining whether and / or how a user is interacting with a virtual object, the computer system optionally determines the distance between the user's hand and the virtual object. In some embodiments, the computer system determines the distance between the user's hand and the virtual object by determining the distance between the location of the hand in the three-dimensional environment and the location of the target virtual object in the three-dimensional environment. For example, one or more hands of a user are positioned at particular positions in the physical world, which the computer system optionally captures and displays at particular corresponding positions in the three-dimensional environment (e.g., positions in the three-dimensional environment at which the hands are displayed, if the hands are virtual rather than physical hands). The positions of the hands in the three-dimensional environment are optionally compared to positions of target virtual objects in the three-dimensional environment to determine a distance between the user's one or more hands and the virtual objects. In some embodiments, the computer system optionally determines the distance between a physical object and a virtual object by comparing positions in the physical world (e.g., as opposed to comparing positions in the three-dimensional environment).For example, when determining the distance between one or more of a user's hands and a virtual object, the computer system optionally determines the corresponding location in the physical world of the virtual object (e.g., the position at which the virtual object would be located in the physical world as if it were a physical object rather than a virtual object), and then determines the distance between the corresponding physical position and the user's one or more hands. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Thus, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system optionally performs any of the above-mentioned techniques to map the location of the physical object to the three-dimensional environment and / or to map the location of the virtual object to the physical environment.
[0119] In some embodiments, the same or similar techniques are used to determine where or what a user's gaze is directed at and / or where or what a physical stylus held by the user is directed at. For example, if a user's gaze is directed at a particular position in the physical environment, the computer system optionally determines a corresponding position in the three-dimensional environment (e.g., a virtual position of the gaze), and if a virtual object is located at that corresponding virtual position, the computer system optionally determines that the user's gaze is directed at that virtual object. Similarly, the computer system can optionally determine where the physical stylus is pointing in the physical environment based on the orientation of the stylus. In some embodiments, based on this determination, the computer system determines a corresponding virtual position in the three-dimensional environment that corresponds to the location in the physical environment where the stylus is pointing, and optionally determines that the stylus is pointing to the corresponding virtual position in the three-dimensional environment.
[0120] Similarly, embodiments described herein may refer to the location of a user (e.g., a user of a computer system) and / or the location of the computer system within a three-dimensional environment. In some embodiments, a user of a computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system is used as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user within the physical environment corresponds to a distinct location within the three-dimensional environment. For example, if a user stands at a location facing a distinct portion of the physical environment displayed by the display generating components, the location of the computer system is the location within the physical environment (and its corresponding location within the three-dimensional environment) where the user would see objects within the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other) as the objects are displayed by the display generating components of the computer system within the three-dimensional environment. Similarly, if the virtual objects displayed in the three-dimensional environment were physical objects in the physical environment (e.g., the virtual objects are located in the same physical environment location and have the same physical environment size and orientation as in the three-dimensional environment), the location of the computer system and / or user is the position at which the user would see the virtual objects in the physical environment in the same position, orientation, and / or size (e.g., absolutely and / or relative to each other and to real-world objects) as they were displayed by the display generation components of the computer system in the three-dimensional environment.
[0121] In this disclosure, various input methods are described with respect to interaction with a computer system. Where one example is provided using one input device or input method and another example is provided using a different input device or input method, it should be understood that each example may be compatible with, and optionally utilize, the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. Where one example is provided using one output device or output method and another example is provided using a different output device or output method, it should be understood that each example may be compatible with, and optionally utilize, the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. Where one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with, and optionally utilize, the method described with respect to the other example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.
[0122] User Interface and Related Processes We now turn our attention to embodiments of user interfaces (“UIs”) and associated processes that may be executed in a computer system, such as a portable multifunction device or a head-mounted device, equipped with display generating components, one or more input devices, and (optionally) one or more cameras.
[0123] 7A-7E illustrate examples of methods for generating virtual lighting effects while an electronic device is presenting a content item, according to some embodiments.
[0124] FIG. 7A shows electronic device 101 displaying a three-dimensional environment 702 via display generating component 120. It should be understood that in some embodiments, electronic device 101 utilizes one or more of the techniques described with reference to FIGS. 7A-7E in a two-dimensional environment without departing from the scope of this disclosure. As described above with reference to FIGS. 1-6 , electronic device 101 optionally includes display generating component 120 (e.g., a touchscreen) and multiple image sensors 314. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that electronic device 101 can use to capture one or more images of a user or a portion of a user while the user interacts with electronic device 101. In some embodiments, display generating component 120 is a touchscreen capable of detecting a user's hand gestures and movements. In some embodiments, the user interfaces described below may also be implemented in a head-mounted display that includes display generation components that display the user interface to the user and sensors that detect the physical environment and / or the movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).
[0125] 7A , electronic device 101 displays content item 704 in a three-dimensional environment 702. Three-dimensional environment 702 further includes representations of real objects in the physical environment of electronic device 101, including a representation of a table 706a, a representation of a sofa 706b, a representation of a wall 708a, and a representation of a ceiling 708b. Electronic device 101 detects, via one or more sensors 314, a user's gaze 713a directed toward content item 704, which is optionally video content currently being played on electronic device 101. In some embodiments, electronic device 101 displays three-dimensional environment 702 with one or more virtual lighting effects as user's gaze 713a is directed toward content item 704 while content item 704 is being played.
[0126] In some embodiments, the virtual lighting effects generated by electronic device 101 include visually highlighting the content item 704, such as by blurring and / or darkening portions of three-dimensional environment 702 that do not include content item 704, and displaying virtual light spill emanating from content item 704. The virtual light spill includes virtual lighting 710a displayed on wall representation 708a, virtual lighting 710b displayed on ceiling representation 708b, virtual lighting 710c displayed on table representation 706a, and virtual lighting 710d displayed on sofa representation 706b. In some embodiments, virtual lighting 710a-d is based on the video content of content item 704. For example, the color, intensity, etc. of the virtual lights 710a-d are optionally based on the color, intensity, etc. of the video content of the content item 704 to simulate the virtual lights 710a-d being reflections of the video content of the content item 704 and / or the virtual lights 710a-d emanating from the content item 704 on various surfaces within the three-dimensional environment 702. In some embodiments, the sizing of the virtual lights 710a, 710b, 710c, and 710d is based on the distance of the content item 704 from the respective surfaces of the virtual lighting effect, and the position of the virtual lights 710a, 710b, 710c, and 710d is based on the position of the content item 704 within the three-dimensional environment 702.
[0127] 7A, the user's hand 703a is in a pose that causes the electronic device 101 to not display one or more selectable options for controlling playback of the content item 704, as described in more detail below with reference to Figures 7B-7E. Thus, the electronic device 101 ceases to display selectable options for controlling playback of the content item 704 of Figure 7A.
[0128] 7B illustrates the display of multiple selectable options 712a-L for controlling playback of the content item 704 and user input corresponding to a request to resize the content item 704 in the three-dimensional environment 702. In some embodiments, the electronic device 101 displays the selectable options 712a-L in response to detecting a user's hand 703b in a ready pose while the user's gaze 713a is directed at the content item 704. In some embodiments, detecting the hand 703b in a ready pose includes detecting the hand 703b in a pre-pinch hand shape with the thumb within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, or 5 centimeters) of but not touching another finger of the hand 703b, or detecting the hand 703b in a pointing hand shape with one or more fingers extended and one or more fingers curled toward the palm.
[0129] 7B , multiple selectable options 712c-L for controlling playback of content item 704 are displayed in a user interface element 711 that is separate from content item 704. In some embodiments, user interface element 711 is angled toward the user's viewpoint at an angle different from the angle at which content item 704 is displayed relative to the user's viewpoint, as described in more detail below with reference to method 1400. For example, user interface element 711 is displayed at a lower height in three-dimensional environment 702 than content item 704, so electronic device 101 orients user interface element 711 upward relative to the angle of content item 704. As shown in FIG. 7B , a portion of user interface element 711 is visually or spatially overlaid on content item 704, and user interface element 711 is, optionally, displayed closer to the user's viewpoint in three-dimensional environment 702 than content item 704.
[0130] Selectable options 712c-L included in user interface element 711 are now described. Option 712c, when selected, causes electronic device 101 to display content item 704 in an immersive content mode according to one or more steps of method 1400. Option 712d, when selected, causes electronic device 101 to display content item 704 in a picture-in-picture element according to one or more steps of method 1000. Option 712e, when selected, causes electronic device 101 to present a content playback queue (e.g., in a user interface element separate from content item 704 and / or in place of content item 704) including one or more content items that electronic device 101 is configured to play next. Option 712f, when selected, causes electronic device 101 to adjust the playback position of content item 704 back a predetermined amount (e.g., 5, 10, 15, 30, or 60 seconds). Option 712g, when selected, causes the electronic device 101 to pause playback of the content item 704. In some embodiments, while the content item 704 is paused, the electronic device 101 ceases displaying the content item 704 with greater visual emphasis relative to the rest of the three-dimensional environment. In some embodiments, while the content item 704 is paused, the electronic device 101 ceases displaying virtual lighting effects, including blurring and / or dimming the three-dimensional environment 702 other than the content item 704, and / or virtual light spill emanating from the content item 704. Option 712h, when selected, causes the electronic device 101 to adjust the playback position of the content item 704 forward a predetermined amount (e.g., 5, 10, 15, 30, or 60 seconds). Option 712i, when selected, causes the electronic device 101 to present subtitle options associated with the content item 704. When selected, option 712j causes the electronic device 101 to adjust virtual lighting effects such as blurring effects and / or dimming effects and / or light spill effects, as described in more detail below with reference to Figures 7C-7E.Option 712k, when selected, causes the electronic device 101 to adjust the playback volume of the audio content included in the content item 704. The user interface element 711 further includes a scrub bar 712L that includes an indication of the current playback position of the content item 704, and in response to an input that moves the indication of the current playback position, causes the electronic device 101 to adjust the playback position of the content item 704 and resume playback of the content item 704 from the adjusted playback position.
[0131] In addition to the selectable options 712c-712L displayed in the user interface element 711, the electronic device 101 further displays a selectable option 712a that, when selected, causes the electronic device 101 to stop displaying the content item 704 (and, optionally, the user interface element 711 and options 712a-712b), and a selectable option 712b that causes the electronic device 101 to resize the content item 704 within the three-dimensional environment 702 when the electronic device 101 detects an input directed at it. The selectable option 712a for closing the content item 704 is displayed overlaid on the content item 704 outside of the user interface element 711. The selectable option 712b for resizing the content item 704 is displayed outside of the content item 704 and the user interface element 711.
[0132] 7B , electronic device 101 receives input provided by hand 703a and gaze 713b directed toward selectable option 712b for resizing content item 704. In some embodiments, detecting the input includes detecting that the user executes a selection pose with hand 703a including a predetermined hand shape, such as a pinch hand shape in which the thumb of hand 703a touches another finger of the hand, or a pointing hand shape in which one or more fingers of hand 703a are extended and one or more fingers of hand 703a are curled toward the palm of hand 703a, while gaze 713b is directed toward option 712b. In some embodiments, while maintaining the predetermined hand shape, the user moves their hand, and in response to detecting the movement, electronic device 101 resizes content item 704 according to the movement (e.g., speed, duration, distance, direction, etc.) of hand 703a. As will be described in more detail below with reference to FIG. 7C, the electronic device 101 does not resize the element 711 when resizing the content item 704 according to input provided by the hand 703a and gaze 713b.
[0133] 7C illustrates how electronic device 101 resizes content item 704 in response to the input shown in FIG. 7B without changing the size of user interface element 711. In response to the input shown in FIG. 7B, electronic device 101 displays content item 704 in FIG. 7C at a smaller size than the size at which content item 704 was displayed in FIG. 7B, while maintaining the size of user interface element 711. Electronic device 101 also updates one or more characteristics (e.g., size) of virtual lights 710a-710c in accordance with the updated size of content item 704, for example, by reducing the size of virtual lights 710a-710c in three-dimensional environment 702 to correspond to the reduced size of content item 704. In FIG. 7C, electronic device 101 optionally continues to display user interface element 711 and other selectable options in response to detecting hand 703b in a ready state as described above.
[0134] 7C also shows electronic device 101 displaying user interface element 714 indicating the functionality of virtual lighting effects option 712j in response to user gaze 713d being directed to option 712j. In some embodiments, electronic device 101 displays user interface element 714 associated with lighting effects option 712j because lighting effects option 712j is particularly related to displaying content items in three-dimensional environment 702. As shown in FIG. 7C, if user gaze 713c is instead directed to option 712f to adjust the playback position of the content item back a predetermined amount, electronic device 101 optionally ceases displaying user interface element 714 associated with option 712f because option 712f is not particularly related to presenting the content item in a three-dimensional environment than to presenting the content item in other environments or user interfaces. In some embodiments, a first plurality of interactive elements in user interface element 711 is associated with a visual indication similar to indication 714, and a second plurality of interactive elements in user interface element 711 is not associated with a visual indication similar to indication 714.
[0135] 7C , the electronic device 101 detects input provided by a hand 703a and gaze 713a directed toward the content item 704. In some embodiments, the input corresponds to a request to update the position of the content item within the three-dimensional environment 702. In some embodiments, the electronic device 101 displays a user interface element other than the content item itself that, when selected, causes the electronic device 101 to begin a process of repositioning the content item 704 within the three-dimensional environment 702. In some embodiments, the reposition user interface element, similar to the resize user interface element 714, is displayed outside the content item 704 and outside the user interface element 711. In some embodiments, the reposition user interface element is a horizontal bar or line aligned along the bottom of the user interface element 711. In some embodiments, the reposition user interface element is either included in the user interface element 711 or is overlaid on the content item 704 without being included in the user interface element 711.
[0136] In some embodiments, detecting an input corresponding to a request to reposition the content item 704 within the three-dimensional environment 702 includes detecting that the user has made a predefined hand shape with their hand 703a, such as the pinch hand shape or pointing hand shape described above, while detecting a gaze 713a directed toward the content item 704. In some embodiments, while the user is making the predefined hand shape, the electronic device 101 detects movement of the hand 703a and, in response, moves the content item 704 and the user interface element 711 according to the movement (e.g., distance, duration, speed, direction, etc.) of the hand 703a, as shown in FIG.
[0137] Figure 7D shows electronic device 101 displaying content item 704 and user interface element 711 at updated locations within three-dimensional environment 702 in accordance with the input shown in Figure 7C. In accordance with the input shown in Figure 7C, electronic device 101 displays content item 704 and user interface element 711 in Figure 7D at a location within three-dimensional environment 702 that is closer to the user's viewpoint from which three-dimensional environment 702 is displayed than the location of content item 704 and user interface element 711 in Figure 7C. In accordance with displaying content item 704 closer to the user's viewpoint in Figure 7D, electronic device 101 displays content item 704 at a larger angular size in Figure 7D than in Figure 7C (e.g., occupying more space in the field of view of the user and / or display generation component 120), although in some embodiments, the size of content item 704 within three-dimensional environment 702 is the same in Figures 7C and 7D (e.g., the size of content item 704 within three-dimensional environment 702 does not change from Figure 7C to Figure 7D). As shown in Figures 7C-7D, the electronic device 101 does not update the angular size of the user interface element 711, even though the user interface element 711 is closer to the user's viewpoint in Figure 7D than in Figure 7C. In some embodiments, refraining from updating the angular size of the user interface element 711 includes updating the size of the user interface element 711 in the three-dimensional environment 702. For example, because the user interface element 711 is closer to the user's viewpoint in Figure 7D than in Figure 7C, but is displayed with the same angular size in Figures 7C and 7D, the electronic device 101 reduces the size of the user interface element 711 in the environment 702 of Figure 7D compared to the size of the user interface element 711 in the environment 702 of Figure 7C. As shown in Figure 7D, as the content item 704 moves, the electronic device 101 updates the positions of the virtual lights 710a, 710b, and 710d on other portions of the three-dimensional environment 702.
[0138] 7D , electronic device 101 detects input directed toward lighting effect option 712j provided by gaze 713e and hand 703a. In some embodiments, detecting the input includes detecting hand 703a forming the pinch hand shape described above or the pointing hand shape described above while detecting gaze 713e directed toward option 712j. In some embodiments, in response to detecting a pinch hand shape or a pointing hand shape that is less than a predetermined time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds), electronic device 101 switches lighting effects (e.g., blur and / or darkening lighting effects and / or light spill lighting effects) on or off. In some embodiments, in response to detecting a pinch hand shape or a pointing hand shape that exceeds the time threshold, electronic device 101 presents slider element 716 that, when interacted with, allows the user to adjust the level (e.g., intensity) of the lighting effect. In some embodiments, both the blur and / or dimming lighting effect and the light spill lighting effect are updated according to inputs directed to the lighting effect options 712j and / or slider 716. In some embodiments, the three-dimensional environment 702 includes separate elements for adjusting each lighting effect. As shown in FIG. 7D , the electronic device 101 detects the user's gaze 713e directed toward the slider 716 and detects the movement of the hand 703a while the hand 703a makes a pinch or pointing hand shape. In response to the inputs shown in FIG. 7E , the electronic device 101 reduces the intensity of the blur and / or dimming lighting effect and the light spill lighting effect, as shown in FIG.
[0139] 7E illustrates a three-dimensional environment 702 in which the visual emphasis of the content item 704 relative to the rest of the three-dimensional environment 702 has been reduced, e.g., the intensity of the blurring and / or darkening lighting effects and light spill effects has been reduced. For example, in FIG. 7E, the amount of darkening and / or blurring applied to areas of the three-dimensional environment 702 other than the content item 704 has been reduced, and the size and / or intensity of the virtual lights 710a, 710b, and 710d has been reduced. In some embodiments, the electronic device 101 reduces the intensity of the virtual lighting effects in response to the input shown in FIG. 7D. In some embodiments, the electronic device 101 reduces the intensity of the virtual lighting effects (or ceases displaying them) in response to detecting a user's gaze 713f directed away from the content item 704. In some embodiments, the electronic device reduces the intensity of the virtual lighting effects (or ceases displaying them) in response to detecting that the content item has been paused. FIG. 7E also shows electronic device 101 ceasing to display selectable options 712a-712L and user interface elements 711 in response to no longer detecting the user's hand in the ready state.
[0140] Additional or alternative details regarding the embodiment shown in FIGS. 7A-7E are provided below in the description of method 800, which is described with reference to FIGS. 8A-8O.
[0141] 8A-8O are flowcharts illustrating a method for generating virtual lighting effects while presenting a content item, according to some embodiments. In some embodiments, method 800 is performed in a computer system (e.g., computer system 101 of FIG. 1 ) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4 ) (e.g., a heads-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing down the user's hand (e.g., color sensors, infrared sensors, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 800 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations of method 800 are, optionally, combined, and / or the order of some operations is, optionally, changed.
[0142] In some embodiments, such as in FIG. 7A , method 800 is performed in an electronic device (e.g., 101) that communicates with a display generation component (e.g., 120) and one or more input devices (e.g., 314) (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users. In some embodiments, the one or more input devices include electronic devices or components that can receive user input (e.g., capture user input, detect user input, etc.) and transmit information related to the user input to the electronic device. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., hand tracking device, hand motion sensor), etc. In some embodiments, the electronic device is in communication with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreen, trackpad). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.
[0143] In some embodiments, such as in FIG. 7B , while presenting a content item (e.g., 704) within a three-dimensional environment (e.g., 702), the electronic device (e.g., 101), via a display generation component (e.g., 120), displays (802a) a user interface (e.g., 711) associated with the content item, the user interface (e.g., 711) including one or more user interface elements (e.g., 712f) for modifying playback of the content item and a separate user interface element (e.g., 712j) for modifying virtual lighting effects affecting the appearance of the three-dimensional environment (e.g., 702). In some embodiments, the content item is video content, and the one or more user interface elements for modifying playback of the content item include play / pause, skip ahead, skip back, subtitles, and audio options. In some embodiments, the one or more user interface elements include an option for displaying the content item in a picture-in-picture user interface element according to one or more steps of method 1000 or for displaying the content item in an immersive (e.g., full-screen) mode according to one or more steps of method 1400.
[0144] In some embodiments, the content item is an item of video content, such as a movie, an episode of a series of episodic content, or a video clip, that is being played / displayed within the three-dimensional environment, or the content item is an item of audio content, such as a music, podcast, or audiobook, that is being played within the three-dimensional environment. In some embodiments, the three-dimensional environment includes virtual objects, such as application windows, operating system elements, representations of other users, and / or representations of content items and physical objects in the physical environment of the electronic device. In some embodiments, representations of physical objects are displayed in the three-dimensional environment through display generation components (e.g., virtual or video pass-through). In some embodiments, the representations of physical objects are views of physical objects in the physical environment of the electronic device seen through transparent portions of the display generation components (e.g., true pass-through or real pass-through). In some embodiments, the electronic device displays the three-dimensional environment from the user's perspective at a location in the three-dimensional environment that corresponds to the physical location of the electronic device in the physical environment of the electronic device. In some embodiments, the three-dimensional environment is generated, displayed, or otherwise made visible by a device (e.g., a computer-generated reality (XR) environment, such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment).
[0145] In some embodiments, such as in FIG. 7D , while displaying a user interface (e.g., 711) associated with a content item (e.g., 704), the electronic device (e.g., 101) receives (802b) user input via one or more input devices directed to a respective user interface element (e.g., 716), the user input corresponding to a request to modify a virtual lighting effect. In some embodiments, the input corresponds to a request to change the amount of the virtual lighting effect to a different (e.g., non-zero) amount. In some embodiments, the input corresponds to a request to display the three-dimensional environment without virtual lighting effects. In some embodiments, the virtual lighting effects are based on the content item, such as light spill effects on representations of virtual and / or real objects in the three-dimensional environment having color, pattern, and / or movement based on the image and / or video content of the content item. In some embodiments, the virtual lighting effects are independent of the content item, such as a degree of dimming and / or blurring applied to areas of the three-dimensional environment other than the content item or the user interface including one or more user interface elements for modifying playback of the content item (discussed below).
[0146] In some embodiments, such as FIG. 7E, in response to receiving user input (802c), the electronic device (e.g., 101) continues to present (802d) a content item (e.g., 704) within the three-dimensional environment (e.g., 702).
[0147] In some embodiments, such as in FIG. 7E , in response to receiving user input (802c), the electronic device (e.g., 101) applies virtual lighting effects to the three-dimensional environment (e.g., 702) (802e). For example, in response to a request to display the three-dimensional environment without dimming virtual lighting effects, the electronic device displays the three-dimensional environment such that (e.g., all) areas of the three-dimensional environment, including areas where content is presented and areas where content is not presented, have the same relative dimming. As another example, in response to a request to display the three-dimensional environment with increasing amounts of light spill virtual lighting effects, the electronic device increases the size and / or brightness of light spill (e.g., from content items) on objects in the three-dimensional environment. In some embodiments, the virtual light spill changes over time according to changes to the content items (e.g., visual content included in the content items).
[0148] Modifying the amount of virtual lighting effects in which a three-dimensional environment is displayed provides an efficient way to switch between an immersive experience and a less distracting virtual environment, thereby reducing the cognitive burden on the user both when engaging with the content item and when engaging with other content or applications within the three-dimensional environment.
[0149] In some embodiments, prior to receiving user input directed to a respective user interface (e.g., 711 in FIG. 7C ), the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (804a) the three-dimensional environment (e.g., 702) without virtual lighting effects. In some embodiments, the virtual lighting effects include one or more of: displaying virtual light spill effects emanating from the content items in areas of the three-dimensional environment other than the content items; dimming areas of the three-dimensional environment other than the content items relative to the content items; and / or blurring areas of the three-dimensional environment other than the content items relative to the content items. In some embodiments, displaying the three-dimensional environment without virtual lighting effects includes withdrawing display of virtual light spill effects emanating from the content items; displaying areas of the three-dimensional environment other than the content items with the same blur level as the content items; and / or displaying areas of the three-dimensional environment other than the content items with the same dimming level as the content items.
[0150] In some embodiments, in response to receiving the user input, the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (804b) the three-dimensional environment (e.g., 702) with virtual lighting effects, such as those in FIG. 7A . In some embodiments, the virtual lighting effects include one or more of: displaying virtual light spill effects emanating from the content items in areas of the three-dimensional environment other than the content items; dimming areas of the three-dimensional environment other than the content items relative to the content items; and / or blurring areas of the three-dimensional environment other than the content items relative to the content items. In some embodiments, in response to detecting a first input directed at a respective user interface element, the electronic device toggles the display of the three-dimensional environment with or without virtual lighting effects. For example, the first input is a selection (e.g., a primary selection such as a “click”) of the respective user interface element. In some embodiments, the first input includes detecting a predetermined portion of the user (e.g., the user's hand in a pinching hand shape) in a predetermined shape for a period of less than a respective threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds) while the user's gaze is directed at the respective user interface element. In some embodiments, in response to a second input (e.g., a secondary selection, an input similar to a "long click") directed at the respective user interface element, the electronic device updates the respective user interface element to a user interface element for adjusting the amount of lighting effects applied to the three-dimensional environment, as described below. For example, detecting the second input includes detecting a predetermined portion of the user (e.g., the user's hand) making a predetermined gesture, such as a pinch hand gesture, that includes touching a thumb to another finger for more than a threshold amount of time (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds) before removing the thumb from the fingers.
[0151] Turning on virtual lighting effects in response to input directed at individual user interface elements provides an efficient way to switch virtual lighting effects on and off, thereby reducing the cognitive burden on the user to switch between an immersive experience and a three-dimensional environment in which other user interface elements are clearly displayed.
[0152] In some embodiments, prior to receiving user input directed to a respective user interface (e.g., 711), such as in FIG. 7iD, the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (806a) the three-dimensional environment (e.g., 702) with a first amount of virtual lighting effects. In some embodiments, displaying the three-dimensional environment with the first amount of virtual lighting effects includes one or more of: displaying virtual light spill effects emanating from the content items at a first size, intensity, sharpness, etc., in areas of the three-dimensional environment other than the content items; dimming areas of the three-dimensional environment other than the content items by a first amount relative to the content items; and / or blurring areas of the three-dimensional environment other than the content items by a first amount relative to the content items. In some embodiments, the first amount of virtual lighting effects is zero (e.g., the electronic device displays the three-dimensional environment without virtual lighting effects).
[0153] In some embodiments, in response to receiving a user input, the electronic device (e.g., 101) displays (806b), via the display generation component (e.g., 120), the three-dimensional environment (e.g., 702) with a second amount of virtual lighting effects, the second amount being different from the first amount, such as in FIG. 7E . In some embodiments, the electronic device modifies multiple virtual lighting effects (e.g., light spill, blur, dimming) by the same amount in response to the input. For example, in response to an input to increase a virtual lighting effect by a discrete amount, the electronic device increases the size, intensity, sharpness, etc. of the light spill by a discrete amount, increases the blur by a discrete amount, and increases the dimming by a discrete amount. As another example, in response to an input to decrease a virtual lighting effect by a discrete amount, the electronic device decreases the size, intensity, sharpness, etc. of the light spill by a discrete amount, decreases the blur by a discrete amount, and decreases the dimming by a discrete amount. In some embodiments, the electronic device adjusts the amount of virtual lighting effects according to a determination that input directed to a respective user interface element satisfies one or more respective criteria, such as a criterion met when the electronic device detects a predetermined portion of the user (e.g., the user's hand) in a predetermined hand shape (e.g., a pinching hand shape) for more than a threshold time period (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds). In some embodiments, in response to detecting the predetermined pose for more than the threshold time period, the electronic device presents an interactive slider that controls the amount of virtual lighting effects with which the three-dimensional environment is displayed. For example, continuation of input directed to a respective user interface element, including movement of a predetermined portion of the user, causes the electronic device to adjust the amount of virtual lighting effects and the position of an indicator on the slider.
[0154] Adjusting the amount of virtual lighting effects applied to the three-dimensional environment in response to inputs directed at individual user interface elements provides an efficient way to select a trade-off between the immersion of a content item and the clarity of elements in the three-dimensional environment other than the content item, which reduces the cognitive burden on the user when interacting with various elements in the three-dimensional environment.
[0155] In some embodiments, such as in FIG. 7C , prior to receiving user input, areas of the three-dimensional environment (e.g., 702) that do not include a content item (e.g., 704) are displayed at a first brightness level (808a). In some embodiments, the content item is displayed at a brightness level higher than the first brightness level (e.g., a dimming visual effect is active). In some embodiments, the content item is displayed at the same brightness level as areas of the three-dimensional environment that do not include a content item (e.g., a dimming visual effect is not active). In some embodiments, the three-dimensional environment includes representations of virtual and / or real objects.
[0156] In some embodiments, such as in FIG. 7E , displaying the three-dimensional environment (e.g., 702) with virtual lighting effects in response to user input includes displaying regions of the three-dimensional environment (e.g., 702) that do not include content items (e.g., 704) at a second brightness level that is different (e.g., lower, higher) from the first brightness level (808b). In some embodiments, the electronic device modifies the brightness level of the regions that do not include content items without modifying the brightness levels of the content items. In some embodiments, in response to the input, the electronic device turns a dimming visual effect on or off. In some embodiments, in response to the input, the electronic device adjusts the amount of the dimming visual effect. In some embodiments, the entire three-dimensional environment is made up of regions that do not include content items and the content items (e.g., the brightness of the entire three-dimensional environment other than the content items is adjusted). In some embodiments, the three-dimensional environment includes a third region that does not include content items that is not affected by the adjustment in the amount of brightness. In some embodiments, the three-dimensional environment includes representations of one or more virtual objects and / or areas outside the content item and / or real objects and / or areas outside the content item that are dimmed when the three-dimensional environment is displayed using virtual lighting effects.
[0157] Adjusting the brightness level of areas of the three-dimensional environment that do not contain content items in response to inputs directed at individual user interface elements provides an efficient way to make a trade-off between an immersive experience with the content items and the ability to see and interact with elements in the three-dimensional environment other than the content items.
[0158] In some embodiments, such as in FIG. 7A , displaying the three-dimensional environment (e.g., 702) with virtual lighting effects in response to user input includes displaying (810) individual virtual lighting effects (e.g., 710a) (e.g., virtual light spill) emanating from the content item (e.g., 704) on one or more objects (e.g., 708a) in the three-dimensional environment (e.g., 704). In some embodiments, the individual virtual lighting effects change over time based on the content. For example, the individual virtual lighting effects include one or more colors, intensities, patterns, animations, etc. currently included in the video and / or image content of the content item. In some embodiments, the individual virtual lighting effects are virtual light spills that simulate reflections of light emanating from the image and / or video content of the content item on one or more (e.g., real, virtual) surfaces and / or objects in the three-dimensional environment outside the content item. In some embodiments, the one or more objects include virtual objects. In some embodiments, the one or more objects include representations of real objects in the physical environment of the electronic device and / or display generating component. In some embodiments, the representation of the real object includes one or more of a true or real pass-through and / or a video or virtual pass-through, as described in more detail above. In some embodiments, the virtual light spill is displayed more intensely (e.g., displayed with higher brightness, sharpness, and / or size) at locations in the three-dimensional environment closer to the content item than at locations in the three-dimensional environment further from the content item.
[0159] In some embodiments, presenting individual virtual lighting effects emanating from a content item on one or more objects within a three-dimensional environment provides an immersive experience with the content item, which reduces distractions and cognitive burden for the user while consuming the content item.
[0160] In some embodiments, such as FIG. 7A , prior to receiving user input directed to a respective user interface (e.g., 711), the electronic device (e.g., 101), via the display generation component (e.g., 120), displays the three-dimensional environment (e.g., 702) with a first amount of virtual lighting effects (e.g., 812a), including displaying areas of the three-dimensional environment (e.g., 702) that do not include a content item at a first brightness level and displaying a first amount of a respective virtual lighting effect (e.g., 710a) (e.g., virtual light spill) emanating from the content item (e.g., 704) on one or more objects (e.g., 708a) in the three-dimensional environment. In some embodiments, the virtual lighting effects include dimming areas of the three-dimensional environment other than the content item relative to the content item and displaying virtual light spill emanating from the content item on other objects in the three-dimensional environment. In some embodiments, the first amount of virtual lighting effects is zero (e.g., the electronic device displays the three-dimensional environment without virtual lighting effects).
[0161] In some embodiments, such as in FIG. 7E , in response to receiving a user input, the electronic device (e.g., 101), via the display generation component (e.g., 120), displays the three-dimensional environment (e.g., 702) with a second amount of virtual lighting effects (812b), including displaying areas of the three-dimensional environment (e.g., 702) that do not include the content item (e.g., 704) at a second brightness level and displaying a second amount of individual virtual lighting effects (e.g., 710a) (e.g., virtual light spill) emanating from the content item (e.g., 704) on one or more objects in the three-dimensional environment (e.g., 702). In some embodiments, the individual user interface element controls both the level of dimming of areas of the three-dimensional environment other than the content item relative to the content item and the level (e.g., brightness, size, translucency, etc.) of virtual light spill emanating from the content item on other objects in the three-dimensional environment. In some embodiments, in response to a first input directed at the individual user interface element, the electronic device switches the dimming and light spill virtual lighting effects on or off. In some embodiments, in response to a second input directed at the respective user interface element, the electronic device adjusts the level and level(s) of dimming (e.g., brightness, size, translucency, etc.) of virtual light spill emanating from the content item onto other objects outside the content item in the three-dimensional environment.
[0162] Adjusting both the amount of individual virtual lighting effects and the amount of dimming emanating from a content item in response to inputs directed at individual user interface elements provides an efficient way to adjust multiple characteristics of the virtual lighting of a three-dimensional environment with a single input, thereby reducing the amount of time and number of inputs required to make the adjustments.
[0163] 7A , displaying the three-dimensional environment (e.g., 702) with virtual lighting effects in response to user input includes displaying (815b) the three-dimensional environment (e.g., 702) with a first amount of virtual lighting effects via a display generation component (e.g., 120) in accordance with detecting, via one or more input devices (e.g., 314), that a user of the electronic device (e.g., 101) has their attention (e.g., gaze 713a) directed toward a first region (e.g., including the content item 704) of the three-dimensional environment (e.g., 702) (e.g., detecting, via an eye-tracking device, that the user's gaze is directed toward the first region). In some embodiments, in response to detecting that the user's attention (e.g., gaze) is directed toward the content item, the electronic device increases the amount of virtual lighting effects with which the three-dimensional environment is displayed.
[0164] In some embodiments, such as in FIG. 7E , displaying the three-dimensional environment (e.g., 702) with virtual lighting effects in response to user input includes: displaying (814c) the three-dimensional environment with a second amount of virtual lighting effects (814a) different from the first amount via a display generation component (e.g., 120) in accordance with detecting, via one or more input devices (e.g., 314), that the user's attention (e.g., gaze 713f) is directed to a second region of the three-dimensional environment (e.g., 702) different from the first region (e.g., detecting, via an eye-tracking device, that the user's gaze is directed to a second region that does not include a content item). In some embodiments, in response to detecting that the user's attention (e.g., gaze) is not directed to the content item, the electronic device reduces the amount of virtual lighting effects with which the three-dimensional environment is displayed. In some embodiments, in response to detecting that the user's attention (e.g., gaze) is not directed to the content item, the electronic device updates the three-dimensional environment to display the three-dimensional environment without the virtual lighting effects.
[0165] Adjusting the amount of virtual lighting effects depending on the area of the three-dimensional environment to which the user focuses their attention provides an efficient way to automatically adjust the level of immersion in a content item based on the user's attention, which reduces the cognitive burden and number of inputs for the user when the user changes the area of the three-dimensional environment to which they focus their attention.
[0166] In some embodiments, while displaying the three-dimensional environment without virtual lighting effects, the electronic device (e.g., 101) receives 816a a first input via one or more input devices (e.g., 314) directed at a distinct user interface element (e.g., 712j in FIG. 7C ), which includes detecting a predefined portion of a user of the electronic device (e.g., hand 703b) in a predefined pose for less than a predetermined time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds) via one or more input devices (e.g., 314). In some embodiments, the predefined pose of the predefined portion of the user is the user's hand making a pinch shape with the thumb touching the fingers of the hand. In some embodiments, the pre-defined pose of the pre-defined portion of the user is making a pointing hand shape with one or more fingers outstretched and one or more fingers curled toward the palm while the user's hand is within a threshold distance (e.g., 1, 2, 3, 5, 10, 15, 30, or 50 centimeters) of a location in the three-dimensional environment corresponding to the individual user interface element. In some embodiments, detecting the first input further includes detecting, via one or more input devices, that the user's attention is directed toward the individual user interface element. In some embodiments, detecting the first input further includes detecting, via eye-tracking devices of the one or more input devices, that the user's gaze is directed toward the individual user interface element.
[0167] In some embodiments, in response to receiving the first input, the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (816b) the three-dimensional environment (e.g., 702) with virtual lighting effects, such as those in Figure 7A. In some embodiments, in response to detecting the first input while displaying the three-dimensional environment without virtual lighting effects, the electronic device switches on the virtual lighting effects.
[0168] In some embodiments, while displaying the three-dimensional environment with virtual lighting effects, the electronic device (e.g., 101) receives (816c) a second input via one or more input devices (e.g., 314) directed at a respective user interface element (e.g., 712j in FIG. 7C ), comprising detecting, via one or more input devices (e.g., 314), a predetermined portion (e.g., hand 703b) of a user of the electronic device (e.g., 101) in a predetermined pose for less than a predetermined time threshold (e.g., 0.1, 0.2, 0.3, 0.5, 1, or 2 seconds). In some embodiments, detecting the second input further comprises detecting, via the one or more input devices, that the user's attention is directed at the respective user interface element. In some embodiments, detecting the second input further comprises detecting, via an eye-tracking device of the one or more input devices, that the user's gaze is directed at the respective user interface element.
[0169] In some embodiments, in response to receiving the second input, the electronic device (e.g., 101) displays (816d) the three-dimensional environment (e.g., 314) without the virtual lighting effects via the display generation component (e.g., 120). In some embodiments, in response to detecting the second input while displaying the three-dimensional environment with the virtual lighting effects, the electronic device switches off the virtual lighting effects.
[0170] Turning virtual lighting effects on or off in response to detecting a predetermined pose of a predetermined portion of the user for less than a threshold amount of time provides an efficient way to switch with less distraction between an immersive experience with a content item and viewing other portions of the three-dimensional environment, thereby reducing the cognitive burden on the user when interacting with the three-dimensional environment.
[0171] In some embodiments, such as in FIG. 7D , while displaying the three-dimensional environment (e.g., 702) with a first amount of virtual lighting effects, the electronic device (e.g., 101) receives input (818a) directed at a respective user interface element (e.g., 712j), including detecting movement of a predetermined portion of the user of the electronic device (e.g., 101) via one or more input devices (e.g., 314) while the predetermined portion of the user (e.g., hand 703a) is in a predetermined pose. In some embodiments, detecting the input further includes detecting that the user's attention (e.g., gaze) is directed at the respective user interface element via one or more input devices (e.g., eye-tracking devices). In some embodiments, the predetermined pose is the user's hands in the pinch hand shape described above. In some embodiments, the predetermined pose is the user's hands in the pointing hand shape described above while the user's hands are within a threshold distance (e.g., 5, 10, 15, 30, or 50 centimeters) of the location of the respective user interface element.
[0172] In some embodiments, in response to input directed at a respective user interface element (e.g., 712j in FIG. 7D ), the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (818b) the three-dimensional environment (e.g., 702) with a second amount of virtual lighting effects, such as in FIG. 7E , where the second amount is based on the movement (e.g., speed, duration, distance, etc.) of a predetermined portion of the user (e.g., hand 703a in FIG. 7D ) while the predetermined portion of the user is in a predetermined pose. In some embodiments, in response to detecting movement of the predetermined portion of the user in a first direction (e.g., downward, leftward), the electronic device decreases the amount of the virtual lighting effects. In some embodiments, in response to detecting movement of the predetermined portion of the user in a second direction (e.g., upward, rightward), the electronic device increases the amount of the virtual lighting effects. In some embodiments, in response to movement having a first magnitude (e.g., speed, duration, distance), the electronic device modifies the amount of the virtual lighting effects by a first amount corresponding to the first magnitude. In some embodiments, in response to a movement having a second magnitude (e.g., of speed, duration, or distance), the electronic device changes the amount of the virtual lighting effect by a second amount corresponding to the first magnitude. For example, in response to a downward movement having a relatively small magnitude, the electronic device decreases the amount of the virtual lighting effect by a relatively small amount. As another example, in response to an upward movement having a relatively large magnitude, the electronic device increases the amount of the virtual lighting effect by a relatively large amount.
[0173] Adjusting the amount of virtual lighting effects based on movement of a predetermined portion of the user during input directed at an individual user interface element provides an efficient way for the user to make a trade-off between an immersive experience with a content item and the clarity of the rest of the three-dimensional environment, which reduces the cognitive burden on the user when interacting with elements in the three-dimensional environment.
[0174] In some embodiments, while the content item is being played (820a), the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (820b) the three-dimensional environment (e.g., 702) using virtual lighting effects, such as those in Figure 7A. In some embodiments, the virtual lighting effects include one or more of blurring and / or dimming areas of the three-dimensional environment that do not include the content item and / or virtual light spill emanating from the content item as described above.
[0175] In some embodiments, while the content item is being played (820a), the electronic device (e.g., 101) receives (820c) user input via one or more input devices (e.g., selection of option 712g in FIG. 7B ) corresponding to a request to pause the content item (e.g., 704). For example, the input corresponding to the request to pause the content item is a selection of a respective one of one or more user interface elements for modifying playback of the content item displayed in a user interface associated with the content item.
[0176] In some embodiments, in response to receiving user input corresponding to a request to pause the content item (820d), the electronic device (e.g., 101) pauses (820e) the content item (e.g., 704 in FIG. 7A). In some embodiments, the electronic device continues to display the paused content item (e.g., displaying a frame of the video content at the playback position where the video content was paused).
[0177] In some embodiments, in response to receiving a user input corresponding to a request to pause a content item (e.g., 704 in FIG. 7A ) (820d), the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (820f) the three-dimensional environment (e.g., 702) without virtual lighting effects (or with lighting effects having a reduced magnitude). In some embodiments, in response to receiving an input to play the content item while the content item is paused and the electronic device is displaying the three-dimensional environment without virtual lighting effects, the electronic device resumes playing the content item and displays the three-dimensional environment with virtual lighting effects. In some embodiments, in response to the input to pause the content item, the electronic device reduces the amount of virtual lighting effects in which the three-dimensional environment is displayed without ceasing the display of the virtual lighting effects. Ceasing the display of virtual lighting effects in response to the input to pause the content item provides an efficient way to improve the readability of elements in the three-dimensional environment other than the content item while the content item is paused, thereby reducing the number of inputs required to switch between engaging with the content item and other elements in the three-dimensional environment.
[0178] In some embodiments, such as FIG. 7C , the electronic device (e.g., 101) receives (822a) a separate user input via one or more input devices (e.g., 314) directed to a second separate interface element of one or more user interface elements (e.g., 712f) for modifying playback of the content item (e.g., 704).
[0179] In some embodiments, in response to receiving a distinct user input (822b), the electronic device (e.g., 101) toggles (822c) the play or pause state of the content item (e.g., 704) in accordance with determining that the second distinct user interface element is a user interface element (e.g., 712g in FIG. 7B ) that, when selected, causes the electronic device (e.g., 101) to toggle between playing and pausing the content item (e.g., 704). In some embodiments, the electronic device selects the user interface element in response to detecting, via one or more input devices (e.g., eye-tracking devices), that the user's attention (e.g., gaze) is directed toward the user interface element while detecting, via one or more input devices (e.g., eye-tracking devices), that a predetermined portion of the user (e.g., hands) is in a predetermined pose. For example, detecting a predetermined portion of the user in a predetermined pose includes detecting the user's hands in a pinching hand configuration, as described above. In some embodiments, in response to detecting an input directed at a user interface element that causes the electronic device to toggle between a play or pause state while the content item is being played, the electronic device pauses the content. In some embodiments, in response to detecting an input directed at a user interface element that causes the electronic device to toggle between a play or pause state while the content item is paused, the electronic device plays the content.
[0180] In some embodiments, in response to receiving the respective user input (822b), following a determination that the second respective user interface element is a user interface element (e.g., 712f, 712h of FIG. 7B ) that, when selected, causes the electronic device (e.g., 101) to update the playback position of the content item (e.g., 704) (e.g., a skip forward option or a skip back option), the electronic device (e.g., 101) updates (822d) the playback position of the content item (e.g., 704) according to the respective user input. In some embodiments, selection of the respective user interface element causes the electronic device to change the playback position of the content item at a rate different from the rate at which the playback position of the content item changes while playing the content item. In some embodiments, the respective user interface element for modifying the virtual lighting effect is displayed in a user interface associated with the content item that includes the first user interface element and the second user interface element. In some embodiments, the user interface further includes selectable options for accessing audio and subtitle settings for a content item, toggling picture-in-picture elements according to one or more steps of method 1000, toggling immersive content modes according to one or more steps of method 1400, and viewing the content item playback queue of the electronic device.
[0181] Displaying separate user interface elements within a user interface having user interface elements for switching between play or pause states and for updating the playback position of a content item provides an efficient way to facilitate modifying the playback of a content item and modifying the three-dimensional environment, which reduces the cognitive burden on the user when interacting with the content item and the three-dimensional environment.
[0182] In some embodiments, the electronic device (e.g., 101) receives (824a) a separate user input via one or more input devices (e.g., 314) directed to a second separate interface element (e.g., 712k in FIG. 7B) of the one or more user interface elements for modifying playback of the content item.
[0183] In some embodiments, in response to receiving the respective user input, in accordance with determining that the second respective user interface element is a user interface element (e.g., 712k of FIG. 7B ) that, when selected, causes the electronic device (e.g., 101) to modify the volume of the audio content of the content item (e.g., 704), the electronic device (e.g., 101) modifies the volume of the audio content in accordance with the respective input (824b). In some embodiments, in response to receiving the first input directed to the user interface element for modifying the volume of the audio content of the content item, the electronic device updates the user interface element to include a slider user interface element that, when the position of an indicator of the slider is changed, causes the electronic device to change the volume of the audio content in accordance with an updated position of the indicator. In some embodiments, the user interface element for modifying the volume of the audio content is displayed in a user interface associated with the content item along with a respective user interface element for modifying the virtual lighting effects.
[0184] Displaying separate user interface elements within a user interface having user interface elements for modifying the volume of audio content provides an efficient way to facilitate modifying the playback of content items and modifying the three-dimensional environment, which reduces the cognitive burden on the user when interacting with the content items and the three-dimensional environment.
[0185] In some embodiments, such as in FIG. 7B , a user interface (e.g., 711) associated with a content item (e.g., 704) is a separate user interface from the content item (e.g., 704) and is displayed (826) between the content item (e.g., 704) and a viewpoint of a user of an electronic device (e.g., 101) within the three-dimensional environment (e.g., 702) via a display generation component (e.g., 120). In some embodiments, the content item and the user interface are displayed in separate windows within the three-dimensional environment. In some embodiments, the user interface associated with the content item is partially overlaid on the content item. In some embodiments, the user interface associated with the content item is not overlaid on the content item.
[0186] Displaying a user interface associated with a content item separately from the content item in a three-dimensional environment between the user's viewpoint and the content item provides an efficient way to facilitate user interaction with the user interface and reduces the cognitive load and time required to modify playback of the content item via interaction with the user interface.
[0187] In some embodiments, a content item (e.g., 704 in FIG. 7B ) is displayed (828a) via a display generation component (e.g., 120) at a first angle relative to a user's viewpoint within the three-dimensional environment. In some embodiments, the first angle comprises a lateral angle within the three-dimensional environment (e.g., tilting toward the user's left or right). In some embodiments, the first angle comprises a vertical angle within the three-dimensional environment (e.g., tilting upward or downward from the user's viewpoint).
[0188] In some embodiments, a user interface (e.g., 711 in FIG. 7B ) associated with a content item (e.g., 704) is displayed (828b) via a display generation component (e.g., 120) at a second angle relative to a user's viewpoint within the three-dimensional environment (e.g., 702) that is different from the first angle. In some embodiments, the second angle comprises a lateral angle within the three-dimensional environment (e.g., tilting toward the user's left or right). In some embodiments, the second angle comprises a vertical angle within the three-dimensional environment (e.g., tilting upward or downward from the user's viewpoint). For example, the user interface associated with the content item is displayed in a location within the three-dimensional environment below the content item and at a more upward angle relative to the user's viewpoint than the angle at which the content item is displayed to the user.
[0189] Displaying content items and user interfaces associated with the content items at different angles relative to a user's viewpoint within a three-dimensional environment provides an efficient way of recognizably presenting content items and user interfaces associated with the content items when the content items and user interfaces associated with the content items are in different positions within the three-dimensional environment, which reduces the time and effort required to interact with the content items and user interfaces associated with the content items.
[0190] In some embodiments, such as in FIG. 7B , the electronic device (e.g., 101) displays (830) one or more user interface elements (e.g., 712f, 712g) for modifying playback of the content item (e.g., 704) in response to detecting, via one or more input devices (e.g., 314) (e.g., hand tracking devices), a pre-defined portion of the user of the electronic device (e.g., hand 703b) in a pose that meets one or more criteria. In some embodiments, detecting a pre-defined portion of the user's pose that meets one or more criteria includes detecting a movement of the user's hand from a position close to the user's body to an elevated location (e.g., within a pre-defined area of the three-dimensional environment). In some embodiments, detecting a pre-defined portion of the user's pose that meets one or more criteria includes detecting the user's hand in a pre-defined hand shape, such as the pointing hand shape described above, or a pre-pinch hand shape in which the thumb of the hand is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, or 3 centimeters) but is not touching another finger of the hand. In some embodiments, the one or more criteria include a criterion that is met when the electronic device detects via one or more input devices (e.g., eye tracking devices) that the user's attention (e.g., gaze) is directed toward the content item. In some embodiments, while the electronic device does not detect a predetermined portion of the user in a pose that meets the one or more criteria, the electronic device ceases displaying one or more user interface elements for modifying playback of the content item. In some embodiments, the electronic device continues to present (and play) the content item.
[0191] Displaying one or more user interface elements for modifying playback of a content item in response to detecting a predetermined portion of the user in a pause that meets one or more criteria provides an efficient way to selectively facilitate interaction with one or more user interface elements, which reduces the time and input required to modify playback of a content item.
[0192] In some embodiments, while displaying one or more user interface elements (e.g., 712f, 712g in FIG. 7B ) for modifying playback of a content item (e.g., 704) (e.g., in response to detecting a predefined portion of the user in a pose that satisfies one or more criteria), the electronic device (e.g., 101) detects (832a) via one or more input devices (e.g., 314) (e.g., hand tracking devices) a predefined portion of the user (e.g., hand 703a) in a pose that does not satisfy one or more criteria, such as in FIG. 7A . In some embodiments, the one or more input devices (e.g., hand tracking devices, eye tracking devices) detect that one or more of the above-described criteria are not satisfied. In some embodiments, the one or more input devices (e.g., hand tracking devices) do not detect a predefined portion of the user (e.g., hand) (e.g., because the predefined portion of the user is out of range of the one or more input devices (e.g., hand tracking devices)). In some embodiments, the electronic device detects a predefined portion of the user in a pose (e.g., shape and / or position) that does not satisfy one or more criteria. For example, the electronic device detects when a user drops their hand on their lap or at their side.
[0193] In some embodiments, in response to detecting a predefined portion (e.g., 703a) of the user pausing that does not meet one or more criteria, such as in FIG. 7A , the electronic device (e.g., 101) reduces (832b), via the display generation component (e.g., 120), the visual emphasis that the electronic device (e.g., 101) displays of one or more user interface elements (e.g., 712f in FIG. 7B ) for modifying playback of the content item. In some embodiments, the electronic device reduces the opacity of the one or more user interface elements for modifying playback of the content item. In some embodiments, the electronic device ceases displaying the one or more user interface elements for modifying playback of the content item. In some embodiments, the electronic device continues to present (and play) the content item.
[0194] Ceasing the display of one or more user interface elements for modifying playback of a content item in response to detecting a predetermined portion of the user in a pause that does not meet one or more criteria provides an efficient way of reducing distraction while the user is consuming a content item without indicating an intention to interact with one or more interactive elements, which reduces the cognitive burden on the user while consuming the content item.
[0195] In some embodiments, such as in FIG. 7B , while displaying a content item (e.g., 704) at a first size and a user interface (e.g., 711) associated with the content item (e.g., 704) at a second size via a display generating component (e.g., 120), the electronic device (e.g., 101) receives input (834a) via one or more input devices (e.g., 314) corresponding to a request to resize the content item (e.g., 704). In some embodiments, the input corresponding to the request to resize the content item is a request to change the virtual size of the content item within the three-dimensional environment, with or without changing the position of the content item within the three-dimensional environment. In some embodiments, the input corresponding to the request to resize the content item is a request to change the angular size (e.g., the portion of the display generating component occupied by the content item), with or without changing the position and / or virtual size of the content item within the three-dimensional environment.
[0196] 7C , in response to receiving an input corresponding to a request to resize the content item (e.g., 704) (834b), the electronic device (e.g., 101), via the display generation component (e.g., 120), displays the content item (e.g., 704) at a third size different from the first size in accordance with the input corresponding to the request to resize the content item (e.g., 704). In some embodiments, the electronic device resizes the content item in a direction corresponding to a direction of movement of a predetermined part of the user (e.g., a hand) and by an amount corresponding to a magnitude (e.g., speed, duration, distance, etc.) of the movement of the predetermined part of the user.
[0197] In some embodiments, such as in FIG. 7C , in response to receiving input corresponding to a request to resize the content item (834b), the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (834d) a user interface (e.g., 711) associated with the content item (e.g., 704) at a second size. In some embodiments, the (e.g., angular) size of the user interface associated with the content item remains constant as the content item is resized. In some embodiments, if the input is a request to resize the content item without changing the position of the content item, the electronic device maintains the angular size and virtual size of the user interface associated with the content item and also maintains the position of the user interface within the three-dimensional environment. In some embodiments, if the input is a request to change the size and position of the content item within the three-dimensional environment, the electronic device maintains the angular size of the user interface and updates the virtual size of the user interface according to the updated position of the user interface to maintain the angular size of the user interface within the three-dimensional environment.
[0198] Maintaining the size of a user interface associated with a content item when updating the size of the content item provides an efficient way to maintain the readability of the user interface in a three-dimensional environment, which reduces the cognitive burden on a user when interacting with the user interface.
[0199] 7C , while displaying, via a display generation component (e.g., 120), a content item (e.g., 704) at a first (e.g., angular) size and a first distance from a user's viewpoint within the three-dimensional environment (e.g., 702) and a user interface (e.g., 711) associated with the content item (e.g., 704) at a second (e.g., angular) size and a second distance from the user's viewpoint within the three-dimensional environment (e.g., 702), the electronic device (e.g., 101) receives, via one or more input devices (e.g., 314), input corresponding to a request to reposition the content item (e.g., 704) (e.g., and the user interface) within the three-dimensional environment (e.g., 702) (836a). In some embodiments, the electronic device repositions the content item and the user interface according to movements of a predetermined part (e.g., hand) of the user while the user is providing the input. For example, the electronic device, while providing the input, moves the content items and the user interface in a direction and amount corresponding to the direction and amount (e.g., speed, distance, duration, etc.) of movement of a predetermined part (e.g., hand) of the user. In some embodiments, detecting the input includes detecting the user's gaze directed toward the content items or a user interface element for repositioning the content items while detecting the user's hand in a predetermined hand shape, such as a pinch hand shape or a pointing hand shape.
[0200] In some embodiments, such as in FIG. 7D , in response to receiving an input corresponding to a request to reposition the content item (e.g., 704) (836b), the electronic device (e.g., 101), via the display generating component (e.g., 120), displays the content item (836c) at a third distance from the user's viewpoint within the three-dimensional environment (e.g., 702) different from the first distance and at a third (e.g., angular) size in accordance with the input corresponding to the request to reposition the content item (e.g., 704). In some embodiments, the electronic device maintains a virtual size of the content item in response to a request to move the content item within the three-dimensional environment, such that the electronic device updates the angular size of the content item (e.g., the portion of the display generating component occupied by the content item) according to a change in the distance between the content item and the user's viewpoint. For example, in response to a request to move the content item farther away from the user, the electronic device displays the content item at a smaller angular size, and in response to a request to move the content item closer to the user, the electronic device displays the content item at a larger angular size. In some embodiments, the electronic device updates the virtual size of the content item at a rate that is less than the change in the angular size of the content item (e.g., to keep the display of the content item within the maximum and minimum sizes if updating the angular size of the content item without changing the virtual size of the content item would cause the angular size of the content item to be outside the maximum or minimum angular size).
[0201] 7D , in response to receiving input corresponding to a request to reposition the content item (e.g., 704) (836b), the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (836d) a user interface (e.g., 711) associated with the content item (e.g., 704) at a second (e.g., angular) size and a fourth distance from the user's viewpoint in accordance with the input corresponding to the request to reposition the content item (e.g., 704). In some embodiments, the electronic device updates the virtual size of the user interface in accordance with updating the distance between the user interface and the user's viewpoint within the three-dimensional environment while maintaining the angular size of the user interface.
[0202] Maintaining the size of the user interface associated with a content item provides an efficient way to maintain the readability of the user interface in a three-dimensional environment, which reduces the cognitive burden on the user when interacting with the user interface.
[0203] In some embodiments, such as in Figure 7B, a content item (e.g., 704) is separate (838a) from a user interface (e.g., 711) associated with the content item (e.g., 704) within the three-dimensional environment (e.g., 702). In some embodiments, the content item and the user interface are displayed in separate containers (e.g., platters, user interface elements, windows, etc.) in separate locations within the three-dimensional environment.
[0204] In some embodiments, such as in FIG. 7B , the electronic device (e.g., 101), via the display generation component (e.g., 120), displays (838b) one or more second user interface elements (e.g., 712a) for modifying playback of the content item (e.g., 704), where the one or more second user interface elements (e.g., 712a) are displayed overlaid on the content item (e.g., 704) in the three-dimensional environment (e.g., 702). In some embodiments, the one or more second user interface elements are displayed within a container that includes the content item. In some embodiments, the one or more second user interface elements visually (e.g., from the user's perspective) or spatially overlay the content item in the three-dimensional environment. For example, the electronic device, via the display generation component, displays a user interface element that, when selected, causes the electronic device to cease displaying the content item overlaid on the content item within the content item's container and display one or more other user interface elements (e.g., for modifying playback of the content item) within a user interface within a container separate from the content item.
[0205] Displaying one or more second user interface elements overlaid on a content item within a three-dimensional environment provides an efficient way to interact with the one or more second user interface elements while viewing the content item, thereby reducing the cognitive burden on the user.
[0206] In some embodiments, such as in FIG. 7C , the electronic device (e.g., 101) detects (840a) via one or more input devices (e.g., 314) (e.g., eye-tracking devices) that a user's attention (e.g., gaze 713c, 713d) directed to an individual user interface element (e.g., 712f, 712j) of the one or more user interface elements satisfies one or more first criteria. In some embodiments, the one or more first criteria include a criterion that is met when the user's attention (e.g., gaze) is directed to the individual user interface element for at least a threshold period (e.g., 0.1, 0.2, 0.3, 0.4, 0.5, 1, 2, or 3 seconds). In some embodiments, the one or more criteria are met the instant the user's attention (e.g., gaze) is directed to the individual user interface element.
[0207] In some embodiments, such as FIG. 7C , in response to detecting 840b that a user's attention (e.g., gaze 713d) directed to an individual user interface element (e.g., 712j) satisfies one or more first criteria, and in accordance with determining that the individual user interface element (e.g., 712j) satisfies one or more second criteria, the electronic device (e.g., 101), via the display generation component (e.g., 120), displays 840c a visual indication (e.g., 714) identifying a function of the individual user interface element (e.g., 712j). In some embodiments, an individual user interface element satisfies the one or more second criteria when the individual user interface element is included in a predetermined subset of one or more user interface elements included in a user interface associated with a content item. For example, one or more of an immersive content option according to one or more steps of method 1400, a picture-in-picture option according to one or more steps of method 1000, and an individual user interface element for modifying virtual lighting effects are included in the predetermined subset. In some embodiments, a distinct user interface element satisfies one or more second criteria if the distinct user interface element is associated with a function related to the presentation of a content item in a three-dimensional environment (e.g., it may not be commonly associated with the presentation of content items in another environment). In some embodiments, the visual indication identifying the function of the distinct user interface element includes text describing the function of the distinct user interface element. For example, a visual indication identifying the function of a distinct option for modifying virtual lighting effects includes text that says "lighting effects," etc. In some embodiments, the electronic device ceases displaying the visual indication in response to detecting the user's attention being directed away from the distinct user interface element. In some embodiments, the electronic device displays the visual indication in response to detecting a user's ready state while the user's gaze is directed at the distinct user interface element via one or more input devices.
[0208] In some embodiments, such as in FIG. 7C , in response to detecting (840b) that a user's attention (e.g., gaze 713c) directed to an individual user interface element (e.g., 712f) satisfies one or more first criteria, and in accordance with determining that the individual user interface element (e.g., 712f) does not satisfy one or more second criteria, the electronic device (e.g., 101) ceases displaying (840d) a visual indication identifying the function of the individual user interface element (e.g., 712f). In some embodiments, an individual user interface element does not satisfy one or more second criteria when the individual user interface element is not included in a predetermined subset of one or more user interface elements included in the user interface associated with the content item. For example, one or more of a playback queue option, an option to skip back or head at the playback position of the content item, an option to play / pause the content item, an option to view subtitle options for the content item, and an option to view audio options for the content item are not included in the predetermined subset of one or more user interface elements. In some embodiments, an individual user interface element does not satisfy one or more second criteria when the individual user interface element is associated with functionality related to the presentation of content items in general (e.g., it may not be specifically associated with the presentation of content items in a three-dimensional environment).
[0209] Displaying a visual indication identifying the functionality of individual user interface elements that meet one or more second criteria provides an efficient way to show a user the actions that will be performed in response to further input directed to the individual user interface elements, which reduces user errors and the number of inputs required to correct the user errors.
[0210] In some embodiments, such as in FIG. 7B , the electronic device (e.g., 101), via a display generation component (e.g., 120), displays (842a) a content item (e.g., 711) and a separate user interface element (e.g., 712b) displayed separately from the user interface (e.g., 711) associated with the content item (e.g., 704). In some embodiments, the content item and the user interface are displayed in separate containers (e.g., platters, user interface elements, windows, etc.) in separate locations within the three-dimensional environment, and the separate user interface elements are displayed outside of these containers in locations within the three-dimensional environment that are different from the locations of the content item and the user interface. In some embodiments, the electronic device, optionally while detecting the user's gaze directed toward the content item, displays the separate user interface element, optionally along with one or more selectable elements for modifying playback of the content item, in response to detecting a predefined portion of the user in a pose that satisfies one or more criteria. For example, in response to detecting a predetermined portion of the user in a pose that satisfies one or more criteria while detecting the user's gaze, optionally directed at the content item, the electronic device displays a respective user interface element and a plurality of selectable elements for modifying playback of the content item.
[0211] In some embodiments, such as FIG. 7B, while displaying an individual user interface element (e.g., 712b), the electronic device (e.g., 101) receives (842b) input directed to the individual user interface element (e.g., 712b) via one or more input devices (e.g., 314).
[0212] In some embodiments, such as in FIG. 7B , in response to detecting input directed at an individual user interface element (e.g., 712b), the electronic device (e.g., 101) initiates a process (842c) to resize a content item (e.g., 704) within the three-dimensional environment (e.g., 702) according to the input directed at the individual user interface element (e.g., 712b). In some embodiments, in response to detecting a selection of an individual user interface element, the electronic device initiates a process to resize the individual user interface element according to a user's predetermined portion (e.g., hand) movement after the selection of the individual user interface element. In some embodiments, in response to detecting a user's predetermined portion (e.g., hand) movement after the selection of the individual user interface element, the electronic device resizes the content item according to the user's predetermined portion (e.g., hand) movement without resizing individual user interface elements, including one or more selectable elements for modifying playback of the content item, as described above.
[0213] Displaying separate user interface elements for resizing content items separately from the content items and the user interfaces associated with the content items provides an efficient way to view the size of a content item while providing input for resizing the content item, thereby improving the ergonomics of the resize input and reducing the time and number of inputs required to resize a content item to a desired size.
[0214] 9A-9E illustrate an exemplary method for displaying media content in a three-dimensional environment.
[0215] FIG. 9A illustrates a three-dimensional environment 904 being displayed by display generating components 120 of electronic device 101 and an overhead view 920 of three-dimensional environment 904. As described above with reference to FIGS. 1-6 , electronic device 101 optionally includes display generating components (e.g., touch screen 120) and multiple image sensors (e.g., image sensor 314 of FIG. 3 ). The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor, and electronic device 101 can be used to capture one or more images of a user or portions of a user while the user interacts with electronic device 101. In some embodiments, the user interfaces illustrated below may also be realized on a head-mounted display that includes display generating components that display the user interface to the user and sensors for detecting the physical environment, the movement of the user's hands (e.g., external sensors facing outward from the user), and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).
[0216] 9A , electronic device 101 is displaying a three-dimensional environment 904 that includes a user interface for application 1 906, a user interface for application 2 910, a user interface for application 3 912, and a representation of a table 918, which is a physical table in the physical environment of device 101. In some embodiments, as described in more detail below, applications 1-3 are, optionally, media applications, gaming applications, social applications, navigation applications, streaming applications, etc. In some embodiments, representation 918 and user interfaces 906, 910, and 912 are displayed by electronic device 101 because these objects are within the field of view of a current viewpoint of user 922 of three-dimensional environment 904. For example, as shown in overhead view 920 of FIG. 9A , user 922's current viewpoint of three-dimensional environment 904 corresponds to a center position in three-dimensional environment 904 and is oriented toward the top / rear of three-dimensional environment 904. For ease of explanation in the remainder of this disclosure, the position / pose of user 922 within three-dimensional environment 904 will be referred to herein as the user's 922 current viewpoint of three-dimensional environment 904, or more simply, the viewpoint of user 922 as shown in overhead view 920.
[0217] 9A , electronic device 101, via display generation component 120, is displaying representation 918 and user interfaces 906, 910, and 912 because these objects are within the field of view of user 922's current viewpoint of three-dimensional environment 904 (as shown in overhead view 920). Conversely, electronic device 101, via display generation component 120, is not displaying sofa representation 924, corner table representation 930, coffee table representation 932, and user interfaces 926 and 928 because these objects are not within the field of view of user 922's current viewpoint of three-dimensional environment 904, as shown in overhead view 920.
[0218] In some embodiments, the viewpoint of user 922 in three-dimensional environment 904 corresponds to the physical location of user 922 in physical environment 902 (e.g., operating environment 100) of electronic device 101. For example, user 922's viewpoint is, optionally, the viewpoint shown in overhead view 920 because user 922 is currently oriented toward a back wall within physical environment 902 and is located at the center of physical environment 902 while holding electronic device 101 (e.g., or wearing device 101 if device 101 is a head-mounted device).
[0219] 9A , electronic device 101 is currently playing TV program A in user interface 906. In some embodiments, TV program A is playing in user interface 906 in response to electronic device detecting a request to start playing TV program A. In some embodiments, as described in more detail below, electronic device 101 can present TV program A in different presentation modes, including a picture-in-picture presentation mode and / or an extended presentation mode (e.g., different from picture-in-picture presentation mode). In the example of FIG. 9A , electronic device 101 is currently presenting TV program A in extended presentation mode in user interface 906. It should be understood that user interface 906 may also be displayed at other locations relative to the user's viewpoint of three-dimensional environment 904, such as other locations within three-dimensional environment 904 that are within the field of view of the user's current viewpoint of three-dimensional environment 904, while TV program A is presented in extended presentation mode.
[0220] In Figure 9B, while the electronic device 101 is playing TV program A in an augmented presentation mode, the electronic device 101 detects that the viewpoint of the user 922 in the three-dimensional environment 904 has moved from the viewpoint shown in Figure 9A to the viewpoint shown in Figure 9B. In some embodiments, the user's viewpoint in the three-dimensional environment 904 has moved to the viewpoint shown in Figure 9B because the user 922 has moved to a corresponding pose and / or position in the physical environment 902. As shown in Figure 9B, in response to the electronic device 101 detecting the movement of the user's 922 viewpoint in the three-dimensional environment 904, the electronic device 101 displays the three-dimensional environment 904 from the user's new viewpoint of the three-dimensional environment 904, which is shown in the overhead view 920 of Figure 9B.
[0221] Specifically, as a result of the movement of user 922's viewpoint from the viewpoint shown in Figure 9A to the viewpoint shown in Figure 9B, electronic device 101 is no longer presenting user interface 912 of application 3 as previously shown in Figure 9A because user interface 912 is no longer in the field of view from the user's current viewpoint of three-dimensional environment 904 (as shown in overhead view 920 of Figure 9B). Additionally, as a result of the movement of user 922's viewpoint, user 922's viewpoint of three-dimensional environment 904 has moved to the right from the viewpoint of user 922 shown in Figure 9A, and therefore electronic device 101 displays representation 918 of the table, user interface 910, and user interface 906 at locations in the three-dimensional environment that are further to the left of the user's field of view compared to Figure 9A.
[0222] In some embodiments, when media content is presented in an augmented presentation mode, the location of the user interface presenting the media content does not change within the three dimensional environment 904 as the user's viewpoint of the three dimensional environment 904 moves. For example, when the user's viewpoint of the three dimensional environment 904 moved from the viewpoint shown in Figure 9A to the viewpoint shown in Figure 9B, the location of the user interface 906 within the three dimensional environment 904 did not change (as shown in the overhead view 920 of Figures 9A and 9B) because the user interface 906 was presenting TV Program A in an augmented presentation mode.
[0223] 9B , electronic device 101 also presents playback control user interface 908. In some embodiments, electronic device 101 displays playback control user interface 908 in response to electronic device 101 detecting that hand 916 of user 922 is in a “pointing” pose (e.g., one or more fingers of hand 1331 are outstretched and one or more fingers of hand 916 are curled toward the palm of hand 916) or a “pre-pinch” pose (e.g., the thumb of hand 916 is within a threshold distance (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, or 5 centimeters) of, but not touching, another finger of hand 916) while the user's gaze is optionally directed at media user interface 906. In some embodiments, user interface elements 908a-908j displayed in playback control user interface 908 are similar to selectable user interface options 712c-L described above in the series of FIGS. 7 .
[0224] 9B , while electronic device 101 is playing TV program A in enhanced presentation mode, electronic device 101 detects a request to start playing TV program A in picture-in-picture presentation mode. In some embodiments, electronic device 101 detected the request to start playing TV program A in picture-in-picture presentation mode because user's hand 916 was in a "pointing" or "pinching" pose (e.g., the thumb and index finger of hand 916 converge to a threshold distance from each other (e.g., 0.2, 0.5, 1, 1.5, 2, or 2.5 centimeters)) while user's gaze 914 was directed at user interface element 908b.
[0225] 9C , in response to detecting a request to begin presenting TV program A in the picture-in-picture presentation mode of FIG. 9B , electronic device 101 stops playing TV program A in user interface 906 and begins presenting TV program A in picture-in-picture user interface 934. In some embodiments, as playback of TV program A is transitioning from media user interface 906 to picture-in-picture user interface 934, electronic device 101 displays an animation of the picture-in-picture user interface fading in and playback of TV program A fading out in user interface 906.
[0226] As shown in the overhead view 920, the picture-in-picture user interface 934 is displayed at a location within the three-dimensional environment 904 that is at a position within the three-dimensional environment 904 that is in front of and to the right of the user's current viewpoint of the three-dimensional environment 904. In some embodiments, the electronic device 101 is displaying the picture-in-picture user interface 934 at the location shown in the overhead view 920 because the location within the three-dimensional environment 904 is within a threshold distance (e.g., 0.1, 0.2, 0.5, 1, 1.5, or 3 feet) from the user's current viewpoint of the three-dimensional environment 904 and / or occupies a predetermined portion (e.g., bottom right, bottom left, top right, top left) of the field of view from the user's 922 viewpoint of the three-dimensional environment 904. Thus, in some embodiments, the user interface 934 is displayed at a location within the three-dimensional environment 904 that is based on the viewpoint of the user 922, and the user interface 906 is displayed at a location within the three-dimensional environment 904 that is not based on the viewpoint of the user 922.
[0227] In some embodiments, while electronic device 101 is presenting media content in picture-in-picture presentation mode, electronic device 101 optionally displays one or more representations of media items selectable for playback. For example, in Figure 9C, in response to electronic device 101 receiving a request to transition the presentation of TV program A from enhanced presentation mode to picture-in-picture presentation mode in Figure 9B, electronic device 101 updates user interface 906 (the user interface that previously presented TV program A during enhanced presentation mode) to include multiple representations 940-958 of individual media content. Multiple representations 940-958 are optionally selectable such that, upon electronic device 101 detecting a selection of one of representations 940-958, the media item corresponding to the selected representation begins playing in user interface 906 (optionally without ceasing playback of TV program A in media user interface 934) and / or in picture-in-picture user interface 934.
[0228] In some embodiments, multiple representations 940-958 of individual media content are displayed in one or more groups (e.g., columns) within user interface 906. For example, in FIG. 9C , representations 940-946 are displayed in a first column of user interface 906 because the corresponding media items were selected for display based on the content consumption history of user 922. Similarly, representations 948-958 are displayed in a second column of user interface 906 because the corresponding media items correspond to popular / currently trending content items (e.g., more users have recently viewed the media content corresponding to representations 948-958 over the past hour, day, week, month, etc.).
[0229] In some embodiments, electronic device 101 updates the type / category of media content displayed in user interface 906. For example, in Figure 9C, electronic device 101 is displaying user interface 936 (which is also optionally displayed in response to electronic device 101 receiving a request to begin presenting TV program A in picture-in-picture presentation mode, as described above). User interface elements 936 include selectable option 936a that, when selected, causes user interface 906 to display representations of currently popular and / or recommended media content based on user 922's content consumption history (as shown in user interface 906 of FIG. 9C ); selectable option 936b that, when selected, causes electronic device 101 to display one or more representations of media content corresponding to one or more TV programs within user interface 906; selectable option 936c that, when selected, causes electronic device 101 to display one or more representations of media content corresponding to one or more movies within user interface 906; selectable option 936d that, when selected, causes electronic device 101 to display one or more representations of media content corresponding to one or more (e.g., live) sports games within user interface 906; and selectable option 936e that, when selected, causes electronic device 101 to present a user interface for searching for specific media content within user interface 906.
[0230] In Figure 9D, electronic device 101 has detected that the viewpoint of user 922 has moved from the viewpoint shown in Figure 9C to the viewpoint shown in Figure 9D. In some embodiments, the viewpoint of user 922 in three-dimensional environment 904 has moved to the viewpoint shown in Figure 9D because user 922 has moved to a corresponding pose and / or location in physical environment 902. As shown in Figure 9D, in response to detecting that the viewpoint of user 922 in three-dimensional environment 904 has moved to the viewpoint shown in Figure 9D, electronic device 101 displays three-dimensional environment 904 from the user's new viewpoint of three-dimensional environment 904. Specifically, display generation component 120 of device 101 is now displaying user interfaces 926 and 928 and representations 924 and 932 because these elements are now within the field of view from the user's viewpoint shown in Figure 9D.
[0231] In some embodiments, as the viewpoint of the user 922 of the three-dimensional environment 904 moves, the electronic device 101 updates the location of the picture-in-picture user interface 934 based on the new viewpoint of the user 922 of the three-dimensional environment 904. For example, as shown in the overhead view 920 of Figures 9C and 9D, in response to the electronic device 101 detecting that the viewpoint of the user of the three-dimensional environment 904 has moved from the viewpoint shown in Figure 9C to the viewpoint shown in Figure 9D, the electronic device 101 moves the location of the picture-in-picture user interface 934 from the location shown in Figure 9C to the location shown in Figure 9D. In some embodiments, the electronic device 101 moved the picture-in-picture user interface 934 from its location in the three-dimensional environment 904 shown in the overhead view 920 of FIG. 9C because, based on the user's current viewpoint of the three-dimensional environment 904, the location of the user interface 934 shown in the overhead view 920 is no longer within a threshold distance (e.g., 0.1, 0.2, 0.5, 1, 1.5, or 3 feet) of the user's new viewpoint of the three-dimensional environment 904 and / or because its location is no longer in a predetermined (e.g., bottom-right) portion of the user's field of view from the user's new viewpoint of the three-dimensional environment 904. In some embodiments, the picture-in-picture user interface 934 is optionally displayed at a location in the three-dimensional environment 904 shown in the overhead view 920 for reasons similar to those described with reference to FIG. 9C . Additionally, as shown in the overhead view 920, the location of the user interface 906 in the three-dimensional environment 904 has not changed for reasons similar to those previously described with reference to FIG. 9B .
[0232] In some embodiments, while electronic device 101 is presenting content in picture-in-picture presentation mode, playback controls are displayed overlaid or integrated on the picture-in-picture user interface. For example, because user interface 934 is currently presenting TV program A in picture-in-picture presentation mode, electronic device 101 is displaying user interface elements 936-948 overlaid on user interface 934 (as opposed to when playback controls were presented in a separate user interface while TV program A was presented in enhanced presentation mode, as described with reference to FIG. 9B ). In some embodiments, as shown in FIG. 9D , user interface elements 936-948 are displayed in media user interface 934 when electronic device detects that hand 916 is in a “pre-pinch” pose and, optionally, when user 922's gaze is directed toward user interface 934. If electronic device 101 does not detect that hand 916 is in a “pre-pinch” pose, user interface elements 936-948 are optionally not displayed.
[0233] Next, functions associated with user interface elements 936-948 will be described. User interface element 936 is optionally selectable and, when selected, causes electronic device 101 to stop presenting TV program A in picture-in-picture presentation mode and begin presenting TV program A in enhanced presentation mode. User interface element 938 is optionally selectable and, when selected, causes electronic device 101 to stop playing TV program A (and optionally stop displaying user interface 934). User interface element 940 is optionally selectable and, when selected, causes electronic device 101 to fast forward TV program A by a predetermined amount (e.g., 10, 15, 20, 30, 40, or 60 seconds). User interface element 942 is optionally selectable and, when selected, causes electronic device 101 to rewind TV program A by a predetermined amount (e.g., 10, 15, 20, 30, 40, or 60 seconds). User interface element 944 is optionally selectable and, when selected, causes electronic device 101 to pause playback of TV program A (e.g., if TV program A is currently playing) or start playback of TV program A (e.g., if TV program A is currently paused). Finally, media user interface 934 includes scrub bar 946 that includes an indication 948 that shows the current playback position of TV program A. Further details of scrub bar 908j and operations associated with scrub bar 908j are described with reference to method 1400 and Figures 13A-13E.
[0234] Additionally, as shown in Figure 9D, while displaying user interface elements 936-946 and while TV program A is playing in picture-in-picture presentation mode, the electronic device receives a request to present TV program A in an enhanced presentation mode (indicated by selection of user interface 936). In some embodiments, the input for selecting user interface element 936 is similar to the input for selecting user interface element 918b in Figure 9B. In some embodiments, in response to receiving the request to begin presenting TV program A in an enhanced presentation mode, electronic device 101 ceases presenting TV program A in media user interface 934 and begins displaying TV program A in the enhanced presentation mode in user interface 906 shown in Figure 9C (and, optionally, ceases displaying representations 940-958 in user interface 906 and user interface 934 shown in Figure 9C).
[0235] In some embodiments, when electronic device 101 receives a request to transition the presentation mode of TV program A from picture-in-picture presentation mode to an augmented presentation mode, if user interface 906 (e.g., a user interface presenting TV program A in the augmented presentation mode) is not within the field of view of the user's current viewpoint of three-dimensional environment 904, electronic device 101 updates the location of user interface 906 to be within the field of view of the user's current viewpoint of three-dimensional environment 904, as shown in Figure 9E. Conversely, if electronic device 101 receives a request to transition TV program A from being presented in picture-in-picture presentation mode to the augmented presentation mode while user interface 906 was in a location in three-dimensional environment 904 that was currently within the user's field of view (as in Figure 9C), electronic device 101 will, optionally, not have updated the location of user interface 906 within three-dimensional environment 904.
[0236] In some embodiments, when playback of a media item in the enhanced presentation mode has finished, electronic device 101 displays one or more representations of one or more suggested media items to view next. For example, in FIG. 9E , electronic device 101 detects that playback of TV program A in user interface 906 has finished or that a distinct position in the playback (e.g., 0.25, 0.5, 1, 2, 3, or 5 minutes from the end of playback) has been reached. In response, electronic device 101 displays user interface 909 including representations 946-950 of the respective media items. In some embodiments, the media items corresponding to representations 946-950 were selected for display in user interface 909 based on the content consumption history of user 922. In some embodiments, representations 946-950 are selectable, and when selected, cause electronic device 101 to play the corresponding media item in user interface 906. For example, electronic device 101 optionally begins playing media item C in user interface 906 when electronic device 101 detects that user 916's hand is in a "pointing" or "pinching" pose (as described above) while user 922's gaze 914 is directed at representation 950.
[0237] Additionally or alternatively, media items corresponding to representations 946-950 are optionally selected for playback based on a user's gaze 914 (without detecting input from the user's hand). For example, as shown in FIG. 9E , user 922's gaze 914 is currently directed at representation 946 of item A. In some embodiments, electronic device 101 begins playback of item A in user interface 906 when electronic device detects that user 922's gaze 914 is directed at representation 914. Alternatively, in some embodiments, electronic device 101 begins playback of media item A only when user's gaze 914 has been directed at representation 946 for at least a threshold amount of time (e.g., 15, 30, 60, 90, or 200 seconds). For example, in FIG. 9E , electronic device 101 has not begun playback of media item A in user interface 906 because user 922's gaze 914 has not been directed at representation 946 for the aforementioned threshold amount of time.
[0238] In some embodiments, the electronic device 101 displays an indication 915 that indicates the amount of time remaining until the user's 922 gaze 914 causes the electronic device 101 to begin playing the media item. For example, in FIG. 9E , the electronic device 101 displays a circular visual indication 915 within the representation 946. In some embodiments, as the user's gaze 914 remains directed toward the representation 946, the electronic device 101 updates (e.g., in real time) the visual indication 915 to occupy a space ranging from 0 degrees (e.g., when the user's gaze 914 is not directed toward the representation 914) to an angular distance of 360 degrees (e.g., when the user's gaze 914 has been directed toward the representation 914 for the threshold amount of time described above). For example, in FIG. 9E , the visual indication 915 indicates that the user's gaze 914 has currently been directed toward the representation 946 for half of the threshold amount of time described above, because the visual indication 915 occupies an angular distance of 180 degrees.
[0239] Additional or alternative details regarding the embodiment illustrated in FIGS. 9A-9E are provided below in the description of method 1000 described with reference to FIGS. 10A-10I.
[0240] 10A-10I are flowcharts illustrating a method for displaying media content in a three-dimensional environment, according to some embodiments. In some embodiments, method 1000 is performed in a computer system (e.g., computer system 101 of FIG. 1 ) that includes display generation components (e.g., display generation components 120 of FIGS. 1, 3, and 4 ) (e.g., a heads-up display, a display, a touchscreen, a projector, etc.) and one or more cameras (e.g., a camera pointing down the user's hand (e.g., color sensors, infrared sensors, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 1000 is governed by instructions stored on a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control unit 110 of FIG. 1A ). Some operations of method 1000 are optionally combined and / or the order of some operations is optionally changed.
[0241] In some embodiments, method 1000 is performed in an electronic device in communication with a display generation component and one or more input devices (e.g., a mobile device (e.g., a tablet, smartphone, media player, or wearable device), or a computer). In some embodiments, the display generation component is a display integrated with the electronic device (optionally a touchscreen display), an external display such as a monitor, projector, television, or a hardware component (optionally built-in or external) for projecting a user interface and making the user interface visible to one or more users, etc. In some embodiments, the one or more input devices include electronic devices or components capable of receiving user input (e.g., capturing user input, detecting user input, etc.) and transmitting information related to the user input to the electronic device. Examples of input devices include a touchscreen, a mouse (e.g., external), a trackpad (optionally integrated or external), a touchpad (optionally integrated or external), a remote control device (e.g., external), another mobile device (e.g., separate from the electronic device), a handheld device (e.g., external), a controller (e.g., external), a camera, a depth sensor, an eye tracking device, and / or a motion sensor (e.g., hand tracking device, hand motion sensor), etc. In some embodiments, the electronic device is in communication with a hand tracking device (e.g., one or more cameras, depth sensors, proximity sensors, touch sensors (e.g., touchscreen, trackpad). In some embodiments, the hand tracking device is a wearable device such as a smart glove. In some embodiments, the hand tracking device is a handheld input device such as a remote control or a stylus.
[0242] In some embodiments, an electronic device (e.g., device 101 of FIGS. 9A-9E ) presents content (e.g., media) via a display generation component and displays (1002 a) a three-dimensional environment (e.g., the three-dimensional environment is a computer-generated reality (XR) environment such as a virtual reality (VR) environment, a mixed reality (MR) environment, or an augmented reality (AR) environment) including a first media user interface (e.g., a user interface such as those described with reference to methods 800, 1200, and / or 1400) located at a first discrete location within the three-dimensional environment. For example, in FIG. 9A , electronic device 101 displays three-dimensional environment 904 including “TV Program A” presented in user interface 906. In some embodiments, while the first media user interface is displayed in the three-dimensional environment, the first media user interface presents a movie, a TV show, a music video, and / or other type of video or audio content. In some embodiments, the first media user interface is located at a first distinct location within the three-dimensional environment because the first distinct location is a default launch location of an application associated with the first media user interface (e.g., in response to launching a distinct application, the first media user interface is displayed at the first distinct location within the three-dimensional environment). In some embodiments, the first media user interface is located at a first distinct location within the three-dimensional environment because a user of the electronic device moved the first media user interface to the first distinct location. In some embodiments, when the three-dimensional environment is displayed while a user's viewpoint is at a first viewpoint (e.g., while the electronic device is oriented at a first area within the physical environment), the three-dimensional environment optionally includes representations of some or all of objects located at the first location within the physical environment and / or representations of virtual objects (e.g., objects that are not in the physical environment but are displayed because the user's viewpoint corresponds to the first viewpoint).In other words, different viewpoints of a user of the electronic device optionally cause representations of different virtual and / or physical objects to be present in the user's field of view while viewing the three-dimensional environment from the respective viewpoints. For example, a first location in the physical environment may include one or more physical objects, such as a chair, a sofa, a table, etc., and the three-dimensional environment may include representations of those one or more chairs, sofas, tables, etc. Similarly, a first media user interface is optionally at a location in the three-dimensional environment such that when the user's viewpoint corresponds to the first viewpoint, the first media user interface is in the user's field of view from the first viewpoint.
[0243] In some embodiments, the electronic device displays the three-dimensional environment with the first media user interface at a first discrete location within the three-dimensional environment having a pose (e.g., position and / or orientation) that is within a discrete range of poses relative to a first perspective of a user of the electronic device (e.g., in some embodiments, the discrete range of poses relative to the user's first perspective includes all (or a subset thereof) of the poses in the three-dimensional environment that are within the user's field of view of the three-dimensional environment from the user's first perspective. In some embodiments, a position of the first media user interface within the three-dimensional environment is not within the discrete range of poses relative to the first perspective if the position is not within the user's field of view from the first perspective). The electronic device then detects (1002b) a movement of the user's viewpoint within the three-dimensional environment from the first perspective to a second perspective that is different from the first perspective (e.g., the electronic device detects that the user of the electronic device has begun to look at a different location within the physical environment (e.g., the orientation of the user's field of view into the three-dimensional environment has changed)). 9B , electronic device 101 detects that the user's viewpoint of three-dimensional environment 904 has changed from the viewpoint shown in overhead view 920 of FIG. 9A to the viewpoint shown in overhead view 920 of FIG. 9B . In some embodiments, a portion of the three-dimensional environment that was within the user's field of view from the user's first perspective, optionally, remains within the user's field of view from the second perspective (e.g., at least a portion of the three-dimensional environment displayed via the display generation components while the three-dimensional environment was presented from the first perspective, is displayed via the display generation components while the three-dimensional environment is presented from the second perspective). Alternatively, in some embodiments, an area / portion of the three-dimensional environment that was within the user's field of view from the first perspective, while the three-dimensional environment is presented from the user's second perspective, is optionally no longer within the user's field of view from the second perspective. In some embodiments, the user's viewpoint changes as the user moves (e.g., walks, runs, etc.) within the physical environment and / or looks toward different areas within the physical environment (e.g., while remaining stationary).
[0244] In some embodiments, in response to detecting a movement of the user's viewpoint from the first viewpoint to the second viewpoint, the electronic device, via the display generation component, displays (1002c) the three-dimensional environment from the second viewpoint (e.g., electronically updates the display of the three-dimensional environment to correspond to the user's new viewpoint of the three-dimensional environment—the second viewpoint). In some embodiments, the display of the three-dimensional environment from the second viewpoint is similar to the display of the three-dimensional environment from the user's first viewpoint, as described above).
[0245] In some embodiments, following a determination that the content is being presented in a first presentation mode (e.g., in some embodiments, the content is being presented in a first presentation mode when the content is not being presented in a picture-in-picture (PiP) UI; in some embodiments, the content is being presented in a first presentation mode when the content is being presented in a default presentation mode (e.g., playing natively in a video player application associated with the first media user interface). In some embodiments, the content is being presented in a first presentation mode when the content is being presented in a video aspect ratio higher than 9:16 or 16:9, and the electronic device 101 maintains (1002d) the first media user interface at a first discrete location within the three-dimensional environment, and the first media user interface is no longer within a discrete range of pose relative to the user's second perspective. For example, when the perspective of the user 922 of the three-dimensional environment 904 is different from that of FIGS. 9A and 9B , 9A and 9B ), the location of the user interface 906 in the three-dimensional environment 904 did not change (as shown in the overhead view 920 of FIGS. 9A and 9B ). For example, if the content in the first media user interface is presented in a non-PiP presentation mode (e.g., presented natively within a video player application), movement of the user's viewpoint does not change the location of the first media user interface within the three-dimensional environment. Thus, in some embodiments, when the three-dimensional environment is displayed from the user's second viewpoint, the first media user interface is optionally no longer located within the user's field of view. In some embodiments, when the three-dimensional environment is presented from the user's second viewpoint, the first media user interface is no longer within the user's individual range of pose for the user's second viewpoint because the location (e.g., position) of the first media user interface within the three-dimensional environment is not within the user's field of view from the user's second viewpoint of the three-dimensional environment.
[0246] In some embodiments, following a determination that the content is being presented in a second presentation mode different from the first presentation mode (e.g., where the content is being presented in a picture-in-picture (PiP) format), the electronic device displays (1002e) the first media user interface at a second discrete location within the three-dimensional environment different from the first discrete location, and by displaying the first media user interface at the second discrete location, the first media user interface is displayed in a pose that is within a discrete range of poses for the user's second viewpoint, such as the location of user interface 934 moving from the location shown in FIG. 9C to the location shown in FIG. 9D as a result of a movement of the user's viewpoint (e.g., in some embodiments, the range of poses for the user's second viewpoint includes all (or a subset thereof) of the poses in the three-dimensional environment that are within the user's field of view from the user's second viewpoint). For example, when content within the first media user interface is presented in a picture-in-picture (PiP) presentation mode, the location of the first media user interface within the three-dimensional environment changes as the user's viewpoint of the three-dimensional environment changes, such that the first media user interface is always displayed in the portion of the three-dimensional environment currently being displayed (e.g., corresponding to the user's current viewpoint). In some embodiments, when the first media user interface is presented in a second presentation mode, the first media user interface is displayed at a location within the three-dimensional environment such that the first media user interface appears within a threshold distance (e.g., 0.5, 1, 2, 4, or 6 feet) of a user (e.g., a predetermined portion) of the electronic device (or within a threshold distance (0.5, 1, 2, 4, or 6 feet) of a discrete body part of the user (e.g., right or left hip, right or left shoulder)). For example, when the first media user interface is presented in the second presentation mode, the first media user interface is displayed in the lower right portion of the user's field of view of the three-dimensional environment, regardless of the location and / or orientation of the viewpoint.In some embodiments, when the first media user interface is not presented in the second presentation mode (e.g., presented in the first presentation mode), the first media user interface is optionally displayed at a location within the three-dimensional environment that is not within a threshold distance (0.5, 1, 2, 4, or 6 feet) of a user of the electronic device (or within a threshold distance (0.5, 1, 2, 4, or 6 feet) of an individual body part of the user (e.g., right or left hip, right or left shoulder)).
[0247] Changing the location of the first media user interface when the user's viewpoint in the three-dimensional environment changes provides an efficient way of providing continuous access to a particular user interface in the three-dimensional environment regardless of the user's current viewpoint of the three-dimensional environment when the content of the first media user interface is presented in the second presentation mode, thereby reducing cognitive burden on the user both when engaging with the first media user interface and when engaging with other content or applications in the three-dimensional environment.
[0248] In some embodiments, the second distinct location is based on the second perspective (1004a) (e.g., the location of the first media user interface is no longer based on the location of the user's first perspective, but rather on the location of the user's second perspective). In some embodiments, the second distinct location is a location in the three-dimensional environment that is within or at a threshold distance (e.g., 0.1, 0.2, 0.5, 1, 1.5, or 3 feet) from the second perspective. In some embodiments, the second distinct location is a location in the user's field of view from the second perspective. In some embodiments, the second distinct location in the three-dimensional environment corresponds to a location in the physical environment that is within or at a threshold distance of a distinct body part (e.g., part) of the user (e.g., within 0.1, 0.2, 0.3, 1, 2, or 3 feet of the user's hips, hand, head, foot, or knee). In some embodiments, the second distinct location corresponds to the lower right portion (or the lower left portion or the upper right portion) of the user's field of view from the user's second perspective. In some embodiments, displaying the first media user interface at the second discrete location may be performed if movement of the user's viewpoint after moving to the second viewpoint satisfies one or more criteria (e.g., in some embodiments, following movement of the user's viewpoint of the three-dimensional environment to the second viewpoint, the user's viewpoint of the three-dimensional environment has not changed by more than a predetermined amount (e.g., the user's viewpoint has not moved by more than a threshold movement amount (e.g., movement less than 1 cm, 2 cm, 5 cm, 10 cm, 50 cm, 100 cm, 300 cm, or 1000 cm)). 9D , if the viewpoint of user 922 of three-dimensional environment 904 satisfies the above-mentioned criteria, user interface 934 is displayed at a location within three-dimensional environment 904 that is within the field of view of user 922.For example, if the user's viewpoint of the three-dimensional environment corresponds to the second viewpoint for at least a threshold amount of time (e.g., 0.5, 1, 3, 7, 10, 20, or 30 seconds) and / or moves less than a threshold amount of movement (e.g., movement less than 1 cm, 2 cm, 5 cm, 10 cm, 50 cm, 100 cm, 300 cm, or 1000 cm), the first media user interface is displayed at a second discrete location within the three-dimensional environment.
[0249] In some embodiments, displaying the first media user interface at the second discrete location includes ceasing to display (1004c) the first media user interface at the second discrete location in accordance with a determination that movement of the user's viewpoint after moving to the second viewpoint does not satisfy one or more criteria (e.g., in some embodiments, one or more criteria are not met if, following movement of the user's viewpoint of the three-dimensional environment to the second viewpoint, the user's viewpoint of the three-dimensional environment has moved more than a threshold amount (e.g., the user's viewpoint has moved more than 1 cm, 2 cm, 5 cm, 10 cm, 50 cm, 100 cm, 300 cm, or 1000 cm), and / or has not corresponded to the second viewpoint for at least a threshold amount (e.g., 0.1, 1, 3, 7, 10, 20, or 30 seconds)). 9D , if the user's viewpoint of the three-dimensional environment 904 does not meet the above-mentioned criteria, the user interface 934 is not displayed at a location within the three-dimensional environment 904 that is within the field of view of the user 922. For example, if the user's viewpoint of the three-dimensional environment has not corresponded to a second viewpoint (and / or has moved more than a threshold amount of movement) for at least a threshold amount (e.g., 0.1, 1, 3, 7, 10, 20, or 30 seconds), the first media user interface is not displayed at a second discrete location until the user's viewpoint of the three-dimensional environment meets one or more criteria. In some embodiments, if one or more criteria are not met (e.g., the user's viewpoint of the three-dimensional environment has not corresponded to the second viewpoint for at least the above-mentioned threshold amount of time), the first media user interface continues to be displayed at a location within the three-dimensional environment based on the user's first viewpoint. In some embodiments, when the user's viewpoint is moved to the second viewpoint, the first media user interface at the first discrete location fades out and fades back in after one or more criteria are met. In some embodiments, the first media user interface does not change location within the three-dimensional environment until one or more criteria are met.Thus, while the electronic device ceases to display the first media user interface at the second discrete location, the electronic device optionally remains displayed at the first discrete location within the three-dimensional environment.
[0250] Displaying or delaying the display of the first media user interface following a movement of the user's viewpoint in the three-dimensional environment provides an efficient way of displaying the first media user interface relative to the user's new viewpoint after the user's movement has settled, thereby reducing cognitive burden on the user both when engaging with the first media user interface and when engaging with other content or applications within the three-dimensional environment.
[0251] In some embodiments, the first media user interface is associated with a separate application (e.g., a video application, a media application, or a streaming application). In some embodiments, the orientation of the first media user interface during the second presentation mode (e.g., whether the first media user interface is displayed in portrait or landscape mode) is defined by the application associated with the first media user interface. In some embodiments, the type of application associated with the first media user interface defines the orientation of the first media user interface during the second presentation mode. In some embodiments, during the second presentation mode, the first media user interface is automatically oriented toward the user's viewpoint (e.g., perpendicular to the user's viewpoint) such that content within the first media user interface is displayed (e.g., angled) toward the user's viewpoint of the three-dimensional environment.
[0252] In some embodiments, while the first media user interface is presenting the content in the second presentation mode, the first media user interface includes one or more user interface elements selectable to modify playback of the content (1006a), such as user interface elements 936-948 in user interface 934 of Figure 9D. In some embodiments, while displaying the one or more user interface elements in the first media user interface, the electronic device receives input via one or more input devices corresponding to a selection of an individual one of the one or more user interface elements, such as a selection of user interface element 936 of Figure 9D (1006b).
[0253] In some embodiments, in response to receiving the input, the electronic device modifies (1006c) the playback of the content according to the selection of a respective user interface element. For example, in response to the electronic device 101 detecting the selection of the user interface element 936, the electronic device 101 transitions the playback of TV program A from a picture-in-picture presentation to an enhanced presentation mode, as described in more detail with reference to FIG. 9D . For example, when the content is presented in a picture-in-picture mode (e.g., the second presentation mode), the user interface elements for modifying the playback of the content are displayed overlaid on the first media user interface. In some embodiments, the user interface elements are integrated within the first media user interface (as opposed to overlaid on the first media user interface) such that the content and the user interface elements presented within the first media user interface are at the same Z-depth within the three-dimensional environment. In some embodiments, the user interface elements for modifying the playback of the content include user interface elements for playing, pausing, fast-forwarding, rewinding, displaying subtitles, and / or modifying audio associated with the content presented in the first media user interface. In some embodiments, the first media user interface also includes user interface elements associated with playing content in a third (e.g., immersive) presentation mode if the content being presented in the first media user interface is immersive content (as described in more detail in method 1400). In some embodiments, the one or more user interface elements are displayed in the first media user interface after the electronic device detects that a user of the electronic device has performed a pinch gesture (e.g., with the thumb and index finger of the user's hand) while the user's gaze is directed at the first media user interface element.In some embodiments, the user interface element is displayed in the first media user interface when only the user's gaze is directed toward the first media user interface and / or when the user's gaze is directed toward the first media user interface while the user's hand is performing the initiation of a pinch gesture (e.g., when the thumb and index finger of the user's hand are more than a threshold distance apart (e.g., 0.5, 1, 1.5, 3, or 6 cm) and have not yet converged within the aforementioned threshold distance of each other).
[0254] Displaying user interface elements in the first media user interface (e.g., overlaid or integrated) provides an efficient way of displaying user interface elements associated with modifying playback of content and interacting with such controls, thereby reducing the cognitive burden on the user when engaging with the first media user interface and modifying playback of the first media user interface.
[0255] In some embodiments, while the first media user interface presents the content in the first presentation mode, the three-dimensional environment includes a playback control user interface separate from the first media user interface, the playback control user interface including one or more user interface elements selectable to modify playback of the content, where the first media user interface does not include one or more user interface elements selectable to modify playback of the content (1008a). For example, in FIG. 9B , the playback control user interface 908 is displayed separately from the user interface 906 during the extended presentation mode. In some embodiments, content is presented in the first presentation mode if the content is not presented in a picture-in-picture presentation mode. In some embodiments, content is presented in the first presentation mode if the content is presented at a display size larger than the display size of the content in the second presentation mode. In some embodiments, content is presented in the first presentation mode when the content is presented in a default presentation mode (e.g., playing natively in a video player application associated with the first media user interface).
[0256] In some embodiments, while displaying the playback control user interface, electronic device 101 receives input via one or more input devices corresponding to a selection of an individual user interface element of the one or more user interface elements (1008b). For example, in FIG. 9B, electronic device 101 detects a selection of user interface element 908b. In some embodiments, in response to receiving the input, electronic device 101 modifies playback of content according to the selection of the individual user interface element (1008c). For example, in FIG. 9C, in response to electronic device 101 detecting a selection of user interface element 908b in FIG. 9B, electronic device 101 displays TV program A in picture-in-picture user interface 934. For example, during presentation of content in the first presentation mode, user interface elements associated with modifying playback of the content being presented in the first media user interface are displayed in a playback control user interface that is separate from (e.g., not integrated with and / or overlaying) the first media user interface. In some embodiments, the one or more user interface elements include options for playing / pausing the content, moving the content forward a predetermined amount (e.g., 15, 30, 60, 90 seconds), and moving the content back a predetermined amount (e.g., 15, 30, 60, 90 seconds) to modify the display of subtitles and / or audio associated with the content being presented in the first media user interface. In some embodiments, the playback control user interface is angled differently toward the user's viewpoint of the three-dimensional environment than the first media user interface (e.g., perpendicular to the user's viewpoint). For example, in some embodiments, the playback control user interface is displayed at an upward angle relative to a fixed frame of reference, and the first media user interface is displayed parallel to the fixed frame of reference. In some embodiments, both the playback control user interface and the first media user interface are perpendicular to the user's viewpoint.In some embodiments, the playback control user interface includes an option to initiate playback of the content in a third presentation mode (e.g., an immersive presentation) if the content is immersive content, as described in more detail with reference to method 1400. In some embodiments, the playback control user interface is displayed in the three-dimensional environment after the electronic device detects that a user of the electronic device has performed a pinch gesture (e.g., with the thumb and index finger of the user's hand) while the user's gaze is directed toward a first media user interface element. In some embodiments, the playback control user interface is displayed in the three-dimensional environment when only the user's gaze is directed toward the first media user interface and / or when the user's gaze is directed toward the first media user interface while the user's hand is performing the initiation of a pinch gesture (e.g., when the thumb and index finger of the user's hand are more than a threshold distance (e.g., 0.5, 1, 1.5, 3, 6 cm) apart and have not yet converged within the aforementioned threshold distance of each other).
[0257] Displaying user interface elements for modifying playback of content in a first media user interface in a separate user interface during a first presentation mode provides an efficient way to access and interact with such user interface elements during the first presentation mode, thereby reducing the cognitive burden on the user when engaging with and modifying playback of the first media user interface.
[0258] In some embodiments, while the content is being displayed in a first presentation mode in a first media user interface (e.g., in some embodiments, the content is being presented in the first presentation mode if the content is not being presented in a picture-in-picture presentation mode; in some embodiments, the content is being presented in the first presentation mode if the content / first media user interface is being presented at a display size that is larger than the display size of the content / first media user interface during the second presentation mode; in some embodiments, the content is being presented in a default presentation mode (e.g., when a video player application associated with the first media user interface is While simultaneously displaying a first media user interface having a first discrete user interface element selectable for transitioning the content from the first presentation mode to a second presentation mode (e.g., displaying a user interface element selectable for transitioning the content from being presented in the first presentation mode to a second presentation mode (e.g., picture-in-picture presentation mode)), electronic device 101 receives (1010a) via one or more input devices a first input corresponding to a selection of the first discrete user interface element, such as a selection of user interface element 908b of FIG. 9B .
[0259] In some embodiments, in response to receiving a first input (e.g., ceasing presentation of content in the media user interface) (1010b), the electronic device, via the display generation component, displays (1010c) a second media user interface (e.g., different from the first media user interface) presenting the content in a second presentation mode. For example, in response to electronic device 101 detecting selection of user interface element 908b of FIG. 9B, electronic device 101 transitions the display of TV Program A from user interface 906 to picture-in-picture user interface 934 of FIG. 9C. For example, after receiving the first input, content transitions from playing in the first media user interface of the media application to the second media user interface (e.g., picture-in-picture user interface). In some embodiments, the second media user interface is smaller in size (e.g., in a three-dimensional environment) than the first media user interface (e.g., has a smaller width and / or height in the three-dimensional environment than the first media user interface). In some embodiments, the portion of the user's field of view occupied by the second media user interface is smaller than the portion of the user's field of view occupied by the first media user interface.
[0260] In some embodiments, in response to receiving the first input (1010b), the electronic device displays (1010d) in the first media user interface one or more selectable representations of one or more content items, including a first selectable representation of the first content item selectable to cause playback of the first content item in the first or second media user interface, such as representations 940-958 of FIG. 9C. For example, after receiving the first input, the first media user interface (e.g., which presented content before the content began playing in the second media user interface) begins displaying representations of the content items selectable to cause playback of the corresponding content items (e.g., within the first media user interface or the second media user interface). In some embodiments, the content items corresponding to the one or more selectable representations correspond to content items that have been recommended based on the user's content consumption history. In some embodiments, the content items corresponding to the one or more selectable representations correspond to trending, popular, and / or newly released content items.
[0261] Updating the first media user interface to include representations of additional content items when content that was presented in the first media user interface begins to be displayed in a different user interface provides an efficient way to access the additional content items at the same time (without requiring additional input) as the content is being presented in the second presentation mode, thereby reducing the cognitive burden on the user when interacting with the first media user interface and when modifying the presentation of the content being presented in the first media user interface.
[0262] In some embodiments, in response to receiving the first input, before presenting the content in the second presentation mode in the second media user interface, the electronic device displays (1012a) an animation of the content transitioning from the first presentation mode in the first media user interface to the second presentation mode in the second media user interface. For example, an animation is displayed when the electronic device 101 transitions the presentation of TV Program A from the augmented presentation mode shown in FIG. 9B to the picture-in-picture presentation mode shown in FIG. 9C. For example, after receiving a request to change the content from being presented in the first presentation mode to the second presentation mode, an animation is displayed indicating the content transitioning from the first presentation mode to the second presentation mode. In some embodiments, the animation includes content that is visually de-emphasized (e.g., fading out) in the first media user interface and / or content that is visually emphasized (e.g., fading in) in the second media user interface. In some embodiments, the animation includes visually highlighting the second media user interface to indicate that the content is currently being presented in the second media user interface. In some embodiments, the second media user interface remains highlighted or visually enhanced until the content has been presented in the second media user interface for a threshold amount of time (e.g., 5, 10, 20, 40, 60, 120 seconds) or until the user's attention is directed to the second media user interface (e.g., the user's gaze is directed toward the second media user interface). In some embodiments, the animation includes the content shrinking and / or moving within the three-dimensional environment from a location in the first media user interface to a location in the second media user interface where it is to be displayed.
[0263] Displaying animation as content transitions from a first presentation mode to a second presentation mode provides an efficient way of indicating the current presentation mode associated with the content, thereby reducing the cognitive burden on the user when engaging with content presented in the first media user interface.
[0264] In some embodiments, while presenting content in a second presentation mode in a second media user interface (e.g., while the content is presented in a picture-in-picture user interface), the electronic device receives a second input via one or more input devices corresponding to a request to change the presentation of the content from the second presentation mode to the first presentation mode (1014a), such as an input to select user interface element 936 in FIG. 9D . In some embodiments, the request to transition the presentation of the content from the second presentation mode to the first presentation mode is received when a user interface element displayed in or with the second media user interface is selected as described above. In some embodiments, in response to receiving the second input (1014b), the electronic device ceases displaying the second media user interface and one or more selectable representations within the second media user interface (1014c). For example, the second user interface (e.g., a picture-in-picture user interface) ceases being displayed in the three-dimensional environment when the presentation mode associated with the content switches from the second presentation mode to the first presentation mode. In some embodiments, the electronic device presents (1014d) the content in a first media user interface, the content being presented in a first presentation mode while the content is displayed in the first media user interface. For example, if the electronic device 101 detects a request to transition playback of TV program A from a picture-in-picture presentation mode to an augmented presentation mode, the electronic device 101 replaces representations 940-958 in the user interface 906 with playback of TV program A. For example, when the presentation mode associated with the content switches to the first presentation mode, the first media user interface in the three-dimensional environment begins presenting the content.In some embodiments, when the presentation mode of the content switches to the first presentation mode, the location of the content within the three-dimensional environment changes from a location corresponding to the second media user interface (e.g., a picture-in-picture user interface) to a location within the three-dimensional environment corresponding to the location of the application currently facilitating playback of the content. In some embodiments, when the content is being presented in the first presentation mode, the content is being presented in a first media user interface that is larger in size compared to the second media user interface (e.g., the content is therefore displayed at a larger size compared to the size presented in the second media user interface). In some embodiments, if a second input is received while the content was in the first playback position, the electronic device begins presenting the content in the first media user interface from the first playback position.
[0265] Presenting content in different media user interfaces based on the presentation mode of that content provides an efficient way of indicating the current presentation mode associated with the content, thereby reducing the cognitive burden on the user when engaging with content being presented in a first media user interface.
[0266] In some embodiments, in response to receiving the second input and prior to presenting the content in the first media user interface in the first presentation mode, the electronic device displays (1016a) an animation of the content transitioning from the second presentation mode in the second media user interface to the first presentation mode in the first media user interface. For example, the animation is displayed when the electronic device 101 is transitioning the presentation of TV Program A from the picture-in-picture presentation mode shown in FIG. 9D to the augmented presentation mode in response to selecting user interface element 936. For example, in response to receiving a request to change the content from being presented in the second presentation mode to the first presentation mode, the animation indicating the content transitioning from the second presentation mode to the first presentation mode is displayed. In some embodiments, the animation includes the content fading out (e.g., visually de-emphasizing) in the second media user interface and / or the content fading in (e.g., visually emphasizing) in the first media user interface. In some embodiments, the animation includes visually highlighting the first media user interface to indicate that the content is currently presented in the first media user interface (rather than the second media user interface). In some embodiments, the first media user interface remains highlighted or visually enhanced until the content has been presented in the first media user interface for a threshold amount of time (e.g., 5, 10, 20, 40, 60, 120 seconds) or until the user's attention is directed to the second media user interface (e.g., the user's gaze is directed to the first media user interface). In some embodiments, the animation includes the content expanding and / or moving within the three-dimensional environment from a location in the second media user interface to a location in the first media user interface.Displaying animation as content transitions from the second presentation mode to the first presentation mode provides an efficient way of indicating the current presentation mode associated with the content, thereby reducing the cognitive burden on the user when engaging with content presented in the first media user interface.
[0267] In some embodiments, while presenting content in a first presentation mode in a first media user interface (e.g., in some embodiments, content is presented in the first presentation mode if it is not presented in a picture-in-picture presentation mode. In some embodiments, content is presented in the first presentation mode if it is presented at a display size that is larger than the display size of the content in a second presentation mode. In some embodiments, content is presented in the first presentation mode when it is presented in a default presentation mode (e.g., playing natively in a video player application associated with the first media user interface), the electronic device receives a first input via one or more input devices (1018a) corresponding to a request to display a first user interface of a first application (e.g., a request to launch a new application within the three-dimensional environment is received). In some embodiments, in response to receiving the first input (1018b), the electronic device displays the first user interface of the first application in the three-dimensional environment (1018c). For example, when the electronic device receives a request to open / launch a first application in the three-dimensional environment, the user interface of the first application is displayed in the three-dimensional environment. In some embodiments, in response to receiving a first input (1018b), the electronic device ceases presenting content in the first presentation mode in the first media user interface (1018d). For example, when an application in the three-dimensional environment is launched while content is being presented in the first presentation mode, the content ceases being presented in the first presentation mode. In some embodiments, the first media user interface also ceases displaying in the three-dimensional environment.In some embodiments, in response to receiving the first input (1018b), the electronic device displays (1018e) a second media user interface presenting the content in the three-dimensional environment, the content being presented in a second presentation mode while the content is being presented in the second media user interface. For example, in FIG. 9A , if electronic device 101 receives a request to launch a new application in three-dimensional environment 904, electronic device 101 automatically transitions the presentation of TV program A from the augmented presentation mode to a picture-in-picture presentation mode. In some embodiments, the second media user interface is displayed simultaneously with the first user interface of the first application. For example, if a request to launch a new application is received in the three-dimensional environment while the content is being presented in the first presentation mode, the content begins playing in a different user interface (e.g., a picture-in-picture user interface) and a different presentation mode (e.g., a second presentation mode). In some embodiments, launching the first application in the three-dimensional environment transitions the content from the first presentation mode to the second presentation mode because the default launch location of the first application corresponds to the current location of the first media user interface in the three-dimensional environment. In some embodiments, launching the first application in the three-dimensional environment transitions the content from the first presentation mode to the second presentation mode because the display location of the first user interface occludes (or partially occludes) the first media user interface. In some embodiments, the content being presented in the second user interface is not occluded by the first user interface of the first application.
[0268] Switching the presentation mode of content from a first presentation mode to a second presentation mode when a request to launch a new application in the three-dimensional environment is received provides an efficient way of continuing to display content in the three-dimensional environment when displaying a new user interface in the three-dimensional environment, thereby reducing the cognitive burden on the user when engaging with content being presented in the first media user interface.
[0269] In some embodiments, the electronic device detects (1020a) that playback of the content has reached a predetermined playback threshold (e.g., playback has completed, playback of the content item is within a threshold amount of time from completion (e.g., playback of the content has finished in 0.5, 1, 1.5, 3, 5, 10, 20 minutes)). In some embodiments, in response to detecting that playback of the content has reached the predetermined playback threshold, the electronic device displays (1020b) a second user interface in the three-dimensional environment that includes one or more representations of recommended content that, when selected, cause the corresponding content to begin playing in the first media user interface. For example, in FIG. 9E , electronic device 101 displays user interface 946 in response to electronic device 101 detecting that playback of TV Program A has completed. For example, when playback of the content has reached the predetermined playback threshold, a second media user interface is displayed in the three-dimensional environment that includes one or more representations of content that can be selected to begin playing new content in the first media user interface. In some embodiments, the second user interface is displayed when the content is presented in the first presentation mode and is not displayed when the content is not presented in the first presentation mode. In some embodiments, the content corresponding to the one or more representations corresponds to content that is recommended based on a content consumption history of a user of the electronic device and / or because the user has previously saved / favorited content. In some embodiments, the playback control user interface is displayed simultaneously with the first media user interface and / or below the first media user interface presenting the content.
[0270] Displaying a second user interface including selectable representations of recommended content when content being played in a first media user interface reaches a predetermined playback position provides an efficient way of providing access to other content that can be played in the first media user interface, thereby reducing the cognitive burden on the user when engaging with content presented in the first media user interface.
[0271] In some embodiments, the one or more representations of the recommended content include a first individual representation of a first recommended content (1022a) (e.g., as described above, a representation corresponding to the first content item is displayed when the content reaches a predetermined playback threshold). In some embodiments, while the user's gaze is directed at the first individual representation (1022b), pursuant to a determination that the user's gaze has been directed at the first individual representation for more than a threshold amount of time (e.g., optionally without considering any other input / gesture being performed by the user of the electronic device, such as without detecting input from the user's hand directed at the first individual representation and / or any other element in the three-dimensional environment), the electronic device 101 begins playback of the first recommended content (e.g., in the first media user interface) (1022c). For example, if the user's gaze is directed at the first individual representation for more than a threshold amount of time (e.g., 5, 7, 9, 10, 20, 30, 60 seconds), the content corresponding to the first individual representation (the first recommended content) begins playback in the three-dimensional environment. In some embodiments, while the user's gaze is directed at the first individual representation (1022b), in accordance with a determination that the user's gaze has not been directed at the first individual representation for more than a threshold amount of time, the electronic device refrains from initiating playback of the first recommended content (e.g., in the first media user interface) (1022d). For example, the electronic device 101 initiates playback of item A if the user's gaze 914 has been directed at the corresponding representation 946 for the aforementioned threshold amount of time, and does not initiate playback of the item if the user's gaze 914 has not been directed at the corresponding representation 946 for the aforementioned threshold amount of time. For example, if the user's gaze has not been directed at the first individual representation for more than a threshold amount of time (e.g., 5, 7, 9, 10, 20, 30, 60 seconds), the content corresponding to the first individual representation (the first recommended content) does not begin playing in the three-dimensional environment until the user's gaze has been directed at the first individual representation for the aforementioned threshold amount of time.Initiating playback of content based on the amount of time the user's gaze has been directed at a corresponding representation displayed in the three-dimensional environment provides an efficient way of initiating playback of content in the three-dimensional environment without requiring the user to perform an additional gesture (e.g., a hand gesture), thereby reducing the cognitive burden on the user when engaging with content presented in the second first media user interface.
[0272] In some embodiments, while the user's gaze is directed at the first individual representation (e.g., and while the user's gaze is not directed at the first individual representation for more than a threshold amount of time), the electronic device displays (1024a) a visual indication in association with the first individual representation, which is updated as the user's gaze remains directed at the first individual representation to indicate progress toward reaching the threshold amount of time, such as visual indication 915 of FIG. 9E. For example, a visual indication is displayed indicating the amount of time remaining until the user's gaze is directed at the first individual representation for at least the aforementioned threshold amount of time (e.g., 5, 7, 9, 10, 20, 30, 60 seconds). In some embodiments, a progress indicator is displayed overlaying the first individual representation when the user's gaze is directed at the first individual representation and is not displayed when the user's gaze is not directed at the first individual representation. In some embodiments, in accordance with a determination that the user's gaze has been directed toward the first individual representation for at least a threshold amount of time, the visual indication stops being updated and content corresponding to the first individual representation (e.g., the first recommended content) begins playing within the three-dimensional environment. In some embodiments, the visual indication is a symbol (e.g., an arrow) that extends in a circular shape when the user's gaze is directed toward the first individual representation.
[0273] Providing an indication of when new content will begin playing within the three-dimensional environment based on the amount of time the user's gaze is directed at the corresponding representation of the content provides an efficient way of indicating when new content will play, thereby reducing the cognitive burden on the user when engaging with content presented within the second media user interface.
[0274] In some embodiments, while presenting the content in the second presentation mode, the pose of the first media user interface at the first discrete location relative to the first viewpoint is the same as the pose of the first media user interface at the second discrete location relative to the user's second viewpoint (1026a). For example, the first media user interface is displayed in a predetermined portion of the user's field of view (e.g., bottom right, top right, bottom left, top left, or bottom center) regardless of the user's current viewpoint of the three-dimensional environment. In some embodiments, the first media user interface is (e.g., always) displayed in the same relative position and / or orientation with respect to the user's viewpoint of the three-dimensional environment.
[0275] Displaying the first media user interface in the same pose (e.g., position and / or orientation) relative to the user's viewpoint provides an efficient way of displaying the first media user interface in a uniform manner regardless of the user's viewpoint of the three-dimensional environment, thereby reducing the cognitive burden on the user when engaging with the first media user interface.
[0276] In some embodiments, the user's first viewpoint corresponds to a first location within the electronic device's physical environment, and the user's second viewpoint corresponds to a second location within the physical environment that is different from the first location (1028a). In some embodiments, the electronic device displays the three-dimensional environment from the user's viewpoint at a location within the three-dimensional environment that corresponds to the electronic device's physical location within the electronic device's physical environment. In some embodiments, detecting movement of the user's viewpoint includes detecting movement of at least a portion of the user (e.g., the user's head, torso, or hands) within the physical environment. In some embodiments, detecting movement of the user's viewpoint includes detecting movement of the electronic device or a display generating component in the physical environment. In some embodiments, displaying the three-dimensional environment from the user's viewpoint includes displaying the three-dimensional environment from a perspective view associated with the user's viewpoint location within the three-dimensional environment. In some embodiments, updating the user's viewpoint causes the electronic device to display the plurality of virtual objects from a perspective view associated with the user's updated viewpoint location. For example, if the electronic device detects a leftward movement in the physical environment, the user's viewpoint moves left in the three-dimensional environment, and the electronic device updates the positions of multiple virtual objects displayed via the display generation component to move them to the right.
[0277] Displaying a three-dimensional environment from a perspective based on the user's physical location provides an efficient way to interact with the three-dimensional environment based on the user's actual pose and / or location in the physical environment, thereby reducing the cognitive burden on the user when interacting with the three-dimensional environment.
[0278] In some embodiments, while the content is being presented in the first media user interface in a second presentation mode and while the second media user interface is displayed at a third separate location within the three-dimensional environment (e.g., while the content in the first media user interface is being presented in a picture-in-picture user interface within the three-dimensional environment and the second media user interface is displaying a representation of content recommendations), the electronic device receives a second input via one or more input devices corresponding to a request to change the presentation of the content from the second presentation mode to the first presentation mode (1030a). In some embodiments, the electronic device receives the input because a user has selected a user interface element associated with changing the presentation mode of the content from the second presentation mode to the first presentation mode, as described above. In some embodiments, in response to receiving the second input (1030b), the electronic device ceases displaying the first media user interface (1032c). In some embodiments, in response to receiving the second input (1030b), in accordance with determining that the second media user interface is within a second discrete range of pose relative to the user's second viewpoint, the electronic device presents the content in the second media user interface at a third discrete location (1032d). For example, if electronic device 101 receives a request to transition playback of TV program A in FIG. 9C , the location of user interface 906 does not change because user interface 906 is currently within the field of view of user 922's current viewpoint of the three-dimensional environment. For example, if a request to transition content from the second presentation mode to the first presentation mode is received while the second media user interface (e.g., a video player / video application user interface) is within the user's field of view, the content begins playing in the second media user interface without a change in the location of the second media user interface within the three-dimensional environment.In some embodiments, the second individual range of poses for the user's second viewpoint includes all (or a subset thereof) of poses in the three-dimensional environment that are within the user's field of view from the user's second viewpoint. In some embodiments, in accordance with determining that the second media user interface is not within the second individual range of poses for the user's second viewpoint (1032e) (e.g., in some embodiments, the second media user interface is not within the second individual range of poses for the user's second viewpoint if the second media user interface is not at a location in the three-dimensional environment that is within the user's field of view from the second viewpoint), the electronic device displays the second media user interface at a fourth individual location within the three-dimensional environment that is different from the third individual location (1032f), wherein displaying the second media user interface at the fourth individual location causes the second media user interface to be displayed in an individual pose that is within the second individual range of poses for the second viewpoint, and wherein the second media user interface includes content. For example, if electronic device 101 detects a request to transition playback of a TV program from a picture-in-picture presentation mode to an augmented presentation mode, electronic device 101 updates the location of user interface 906 to be within the field of view of the user's current viewpoint of three-dimensional environment 904 and presents TV program A at the new location of user interface 906 within three-dimensional environment 904. For example, if a request to transition content from the second presentation mode to the first presentation mode is received while a second media user interface (e.g., a video player / video application user interface) is not within the user's field of view from the second viewpoint, the location of the media user interface moves to a location within the user's field of view from the user's second viewpoint.Moving the location of the second media user interface within the three-dimensional environment when the second user interface is not within the user's field of view provides an efficient way of moving the second media user interface within the user's field of view (when the first media user interface is not currently within the user's field of view) when a request to present content within the second user interface is received, thereby reducing the user's cognitive burden when engaging with the second media user interface.
[0279] 11A-11E show examples of how an electronic device can enhance navigation to individual playback positions of a content item, according to some embodiments.
[0280] FIG. 11A shows electronic device 101 displaying a three-dimensional environment 1102 via display generating component 120. It should be understood that in some embodiments, electronic device 101 utilizes one or more of the techniques described with reference to FIGS. 11A-11E in a two-dimensional environment without departing from the scope of this disclosure. As described above with reference to FIGS. 1-6 , electronic device 101 optionally includes display generating component 120 (e.g., a touchscreen) and multiple image sensors 314. The image sensors optionally include one or more of a visible light camera, an infrared camera, a depth sensor, or any other sensor that electronic device 101 can use to capture one or more images of a user or a portion of a user while the user interacts with electronic device 101. In some embodiments, display generating component 120 is a touchscreen capable of detecting a user's hand gestures and movements. In some embodiments, the user interfaces described below may also be implemented in a head-mounted display that includes display generation components that display the user interface to the user and sensors that detect the physical environment and / or the movement of the user's hands (e.g., external sensors facing outward from the user) and / or the user's line of sight (e.g., internal sensors facing inward toward the user's face).
[0281] 11A , electronic device 101 presents content item 1104 in three-dimensional environment 1102. In some embodiments, content item 1104 is an item of video content. In addition to content item 1104, three-dimensional environment 1102 includes representations of real objects in the physical environment of electronic device 101, such as wall representation 1108a, ceiling representation 1108b, table representation 1106a, and sofa representation 1106b. Electronic device 101, according to one or more steps of method 800, displays content item 1104 with increased visual emphasis relative to the rest of three-dimensional environment 1102, such as displaying areas of three-dimensional environment 1102 that do not include content item 1104 with a greater amount of blur and / or darkening than content item 1104, and displays virtual lighting effects 1110a-d to simulate light spill emanating from content item 1104.
[0282] 11A , a user's gaze 1113a is directed toward content item 1104 while the content item is being played. The playback position 1106 of content item 1104 in FIG. 11A advances as playback of the content item continues. In some embodiments, if the user shifts their attention away from content item 1104, electronic device 101 continues to play the content item. In some embodiments, when the user shifts their attention away from content item 1104 and then shifts their attention back to content item 1104, electronic device 101 presents a selectable option that, when selected, causes electronic device 101 to update the playback position to the playback position associated with the playback position that was playing at the time the user shifted their attention away from content item 1104.
[0283] 11B , the user shifts their attention away from the content item 1104 while the content item playback position 1106 is at the playback position shown in the figure. In some embodiments, detecting that the user has shifted their attention away from the content item 1104 includes detecting the user's gaze 1113b directed toward a location in the three-dimensional environment 1102 other than the content item 1104. In some embodiments, in response to detecting the user's gaze 1113b directed away from the content item 1104, the electronic device 101 reduces the amount of visual emphasis of the content item 1104 relative to the remainder of the three-dimensional environment 1102. In some embodiments, the user's gaze 1113b must be directed away from the content item 1104 for a threshold period of time (e.g., 1, 2, 3, 5, 10, 15, 30, or 45 seconds, 1, 2, 3, or 5 minutes) for the electronic device 101 to determine that the user's attention is being directed away from the content item 1104. In some embodiments, the moment the user's gaze 1113b is directed away from the content item 1104, the electronic device 101 determines that the user's attention is directed away from the content item 1104. In some embodiments, detecting that the user has shifted their attention away from the content item 1104 optionally includes detecting that the user closes their eyes, as shown in legend 1123, for at least a threshold period of time (e.g., 0.1, 0.2, 0.3, 0.5, 1, 2, 3, 5, 10, 15, 30, or 45 seconds, 1, 2, 3, or 5 minutes), for example, corresponding to the user falling asleep. In some embodiments, the electronic device 101 continues playing the content item 1104 after detecting the user's attention being directed away from the content item 1104 and advances the playback position 1106 beyond the point shown in FIG. 11B , which is the playback position of the content item at the time the electronic device 101 determined that the user's attention was directed away from the content item 1104.
[0284] 11C shows electronic device 101 presenting selectable options 1112a and 1112b that, when selected, cause electronic device 101 to resume playback of content item 1104 from playback position 1106 corresponding to the moment the user shifted their attention away from content item 1104. In some embodiments, electronic device 101 presents selectable options 1112a and / or selectable options 1112b in response to detecting the user's attention directed toward content item 1104 and / or the user's hand 1103b in a ready pose. In some embodiments, detecting the user's attention directed toward content item 1104 includes detecting the user's gaze 1103d directed toward content item 1104. In some embodiments, in response to detecting gaze 1103d directed toward content item 1104, electronic device 101 increases the visual prominence of content item 1104 relative to the rest of three-dimensional environment 1102. 11C , the electronic device 101 also displays options 1114a and 1114b and user interface element 1116, including additional options for modifying playback of the content item 1104, in response to detecting a gaze 1103d directed toward the content item 1104 and / or the hand 1103b in the ready pose, as described in more detail above with reference to method 800.
[0285] In some embodiments, electronic device 101 presents both options 1112a and 1112b. In some embodiments, electronic device 101 presents option 1112a or option 1112b, but not both. Option 1112a is displayed outside of user interface element 1116 overlaid on content item 1104. Option 1112b is displayed as part of scrub bar 1111 included in user interface element 1116. Scrub bar 1111 includes an indication 1113 of the current playback position of content item 1104 (which optionally continues playback while options 1112a and / or 1112b are displayed). Electronic device 101 displays option 1112b in a location on scrub bar 1111 corresponding to the playback position where playback of content item 1104 will resume in response to selection of option 1112b. In some embodiments, the playback position at which playback of content item 1104 resumes is the playback position in Figure 1 IB at which the user shifts their attention away from content item 1104. In some embodiments, the playback position at which playback of content item 1104 resumes is a playback position that is a predetermined time (e.g., 1, 2, 3, 5, 10, 15, or 30 seconds) before or after the playback position in Figure 1 IB at which the user shifts their attention away from content item 1104.
[0286] As shown in FIG. 11C , the user selects option 1112a with gaze 1103c and hand 1103a, for example, via indirect input. In some embodiments, detecting the selection of option 1112a via indirect input includes detecting that hand 1103a makes a pinch gesture in which the thumb of hand 1103a touches another finger of the hand while gaze 1103c is directed at option 1112a. While FIG. 11C illustrates a first input state of hand 1103a and gaze 1103d and a second input state of hand 1103b and gaze 1103c, it should be understood that in some embodiments, these input states are detected at different times. In some embodiments, other selection inputs are possible. As described in more detail below with reference to FIG. 11E , in response to detecting the selection of option 1112a, electronic device 101 updates content item playback position 1106 to a playback position associated with the time the user shifted their attention away from content item 1104.
[0287] FIG. 11D , for example, illustrates selection of selectable option 1112b via direct input. In some embodiments, detecting selection of selectable option 1112b via direct input includes detecting a user's hand 1103a within a predetermined threshold distance of option 1112b while the hand 1103a is in a predetermined shape. In some embodiments, the predetermined shape is a pinch hand shape with the thumb touching the other fingers of the hand. In some embodiments, the predetermined shape is a pointing hand shape with one or more fingers extended and one or more fingers curled toward the palm. In some embodiments, direct input includes detecting a user's gaze 1103e directed toward option 1112b, and in some embodiments, direct input does not include detecting a user's gaze 1103e directed toward option 1112b. In some embodiments, electronic device 101 detects an indirect input selecting option 1112b similar to the indirect input or another type of selection input described above with reference to FIG. 11C . In response to the input shown in FIG. 11D, the electronic device 101 updates the playback position 1106 of the content item 1104 to the pl...
Claims
1. 1. A method comprising:
1. An electronic device in communication with a display generating component and one or more input devices, comprising: displaying a three-dimensional environment via said display generation component; a media user interface object including a respective piece of content of a first content item being presented in a first presentation mode; displaying the three-dimensional environment via the display generation component, the first user interface element transitioning the presentation of the first content item to a second presentation mode, wherein during presentation of the first content item in the first presentation mode, the first content item is displayed in a frame occupying a first portion of a field of view of the three-dimensional environment from a viewpoint of a user of the electronic device, and a second portion of the field of view from the viewpoint of the user is occupied by other elements of the three-dimensional environment representing a physical environment surrounding the electronic device; receiving, via the one or more input devices, a first input corresponding to a selection of the first user interface element while displaying the three-dimensional environment including the media user interface object and the first user interface element; and displaying the first content item in the second presentation mode in the three-dimensional environment in response to receiving the first input, wherein displaying the first content item in the second presentation mode simultaneously displaying the individual pieces of content of the first content item; and replacing at least a portion of the other elements of the three-dimensional environment representing the physical environment surrounding the electronic device with virtual three-dimensional content that extends to at least one edge of the field of view from the viewpoint of the user of the electronic device.
2. 2. The method of claim 1, wherein during the presentation of the first content item in the second presentation mode, the first content item extends to at least a plurality of respective edges of the field of view from the viewpoint of the user of the electronic device.
3. 10. The method of claim 1, wherein during presentation of the first content item in the second presentation mode, the first content item extends beyond at least one edge of the field of view from the viewpoint of the user of the electronic device.
4. 2. The method of claim 1, further comprising displaying a playback control user interface including one or more user interface elements, including the first user interface element, for modifying playback of the first content item in the three-dimensional environment while the media user interface object is presenting the first content item in the first presentation mode, the playback control user interface being displayed at a first location within the three-dimensional environment based on a location of the media user interface object.
5. During presentation of the first content item in the first presentation mode, the viewpoint of the user corresponds to a first viewpoint, and the method further comprises: detecting a movement of the user's viewpoint from the first viewpoint to a second viewpoint while presenting the first content item in the three-dimensional environment from the first viewpoint in the first presentation mode; In response to detecting the movement of the user's viewpoint to the second viewpoint, displaying the three-dimensional environment from the second viewpoint of the user; The method of claim 4 , further comprising: maintaining a display of the playback control user interface at the first location within the three-dimensional environment.
6. displaying the media user interface object within the three-dimensional environment and, while displaying the playback control user interface at the first location, receiving a second input via the one or more input devices corresponding to a request to move the media user interface object to a different location within the three-dimensional environment; In response to receiving the second input, moving the media user interface object to the different location within the three-dimensional environment; 5. The method of claim 4, further comprising: displaying the playback control user interface at a second location within the three-dimensional environment that is different from the first location, the second location within the three-dimensional environment being based on the different location within the three-dimensional environment.
7. While displaying the first content item in the first presentation mode, displaying a playback control user interface at a first location within the three-dimensional environment, the first location being a first distance from the viewpoint of the user, the playback control user interface including one or more user interface elements for modifying playback of the first content item; 2. The method of claim 1, further comprising: while presenting the first content item in the second presentation mode, displaying the playback control user interface at a second location within the three-dimensional environment, the second location within the three-dimensional environment being a second distance closer than the first distance from the user's viewpoint within the three-dimensional environment.
8. The method of claim 7 , wherein the second location in the three-dimensional environment is based on a location of the viewpoint of the user.
9. The method comprises: detecting, while the user's viewpoint is a first viewpoint and the media user interface object is displayed in the second presentation mode and the playback control user interface is displayed at the second location within the three-dimensional environment, a movement of the user's viewpoint from the first viewpoint to a second viewpoint different from the first viewpoint; In response to detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, displaying the three-dimensional environment from the second viewpoint of the user of the electronic device via the display generation component; 8. The method of claim 7, further comprising: displaying the playback control user interface at a third location within the three-dimensional environment that is different from the second location, the third location being based on a location of the second viewpoint of the user.
10. receiving, while presenting the first content item in a discrete presentation mode, a second input via the one or more input devices, the second input comprising a movement of a discrete portion of the user, the second input corresponding to a request to scrub the first content item; 10. The method of claim 1, further comprising: in response to receiving the second input, scrubbing the first content item according to the movement of the discrete portion of the user.
11. Scrubbing the first content item includes:
11. The method of claim 10, further comprising: displaying, in the media user interface object, content corresponding to a current scrub position within the first content item that changes as the individual portion of the user moves, in accordance with a determination that one or more second criteria are met while detecting the second input.
12. Scrubbing the first content item includes: in response to a determination that one or more second criteria are not met while detecting the second input and the content is displayed as immersive content; pausing playback of the first content item in the media user interface object; displaying in the three-dimensional environment a visual indication of separate content within the first content item that is separate from the immersive content, the separate content corresponding to a current scrub position within the first content item that changes as the separate portion of the user moves without changing the appearance of the immersive content as the separate portion of the user moves; responsive to detecting an end of the second input, ceasing to display the visual indication of the respective content and playing the first content item within the media user interface object as immersive content starting from the respective scrub position where the end of the second input was detected.
13. 12. The method of claim 11, wherein the one or more second criteria include a criterion that is met when the second input is received while the first content item occupies less than a threshold portion of the user's field of view and is not met when the second input is received while the first content item occupies more than the threshold portion of the user's field of view.
14. 2. The method of claim 1, wherein during presentation of the first content item in the first presentation mode, the first content item is displayed in the three-dimensional environment at a first size, and during presentation of the first content item in the second presentation mode, the first content item is displayed in the three-dimensional environment at a second size that is larger than the first size.
15. presenting the first content item in the first presentation mode includes presenting the first content item in the three-dimensional environment while a representation of a first portion of the physical environment surrounding the electronic device has a first level of visual emphasis; 2. The method of claim 1 , wherein presenting the first content item in the second presentation mode includes reducing visual emphasis of the first portion of the physical environment to a second level of visual emphasis that is lower than the first level of visual emphasis.
16. 10. The method of claim 1, further comprising, while presenting the first content item in the second presentation mode, displaying a second user interface element in the three-dimensional environment that transitions presentation of the first content item from the second presentation mode to the first presentation mode.
17. The method comprises: detecting a movement of the user's viewpoint from the first viewpoint to a second viewpoint different from the first viewpoint while presenting the first content item in the first presentation mode, and while the viewpoint of the user corresponds to a first viewpoint and a first portion of the first content item but not a second portion of the first content item is displayed in the media user interface object; 2. The method of claim 1, further comprising: in response to detecting the movement of the user's viewpoint from the first viewpoint to the second viewpoint, displaying the first content item within the media user interface object from the user's second viewpoint, wherein displaying the first content item within the media user interface object from the user's second viewpoint comprises displaying the second portion of the first content item within the media user interface object.
18. 1. An electronic device comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 1 to 17.
19. 18. One or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any one of claims 1 to 17.
20. The method of claim 1, wherein the individual content of the first content item includes one or more portions of the first content item that were not displayed in the first presentation mode.
Citation Information
Patent Citations
Video playing control method and device, augmented reality equipment and storage medium
CN111580652A
Method for modifying a virtual reality scene and computer-readable storage medium
JP2019527881A
Controls and interfaces for user interaction in virtual spaces
JP2019536131A
Method and system for viewing set top box content in a virtual reality device
US20160227267A1
Devices, Methods, and Graphical User Interfaces for Interacting with Three-Dimensional Environments
US20210097776A1