Devices, methods, and graphical user interfaces for interacting with media and three-dimensional environments
Gaze-based interaction methods in computer systems enhance virtual and augmented reality experiences by reducing input requirements and conserving power, addressing inefficiencies in existing interaction methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2022-09-24
- Publication Date
- 2026-05-15
Smart Images

Figure 0007860227000001 
Figure 0007860227000002 
Figure 0007860227000003
Abstract
Description
Technical Field
[0001] (Cross - reference to Related Applications) This application claims priority to U.S. Patent Application No. 17 / 952,206, filed on September 23, 2022, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR INTERACTING WITH MEDIA AND THREE - DIMENSIONAL ENVIRONMENTS"; U.S. Provisional Patent Application No. 63 / 409,695, filed on September 23, 2022, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR INTERACTING WITH MEDIA AND THREE - DIMENSIONAL ENVIRONMENTS"; and U.S. Provisional Patent Application No. 63 / 248,222, filed on September 24, 2021, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR INTERACTING WITH MEDIA AND THREE - DIMENSIONAL ENVIRONMENTS". The entire contents of each of these applications are hereby incorporated by reference into this specification.
[0002] The present disclosure generally relates to a computer system in communication with a display generation component, including but not limited to, an electronic device that provides virtual reality experiences and mixed reality experiences, and one or more input devices that provide computer - generated experiences.
Background Art
[0003] The development of computer systems for augmented reality has progressed significantly in recent years. An exemplary augmented reality environment includes at least several virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with the virtual / augmented reality environment. Exemplary virtual elements include virtual objects such as digital images, videos, text, icons, and control elements such as buttons and other graphics. [Overview of the project]
[0004] Some methods and interfaces for interacting with media items and environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and restrictive. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in augmented reality environments, and systems where manipulating virtual objects is complex and error-prone impose a significant cognitive burden on the user and detract from the virtual / augmented reality experience. In addition, these methods are unnecessarily time-consuming, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.
[0005] Therefore, there is a need for computer systems with improved methods and interfaces to provide users with computer-generated experiences that make interaction with the computer system more efficient and intuitive for the user. Such methods and interfaces optionally complement or replace conventional methods for interacting with media items and providing users with augmented reality experiences. Such methods and interfaces reduce the number, extent, and / or types of user input by helping the user understand the connection between the input provided and the device response to that input, thereby generating a more efficient human-machine interface.
[0006] The drawbacks and other problems associated with the user interface of a computer system described above are mitigated or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, tablet computer, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to display-generating components, the output devices include one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing multiple functions. In some embodiments, the user interacts with the GUI (and / or computer system) through stylus and / or finger touch and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the user's body as captured by a camera and other motion sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, gameplay, making phone calls, video conferencing, sending emails, instant messaging, training support, digital photography, digital videography, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing those functions optionally reside in a primary computer-readable storage medium and / or a non-primary computer-readable storage medium, or in other computer program products configured to be executed by one or more processors.
[0007] There is a need for electronic devices with improved methods and interfaces for interacting with media items in a three-dimensional environment. Such methods and interfaces can complement or replace conventional methods for interacting with media items in a three-dimensional environment. Such methods and interfaces reduce the number, degree, and / or type of user input, resulting in a more efficient human-machine interface. In the case of battery-operated computing devices, such methods and interfaces conserve power and extend the interval between battery charging. Such methods and interfaces also improve the usability of the device, making the user device interface more efficient by, for example, reducing the number of unnecessary and / or irrelevant received inputs and providing the user with improved visual feedback.
[0008] A method is described according to several embodiments. The method includes, in a computer system communicating with a display generation component and one or more input devices, displaying a media library user interface including representations of multiple media items, including a representation of a first media item, via the display generation component; detecting a user gaze corresponding to a first position in the media library user interface via one or more input devices at a first point in time while the media library user interface is being displayed, changing the appearance of the representation of the first media item via the display generation component from being displayed in a first manner to being displayed in a second manner different from the first manner, at a second point in time after the first point in time, detecting a user gaze corresponding to a second position in the media library user interface different from the first position via one or more input devices, and displaying the representation of the first media item via the display generation component in a third manner different from the second manner, in response to the detection of the user gaze corresponding to the second position in the media library user interface.
[0009] According to some embodiments, non-temporary computer-readable storage media are described. In some embodiments, a non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs include instructions to display a media library user interface including representations of multiple media items, including a representation of a first media item, via the display generation component, while the media library user interface is being displayed, to detect a user gaze corresponding to a first position in the media library user interface via one or more input devices at a first time point, and in response to the detection of the user gaze corresponding to the first position in the media library user interface, to change the appearance of the representation of the first media item via the display generation component from being displayed in a first manner to being displayed in a second manner different from the first manner, and at a second time point after the first time point, to detect a user gaze corresponding to a second position in the media library user interface different from the first position via one or more input devices, and in response to the detection of the user gaze corresponding to the second position in the media library user interface, to display the representation of the first media item via the display generation component in a third manner different from the second manner.
[0010] According to some embodiments, a temporary computer-readable storage medium is described. In some embodiments, a temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs include instructions to display a media library user interface including representations of multiple media items, including a representation of a first media item, via the display generation component, while the media library user interface is being displayed, to detect a user gaze corresponding to a first position in the media library user interface via one or more input devices at a first time point, and in response to the detection of the user gaze corresponding to the first position in the media library user interface, to change the appearance of the representation of the first media item via the display generation component from being displayed in a first manner to being displayed in a second manner different from the first manner, and at a second time point after the first time point, to detect a user gaze corresponding to a second position in the media library user interface different from the first position via one or more input devices, and in response to the detection of the user gaze corresponding to the second position in the media library user interface, to display the representation of the first media item via the display generation component in a third manner different from the second manner.
[0011] According to several embodiments, a computer system is described. In some embodiments, the computer system communicates with a display generation component and one or more input devices, and includes one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors, wherein one or more programs include instructions to display a media library user interface including representations of a plurality of media items, including a representation of a first media item, via the display generation component, and while the media library user interface is being displayed, at a first time point, detect a user gaze corresponding to a first position in the media library user interface via one or more input devices, and in response to the detection of the user gaze corresponding to the first position in the media library user interface, change the appearance of the representation of the first media item via the display generation component from being displayed in a first manner to being displayed in a second manner different from the first manner, and at a second time point after the first time point, detect a user gaze corresponding to a second position in the media library user interface different from the first position via one or more input devices, and in response to the detection of the user gaze corresponding to the second position in the media library user interface, display the representation of the first media item via the display generation component in a third manner different from the second manner.
[0012] In some embodiments, a computer system is described. In some embodiments, the computer system includes means for communicating with a display generation component and one or more input devices to display a media library user interface including representations of a plurality of media items, including a representation of a first media item, via the display generation component; and means for detecting, at a first time point, a user gaze corresponding to a first position in the media library user interface via one or more input devices, while the media library user interface is being displayed, and in response to the detection of the user gaze corresponding to the first position in the media library user interface, changing the appearance of the representation of the first media item via the display generation component from being displayed in a first manner to being displayed in a second manner different from the first manner; and at a second time point after the first time point, detecting, via one or more input devices, a user gaze corresponding to a second position in the media library user interface different from the first position, and in response to the detection of the user gaze corresponding to the second position in the media library user interface, displaying the representation of the first media item via the display generation component in a third manner different from the second manner.
[0013] In some embodiments, a computer program product is described. In some embodiments, the computer program product stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs include instructions to display a media library user interface including representations of multiple media items, including a representation of a first media item, via the display generation component, while the media library user interface is being displayed, at a first time point, detect a user gaze corresponding to a first position in the media library user interface via one or more input devices, and in response to the detection of the user gaze corresponding to the first position in the media library user interface, change the appearance of the representation of the first media item via the display generation component from being displayed in a first manner to being displayed in a second manner different from the first manner, at a second time point after the first time point, detect a user gaze corresponding to a second position in the media library user interface different from the first position via one or more input devices, and in response to the detection of the user gaze corresponding to the second position in the media library user interface, display the representation of the first media item via the display generation component in a third manner different from the second manner.
[0014] The method is described according to several embodiments. The method includes, in a computer system communicating with a display generation component and one or more input devices, displaying a user interface at a first zoom level via the display generation component; detecting one or more user inputs corresponding to a zoom-in user command via one or more input devices while the user interface is being displayed; and, in response to the detection of one or more user inputs corresponding to a zoom-in user command, displaying the user interface at a second zoom level greater than the first zoom level via the display generation component according to a determination that the user's gaze corresponds to a first position in the user interface, wherein displaying the user interface at the second zoom level includes zooming the user interface using a first zoom center selected based on the first position; and, according to a determination that the user's gaze corresponds to a second position in the user interface that is different from the first position, displaying the user interface at the third zoom level includes zooming the user interface using a second zoom center selected based on the second position, wherein the second zoom center is located at a different location from the first zoom center.
[0015] According to some embodiments, a non-temporary computer-readable storage medium is described. In some embodiments, the non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs display a user interface at a first zoom level via the display generation component, and while displaying the user interface, detect one or more user inputs corresponding to zoom-in user commands via one or more input devices, and in response to the detection of one or more user inputs corresponding to zoom-in user commands, determine that the user's gaze corresponds to a first position in the user interface, and then, via the display generation component, zoom to a level greater than the first zoom level The instructions include displaying the user interface at a second zoom level, which includes zooming the user interface using a first zoom center selected based on a first position, and displaying the user interface at a second zoom level greater than the first zoom level via a display generating component, in accordance with the determination that the user's gaze corresponds to a second position in the user interface that is different from the first position, and displaying the user interface at a third zoom level that is greater than the first zoom level, which includes zooming the user interface using a second zoom center selected based on a second position, where the second zoom center is in a different location from the first zoom center.
[0016] According to some embodiments, a temporary computer-readable storage medium is described. In some embodiments, the temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs display a user interface at a first zoom level via the display generation component, and while the user interface is being displayed, detect one or more user inputs corresponding to zoom-in user commands via one or more input devices, and in response to the detection of one or more user inputs corresponding to zoom-in user commands, determine that the user's gaze corresponds to a first position in the user interface, and then, via the display generation component, zoom to a level greater than the first zoom level The instructions include displaying the user interface at a second zoom level, which includes zooming the user interface using a first zoom center selected based on a first position, and displaying the user interface at a second zoom level greater than the first zoom level via a display generating component, in accordance with the determination that the user's gaze corresponds to a second position in the user interface that is different from the first position, and displaying the user interface at a third zoom level that is greater than the first zoom level, which includes zooming the user interface using a second zoom center selected based on a second position, where the second zoom center is in a different location from the first zoom center.
[0017] A computer system is described according to several embodiments. In some embodiments, the computer system communicates with a display generation component and one or more input devices, and the computer system comprises one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs display a user interface at a first zoom level via the display generation component, and while displaying the user interface, detect one or more user inputs corresponding to a zoom-in user command via one or more input devices, and in response to the detection of one or more user inputs corresponding to a zoom-in user command, according to the determination that the user's gaze corresponds to a first position in the user interface, the computer system displays a first zoom level via the display generation component. Displaying the user interface at a second zoom level greater than the bell, and displaying the user interface at the second zoom level includes zooming the user interface using a first zoom center selected based on the first position, and according to the determination that the user's gaze corresponds to a second position in the user interface that is different from the first position, the display generating component includes an instruction to display the user interface at a third zoom level greater than the first zoom level, and displaying the user interface at the third zoom level includes zooming the user interface using a second zoom center selected based on the second position, where the second zoom center is in a different location from the first zoom center.
[0018] In some embodiments, a computer system is described. In some embodiments, the computer system communicates with a display generation component and one or more input devices, and the computer system comprises means for displaying a user interface at a first zoom level via the display generation component; means for detecting one or more user inputs corresponding to zoom-in user commands via one or more input devices while the user interface is being displayed; and means for displaying the user interface at a second zoom level greater than the first zoom level via the display generation component in accordance with the determination that the user's gaze corresponds to a first position in the user interface in response to the detection of one or more user inputs corresponding to a zoom-in user command, and displaying the user interface at the second zoom level includes zooming the user interface using a first zoom center selected based on the first position; and displaying the user interface at a third zoom level greater than the first zoom level via the display generation component in accordance with the determination that the user's gaze corresponds to a second position in the user interface different from the first position, and displaying the user interface at the third zoom level includes zooming the user interface using a second zoom center selected based on the second position, the second zoom center being located at a different location from the first zoom center.
[0019] In some embodiments, a computer program product is described. In some embodiments, a computer program product includes one or more programs configured to run by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs including a command to display a user interface at a first zoom level via the display generation component, to detect one or more user inputs corresponding to a zoom-in user command via one or more input devices while displaying the user interface, and, in response to the detection of one or more user inputs corresponding to a zoom-in user command, to display the user interface at a second zoom level greater than the first zoom level via the display generation component according to a determination that the user's gaze corresponds to a first position in the user interface, and displaying the user interface at the second zoom level includes zooming the user interface using a first zoom center selected based on the first position, and displaying the user interface at a third zoom level greater than the first zoom level via the display generation component according to a determination that the user's gaze corresponds to a second position in the user interface different from the first position, and displaying the user interface at the third zoom level includes zooming the user interface using a second zoom center selected based on the second position, the second zoom center being at a different location from the first zoom center.
[0020] A method is described according to several embodiments. The method includes, in a computer system communicating with a display generation component and one or more input devices, detecting one or more user inputs corresponding to the selection of a first media item via one or more input devices; displaying the first media item in a first manner via the display generation component according to a determination that the first media item is a media item containing depth information of a particular type, in response to the detection of one or more user inputs corresponding to the selection of a first media item; and displaying the first media item in a second manner different from the first manner via the display generation component according to a determination that the first media item is a media item not containing depth information of a particular type.
[0021] According to some embodiments, a non-temporary computer-readable storage medium is described. In some embodiments, the non-temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs include instructions for detecting one or more user inputs corresponding to the selection of a first media item via one or more input devices, and, in response to detecting one or more user inputs corresponding to the selection of a first media item, displaying the first media item in a first manner via the display generation component according to a determination that the first media item is a media item containing a distinct type of depth information, and displaying the first media item in a second manner different from the first manner via the display generation component according to a determination that the first media item is a media item that does not contain a distinct type of depth information.
[0022] According to some embodiments, a temporary computer-readable storage medium is described. In some embodiments, the temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs include instructions for detecting one or more user inputs corresponding to the selection of a first media item via one or more input devices, and, in response to detecting one or more user inputs corresponding to the selection of a first media item, displaying the first media item in a first manner via the display generation component according to a determination that the first media item is a media item containing a distinct type of depth information, and displaying the first media item in a second manner different from the first manner via the display generation component according to a determination that the first media item is a media item that does not contain a distinct type of depth information.
[0023] A computer system is described according to several embodiments. In some embodiments, the computer system communicates with a display generation component and one or more input devices, and the computer system comprises one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for detecting one or more user inputs corresponding to the selection of a first media item via one or more input devices, and, in response to detecting one or more user inputs corresponding to the selection of a first media item, displaying the first media item in a first manner via the display generation component according to a determination that the first media item is a media item containing depth information of a distinct type, and displaying the first media item in a second manner different from the first manner via the display generation component according to a determination that the first media item is a media item that does not contain depth information of a distinct type.
[0024] In some embodiments, a computer system is described. In some embodiments, the computer system communicates with a display generation component and one or more input devices, and the computer system includes means for detecting one or more user inputs corresponding to the selection of a first media item via one or more input devices, and means for displaying the first media item in a first manner via the display generation component in accordance with the determination that the first media item is a media item containing a distinct type of depth information, in response to the detection of one or more user inputs corresponding to the selection of a first media item, and for displaying the first media item in a second manner different from the first manner via the display generation component in accordance with the determination that the first media item is a media item that does not contain a distinct type of depth information.
[0025] In some embodiments, a computer program product is described. In some embodiments, the computer program product stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more input devices, the one or more programs include instructions for detecting one or more user inputs corresponding to the selection of a first media item via one or more input devices, and, in response to detecting one or more user inputs corresponding to the selection of a first media item, displaying the first media item in a first manner via the display generation component according to a determination that the first media item is a media item containing depth information of a distinct type, and displaying the first media item in a second manner different from the first manner via the display generation component according to a determination that the first media item is a media item that does not contain depth information of a distinct type.
[0026] In some embodiments, a method is described that is executed in a computer system that communicates with a display generation component. The method includes displaying, via the display generation component, a user interface that includes a first representation of a stereoscopic media item, the first representation of the stereoscopic media item including at least a first edge, and a visual effect that obscures at least a first portion of the stereoscopic media item and extends inwardly from at least the first edge of the first representation of the stereoscopic media item toward the interior of the first representation of the stereoscopic media item.
[0027] In some embodiments, a non - transient computer - readable storage medium is described. The non - transient computer - readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with a display generation component, and the one or more programs include instructions to display, via the display generation component, a user interface that includes a first representation of a stereoscopic media item, the first representation of the stereoscopic media item including at least a first edge, and a visual effect that obscures at least a first portion of the stereoscopic media item and extends inwardly from at least the first edge of the first representation of the stereoscopic media item toward the interior of the first representation of the stereoscopic media item.
[0028] In some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system, the computer system communicates with a display generation component, and one or more programs, via the display generation component, include instructions to display a user interface which includes a first representation of a stereoscopic media item, the first representation of the stereoscopic media item including at least a first edge, and a visual effect, the visual effect obscures at least a first portion of the stereoscopic media item and extends inward from at least a first edge of the first representation of the stereoscopic media item toward the interior of the first representation of the stereoscopic media item.
[0029] In some embodiments, a computer system is described. The computer system comprises one or more processors, the computer system communicating with a display generation component, and a memory for storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for displaying a user interface via the display generation component, the user interface including a first representation of a stereoscopic media item, the first representation of a stereoscopic media item including at least a first edge, and a visual effect, the visual effect obscuring at least a first portion of the stereoscopic media item and extending inward from at least a first edge of the first representation of the stereoscopic media item toward the interior of the first representation of the stereoscopic media item.
[0030] In some embodiments, a computer system is described. The computer system communicates with a display generation component, and the computer system, via the display generation component, displays a user interface that includes a first representation of a stereoscopic media item, the first representation of the stereoscopic media item including at least a first edge, and a visual effect that obscures at least a first portion of the stereoscopic media item and extends inwardly from at least the first edge of the first representation of the stereoscopic media item into the interior of the first representation of the stereoscopic media item, for displaying the user interface.
[0031] In some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs including instructions to display, via the display generation component, a user interface that includes a first representation of a stereoscopic media item, the first representation of the stereoscopic media item including at least a first edge, and a visual effect that obscures at least a first portion of the stereoscopic media item and extends inwardly from at least the first edge of the first representation of the stereoscopic media item into the interior of the first representation of the stereoscopic media item.
[0032] Note that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and in particular, many additional features and advantages will be apparent to those skilled in the art in view of the drawings, the specification, and the claims. Further, note that the language used herein has been selected solely for readability and for the purpose of explanation, and not for the purpose of defining or limiting the subject matter of the invention.
Brief Description of the Drawings
[0033] To better understand the various embodiments described, the following “Modes for Carrying Out the Invention” should be referenced in conjunction with the following drawings, and similar reference numbers throughout the following drawings refer to the corresponding parts.
[0034] [Figure 1] A block diagram showing the operating environment of a computer system for providing an augmented reality experience, according to several embodiments.
[0035] [Figure 2] A block diagram showing a controller for a computer system configured to manage and adjust a user's augmented reality experience, according to several embodiments.
[0036] [Figure 3] This is a block diagram showing display generation components of a computer system configured to provide a user with visual components of an augmented reality experience, according to several embodiments.
[0037] [Figure 4] This is a block diagram showing a hand tracking unit for a computer system configured to capture user gesture input, according to several embodiments.
[0038] [Figure 5] This is a block diagram showing an eye-tracking unit for a computer system configured to capture user eye-gaze input, according to several embodiments.
[0039] [Figure 6] This is a flowchart illustrating a Glint-assisted eye-tracking pipeline in several embodiments.
[0040] [Figure 7A]This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7B] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7C] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7D] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7E] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7F] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7G] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7H] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7I] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7J] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7K] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7L] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7M] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments. [Figure 7N] This document illustrates exemplary techniques for interacting with media items and user interfaces, according to several embodiments.
[0041] [Figure 8] This is a flowchart illustrating various embodiments of how to interact with media items and user interfaces.
[0042] [Figure 9] This is a flowchart illustrating various embodiments of how to interact with media items and user interfaces.
[0043] [Figure 10] This is a flowchart illustrating various embodiments of how to interact with media items and user interfaces.
[0044] [Figure 11A] This document illustrates exemplary techniques for displaying media items, according to several embodiments. [Figure 11B] This document illustrates exemplary techniques for displaying media items, according to several embodiments. [Figure 11C] This document illustrates exemplary techniques for displaying media items, according to several embodiments. [Figure 11D] This document illustrates exemplary techniques for displaying media items, according to several embodiments. [Figure 11E] This document illustrates exemplary techniques for displaying media items, according to several embodiments. [Figure 11F] This document illustrates exemplary techniques for displaying media items, according to several embodiments.
[0045] [Figure 12] This is a flowchart of a method for displaying media items according to several embodiments. [Modes for carrying out the invention]
[0046] This disclosure relates to user interfaces that provide users with augmented reality (XR) experiences, in several embodiments.
[0047] The systems, methods, and GUIs described herein improve user interface interactions with virtual / augmented reality environments in multiple ways.
[0048] In some embodiments, one or more visual characteristics of displayed content are modified based on the location of the user's gaze. For example, if the user gazes at a first media item in the media library, the visual characteristics of the first media item are modified; if the user gazes at a second media item in the media library, the visual characteristics of the second media item are modified. The location and / or position of the user's gaze is detected using sensors and / or cameras (e.g., sensors and / or cameras integrated with a head-mounted device or installed away from the user (e.g., in an XR room)) as opposed to, for example, a touch-sensitive surface or other physical controller. Modifying the visual characteristics of displayed content based on the user's gaze provides the user with visual feedback regarding the state of the computer system (e.g., that the computer system has detected the user's gaze at a particular position in the user interface). Modifying the visual characteristics of displayed content based on the user's gaze also allows the user to efficiently interact with displayed content in two or more contexts without visually cluttering a display with multiple controls, improving the interactivity of the user interface (e.g., reducing the number of inputs required to achieve a desired result).
[0049] In some embodiments, a computer system allows a user to zoom in on a user interface based on where the user's gaze is directed. For example, if the user gazes at a first position on the user interface while giving a zoom-in command (e.g., one or more hand gestures), the computer system zooms in on the user interface using the first position as the center point of the zoom operation; and if the user gazes at a second position on the user interface while giving a zoom-in command, the computer system zooms in on the user interface using the second position as the center point of the zoom operation. Zooming in on a user interface based on the user's gaze provides the user with visual feedback about the state of the computer system (e.g., that the computer system has detected the user's gaze at a particular position on the user interface), assisting the user in providing appropriate and / or correct input. Zooming in on a user interface based on the user's gaze also allows the user to efficiently interact with content displayed in two or more contexts without visually cluttering a display with multiple controls, improving the interactivity of the user interface (e.g., reducing the number of inputs required to achieve a desired result).
[0050] In some embodiments, media items are displayed differently (e.g., with different visual effects) based on whether the media item contains a particular type of depth information (e.g., whether the media item is a stereoscopic capture). For example, if the media item is a stereoscopic capture, it is displayed with a first set of visual characteristics (e.g., displayed within a three-dimensional shape, displayed with multiple layers, displayed with refraction and / or blurred edges). If the media item is not a stereoscopic capture, it is displayed using a second set of visual characteristics (e.g., displayed within a two-dimensional shape, displayed as a single layer, displayed without refraction and / or blurred edges). Displaying media items differently based on whether they contain a particular type of depth information provides the user with visual feedback about the state of the computer system (e.g., indicating to the user whether the currently displayed media item contains a particular type of depth information).
[0051] In some embodiments, the visual effect is displayed as an overlay on a representation of a previously captured media item (e.g., a previously captured stereoscopic media item). For example, the visual effect may be displayed on one or more edges of a representation of a previously captured media item and may extend inward toward the center of the representation of the previously captured media item. The visual effect has a visual property (e.g., amount of blur) whose value and / or intensity decreases as the visual effect extends toward the center of the representation of the media item. Displaying the visual effect on one or more edges of a representation of a media item helps reduce the amount of visual discomfort (e.g., window violation) that the user may experience while viewing a representation of a previously captured media item.
[0052] Figures 1 to 6 illustrate exemplary computer systems for providing an XR experience to a user. Figures 7A to 7N show exemplary techniques for interacting with media items and user interfaces in several embodiments. Figure 8 is a flowchart of how to interact with media items and user interfaces in various embodiments. Figure 9 is a flowchart of how to interact with media items and user interfaces in various embodiments. Figure 10 is a flowchart of how to interact with media items and user interfaces in various embodiments. The user interfaces in Figures 7A to 7N are used to illustrate the processes in Figures 8 to 10. Figures 11A to 11F show exemplary techniques for displaying media items in several embodiments. Figure 12 is a flowchart of how to display media items in various embodiments. The user interfaces in Figures 11A to 11F are used to illustrate the process in Figure 12.
[0053] The processes described below enhance the usability of the device and make the user device interface more efficient (for example, by helping the user provide appropriate input and reducing user errors when operating / interacting with the device) through various technologies, including providing the user with improved visual feedback, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional controls displayed, performing an operation without requiring further user input when a set of conditions is met, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving memory space, reducing the amount of window violations the user experiences, and / or additional technologies. These technologies also reduce power consumption and improve the battery life of the device by enabling the user to use the device more quickly and efficiently. Saving battery power, and therefore weight, improves the ergonomics of the device. These technologies also enable real-time communication, allow the use of fewer and / or less accurate sensors, resulting in more compact, lighter, and less expensive devices, and enabling the device to be used in a variety of lighting conditions. These technologies reduce energy consumption and thereby reduce the heat emitted by the device, which is especially important for wearable devices that can become uncomfortable for the user to wear if they generate excessive heat, even if the device is well within the operating parameters for its components.
[0054] Furthermore, in any method described herein that is conditional on one or more conditions being met in one or more steps, it should be understood that the method described can be repeated in multiple iterations such that all the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires that a first step be performed if a condition is met, and a second step be performed if the condition is not met, a person skilled in the art will understand that the steps described in the claim are repeated in a specific order until the conditions are met and then not met. Thus, a method described in one or more steps that depends on one or more conditions being met can be rewritten as a method that is repeated until each of the conditions described in the method is met. However, this is not required for a claim of a system or computer-readable medium that includes instructions for performing a conditional operation based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency has been met without explicitly repeating the steps of the method until all the conditions that the steps of the method are conditional on are met. Those skilled in the art will also understand that, as with a method having conditional steps, a system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.
[0055] In some embodiments, as shown in Figure 1, the XR experience is provided to the user via an operating environment 100 which includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or remote server), display generation components 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the display generation component 120 (for example, within a head-mounted device or handheld device).
[0056] When describing an XR experience, various terms are used to refer individually to several related but distinct environments that the user perceives and / or interacts with (for example, using inputs detected by the computer system 101, which causes the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101 that generates the XR experience). The following is a subset of these terms.
[0057] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the help of electronic systems. Examples of physical environments, such as a physical park, include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses of sight, touch, hearing, taste, and smell.
[0058] Augmented Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through an electronic system. In XR, a subset of a person's bodily movements or their representations are tracked, and accordingly, one or more properties of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics. For example, an XR system may detect a person's head rotation and, accordingly, adjust the graphic content and sound field presented to the person in a similar manner to how such views and sounds would change in a physical environment. In some circumstances (e.g., for reasons of accessibility), adjustments to the properties(s) of virtual objects(s) in the XR environment may be made in response to representations of bodily movements (e.g., voice commands). A person may perceive and / or interact with XR objects using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person can perceive and / or interact with audio objects that create a 3D or spatial audio environment, providing the perception of point audio sources in 3D space. In another example, audio objects may enable audio transparency, selectively incorporating ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may perceive and / or interact with only audio objects.
[0059] Examples of XR include virtual reality and mixed reality.
[0060] Virtual reality: A virtual reality (VR) environment refers to a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with virtual objects in a VR environment through a simulation of their presence within the computer-generated environment and / or through a simulation of a subset of their physical movement within the computer-generated environment.
[0061] Mixed Reality: A mixed reality (MR) environment is a simulated environment designed to incorporate sensory input or its representation from a physical environment, in addition to including computer-generated sensory input (e.g., virtual objects), in contrast to a virtual reality (VR) environment designed to rely entirely on computer-generated sensory input. On a virtual continuum, a mixed reality environment is any location between, but not encompassing, the complete physical environment at one end and the virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Also, some electronic systems for presenting an MR environment may track location and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical articles or their representations from the physical environment). For example, the system may take movement into account so that a virtual tree appears stationary relative to the physical ground.
[0062] Examples of mixed reality include augmented reality and augmented virtual reality.
[0063] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on or onto a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display that allows a person to directly view the physical environment. The system may also be configured to present virtual objects on the transparent or translucent display, thereby allowing a person to use the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system composites the image or video with the virtual objects and presents the composite on the opaque display. A person uses this system to perceive the virtual objects superimposed on the physical environment by indirectly viewing the physical environment through the image or video of the physical environment. As used herein, a video of the physical environment shown on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into or onto the physical environment, so that a person can use the system to perceive the virtual objects superimposed on the physical environment. An augmented reality environment also refers to an imitation environment in which the representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may transform one or more sensor images to plane a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be transformed by graphically modifying (e.g., enlarging) a portion of it, so that the modified portion is a non-photorealistic altered version of the original captured image. As yet another example, the representation of the physical environment may be transformed by graphically removing or obscuring a portion of it.
[0064] Augmented Virtuality (AV): An augmented virtuality (AV) environment refers to a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. These sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park might have virtual trees and virtual buildings, but people with faces might be realistically reproduced from images of real people. Another example is that a virtual object might adopt the shape or color of a physical article captured by one or more imaging sensors. A further example is that a virtual object might adopt shadows that correspond to the position of the sun in the physical environment.
[0065] Viewpoint-locked virtual objects: A virtual object is viewpoint-locked when the computer system displays the virtual object in the same location and / or position within the user's view, even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked in the forward direction of the user's head (e.g., the user's viewpoint is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's viewpoint remains fixed even if the user's gaze moves, without moving the user's head. In embodiments where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object displayed in the upper-left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) will continue to be displayed in the upper-left corner of the user's viewpoint even if the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position in which a viewpoint-locked virtual object is displayed from the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so that the virtual object is also referred to as a "head-locked virtual object."
[0066] Environment-Locked Virtual Objects: A virtual object is environment-locked (or "world-locked") when a computer system displays it at a location and / or position in the user's viewpoint that is based on (e.g., selected by reference to and / or fixed to) a location and / or object in a three-dimensional environment (e.g., a physical or virtual environment). As the user's viewpoint shifts, the location and / or object in the environment relative to the user's viewpoint changes, and as a result, the environment-locked virtual object will appear at a different location and / or position in the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear centered in the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes left-leaning in the user's viewpoint (e.g., the tree's position in the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear left-leaning in the user's viewpoint. In other words, the location and / or position in which an environment-locked virtual object is displayed in the user's viewpoint depends on the location and / or object's position and / or orientation in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a fixed location in the physical environment and / or a coordinate system fixed to an object) to determine the position in which the environment-locked virtual object is displayed from the user's viewpoint. The environment-locked virtual object can be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object) or to a moving part of the environment (e.g., a vehicle, animal, person, or a representation of a part of the user's body that moves independently of the user's viewpoint, such as the user's hands, wrists, arms, or feet), so that the virtual object moves as the viewpoint or the part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.
[0067] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed tracking behavior, reducing or delaying the movement of the environment-locked or viewpoint-locked virtual object in response to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed tracking behavior, the computer system detects movement of the reference point that the virtual object is following (e.g., a part of the environment, a viewpoint, or a point fixed to the viewpoint, such as a point between 5 and 300 cm from the viewpoint) and intentionally delays the movement of the virtual object. For example, when the reference point (e.g., a part of the environment or viewpoint) moves at a first velocity, the virtual object is moved by the device so as to remain locked to the reference point, but at a second velocity slower than the first velocity (e.g., the virtual object begins to catch up to the reference point until the reference point stops or slows down). In some embodiments, when a virtual object exhibits delayed tracking behavior, the device ignores small amounts of movement of the reference point (e.g., ignoring movement of the reference point that is below a threshold amount, such as movement between 0 and 5 degrees or movement between 0 and 50 cm). For example, when the reference point (e.g., the part of the environment or viewpoint from which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a different viewpoint or part of the environment from which the virtual object is locked), and when the reference point (e.g., the part of the environment or viewpoint from which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a different viewpoint or part of the environment from which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a "delayed tracking" threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object that maintains a substantially fixed position with respect to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or forward / behind the position of the reference point).
[0068] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system to provide audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. The head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to the person's eye. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto the person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface.In some embodiments, the controller 110 is configured to manage and coordinate the XR experience for the user. In some embodiments, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to Figure 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some embodiments, the controller 110 is communicably coupled to a display generation component 120 (e.g., an HMD, display, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within a housing (e.g., a physical housing) of one or more of the display generation components 120 (e.g., a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0069] In some embodiments, the display generation component 120 is configured to provide the user with an XR experience (e.g., at least the visual components of an XR experience). In some embodiments, the display generation component 120 includes a preferred combination of software, firmware, and / or hardware. The display generation component 120 is described in more detail below with reference to Figure 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.
[0070] According to some embodiments, the display generation component 120 provides the user with an XR experience while the user is virtually and / or physically present in the scene 105.
[0071] In some embodiments, the display generation component is mounted on a part of the user's body (e.g., their head or hand). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some embodiments, the handheld device is optionally placed in a housing mounted on the user's head. In some embodiments, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content when the user is not wearing or holding the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interaction with XR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD where the interaction occurs in the space in front of the HMD and the XR content response is displayed through the HMD. Similarly, a user interface showing interaction with XR content triggered based on the movement of a handheld or tripod-mounted device relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may be implemented similarly to an HMD where the movement is triggered by the movement of the HMD relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).
[0072] While relevant features of the operating environment 100 are shown in Figure 1, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the exemplary embodiments disclosed herein.
[0073] Figure 2 is a block diagram of an example of the controller 110 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0074] In some embodiments, one or more communication buses 204 include circuits that interconnect system components and control communication between system components. In some embodiments, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0075] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some embodiments, memory 220, or the non-temporary computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and XR experience module 240.
[0076] The operating system 230 handles various basic system services and includes instructions for performing hardware-dependent tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for each group of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0077] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1, and optionally from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0078] In some embodiments, the tracking unit 242 is configured to map scene 105 and track the position / location of at least the display generation component 120 relative to scene 105 in Figure 1, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for this purpose, as well as heuristics and metadata for this purpose. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand relative to scene 105 in Figure 1, relative to the display generation component 120, and / or relative to a coordinate system defined for the user's hand. The hand tracking unit 244 is described in more detail below with respect to Figure 4. In some embodiments, the eye-tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hands)) or to XR content displayed via the display generation component 120. The eye-tracking unit 243 is described in more detail below with reference to Figure 5.
[0079] In some embodiments, the adjustment unit 246 is configured to manage and adjust the XR experience presented to the user by the display generation component 120 and optionally by one or more of the output devices 155 and / or peripheral devices 195. For this purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0080] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 248 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0081] While the data acquisition unit 241, tracking unit 242 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 are shown as residing on a single device (e.g., a controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (including, for example, an eye-tracking unit 243 and a hand-tracking unit 244), adjustment unit 246, and data transmission unit 248 may be located in separate computing devices.
[0082] Furthermore, Figure 2 is intended to illustrate the functionality of various features that may be present in a particular embodiment, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules separately shown in Figure 2 can be implemented in a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0083] Figure 3 is a block diagram of an example of a display generation component 120 according to several embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more suitable embodiments of the embodiments disclosed herein. For that purpose, in some non-limiting examples, the display generation component 120 (e.g., HMD) may include one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, Bluetooth, ZiGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional in-facing and / or out-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0084] In some embodiments, one or more communication buses 304 include circuits for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).
[0085] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to the user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), micro-electromechanical system (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, a display generation component 120 (e.g., HMD) includes a single XR display. In another embodiment, the display generation component 120 includes an XR display for each of the user's eyes. In some embodiments, one or more XR displays 312 can present MR or VR content. In some embodiments, one or more XR displays 312 can present MR or VR content.
[0086] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally, at least a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would view if a display generation component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.
[0087] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some embodiments, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and XR presentation module 340.
[0088] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to the user via one or more XR displays 312. For this purpose, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.
[0089] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1. For this purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0090] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. For this purpose, in various embodiments, the XR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0091] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (for example, a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate augmented reality) based on media content data. For this purpose, in various embodiments, the XR map generation unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0092] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral devices 195. For this purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.
[0093] While the data acquisition unit 342, XR presentation unit 344, XR map generation unit 346, and data transmission unit 348 are shown as existing on a single device (e.g., the display generation component 120 in Figure 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, XR presentation unit 344, XR map generation unit 346, and data transmission unit 348 may be located in separate computing devices.
[0094] Furthermore, Figure 3 is intended to illustrate the functionality of various features that may be present in a particular implementation, in contrast to the structural schematics of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 3 can be realized within a single module, and the various functions of a single functional block can be performed by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some embodiments, will partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0095] Figure 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (Figure 1) is controlled by the hand tracking unit 244 (Figure 2) to track the location / position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to the scene 105 of Figure 1 (e.g., relative to a part of the physical environment surrounding the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system defined for the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0096] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures a hand image with sufficient resolution to allow for the distinction of fingers and their respective positions. The image sensor 404 can typically capture images of other parts of the user's body, or images of the entire body, and may have either a zoom function or a dedicated sensor with high magnification to capture an image of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures a 2D color video image of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors that capture the physical environment of the scene 105, or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as input to the controller 110.
[0097] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), which drives the display generation components 120 accordingly. For example, a user can interact with the software running on the controller 110 by moving their hand 406 to change the orientation of their hand.
[0098] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spot in the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a given reference plane at a specific distance from the image sensor 404. In this disclosure, it is assumed that the image sensor 404 defines a set of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods such as stereoscopic imaging or time-of-flight measurement based on one or more cameras or other types of sensors.
[0099] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software running on the processor in the image sensor 404 and / or controller 110 processes the 3D map data to extract patch descriptors of the hand within these depth maps. Based on a previous learning process, the software matches these descriptors against patch descriptors stored in the database 408 to estimate the hand pose in each frame. The pose typically includes the 3D location of the user's wrist and fingertips.
[0100] The software can also analyze the trajectory of the hand and / or fingers across multiple frames in a sequence to identify gestures. The posture estimation function described herein may be interleaved with the motion tracking function, so that patch-based posture estimation is performed only once every two (or more) frames, while tracking is used to detect changes in posture that occur over the remaining frames. Posture, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify the image presented on the display generation component 120, or perform other functions, depending on the posture and / or gesture information.
[0101] In some embodiments, the gesture includes an air gesture. An air gesture is a gesture detected by the user without (or independently of) touching an input element that is part of a device (e.g., a computer system 101, one or more input devices 125, and / or a hand tracking device 140), and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined posture by a predetermined amount and / or speed, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body).
[0102] In some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures, as in some embodiments, performed by moving one or more of the user's fingers relative to other fingers or parts of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, an air gesture is a gesture detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device), and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture involving movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture involving rotation of a part of the user's body by a predetermined speed or amount).
[0103] In some embodiments where the input gesture is an air gesture (i.e., without physical contact with an input device that provides the computer system with information about which user interface element is the target of user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor over a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of user input (e.g., in the case of direct input, as described below). Thus, in implementations involving air gestures, the input gesture is the detected attention (e.g., gaze) to the user interface element in combination (e.g., simultaneously) with the movement of the user's fingers (one or more) and / or hand to perform pinch and / or tap input, as described in more detail below.
[0104] In some embodiments, input gestures directed towards a user interface object are performed directly or indirectly by reference to the user interface object. For example, user input is performed directly towards the user interface object in response to the user performing an input gesture with their hand at a position corresponding to the user interface object's position in a three-dimensional environment (e.g., determined based on the user's current viewpoint). In some embodiments, the input gesture is performed indirectly towards the user interface object according to the user performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in a three-dimensional environment, while detecting the user's attention (e.g., gaze) to the user interface object. For example, in the case of a direct input gesture, the user can direct their input towards the user interface object by initiating the gesture at or near a position corresponding to the user interface object's display position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm from the optional outer edge or optional central portion). In the case of indirect input gestures, the user can direct their input towards the user interface object by paying attention to the user interface object (for example, by gazing at the user interface object), and while paying attention to the options, the user initiates the input gesture (for example, at any position detectable by the computer system) (for example, at a position that does not correspond to the display position of the user interface object).
[0105] In some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch and tap inputs for interacting with virtual or mixed reality environments, as in some embodiments. For example, the pinch and tap inputs described later are performed as air gestures.
[0106] In some embodiments, a pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture involves moving two or more fingers of a hand to touch each other, i.e., including an optional interruption (e.g., within 0 to 1 second) immediately after the touch. A long pinch gesture that is an air gesture involves moving two or more fingers of a hand to touch each other for at least a threshold time amount (e.g., at least 1 second) before detecting an interruption of contact between them. For example, a long pinch gesture includes the user holding a pinch gesture (e.g., if two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected directly and consecutively (e.g., within a predetermined period of time) to each other. For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks contact between two or more fingers), and then performs a second pinch input within a predetermined period (e.g., within 1 second or 2 seconds) after releasing the first pinch input.
[0107] In some embodiments, an air gesture, a pinch-and-drag gesture, includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in relation to (e.g., after) a drag input that changes the user's hand position from a first position (e.g., a drag initiation position) to a second position (e.g., a resistance termination position). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers) to terminate the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together and touches them to each other, and then moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by the user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from the first position to the second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of the user's hands. For example, an input gesture includes two (e.g., or more) pinch inputs performed in relation to each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using the user's first hand, and a second pinch input performed using the other hand (e.g., a second hand of the user's hands) in relation to performing a pinch input using the first hand. In some embodiments, movement between the user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).
[0108] In some embodiments, a tap input performed as an air gesture (e.g., directed towards a user interface element) includes the movement of one or more of the user's fingers toward the user interface element, the movement of the user's hand toward the user interface element with the user's fingers (one or more) optionally extended toward the user interface element, a downward movement of the user's fingers (e.g., mimicking a mouse click or a tap on a touchscreen), or other default movements of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand that performs the tap gesture movement away from the user's viewpoint and / or toward the object that is the target of the tap input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand that performs the tap gesture (e.g., away from the user's viewpoint and / or the end of the movement toward the object that is the target of the tap input, a reversal of the direction of the finger or hand movement, and / or a reversal of the direction of acceleration of the finger or hand movement).
[0109] In some embodiments, the user's attention is determined to be directed towards a part of the three-dimensional environment based on the detection of a gaze directed towards that part of the three-dimensional environment (optionally, without requiring any other conditions). In some embodiments, for the device to determine that the user's attention is directed towards a part of the three-dimensional environment, the device determines that the user's attention is directed towards a part of the three-dimensional environment based on the detection of a gaze directed towards a part of the three-dimensional environment, with one or more additional conditions such as the gaze being directed towards the part of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the part of the three-dimensional environment, and / or the gaze being directed towards a part of the three-dimensional environment. If one of the additional conditions is not met, the device determines that the user's attention is not directed towards the part of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).
[0110] In some embodiments, the detection of a ready state configuration of the user or a part of the user is detected by the computer system. The detection of a ready state configuration of the hand is used by the computer system as an indication that the user is likely to be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap shape where one or more fingers are extended and the palm is facing away from the user), whether the hand is in a predetermined position relative to the user's line of sight (e.g., below the user's head, above the user's waist, or extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular way (e.g., moved towards the area in front of the user above the user's waist, below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of the user interface is responsive to attention (e.g., gaze) input.
[0111] In some embodiments, the software may be downloaded electronically to the controller 110, for example, over a network, or instead, it may be provided on a tangible non-temporary medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively or additionally, some or all of the computer's described functions may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although the controller 110 is shown in Figure 4, for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or in other ways. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device), or by any other suitable computerized device such as a game console or media player. The sensing function of the image sensor 404 can also be integrated into a computer or other computerized device controlled by the sensor output.
[0112] Figure 4 further includes schematic diagrams of depth maps 410 captured by image sensor 404 according to several embodiments. The depth map includes a matrix of pixels, each having a depth value, as described above. Pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The brightness of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from image sensor 404, with the gradation becoming richer as the depth increases. Controller 110 processes these depth values to identify and segment image components (i.e., adjacent pixel groups) that have the characteristics of a human hand. These characteristics may include, for example, the overall size, shape, and frame-to-frame movement of the depth map sequence.
[0113] Figure 4 also schematically shows the hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to several embodiments. In Figure 4, the hand skeleton 414 is superimposed on the hand background 416, which has been segmented from the original depth map. In some embodiments, the hand (e.g., finger joints, fingertips, center of the palm, end of the hand connected to the wrist), and optionally major feature points on the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these major feature points across multiple image frames are used by the controller 110 to determine, according to several embodiments, a hand gesture performed by the hand or the current state of the hand.
[0114] Figure 5 shows an exemplary embodiment of the eye-tracking device 130 (Figure 1). In some embodiments, the eye-tracking device 130 is controlled by an eye-tracking unit 243 (Figure 2) to track the position and movement of the user's gaze toward the scene 105 or toward the XR content displayed via the display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device positioned in a wearable frame, the head-mounted device includes both a component for generating XR content for user viewing and a component for tracking the user's gaze toward the XR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, the eye-tracking device 130 is optionally a separate device from the handheld device or XR chamber. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used with a display generation component that is mounted on the head or a display generation component that is not mounted on the head. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye-tracking device 130 is not a head-mounted device, but is optionally part of a non-head-mounted display generation component.
[0115] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames containing left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include, or be coupled to, one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects on a transparent or translucent display. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as holograms, so that the individual can use the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0116] As shown in Figure 5, in some embodiments, the eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) camera or a near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be directed toward the user's eye to receive reflected IR or NIR light from the light source directly from the eye, or alternatively, it may be directed toward a "hot" mirror positioned between the user's eye and a display panel that reflects IR or NIR light from the eye to the eye-tracking camera while allowing visual light to pass through. The eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by a separate eye-tracking camera and light source.
[0117] In some embodiments, the eye-tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye-tracking device for a specific operating environment 100, e.g., the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at the factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automated calibration process or a manual calibration process. The user-specific calibration process may include estimating the eye parameters of a particular user, e.g., pupil location, central visual location, optical axis, visual axis, interpupillary distance. According to some embodiments, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, the images captured by the eye-tracking camera can be processed using a glint-assisted method to determine the user's current visual axis and viewpoint relative to the display.
[0118] As shown in Figure 5, the eye-tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and an eye-tracking system which includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eyes(s) 592. The eye-tracking camera 540 is positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display or projector of a handheld device) and may be directed towards a mirror 550 that transmits visible light while reflecting IR or NIR light from the eye(s) 592 (e.g., as shown at the top of Figure 5), or may be directed towards the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of Figure 5).
[0119] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames of left and right display panels) and provides the frames 562 to the display 510. For various purposes, for example, when processing the frames 562 for display, the controller 110 uses gaze tracking input 542 from the eye-tracking camera 540. The controller 110 optionally uses a glint-assisted method or other appropriate method to estimate the user's viewpoint on the display 510 based on the gaze tracking input 542 obtained from the eye-tracking camera 540. The viewpoint estimated from the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0120] The following describes, but is not intended to be limiting, several possible use cases of the user's current gaze direction. As an exemplary use case, the controller 110 may render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content within the view based at least partially on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may capture the physical environment of the XR experience and orient an external camera to focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface in the environment that the user is currently viewing on the display 510. In another exemplary use case, the eyepiece 520 may be a focusing lens, and the controller uses eye-tracking information to adjust the focus of the eyepiece 520 so that the virtual object currently being viewed by the user has appropriate binocular coordination to match the convergence of the user's eye 592. The controller 110 can use the eye-tracking information to orient and adjust the focus of the eyepiece 520 so that the nearby object being viewed by the user appears at the correct distance.
[0121] In some embodiments, the eye-tracking device is part of a head-mounted device mounted on a wearable housing, which includes a display (e.g., display 510), two eyepieces (e.g., one or more eyepieces 520), an eye-tracking camera (e.g., one or more eye-tracking cameras 540), and a light source (e.g., a light source 530 (e.g., an IR LED or NIR LED)). The light source emits light (e.g., IR light or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in Figure 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.
[0122] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, thus not introducing noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given as examples and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is positioned on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0123] Embodiments of eye-tracking systems, such as those shown in Figure 5, can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0124] Figure 6 shows glint-assisted eye-tracking pipelines according to several embodiments. In some embodiments, the eye-tracking pipeline is implemented by a glint-assisted eye-tracking system (e.g., an eye-tracking device 130 as shown in Figures 1 and 5). The glint-assisted eye-tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the glint-assisted eye-tracking system tracks the pupil contour and glint in the current frame by using prior information from previous frames when analyzing the current frame. When not in tracking state, the glint-assisted eye-tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues in tracking state for the next frame.
[0125] As shown in Figure 6, the eye-tracking camera can capture left and right images of the user's left and right eyes. The captured images are then fed into the eye-tracking pipeline for processing, which begins at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images may be fed into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0126] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as shown in 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. If they are not successfully detected, the method returns to element 610 and processes the next image of the user's eyes.
[0127] At 640, if the process proceeds from element 610, the current frame is analyzed to track the pupil and glint based in part on previous information from the previous frame. At 640, if the process proceeds from element 630, the tracking state is initialized based on the detected pupil and glint in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results may be checked to determine whether a sufficient number of glints for pupil and gaze estimation are successfully tracked or detected in the current frame. At 650, if the results are unreliable, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze.
[0128] Figure 6 is intended to serve as an example of an eye-tracking technology that may be used in a particular implementation. As will be recognized by those skilled in the art, other eye-tracking technologies that currently exist or may be developed in the future may be used in computer system 101 to provide users with XR experiences in various embodiments, either in place of or in combination with the Glint-assisted eye-tracking technology described herein.
[0129] This disclosure describes various input methods for interaction with computer systems. Where one example is provided using one input device or method, and another example is provided using a different input device or method, each example may be compatible with the input device or method described in the other example, and their use should be considered optional. Similarly, various output methods for interaction with computer systems are described. Where one example is provided using one output device or method, and another example is provided using a different output device or method, each example may be compatible with the output device or method described in the other example, and their use should be considered optional. Similarly, various methods for interaction with virtual or mixed reality environments via computer systems are described. Where one example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, each example may be compatible with the method described in the other example, and their use should be considered optional. Therefore, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment. User interface and related processes
[0130] Here, we focus on embodiments of user interfaces ("UI"), as well as related processes that can be implemented on a computer system such as a portable multifunction device or head-mounted device communicating with display generation components and one or more input devices.
[0131] Figures 7A to 7N illustrate exemplary techniques for interacting with media items and user interfaces according to several embodiments. Figures 8 to 10 are flowcharts of exemplary methods 800, 900, and 1000 for interacting with media items and user interfaces, respectively. The user interfaces in Figures 7A to 7N are used to illustrate the processes described below, including the processes in Figures 8 to 10.
[0132] Figure 7A shows a computer system 700 which is a tablet including a display device (e.g., a touch-sensitive display) 702 and one or more input sensors 715 (e.g., one or more cameras, eye-tracking devices, hand movement tracking devices, and / or head movement tracking devices). In some embodiments described below, the computer system 700 is a tablet. In some embodiments, the computer system 700 is a smartphone, a wearable device, a wearable smartwatch device, a head-mounted system (e.g., a headset), or another computer system including and / or communicating with a display device (e.g., a display screen, a projection device, etc.). The computer system 700 is a computer system (e.g., computer system 101 in Figure 1).
[0133] In Figure 7A, the computer system 700 displays the media window 704 and the three-dimensional environment 706 via a display device (e.g., a touch-sensitive display or the display of an HMD) 702. In some embodiments, the three-dimensional environment 706 is displayed by the display (as shown in Figure 7A). In some embodiments, the three-dimensional environment 706 is a virtual environment image or picture (or video) of a physical environment captured by one or more cameras. In some embodiments, the three-dimensional environment 706 is visible to the user behind the media window 704 but is not displayed by the display. For example, in some embodiments, the three-dimensional environment 706 is a physical environment that is visible to the user behind the media window 704 (e.g., through a transparent display) but is not displayed by the display.
[0134] In Figure 7A, the electronic device 700 displays a media library user interface 708 within a media window 704. The media library user interface 708 includes representations of a plurality of media items 710A to 710E. In some embodiments, the plurality of media items include photographs and / or videos (e.g., stereoscopic photographs and / or videos and / or non-stereoscopic photographs and / or videos). In some embodiments, the plurality of media items include a plurality of media items and / or a subset of media items from a media library (e.g., device 700 and / or a media library associated with the user of device 700). In some embodiments, the plurality of media items 710A to 710E include one or more stereoscopic media items (e.g., stereoscopic photographs and / or stereoscopic videos) and one or more non-stereoscopic media items (e.g., non-stereoscopic photographs and / or non-stereoscopic videos). In some embodiments, the stereoscopic media items include a specific type of depth information, while the non-stereoscopic media items do not include a specific type of depth information. In some embodiments, a stereoscopic media item is a media item comprising at least two images captured simultaneously using two different cameras (or two different sets of cameras). In some embodiments, the stereoscopic media item is displayed to the user's first eye by displaying a first image captured by a first camera (or a first set of cameras) (e.g., without the first image being displayed to the user's second eye), and simultaneously displaying to the user's second eye a second image, different from the first image but captured simultaneously with the first image by a second camera (or a second set of cameras) (e.g., without the second image being displayed to the user's first eye). Non-stereoscopic images include, for example, still images and / or videos captured by a single camera.In some embodiments, a non-stereoscopic media item includes multiple images captured by multiple cameras (e.g., simultaneously by multiple cameras), but the different images are not presented simultaneously to different eyes of the user (e.g., the multiple images are combined into a single image presented to both eyes of the user). In some embodiments, a stereoscopic media item and / or a non-stereoscopic media item is captured by device 700 (e.g., using one or more cameras that are part of device 700 and / or communicate with device 700).
[0135] The media library user interface 708 includes options 712A to 712E. Option 712A is selectable to display a set of stereoscopic media items (e.g., without displaying non-stereoscopic media items). In some embodiments, option 712A is selectable to display a user interface that includes representations of one or more stereoscopic media items, without including representations of non-stereoscopic media items. Option 712B is selectable to display a set of aggregated content items that are generated (e.g., automatically) by aggregating multiple media items from the media library (e.g., a collection of media items associated with a particular time, event, and / or location). Option 712C is selectable to display the media library user interface 708. Option 712D is selectable to display albums of media items (e.g., collections of media items) (e.g., collections of manually curated media items and / or collections of automatically generated media items). Option 712E is selectable to initiate a process of performing a text search of media items (e.g., selectable to display a text search user interface).
[0136] In Figure 7A, the computer system 700 displays the media item 710A using a first set of visual characteristics. In Figure 7A, the computer system 700 detects (e.g., via sensor 715) that the user is gazing at the media item 710A, as indicated by gaze indication 714. In some embodiments, gaze indication 714 is not displayed by the computer system 700.
[0137] In Figure 7B, in response to the determination that the user has gazed at media item 710A for a threshold duration (for example, continuously without gazing at another media item), the computer system 700 modifies one or more visual characteristics of media item 710A and displays media item 710A using a second set of visual characteristics different from a first set of visual characteristics.
[0138] In some embodiments, modifying one or more visual properties of media item 710A and / or displaying media item 710A using a second set of visual properties includes separating and / or expanding elements of media item 710A along predetermined axes (e.g., changing the display of media item 710A from a two-dimensional object to a three-dimensional object). For example, in some embodiments, as part of modifying one or more visual properties of media item 710A, the computer system 700 expands the display of media item 710A along a separate axis of media item 710A (e.g., the z-axis of media item 710A) so that the user perceives media item 710A as a three-dimensional object with depth (e.g., thickness). In some embodiments, media item 710A is expanded along a separate axis of media item 710A based on a determination that media item 710A contains depth information (e.g., a particular type of depth information) (e.g., media item 710A is a stereoscopic media item) (e.g., depending on the determination). For example, if media item 710A is extended along its individual axes, a first set of elements within media item 710A will be displayed at a first depth (e.g., a first z-position along the z-axis of the extended media item 710A), and a second set of elements within media item 710A will be displayed at a second depth different from the first depth (e.g., a second z-position on the z-axis of the extended media item 710A).
[0139] In some embodiments, when a user shifts their viewpoint while gazing at a media item 710A (for example, by moving and / or rotating their head while wearing an HMD, and / or by moving their body relative to the display device 702), the content within the media item 710A shifts in accordance with the user's shift in viewpoint. In some embodiments, the parallax effect is achieved, for example, such that layers of media item 710A further from the user move more slowly (or by less) than layers of media item 710A closer to the user.
[0140] In some embodiments, modifying one or more visual characteristics of media item 710A and / or displaying media item 710A using a second set of visual characteristics includes pushing media item 710A backward (e.g., away from the user). For example, in some embodiments, in Figure 7B, in response to a determination that the user has gazed at media item 710A for a threshold duration, the computer system 700 maintains the displayed depth of media items 710B-710E while pushing media item 710A back. In some embodiments, media item 710A is pushed backward based on a determination that media item 710A contains a particular type of depth information (e.g., is a stereoscopic media item).
[0141] In some embodiments, modifying one or more visual characteristics of media item 710A and / or displaying media item 710A using a second set of visual characteristics includes autoplaying media item 710A (e.g., automatically playing it without further user input other than user gaze). For example, in some embodiments, if media item 710A is a video (e.g., stereoscopic or non-stereoscopic video), the computer system 700 starts playing the video content of media item 710A within the media library user interface 708 in response to a determination that the user has gazed at media item 710A for a threshold duration. In some embodiments, autoplaying media item 710A (e.g., in response to user gaze) includes applying a low-pass filter (e.g., a blur filter and / or a smoothing filter) when playing media item 710A within the media library user interface 708. In some embodiments, the low-pass filter is removed when playing a selected media item 710A (for example, within the selected media user interface 719, as described with reference to Figures 7M to 7N below). In some embodiments, autoplaying a media item 710A within the media library user interface 708 includes outputting the audio content of the media item 710A at a lower volume than when it is played in the selected media user interface 719 (for example, within the selected media user interface 719, as described with reference to Figures 7M to 7N below).
[0142] In some embodiments, when the user's gaze moves away from media item 710A (for example, to another media item 710B-710E), the computer system 700 stops displaying media item 710A using a second set of visual characteristics (for example, stops automatic playback of media item 710A) (for example, returns from displaying media item 710A using a second set of visual characteristics to displaying media item 710A using a first set of visual characteristics and / or displaying media item 710A using a third set of visual characteristics).
[0143] In Figure 7C, while displaying media item 710A using a second set of visual characteristics, the computer system 700 detects that the user's gaze has moved to a different position within the media library user interface 708 corresponding to media item 710C, as indicated by the gaze indication 714. In Figure 7C, upon determining that the user's gaze has moved to media item 710C, the computer system 700 stops displaying media item 710A using the second set of visual characteristics and displays media item 710A using the first set of visual characteristics (as shown in Figure 7A). Furthermore, upon determining that the user's gaze has moved to media item 710C (and optionally maintained over a threshold duration), the computer system 700 transitions media item 710C from a state where it is displayed using a third set of visual characteristics (e.g., Figures 7A-7B) to a state where it is displayed using a fourth set of visual characteristics (e.g., Figure 7C). In some embodiments, the third set of visual characteristics is the same as the first set of visual characteristics. In some embodiments, the fourth set of visual characteristics is the same as the second set of visual characteristics. In some embodiments, the examples described above regarding the transition from the first set of visual characteristics to the second set of visual characteristics can be applied to the transition from the third set of visual characteristics to the fourth set of visual characteristics.
[0144] In Figure 7C, while displaying media item 710C using a fourth set of visual characteristics within the media library user interface 708, the computer system 700 detects a user gesture 716 (e.g., via a sensor 715). In some embodiments, user gesture 716 and the various user gestures described herein include one or more touch inputs (e.g., a display device 702 such as a touch-sensitive display). In some embodiments, user gesture 716 and the various user gestures described herein include one or more air gestures (e.g., movement of one or more parts of the user's body in a predetermined manner) (e.g., captured by a camera and / or body movement sensors (e.g., a hand movement sensor and / or a head movement sensor)). In some embodiments, an air gesture is a gesture detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device), and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture involving movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture involving rotation of a part of the user's body by a predetermined speed or amount). In some embodiments, the user gesture 716 includes a pinch-and-drag gesture (e.g., a pinch-and-drag air gesture) in which the user forms a predetermined pinch shape with their hand and drags (e.g., moves) the hand forming the predetermined pinch shape in a particular direction (e.g., an upward pinch-and-drag gesture) (e.g., while the hand maintains the predetermined pinch shape). In some embodiments, the predetermined pinch shape is a shape in which the user brings a region adjacent to the tip of the thumb of one hand into contact with the tips of one or more other fingers of the same hand.In some embodiments, as shown in Figure 7C, while the media library user interface 708 is being displayed, the computer system 700 detects the user's hand making a default shape (e.g., a default pinch shape) (e.g., via sensor 715) and detects the user's hand moving in a default direction (e.g., upward) (e.g., while maintaining the default shape).
[0145] In Figure 7D, in response to the detection of a user gesture 716, the computer system 700 scrolls the media library user interface 708 upward. In some embodiments, the media library user interface 708 scrolls upward based on the direction and magnitude of the user gesture 716 (for example, based on the direction and magnitude of the movement of a pinch-and-drag gesture (for example, based on the direction of the user's pinched hand and how far the user's pinched hand moves)). As a result, the media library user interface 708 no longer contains media items 710A-710B, moves media items 710C-710E upward, and displays media item 710F. Furthermore, in Figure 7D, the computer system 700 detects that the user's gaze is directed towards media item 710F, as indicated by the gaze indication 714. In Figure 7D, in response to the determination that the user's gaze has been maintained over the media item 710F for a threshold duration, the computer system 700 transitions from displaying the media item 710F having a fifth set of visual characteristics (e.g., two-dimensional and / or unextended) to displaying the media item 710F having a sixth set of visual characteristics (e.g., three-dimensional and / or extended), which is different from the fifth set of visual characteristics. In some embodiments, the fifth set of visual characteristics is the same as the first and / or third set of visual characteristics. In some embodiments, the sixth set of visual characteristics is the same as the second and / or fourth set of visual characteristics. In some embodiments, the examples described above regarding the transition from the first set of visual characteristics to the second set of visual characteristics may also apply to the transition from the fifth set of visual characteristics to the sixth set of visual characteristics.
[0146] In Figure 7D, while the user's gaze is maintained over the media item 710F, the computer system 700 detects a user gesture 718. In some embodiments, the user gesture 718 corresponds to a zoom-in command. In some embodiments, the user gesture 718 includes a two-handed pinch-out gesture (e.g., a two-handed pinch-out air gesture) (e.g., a gesture in which two pinched hands are moved away from each other). In some embodiments, detecting a two-handed pinch-release gesture includes the computer system 700 detecting that the user's first hand has formed a predetermined shape (e.g., a predetermined pinch shape), the computer system 700 detecting that the user's second hand has formed a predetermined shape (e.g., a predetermined pinch shape), the computer system 700 detecting that at a first time point, the first hand forming the predetermined shape and the second hand forming the predetermined shape are at a first distance from each other, and subsequently, the computer system 700 detecting that at a second time point, the first hand forming the predetermined shape and the second hand forming the predetermined shape have moved to a second distance from each other, where the second distance is greater than the first distance. In some embodiments, the user gesture 718 includes a one-handed pinch gesture or a one-handed double pinch gesture (e.g., one hand of the user quickly performs two pinch gestures in succession (e.g., two pinch gestures within a threshold duration of each other)). In some embodiments, the computer system 700 detecting a one-handed pinch gesture includes the computer system 700 detecting a pinch in which two or more of the user's fingers move closer to each other until the two fingers are within a threshold distance of each other (for example, until the two fingers are closer to each other than the threshold distance (for example, until the two fingers touch each other)). In some embodiments, the computer system 700 detecting a one-handed double pinch gesture includes the computer system 700 detecting a first pinch in which two or more of the user's fingers move closer to each other until the two fingers are within a threshold distance of each other (for example, until the two fingers are closer to each other than the threshold distance (for example, until the two fingers touch each other)).After detecting that two fingers are separated, the computer system 700 detects a second pinch, which involves two or more of the user's fingers moving closer to each other until the two fingers are within a threshold distance of each other (for example, until the two fingers are closer to each other than the threshold distance (for example, until the two fingers touch each other)). In some embodiments, the one-handed double pinch gesture includes detecting the second pinch within a threshold period of the first pinch (for example, the second pinch is initiated and / or completed within a threshold period after the first pinch has been initiated and / or completed).
[0147] In Figure 7E, in response to the detection of a user gesture 718 corresponding to a zoom-in command while the user's gaze is directed towards media item 710F, the computer system 700 zooms in on the media library user interface 708 with media item 710F as the center and / or focus of the zoom-in operation. Therefore, in Figure 7E, media item 710F is displayed at a larger size than in Figure 7D. In some embodiments, if the user's gaze is directed towards a different media item, the zoom-in user command will zoom in on the different media item instead of media item 710F. In some embodiments, the magnitude of the zoom operation (e.g., how much the computer system 700 zooms in on the media library user interface 708) is determined based on the direction and / or magnitude of a two-handed pinch-out gesture (e.g., an air gesture) (e.g., how far the user's first pinched hand moves from the user's second pinched hand). In some embodiments, if the user gesture 718 is a one-handed pinch and / or one-handed double pinch gesture (e.g., an air gesture), the computer system 700 zooms in on the media library user interface 708 by a predetermined amount (e.g., a default amount). In some embodiments, when the computer system 700 is a head-mounted device, the computer system 700 separately increases the zoom level of the display of media 710F on two separate display devices of the computer system 700, with each display device of the computer system 700 corresponding to an individual eye of the user (e.g., each display device is visible to an individual eye of the user).
[0148] Returning to Figure 7D, while the user's gaze is maintained over the media item 710F, the computer system 700 detects a user gesture 718. In some embodiments, as described above with reference to Figure 7E, the user gesture 718 corresponds to a media item selection command rather than a zoom-in user command. In some embodiments, the media item selection command includes a one-handed pinch gesture (e.g., an air gesture) (e.g., a one-handed pinch gesture). In some embodiments, the media item selection command includes a two-handed pinch-out gesture (e.g., an air gesture) (e.g., a two-handed pinch-out gesture). In some embodiments, multiple gestures correspond to a single command (e.g., different gestures produce the same result).
[0149] In Figure 7F, in response to detecting a user gesture 718 corresponding to a media item selection command while the user's gaze is directed towards media item 710F, the computer system 700 discontinues displaying the media library user interface 708 in the media window 704 and displays the selected media item 710F in the media user interface 719 within the media window 704. In some embodiments, when the computer system 700 is a head-mounted device, the computer system 700 displays a first representation of media item 710F (e.g., a first perspective view of the environment contained within the media item 710F) to the user's first eye, and the computer system 700 displays a second representation of media item 710F (e.g., a second perspective view of the environment contained within the media item 710F) to the user's second eye, thereby allowing the user to view media item 710F with a stereoscopic depth effect.
[0150] In some embodiments, the user gesture 718 in Figure 7D corresponds to a media item selection command in the scenario shown in Figure 7F and includes one or more intermediate states. For example, if the media item selection command is a one-handed pinch gesture, the user can slowly and / or gradually bring their index finger closer to their thumb to perform the pinch gesture. In some embodiments, as the user progresses through the media item selection command, one or more isolated elements of the media item 710F in Figure 7D move gradually closer to each other (for example, along a default axis (for example, along the z-axis)). In some embodiments, once the user gesture exceeds a completion threshold, the media item 710F is displayed in the selected media user interface 719 with one or more elements re-expanded (for example, as seen in Figure 7F). In some embodiments, if the user stops executing a media item selection gesture before reaching a completion threshold (e.g., the user stops a pinch gesture before reaching a completion threshold, and / or separates their index finger and thumb before reaching a completion threshold), the elements of media item 710F in Figure 7D are re-expanded and / or separated within the media library user interface 708. In some embodiments, as the user progresses through the media item selection command, other media items in Figure 7D (e.g., media items 710C-710E) are gradually blurred (e.g., they become more blurred as the media item selection command progresses towards a completion threshold). In some embodiments, certain types of user gestures result in the display of an intermediate stage between Figure 7D and Figure 7F (e.g., gradual blurring of media items, and / or gradual movement of media item layers and / or elements moving closer to each other), while other types of user gestures immediately trigger a transition from the media library user interface 708 in Figure 7D to the selected media user interface 719 in Figure 7F without displaying an intermediate stage.For example, in some embodiments, a fast one-handed pinch gesture (e.g., a one-handed pinch gesture completed within a threshold time) results in an immediate transition from the media library user interface 708 to the selected media user interface 719 without displaying an intermediate state, while a slow one-handed pinch gesture or a two-handed pinch-out gesture displays an intermediate state during the transition from the media library user interface 708 to the selected media user interface 719 based on the user gesture.
[0151] In Figure 7F, it can be seen that the three-dimensional environment 706 darkens when the selected media user interface 719 is displayed. Furthermore, the media window 704 is displayed along with one or more lighting effects 726-1 extending from the media window 704 (for example, into the three-dimensional environment 706). In some embodiments, when opening a media item (for example, when transitioning from the media library user interface 708 to the selected media user interface 719), the lighting effects extending from the media window 704 (e.g., 726-1) and the lighting effects applied to the three-dimensional environment 706 are gradually realized and / or displayed. For example, in some embodiments, from Figure 7D to Figure 7F, in response to a user gesture 718 (media item selection command), the size of the media item 710F is gradually increased to fill the media window 704, the three-dimensional environment 706 is gradually darkened, and the lighting effects 726-1 gradually extend from the media window 704 and their intensity gradually increases. Similarly, in some embodiments, when a media item is closed (for example, to transition from a selected media user interface 719 to a media library user interface 708), the lighting effect extending from the media window 704 and applied to the three-dimensional environment 706 is gradually removed (for example, the intensity and size of the lighting effect 726-1 gradually decrease, and the three-dimensional environment 706 gradually brightens). In some embodiments, if the computer system 700 is a head-mounted device, the computer system 700 changes the appearance of the lighting effect 726-1 in response to the user rotating their head and / or walking around the physical environment (for example, while wearing the computer system 700). In some embodiments, the lighting effect 726-1 includes a plurality of rays extending from the media window 704. In some embodiments, the visual characteristics of the plurality of rays (e.g., the length of one or more rays, the color (one or more) of one or more rays, and / or the brightness and / or intensity of one or more rays) are determined based on the visual content of the displayed media item (e.g., a media item displayed within a selected media user interface 719).In some embodiments, visual content at the edges of the media window 704 (e.g., visual content within a threshold distance of the edges and / or boundaries of the media window 704) is weighted more heavily than content inside the edges of the media window 704 when determining the visual characteristics of the multiple rays. In some embodiments, one or more rays have variable length, variable color, variable brightness, and / or variable intensity (e.g., based on the visual content of the displayed media item). Thus, in some embodiments, when the visual content of the displayed media item changes (e.g., the displayed media item changes from one media item to another, the displayed media item is a video with visual content that changes as the video is played, and / or the user zooms in or out within the media item), the lighting effect 726-1, including the multiple rays, changes accordingly. In some embodiments, if the displayed media item is a video, the changes in the lighting effect 726-1 (e.g., the multiple rays) are smoothed over time. In some embodiments, the lighting effect 726-1 includes rays extending from the front of the media window 704 (e.g., the front layer 720) (e.g., extending forward) and rays extending from the back of the media window 704 (e.g., the back layer 722) (e.g., extending backward).
[0152] In some embodiments, the lighting effect 726-1 extends from the media window 704 regardless of whether the displayed media item (e.g., a media item displayed within a selected media user interface 719) is a stereoscopic or non-stereoscopic media item. In some embodiments, the rays extending from the media window 704 differ based on whether the displayed media item is a stereoscopic media item (e.g., as shown in Figures 7F, 7G, 7H, and 7N) or a non-stereoscopic media item (e.g., as shown in Figures 7J to 7M). For example, one or more algorithms used to determine the visual characteristics of the rays differ based on whether the displayed media item is a stereoscopic or non-stereoscopic media item. In some embodiments, the rays (or algorithms used to determine the visual characteristics of the rays) do not differ based on whether the displayed media item is a stereoscopic or non-stereoscopic media item.
[0153] In some embodiments, the visual characteristics of the media item, media window 704, and / or the three-dimensional environment 706 depend on whether the displayed media item is a stereoscopic or non-stereoscopic media item. In Figure 7F, the displayed media item is a stereoscopic media item. Based on the determination that the displayed media item is a stereoscopic media item, the computer system 700 displays the media item along with a set of visual characteristics corresponding to a stereoscopic media item. For example, in the illustrated embodiment, the media window 704 displays the displayed media item three-dimensionally, and the elements of the media item are unfolded along the axes. In the illustrated embodiment, the media window 704 is shown with a front layer 720, a back layer 722, and one or more intermediate layers 724 between the front layer 720 and the back layer 722. In some embodiments, when the computer system 700 is a head-mounted device, the computer system 700 displays different perspectives of the media item displayed in the media window 704 in response to the user repositioning themselves in the physical environment and / or rotating their head (for example, while wearing the computer system 700 on their head).
[0154] In some embodiments, according to the determination that the displayed media item is a stereoscopic media item, the displayed media item is displayed within a continuous three-dimensional shape having continuous edges (for example, the media window 704 is a continuous three-dimensional shape having continuous edges). Figure 7I shows a side profile view of an exemplary embodiment of a continuous three-dimensional shape with continuous edges, where the front layer 720 is shown as the left surface of the shape in Figure 7I, and the back layer 722 is shown as the right surface of the shape in Figure 7I. In some embodiments, as shown in Figure 7I, the three-dimensional shape of the media window 704 has a curved front and / or curved back. In some embodiments, the three-dimensional shape of the media window 704 has a refraction edge (for example, a blurred edge and / or a glassy edge). In contrast, in some embodiments, according to the determination that the displayed media item is a non-stereoscopic media item, the displayed media item is displayed within and / or as a two-dimensional object (as shown in Figure 7J) (for example, the media window 704 is a two-dimensional object and / or shape). In some embodiments, when the computer system 700 is a head-mounted device, the computer system 700 displays different perspective views of the three-dimensional shape of a stereoscopic media item in response to the user repositioning themselves in the physical environment and / or rotating their head (for example, while the user is wearing the computer system 700 on their head).
[0155] In Figure 7F, the computer system 700 also displays a sharing option 728A, time and location information 728B, and a close option 728C. The sharing option 728A is selectable to display one or more options for sharing the displayed media item. The time and location information 728B displays time and location information 728B corresponding to the displayed media item (e.g., the time and location of the capture of the displayed media item). The close option 728C is selectable to close the displayed media item (e.g., it is selectable to stop displaying the selected media user interface 719 and redisplay the media library user interface 708). In some embodiments, one or more controls, such as the sharing option 728A, date and time information 728B, and close option 728C, are displayed according to the determination that the displayed media item is a stereoscopic media item (e.g., outside the media window 704 and / or in a manner that does not overlap and / or overlay the stereoscopic media item 710F on which one or more controls are displayed).
[0156] In some embodiments, when a media item is displayed in a selected media user interface 719, the media item is displayed with vignetting applied to the media item (e.g., darkened corners and / or edges) regardless of whether the displayed media item is a stereoscopic or non-stereoscopic media item. In some embodiments, when a media item is displayed in a selected media user interface 719, the media item is optionally displayed in an immersive state. In some embodiments, if a first set of conditions is met, the media item is displayed in an immersive state, and if the first set of conditions is not met, the media item is not displayed in an immersive state. In various embodiments, the first set of conditions includes, for example, one or more of the following: a determination that a first user setting (e.g., an immersive viewing setting) is enabled and / or disabled; a determination that a first media item is of a particular type (e.g., stereoscopic compared to non-stereoscopic) (e.g., according to the determination that the first media item is a stereoscopic media item); and / or a determination that one or more user inputs of a particular type (e.g., one or more gestures (e.g., one or more air gestures)) have been detected. In some embodiments, the non-immersive state corresponds to a first angular size, and the immersive state corresponds to a second angular size that is larger than the first angular size. In other words, in the immersive state, the displayed media item is displayed at a larger angular display size, and in the non-immersive state, the displayed media item is displayed at a smaller angular display size. In some embodiments, if the media item is displayed in an immersive state and the computer system 700 is a head-mounted device, the computer system 700 displays different perspectives of the media item based on the user repositioning themselves in the physical environment and / or rotating their head (for example, while the user is wearing the computer system 700 on their head).
[0157] In Figure 7F, while displaying media item 710F on the selected media user interface 719, the computer system 700 detects a user gesture 732. Various different scenarios corresponding to different types of user gestures for user gesture 732 are described below in turn.
[0158] In the first scenario, user gesture 732 corresponds to a media item close command, for example, to stop displaying the selected media user interface 719 and / or to redisplay the media library user interface 708. In some embodiments, the media item close command includes a pinch-and-drag gesture (e.g., an air gesture) (e.g., a pinch-and-drag gesture in a default direction (e.g., a pinch-and-drag gesture downwards)). In some embodiments, a pinch-and-drag gesture in a first direction (e.g., left) corresponds to a next media item command to display a subsequent media item in the selected media user interface 719 (e.g., the immediately following media item in an ordered sequence of media items), a pinch-and-drag gesture in a second direction (e.g., a second direction opposite to the first direction) (e.g., right) corresponds to a previous media item command to display a previous media item in the selected media user interface 719 (e.g., the immediately preceding media item in an ordered sequence of media items), and a pinch-and-drag gesture in a third direction (e.g., down) corresponds to a command to close a media item. As described above, in some embodiments, the detection of a pinch-and-drag gesture (e.g., a pinch-and-drag air gesture) by the computer system 700 includes the computer system 700 detecting that the user's hand forms a predetermined shape (e.g., a predetermined pinch shape), and the computer system 700 detecting movement of the hand forming the predetermined shape in a particular direction (e.g., while the user's hand is forming and / or maintaining the predetermined shape). In some embodiments, the detection of a pinch-and-drag gesture by the computer system 700 further includes the computer system 700 detecting movement of the hand forming the predetermined shape in a particular direction over at least a threshold distance.
[0159] In the second scenario, user gesture 732 corresponds to a zoom-in command received while the user's gaze is detected at the lower right corner of the media window 704, as indicated by gaze indication 714-1 (e.g., a one-handed double-pinch gesture or a two-handed pinch-out gesture, various embodiments of which are described above with reference to Figure 7D). In the third scenario, user gesture 732 corresponds to a zoom-in command received while the user's gaze is detected at the upper left corner of the media window 704, as indicated by gaze indication 714-2. In the second scenario, in response to user gesture 732 corresponding to the zoom-in command while the user's gaze is directed at the lower right corner of the media window 704, the computer system 700 zooms in to the lower right corner of the media window 704, as shown in Figure 7G. In the third scenario, in response to a user gesture 732 corresponding to a zoom-in command while the user's gaze is directed towards the upper-left corner of the media window 704, the computer system 700 zooms in to the upper-left corner of the media window 704, as shown in Figure 7H.
[0160] Figure 7J shows a selected media user interface 719 displaying a non-stereoscopic media item 731. In Figure 7J, the media item 731 is a panoramic non-stereoscopic image. In Figure 7J, the three-dimensional environment 706 is darkened as in Figure 7F, and the lighting effect 730-1 extends from the media window 704. Similar to the lighting effect 726-1 described above with reference to Figure 7F, in some embodiments, the lighting effect 730-1 is determined based on the visual content displayed in the media window 704 (e.g., the visual content of the displayed media item 731). In some embodiments, the lighting effect 730-1 differs from the lighting effect 726-1 based on the difference in the visual content displayed within the media window 704, but the algorithm used to generate and / or determine the lighting effect 730-1 is the same as the algorithm used to generate and / or determine the lighting effect 726-1. In some embodiments, based on the fact that media item 731 is a non-stereoscopic media item and media item 710F was a stereoscopic media item, the algorithm used to generate and / or determine lighting effect 730-1 is different from the algorithm used to generate and / or determine lighting effect 726-1. In some embodiments, the various characteristics of lighting effect 726-1 described above with reference to Figure 7F are applicable to lighting effect 730-1 and other lighting effects described herein (e.g., 730-2, 730-3, 730-4, 726-2, 726-3, and / or 726-4). In some embodiments, when the computer system 700 is a head-mounted device, the computer system 700 does not display various perspective views of the content contained in media item 731 in response to the user walking around the physical environment and / or turning the user's head (e.g., while the user is wearing the computer system 700 on their head) as a result of media item 731 being a non-stereoscopic media item.
[0161] In some embodiments, the specific visual characteristics of the media window 704 and the selected media user interface 719 in Figure 7J differ from those in Figure 7F, based on the determination that the displayed media item 731 is a non-stereoscopic media item. For example, in some embodiments, based on the determination that the displayed media item 731 in Figure 7J is a non-stereoscopic media item, one or more controls, such as the sharing option 728A, time and location information 728B, and the close option 728C, are overlaid on the displayed media item 731. In another example, in some embodiments, based on the determination that the displayed media item 731 in Figure 7J is a non-stereoscopic media item, the media item 731 is displayed within and / or as a two-dimensional object (whereas the media item 710F in Figure 7F is displayed within a three-dimensional object). Other differences between stereoscopic and non-stereoscopic media items are described in more detail above with reference to Figure 7F.
[0162] In Figure 7J, while displaying a non-stereoscopic panoramic media item 731 within the selected media user interface 719, the computer system 700 detects a user gesture 733 (e.g., via the sensor 715). In the illustrated scenario, the user gesture 733 corresponds to a zoom-in command. In some embodiments, the zoom-in command includes a one-handed double-pinch gesture and / or a two-handed pinch-out gesture, as described above (see, for example, Figure 7D). Furthermore, in Figure 7J, while detecting the user gesture 733, the computer system 700 detects a user gaze directed towards a first position within the displayed media item 731, as indicated by the gaze indication 714.
[0163] In Figure 7K, in response to the detection of a user gesture 733 (e.g., a zoom-in command) while the user's gaze is directed to a first position within the displayed media item 731, the computer system 700 zooms in on the media item 731, using the first position as the center point of the zoom-in operation. Furthermore, in Figure 7K, in response to the detection of the user gesture 733 (e.g., a zoom-in command) and the determination that the displayed media item 731 is a panoramic media item, the computer system 700 enlarges the size of the media window 704 and also curves the media window 704. In Figure 7K, the lighting effect 730-2 extends from the media window 704. In some embodiments, the lighting effect 730-2 differs from the lighting effect 730-1 based on changes in the visual content displayed in the media window 704. In some embodiments, if the computer system 700 is a head-mounted device, the display of the media window 704 occupies a wider field of view of the user while the media window 704 is displayed with a curved appearance (as opposed to, for example, when the media window 704 is displayed without a curved appearance).
[0164] In Figure 7K, the computer system 700 detects a user gesture 734 (for example, via a sensor 715). In some embodiments, user gesture 734 is a continuation of the zoom-in command of user gesture 733 (for example, a continuation of a two-handed pinch-out gesture (for example, the various embodiments described above, see, for example, Figure 7D)). Furthermore, in Figure 7K, while detecting user gesture 734, the computer system 700 detects a user gaze directed towards a second position within the displayed media item, as indicated by a gaze indication 714.
[0165] In Figure 7L, in response to detecting a user gesture 734 (e.g., a zoom-in command (e.g., the various embodiments described above (see, for example, Figure 7D)) while the user's gaze is directed to a second position within the displayed media item, the computer system 700 further zooms in on the panoramic media item 731, using the second position as the center point of the zoom-in operation. Furthermore, in Figure 7L, in response to detecting the user gesture 734 (e.g., a zoom-in command) and determining that the displayed media item 731 is a panoramic media item, the computer system 700 further enlarges the size of the media window 704 and further curves the media window 704. In Figure 7L, the lighting effect 730-3 extends from the media window 704. In some embodiments, the lighting effect 730-3 differs from lighting effects 730-2 and 730-1 based on the changing visual content displayed in the media window 704.
[0166] Furthermore, in Figure 7L, as media item 731 is zoomed in further, it is zoomed in so sufficiently that certain content on the left and right sides of media item 731 is no longer visible within media window 704. In Figure 7L, according to the determination that the additional content not visible within media window 704 is to the left of media item 731, computer system 700 displays visual effect 736A on the left side of media window 704. Similarly, according to the determination that the additional content not visible within media window 704 is to the right of media item 731, computer system 700 displays visual effect 736B on the right side of media window 704. In some embodiments, visual effect 736A includes blurring and / or visually obscuring the left edge of media window 704, and visual effect 736B includes blurring and / or visually obscuring the right edge of media window 704.
[0167] In Figure 7L, while displaying the media item 731 in the selected media user interface 719, the computer system 700 detects a user gesture 738. In some embodiments, the user gesture 738 corresponds to a previous content item command (e.g., a pinch-and-drag gesture in a first (e.g., right) direction (e.g., the various embodiments described above (see, for example, Figures 7D and 7F))) and / or a next content item command (e.g., a pinch-and-drag gesture in a second (e.g., left) direction (e.g., the various embodiments described above (see, for example, Figures 7D and 7F))).
[0168] In Figure 7M, upon detecting a user gesture 738, the computer system 700 ceases displaying the media item 731 of Figure 7L within the selected media user interface 719 and displays the non-stereoscopic video media item 741 within the selected media user interface 719. In Figure 7M, one or more controls are displayed overlaid on the media item, according to the determination that the displayed media item 741 is a non-stereoscopic media item. One or more controls include, for example, a sharing option 728A, time and location information 728B, and a close option 728C, as well as a scrubber 740. The scrubber 740 can be operated by the user to navigate within the video media item 741 (for example, the user can interact with the scrubber 740) and also includes one or more playback controls such as a pause button, a play button, a fast forward button, and / or a rewind button. In Figure 7M, lighting effects 730-4 extend from the media window 704. As described above with reference to Figure 7F, in some embodiments, the lighting effects of the video are smoothed over time so that the lighting effects at individual points in time within the video are determined based on multiple video frames (e.g., multiple video frames immediately before and / or immediately after an individual point in time).
[0169] In Figure 7M, while displaying a non-stereoscopic video media item 741 on the selected media user interface 719, the computer system 700 detects a user gesture 742. In some embodiments, the user gesture 742 corresponds to a previous content item command (e.g., a pinch-and-drag gesture in a first (e.g., right) direction (e.g., the various embodiments described above (referring to Figures 7D and 7F))) and / or a next content item command (e.g., a pinch-and-drag gesture in a second (e.g., left) direction (e.g., the various embodiments described above (referring to Figures 7D and 7F))).
[0170] In Figure 7N, upon detecting a user gesture 742, the computer system 700 ceases displaying the non-stereoscopic video media item 741 of Figure 7M within the selected media user interface 719 and displays the stereoscopic video media item 743 within the selected media user interface 719. In Figure 7N, according to the determination that the displayed media item 743 is a stereoscopic media item, one or more controls (e.g., a sharing option 728A, time and location information 728B, a close option 728C, a scrubber 740) are displayed in the media window 704 (e.g., outside and / or without overlapping). In some embodiments, one or more controls (e.g., 728A, 728B, 728C, and / or 740) are displayed when the user's hand is in a first position (e.g., the up position) and hidden / not displayed when the user's hand is not in the first position (e.g., when the user's hand is in the down position). In some embodiments, one or more controls are displayed when the user's hands are in a first position, and are hidden when the user's hands are not in the first position, regardless of whether the displayed media item is a stereoscopic or non-stereoscopic media item. In some embodiments, one or more controls are displayed when the user's hands are in a first position, and are hidden when the user's hands are not in the first position, regardless of whether the displayed media item is a still image or a video.
[0171] Furthermore, according to the determination that the indicated media item 743 is a stereoscopic media item, the media item 743 is displayed in a three-dimensional shape having multiple layers (e.g., 720, 722, and 724). In Figure 7N, the lighting effect 726-4 extends from the media window 704. As described above with reference to Figures 7F and 7M, in some embodiments, the lighting effect of the video is smoothed over time such that the lighting effect at individual times in the video is determined based on multiple video frames (e.g., multiple video frames immediately before and / or immediately after individual times).
[0172] Further explanations regarding Figures 7A to 7N are given below with reference to methods 800, 900, and 1000 described with reference to Figures 8 to 10.
[0173] Figure 8 is a flowchart of exemplary method 800 for interacting with media items and user interfaces according to several embodiments. In some embodiments, method 800 is performed with a computer system (e.g., computer system 101 in Figure 1) (e.g., smartphone, smartwatch, tablet, and / or wearable device), mouse, keyboard, remote control, visual input device (e.g., camera), audio input device (e.g., microphone), and / or biometric sensor (e.g., fingerprint sensor, face recognition sensor, and / or iris recognition sensor)) that communicates with display generating components (e.g., 702) (e.g., display controller, touch-sensitive display system, display (e.g., integrated and / or connected), 3D display, transparent display, projector, and / or head-up display) and one or more input devices (e.g., 715) (e.g., touch-sensitive surface (e.g., touch-sensitive display)) (e.g., computer system 101 in Figure 1) (e.g., smartphone, smartwatch, tablet, and / or wearable device), mouse, keyboard, remote control, visual input device (e.g., camera), audio input device (e.g., microphone), and / or biometric sensor (e.g., fingerprint sensor, face recognition sensor, and / or iris recognition sensor)). In some embodiments, the method 800 is stored in a non-temporary (or temporary) computer-readable storage medium and on one or more processors 202 of the computer system 101 (for example, Figure 1It is controlled by instructions executed by one or more processors of a computer system, such as control 110. Some operations of method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0174] In some embodiments, a computer system (e.g., 700) displays a media library user interface (e.g., 708) (802) via a display generation component (e.g., 702) which includes representations (e.g., thumbnails and / or previews) of multiple media items (e.g., 710A to 710F) (e.g., images, photographs, and / or videos) including a representation of a first media item (e.g., 710A to 710F).
[0175] While the media library user interface is being displayed (804), the computer system detects, at a first time point, via one or more input devices, a user gaze (e.g., 714) corresponding to a first position in the media library user interface (806) (e.g., detecting and / or determining that the user is gazing at the first position in the media library user interface). In some embodiments, the first position in the media library user interface is either a position corresponding to a representation of a first media item or a position in the media library user interface that does not correspond to a representation of a first media item (e.g., a position corresponding to a representation of a second media item different from the first media item).
[0176] In response to detecting a user gaze corresponding to a first position within the media library user interface (808), the computer system changes the appearance of the representation of the first media item via a display generation component from being displayed in a first manner (e.g., using a first set of visual characteristics) to being displayed in a second manner different from the first manner (e.g., using a second set of visual characteristics) (810) (e.g., media item 710A in Figures 7A-7B and / or media item 710C in Figures 7B-7C).
[0177] The computer system detects (812) a user gaze corresponding to a second position in the media library user interface (e.g., 714 in Figure 7C) via one or more input devices (e.g., detect and / or determine that the user is gazing at the second position in the media library user interface) at a second time point following a first time point (e.g., after or while displaying a representation of the first media item in a second manner). In some embodiments, the second position in the media library user interface is a position corresponding to a representation of the first media item, or a position in the media library user interface that does not correspond to a representation of the first media item different from the first position (e.g., a position corresponding to a representation of a second media item different from the first media item).
[0178] In response to detecting a user gaze corresponding to a second position within the media library user interface (814), the computer system displays a representation of the first media item via a display generation component in a third manner different from the second manner (for example, using a third set of visual characteristics different from the second set of visual characteristics) (816) (for example, media item 710A in Figures 7B to 7C). In some embodiments, the third manner is the same as the first manner.
[0179] Detecting a user gaze corresponding to a first position within the media library user interface and changing the appearance of the representation of a first media item from being displayed in a first manner to being displayed in a second manner provides the user with visual feedback about the system state (e.g., that a user gaze corresponding to a first position has been detected), which provides improved visual feedback.
[0180] In some embodiments, the computer system displays a media library user interface (e.g., 708) (e.g., at a third time point after a second time point) (e.g., at a third time point after a second time point) via a display generation component, which includes representations (e.g., thumbnails and / or previews) of a plurality of media items (e.g., 710A to 710F) (e.g., images, photographs, and / or videos) including representations of a second media item different from a first media item. While displaying the media library user interface including a representation of the second media item, the computer system detects a user gaze (for example, at a fourth time point after the third time point) via one or more input devices that corresponds to a third position in the media library user interface (e.g., 714 in Figure 7C) (e.g., the second position, or a third position different from the first and / or second position) (e.g., a third position in the media library user interface corresponding to a representation of the second media item, or a third position in the media library user interface that does not correspond to a representation of the second media item (e.g., corresponding to a representation of a third media item different from the second media item)) (e.g., detect and / or determine that the user is gazing at the third position in the media library user interface). In response to detecting a user gaze corresponding to a third position within the media library user interface, the computer system changes the appearance of the representation of a second media item (e.g., 710C) from being displayed in a fourth manner (e.g., using a fourth set of visual characteristics) (in some embodiments, the fourth manner is the same as the first, second, and / or third manner) to being displayed in a fifth manner different from the fourth manner (e.g., media item 710C in Figures 7B-7C) (e.g., using a fifth set of visual characteristics different from the fourth set of visual characteristics). In some embodiments, the fifth manner is the same as the first, second, and / or third manner.
[0181] In some embodiments, the computer system detects a user gaze corresponding to a fourth position in the media library user interface that is different from the third position (e.g., a fourth position in the media library user interface that corresponds to a representation of a second media item, or a fourth position in the media library user interface that does not correspond to a representation of a second media item (e.g., a representation of a third media item different from the second media item)) via one or more input devices (e.g., at a fifth time point after a fourth time point) (e.g., after or while displaying a representation of the second media item in a fifth manner). (e.g., detecting and / or determining that the user is gazing at the fourth position in the media library user interface). In response to detecting a user gaze corresponding to the fourth position in the media library user interface, the computer system displays a representation of the second media item in a sixth manner different from the fifth manner (e.g., using a sixth set of visual characteristics different from the fifth set of visual characteristics) via a display generation component. In some embodiments, the sixth manner is the same as the fourth manner.
[0182] In response to detecting a user gaze corresponding to a third position within the media library user interface, changing the appearance of the representation of a second media item from being displayed in a fourth manner to being displayed in a fifth manner provides the user with visual feedback about the system's state (e.g., that a user gaze corresponding to a third position has been detected), which provides improved visual feedback.
[0183] In some embodiments, a first media item (e.g., 710A) includes a plurality of elements, including a first element and a second element. In some embodiments, representing the first media item in a first manner includes displaying the first and second elements within a first two-dimensional plane (e.g., media item 710A in Figure 7A) (e.g., a two-dimensional plane including a first dimension and a second dimension different from the first dimension (e.g., perpendicular to the first dimension)) (e.g., displaying the first and second elements within a two-dimensional object and / or as part of a single two-dimensional object). In some embodiments, representing the first media item in a second manner includes separating the first and second elements along a third dimension (e.g., a third dimension perpendicular to the first two-dimensional plane) that extends outside the first two-dimensional plane (e.g., media item 710A in Figure 7B).
[0184] In some embodiments, the media library user interface defines a first plane having an x-axis and a y-axis perpendicular to the x-axis (for example, the media library user interface includes at least one plane which defines the x-axis and y-axis) (in some embodiments, representations of multiple media items are displayed on the first plane (for example, within the first plane)), the first media item includes multiple elements including a first element and a second element, and displaying the representation of the first media item in a first manner includes displaying the multiple elements at a first position on the z-axis, where the z-axis is perpendicular to both the x-axis and the y-axis (for example, the representation of the first media item is displayed as a two-dimensional object) (for example, multiple within the first media item The elements are displayed on a single two-dimensional plane (e.g., a first plane) (in some embodiments, when displayed in the first manner, the representation of the first media item is displayed as a two-dimensional planar object presented within the first plane) (in some embodiments, the first plane, defined by the media library user interface, is positioned at a first position on the z axis). Displaying the representation of the first media item in the second manner involves simultaneously displaying the first element at a second position on the z axis (in some embodiments, the second position on the z axis is the same as or different from the first position on the z axis), and the second element is at a third position different from the second position on the z axis.
[0185] In some embodiments, a first media item includes a third element, and displaying a representation of the first media item in a second manner includes simultaneously displaying the first element at a second position on the z-axis, the second element at a third position on the z-axis different from the second position, and the third element at a fourth position on the z-axis different from the second and third positions.
[0186] In response to detecting a user gaze corresponding to a first position within the media library user interface, extending the element to a third dimension, thereby changing the appearance of the representation of the first media item from being displayed in a first manner to being displayed in a second manner, provides the user with visual feedback about the state of the system (e.g., that a user gaze corresponding to a first position has been detected), which provides improved visual feedback.
[0187] In some embodiments, displaying a representation of a first media item in a first manner includes displaying the representation of the first media item in a first position relative to the user. In some embodiments, displaying a representation of a first media item in a second manner includes displaying the representation of the first media item in a second position relative to the user that is further away from the user than the first position (for example, as described with reference to media items 710A to 710F in Figures 7A to 7D) (for example, the representation of the first media item is pushed backward and / or away from the user).
[0188] In some embodiments, displaying a media library user interface includes displaying a representation of a first media item and a representation of a third media item that is different from the representation of the first media item, and displaying a representation of the first media item in a second manner includes displaying a representation of the first media item in a second position relative to the user while maintaining the position of the third media item relative to the user (for example, moving the representation of the first media item backward and / or away from the user while maintaining the position of the third media item relative to the user).
[0189] In some embodiments, the media library user interface defines a first plane having an x-axis and a y-axis perpendicular to the x-axis (for example, the media library user interface includes at least one plane which defines the x-axis and y-axis) (in some embodiments, representations of multiple media items are displayed on the first plane (for example, within the first plane)), and displaying the representation of the first media item in a first manner includes displaying the representation of the first media item at a first position on the z-axis, where the z-axis is perpendicular to both the x-axis and the y-axis, and furthermore, the front of the media library user interface defines the positive direction of the z-axis, and the media library user The back of the interface defines the negative direction of the z-axis (in some embodiments, the positive direction of the z-axis extends toward the user, and the negative direction of the z-axis extends toward the user), and displaying a representation of a first media item in a second manner involves displaying a representation of the first media item at a second position on the z-axis that is different from the first position, where the second position is a more negative z-axis position than the first position (for example, the first media item is pushed backward (e.g., further away from the user) in the second manner compared to the first manner) (in some embodiments, the representation of the first media item is displayed at the second position on the z-axis parallel to the first plane).
[0190] In some embodiments, while a representation of a first media item is displayed in a first manner, one or more (e.g., two or more) representations of other media items are displayed at a first position on the z-axis, and while a representation of the first media item is displayed in a second manner, one or more (e.g., two or more) representations of other media items are displayed (e.g., maintained) at the first position on the z-axis.
[0191] In response to detecting the user's gaze corresponding to a first position within the media library user interface, moving the representation of the first media item away from the user, thereby changing the appearance of the representation of the first media item from being displayed in a first manner to being displayed in a second manner, provides the user with visual feedback regarding the system state (e.g., that the user's gaze corresponding to the first position has been detected), which provides improved visual feedback.
[0192] In some embodiments, the first media item includes video content (e.g., movie content and / or video content). Displaying a representation of the first media item in a first manner includes displaying a static representation of the video content of the first media item (e.g., displaying a static thumbnail image representation of the first media item), and displaying a representation of the first media item in a second manner includes displaying playback of the video content of the first media item (e.g., playback of at least a subset of the video content of the first media item) (e.g., as described with reference to media items 710A to 710F in Figures 7A to 7D) (e.g., displaying video content and / or movie content of the first media item) (e.g., displaying playback of video content of the first media item within the media library user interface) (e.g., displaying playback of video content of the first media item within the media library user interface without displaying playback of video content of any other media item within the media library user interface). By detecting the user's gaze corresponding to a first position within the media library user interface and displaying playback of the video content of the first media item, the appearance of the representation of the first media item is changed from being displayed in a first manner to being displayed in a second manner, which provides the user with visual feedback regarding the system state (e.g., that the user's gaze corresponding to the first position has been detected), and this provides improved visual feedback.
[0193] In some embodiments, displaying a representation of the first media item in a third manner different from the second manner includes displaying a second static representation of the video content of the first media item (e.g., displaying a static thumbnail image representation of the first media item) (as described with reference to media items 710A to 710F in Figures 7A to 7D). In some embodiments, the second static representation is the same as or different from the static representation. In some embodiments, in response to detecting a user gaze corresponding to a second position in the media library user interface, the computer system stops playing the video content of the first item. In some embodiments, in response to detecting a user gaze corresponding to a second position in the media library user interface, the computer system changes the appearance of the representation of the second media item, which is different from the first media item, from displaying a static representation of the video content of the second media item (e.g., displaying a static thumbnail image representation of the video content of the first item) to displaying playback of the video content of the second media item (e.g., displaying the video content and / or movie content of the second media item). Detecting a user gaze corresponding to a second position within the media library user interface and changing the appearance of the representation of the first media item from being displayed in a second manner to being displayed in a third manner (for example, by displaying a static representation of the video content of the first media item) provides the user with visual feedback about the system state (e.g., that a user gaze corresponding to a second position has been detected), which provides improved visual feedback.
[0194] In some embodiments, displaying a representation of a first media item in a second manner includes applying a low-pass filter (e.g., a blur filter and / or a smoothing filter) to the video content of the first media item (for example, as described with reference to media items 710A to 710F in Figures 7A to 7D) (in some embodiments, displaying playback of the video content of the first media item includes this). In some embodiments, while displaying the media library user interface, the computer system detects one or more user inputs corresponding to the selection of a first media item (e.g., one or more touch inputs, one or more non-touch inputs, one or more gestures (e.g., one or more air gestures)) (e.g., a one-handed pinch gesture and / or a two-handed pinch-out gesture while the user's gaze is maintained over and / or directed towards the first media item), and in response to detecting one or more user inputs corresponding to the selection of a first media item, the computer system displays playback of the video content of the first media item without the low-pass filter applied. In some embodiments, upon detecting one or more user inputs corresponding to the selection of a first media item, the computer system displays a media player user interface different from the media library user interface (in some embodiments, the computer system discontinues displaying the media library user interface and / or discontinues displaying at least a portion of the media library user interface) and displays within the media player user interface the playback of the video content of the first media item without a low-pass filter applied.Detecting a user gaze corresponding to a first position within the media library user interface and changing the appearance of the representation of a first media item from being displayed in a first manner to being displayed in a second manner (for example, by displaying playback of the video content of the first media item with a low-pass filter applied) provides the user with visual feedback about the system state (for example, that a user gaze corresponding to a first position has been detected), which provides improved visual feedback.
[0195] In some embodiments, while displaying a representation of the first media item in a second manner, including displaying playback of the video content of the first media item, the computer system outputs the audio content of the first media item (e.g., audio content corresponding to the video content of the first media item) at a first volume level. While displaying the media library user interface (in some embodiments, while displaying playback of the video content of the first media item (e.g., within the media library user interface) and / or while outputting the audio content of the first media item at a first volume level), the computer system detects one or more user inputs corresponding to the selection of the first media item via one or more input devices (e.g., one or more touch inputs, one or more non-touch inputs, and / or one or more gestures (e.g., one or more air gestures)) (e.g., one-handed pinch gesture and / or two-handed pinch-out gesture while the user's gaze is maintained over and / or directed towards the first media item). Upon detecting one or more user inputs corresponding to the selection of a first media item, the computer system outputs the audio content of the first media item at a second volume level greater than the first volume level (for example, as described with reference to media items 710A to 710F in Figures 7A to 7D). In some embodiments, the one or more user inputs corresponding to the selection of the first media item include one or more air gestures.In some embodiments, detecting one or more user inputs corresponding to the selection of a first media item includes detecting a one-handed pinch gesture (e.g., a one-handed pinch air gesture) and / or a two-handed pinch-out gesture (e.g., a two-handed pinch-out air gesture) while the user's gaze is maintained over and / or directed towards the first media item (e.g., detecting a one-handed pinch gesture and / or a two-handed pinch-out gesture while detecting that the user's gaze is maintained over and / or directed towards the first media item). In some embodiments, the computer system outputs the audio content of the first media item at a second volume level while continuing to display playback of the video content of the first media item. In some embodiments, upon detecting one or more user inputs corresponding to the selection of a first media item, the computer system displays a media player user interface different from the media library user interface (in some embodiments, the computer system discontinues displaying the media library user interface and / or discontinues displaying at least a portion of the media library user interface), and displays playback of the video content of the first media item within the media player user interface, while outputting the audio content of the first media item at a second volume level. Automatically increasing the volume at which the audio content of the first media item is output in response to the opening and / or selection of the first media item allows the playback volume to be increased without the user providing further user input, thereby reducing the number of inputs required to perform the operation.
[0196] In some embodiments, displaying playback of the video content of a first media item is performed (e.g., without additional user input) in response to the detection of a user gaze corresponding to a first position in the media library user interface (for example, as described with reference to media items 710A to 710F in Figures 7A to 7D). In some embodiments, playback of the video content of a first media item is displayed according to the determination that the user gaze corresponding to the first position in the media library user interface has been maintained for a threshold duration. By displaying playback of the video content of a first media item in response to the detection of a user gaze corresponding to a first position in the media library user interface, the user can play the video content of the first media item without additional user input, thereby reducing the number of inputs required to perform the action.
[0197] In some embodiments, the first media item includes a plurality of elements, including a first element and a second element (for example, a plurality of elements positioned in a plurality of layers, including a first element in a first layer and a second element in a second layer (in some embodiments, each layer is positioned at a different position along the z-axis)). In some embodiments, while continuously detecting the user's gaze corresponding to a first position in the media library user interface, the computer system moves the first element relative to the second element via a display generation component when the user's viewpoint shifts relative to the first media item (for example, as described with reference to media items 710A to 710F in Figures 7A to 7D) (for example, moving the first element relative to the second element in response to a movement of the user's viewpoint, or in response to a movement of the first media item while the user's viewpoint remains in the same place). In some embodiments, the first element shifts relative to the second element by an amount based on the amount of change in the user's viewpoint relative to the first media item. In some embodiments, the first element shifts relative to the second element in a direction based on the direction of the user's gaze shift relative to the first media item. In some embodiments, displaying the first element moving relative to the second element is performed in response to the detection of one or more user inputs (e.g., one or more gestures (e.g., user's hand and / or head movements) (e.g., one or more air gestures) and / or one or more non-gesture inputs). Displaying the first element moving relative to the second element while detecting the user gaze corresponding to the first position in the media library user interface provides the user with visual feedback about the state of the system (e.g., that the user gaze corresponding to the first position has been detected), which provides improved visual feedback.
[0198] In some embodiments, the computer system detects a user gaze (e.g., 714) corresponding to a first position in the media library user interface (e.g., 708) via one or more input devices (e.g., after and / or while displaying a representation of a first media item in a third manner) (e.g., detecting and / or determining that the user is gazing at a first position in the media library user interface (e.g., a first position in the media library user interface corresponding to a representation of a first media item)). While continuously detecting the user's gaze corresponding to a first position within the media library user interface (for example, at a fourth time point after a third time point), the computer system detects one or more user gestures (e.g., 716 or 718) (e.g., movement of the user's hands, head, and / or body) (e.g., one or more air gestures (e.g., one-handed pinch air gesture, one-handed double pinch air gesture, pinch-and-drag air gesture, two-handed pinch air gesture, and / or two-handed pinch-out air gesture)) (in some embodiments, one or more non-gesture inputs) (in some embodiments, one or more user gestures corresponding to the selection of a first media item). In response to detecting one or more user gestures while continuously detecting the user gaze corresponding to a first position in the media library user interface, the computer system changes the appearance of the representation of the first media item via a display generation component from being displayed in a seventh manner (e.g., using a seventh set of visual characteristics) to being displayed in an eighth manner different from the seventh manner (e.g., using an eighth set of visual characteristics) (for example, changing the appearance of the media item 710F in response to a user gesture 718, as described with reference to Figures 7D to 7F). In some embodiments, the seventh manner is the same as the second manner.In some embodiments, in response to detecting one or more user gestures while continuously detecting a user gaze corresponding to a first position in the media library user interface, the computer system changes the appearance of the representation of the first media item from being displayed in a second manner to being displayed in an eighth manner different from the second manner. Changing the appearance of the representation of the first media item in response to detecting one or more user gestures while continuously detecting a user gaze corresponding to a first position in the media library provides the user with visual feedback about the state of the system (e.g., that the system has detected one or more user gestures while continuously detecting a user gaze corresponding to a first position), which provides improved visual feedback.
[0199] In some embodiments, one or more user gestures include pinch gestures (e.g., pinch air gestures) (e.g., one-handed pinch gestures or two-handed pinch gestures) (e.g., two fingers (e.g., two fingers on one hand or both hands) move from a first distance to a second distance relative to each other, where the second distance is smaller than the first distance). In some embodiments, the second distance is smaller than a threshold distance (e.g., the two fingers are moved to a position close enough to satisfy a distance threshold). In some embodiments, a pinch gesture corresponds to a selection gesture indicating a user selection of a first media item (e.g., a user selection of the first media item only, without the selection of any other media items), the display of the representation of the first media item in the eighth method indicates a user selection of the first media item (e.g., a user selection of the first media item only, without the selection of any other media items among the multiple media items), and the display of the representation of the first media item in the seventh method does not indicate a user selection of the first media item.
[0200] In some embodiments, changing the appearance of the representation of a first media item from being displayed in the seventh manner to being displayed in the eighth manner includes enlarging the size of the representation of the first media item. In some embodiments, changing the appearance of the representation of a first media item from being displayed in the seventh manner to being displayed in the eighth manner includes starting playback of the first media item (e.g., starting video playback of the first media item). In some embodiments, changing the appearance of the representation of a first media item from being displayed in the seventh manner to being displayed in the eighth manner includes displaying the representation of the first media item in a selected media user interface different from the media library user interface. In some embodiments, in response to detecting one or more user gestures, including a pinch gesture, the computer system stops displaying one or more representations of media items different from the representation of the first media item (e.g., representations of one or more media items that were displayed in the media library user interface).
[0201] Changing the appearance of the representation of a first media item in response to the detection of one or more user gestures while continuously detecting the user's gaze corresponding to a first position in the media library provides the user with visual feedback about the system state (e.g., that the system has detected one or more user gestures while continuously detecting the user's gaze corresponding to a first position), which provides improved visual feedback.
[0202] In some embodiments, a first media item (e.g., 710A-710F) includes a plurality of elements, including a first element and a second element. In some embodiments, displaying a representation of the first media item in a seventh manner includes displaying the first element in a first position and displaying the second element in a second position closer to the user of the computer system than the first position. In some embodiments, one or more user gestures (e.g., one or more air gestures) include a first gesture (e.g., a first air gesture). In some embodiments, in response to detecting a first gesture while continuously detecting the user's gaze corresponding to a first position in the media library user interface, the computer system moves the second element relative to the first element (as described with reference to Figures 7D-7F, for example) to reduce the distance between the first and second elements. In some embodiments, the first gesture is a pinch gesture (e.g., a pinch air gesture) (e.g., a gesture in which the first finger moves closer to the second finger), and the second element is moved along the axis relative to the first element to reduce the distance between the first and second elements along the axis, based on the magnitude of the pinch gesture (e.g., based on the amount of movement of one finger relative to another finger).
[0203] In some embodiments, the media library user interface defines a first plane having an x-axis and a y-axis perpendicular to the x-axis (for example, the media library user interface includes at least one plane which defines the x-axis and y-axis) (in some embodiments, representations of multiple media items are displayed on the first plane (for example, within the first plane)), the first media item includes multiple elements including a first element and a second element, and displaying a representation of the first media item in the seventh manner includes simultaneously displaying the first element at a first position on the z-axis and the second element at a second position different from the second position on the z-axis, one or more user gestures include a first gesture, and in response to detecting a first gesture while continuing to detect a user gaze corresponding to a first position in the media library user interface, the computer system moves the second element along the z-axis relative to the first element based on the magnitude of the first gesture to reduce the distance between the first and second elements along the z-axis. In some embodiments, the first gesture is a pinch gesture (e.g., a pinch air gesture) (e.g., a gesture in which the first finger moves closer to the second finger), and the second element is moved along the z-axis relative to the first element so as to decrease the distance between the first and second elements along the z-axis, based on the magnitude of the pinch gesture (e.g., based on the amount of movement of one finger relative to the other finger).
[0204] Moving a second element relative to a first element in response to detecting a first gesture while continuously detecting the user's gaze corresponding to a first position provides the user with visual feedback about the system's state (e.g., the system detecting a first gesture while continuously detecting the user's gaze corresponding to a first position), which provides improved visual feedback.
[0205] In some embodiments, one or more user gestures include a second gesture (e.g., a second air gesture) performed following a first gesture. In some embodiments, in response to detecting a second gesture while continuing to detect a user gaze corresponding to a first position in the media library user interface, the computer system moves the second element relative to the first element (e.g., moving the second element along a given axis) to increase the distance between the first and second elements (e.g., based on the magnitude of the second gesture), as described with reference to Figures 7D-7F. In some embodiments, moving the second element relative to the first element to increase the distance between the first and second elements includes moving the first element to a first position on an axis (e.g., a given axis (e.g., an axis extending toward the user of the computer system)) and / or moving the second element to a second position on the axis. In some embodiments, the second gesture is a pinch-out gesture (e.g., an air pinch-out gesture) (e.g., a gesture in which the first finger moves further away from the second finger). In some embodiments, the second element is moved relative to the first element (e.g., moved along an axis) to increase the distance between the first and second elements based on the magnitude of the pinch-out gesture (e.g., based on the amount one finger moves relative to the other finger). In some embodiments, moving the second element relative to the first element to increase the distance between the first and second elements is performed in accordance with the determination that the first gesture did not meet a completion threshold. Moving the second element relative to the first element in response to the detection of the second gesture while continuing to detect the user gaze corresponding to the first position in the media library provides the user with visual feedback regarding the state of the system (e.g., that the system detected the second gesture while continuing to detect the user gaze corresponding to the first position), which provides improved visual feedback.
[0206] In some embodiments, one or more user gestures (e.g., 718) include a third gesture (e.g., a third air gesture) performed following a first gesture, the third gesture indicating a user selection of a first media item (e.g., a user selection of the first media item without selecting any other media items (e.g., selection of only the first media item)). In some embodiments, the third gesture is a continuation of the first gesture (e.g., a continuation of the first gesture exceeding a completion threshold). In some embodiments, upon detection of a third gesture, the computer system increases the distance between the first and second elements (e.g., along a predetermined axis) (as described with reference to Figures 7D-7F, for example). In some embodiments, upon detection of a first gesture, the distance between the first and second elements is decreased to a first distance, and upon detection of a third gesture, the distance between the first and second elements is increased from the first distance to a second distance. In some embodiments, upon detecting a first gesture while continuously detecting the user's gaze corresponding to a first position in the media library user interface, the computer system moves the second element relative to the first element based on the magnitude of the first gesture, reducing the distance between the first and second elements to a first distance, and after moving the second element relative to the first element and reducing the distance between the first and second elements to a first distance, the computer system increases the distance between the first and second elements to a second distance, according to the determination that the first gesture meets a completion threshold. Increasing the distance between the first and second elements upon detecting a third gesture provides the user with visual feedback regarding the state of the system (e.g., that the system has detected a third gesture), which provides improved visual feedback.
[0207] In some embodiments, displaying a media library user interface (e.g., 708) includes displaying a representation of a first media item (e.g., 710A-710F) and simultaneously displaying a representation of a second media item that is different from the representation of the first media item (e.g., 710A-710F). One or more user gestures include a first selection gesture (e.g., an air gesture) (e.g., 718) (e.g., an air tap gesture (e.g., a tap gesture that does not touch any surface or object), a two-finger pinch gesture, a two-finger pinch-out gesture, a one-handed pinch gesture, and / or a one-handed pinch-out gesture corresponding to the selection of a first media item (e.g., the selection of media item 710F in Figure 7D) (e.g., the selection of a first media item without the selection of any other media items (e.g., the selection of only the first media item))). In some embodiments, detecting a first selection gesture includes detecting a default gesture (e.g., an air gesture) as well as detecting that the user's gaze is directed towards and / or maintained over the first media item (e.g., a one-handed pinch gesture and / or a two-handed pinch-out gesture while the user's gaze is directed towards and / or maintained over the first media item). In some embodiments, in response to detecting a first selection gesture corresponding to the selection of the first media item, the computer system visually obscures the representation of the second media item (e.g., blurs the representation of the second media item) (as described with reference to Figures 7D-7F). Visually obscuring the representation of the second media item in response to detecting a first selection gesture provides the user with visual feedback regarding the state of the system (e.g., that the system has detected a first selection gesture corresponding to the selection of the first media item), which provides improved visual feedback.
[0208] In some embodiments, while detecting a user gaze (e.g., 714) corresponding to a first position in the media library user interface (and optionally, while displaying a representation of the first media item in a second manner), the computer system detects a second set of one or more user gestures (e.g., 718) (e.g., movement of the user's hands, head (e.g., left and right), and / or body parts (e.g., one or more air gestures)) via one or more input devices. In response to detecting the second set of one or more user gestures, the computer system shifts the visual content of the representation of the first media item based on the second set of one or more user gestures (e.g., based on the direction of movement of the second set of one or more user gestures) (e.g., Figures 7C-7D). In some embodiments, in response to detecting the second set of one or more user gestures, the computer system displays additional content (e.g., additional content of the first media item and / or additional content of the media library user interface). Shifting the visual content of a representation of a first media item in response to the detection of a second set of one or more user gestures provides the user with visual feedback about the state of the system (e.g., that the system has detected a second set of one or more user gestures), which provides improved visual feedback. In some embodiments, when the computer system is a head-mounted device, the second set of one or more user gestures corresponds to the rotation of the user's head.
[0209] In some embodiments, while displaying the media library user interface (e.g., 708), the computer system detects a third set of one or more user gestures (e.g., 716) (e.g., one or more air gestures) via one or more input devices, including pinch gestures (e.g., pinch air gestures) (e.g., placing two fingers adjacent to each other and / or moving two fingers closer to each other) and drag gestures (e.g., drag air gestures) (e.g., moving a hand in a certain direction (e.g., moving a pinched hand)) (e.g., pinch and drag gestures). In response to detecting a third set of one or more user gestures, the computer system displays a scroll of the media library user interface (e.g., Figures 7C-7D). In some embodiments, before detecting a third set of one or more user gestures, the computer system displays a representation of a first set of media items, and displaying a scroll of the media library user interface includes displaying a representation of a second set of media items different from the first set of media items after detecting a third set of one or more user gestures. In some embodiments, displaying a scroll in the media library user interface includes displaying a scroll of a first set of media items in a certain direction until at least a subset of the first set of media items is no longer visible. Displaying a scroll in the media library user interface in response to the detection of a third set of one or more user gestures provides the user with visual feedback about the state of the system (e.g., that the system has detected a third set of one or more user gestures), which provides improved visual feedback.
[0210] In some embodiments, while displaying a media library user interface (e.g., 708), the computer system detects one or more user inputs (e.g., 718) via one or more input devices corresponding to the selection of a first media item (e.g., 710F) (e.g., one or more gesture inputs (e.g., one or more air gestures) and / or one or more non-gesture inputs (e.g., a one-handed pinch gesture and / or a two-handed pinch-out gesture while the user's gaze is directed at and / or maintained over the first media item) (e.g., selecting the first media item without selecting any other media items (e.g., selecting only the first media item)). In response to detecting one or more user inputs corresponding to the selection of the first media item, the computer system, via a display generation component, displays the first media item to a selected media user interface (e.g., 719) that is different from the media library user interface. Display. In some embodiments, the selected media user interface overlays the media library user interface. In some embodiments, in response to detecting one or more user inputs corresponding to the selection of a first media item, the computer system stops displaying the media library user interface or stops displaying at least a portion of the media library user interface. While displaying the first media item within the selected media user interface, the computer system detects a fourth set (e.g., 732) (e.g., one or more air gestures) of one or more user gestures via one or more input devices, including pinch gestures (e.g., pinch air gestures) (e.g., the arrangement of two fingers adjacent to each other and / or the movement of two fingers approaching each other) and drag gestures (e.g., drag air gestures) (e.g., the movement of a hand in a certain direction (e.g., the movement of a pinched hand)) (e.g., pinch and drag gestures in a first direction).In response to detecting a fourth set of one or more user gestures, the computer system discontinues displaying the first media item in the selected media user interface and, via a display generation component, displays a second media item different from the first media item in the selected media user interface (as described, for example, with reference to Figure 7F). In some embodiments, the computer system replaces the display of the first media item in the selected media user interface with the display of the second media item in the selected media user interface. In some embodiments, displaying the second media item in the selected media user interface is performed according to a determination that a drag gesture (e.g., a pinch-and-drag gesture) corresponds to a particular direction (e.g., left, right, up, and / or down). In some embodiments, a pinch-and-drag gesture in a direction different from a particular direction results in a different action (e.g., discontinuing the display of the selected media user interface and / or displaying the media library user interface). In response to detecting a fourth set of one or more user gestures, ceasing to display the first media item in the selected media user interface and displaying a second media item different from the first media item in the selected media user interface provides the user with visual feedback about the system state (e.g., that the system has detected a fourth set of one or more user gestures), which provides improved visual feedback.
[0211] In some embodiments, while displaying a media library user interface (e.g., 708), the computer system detects one or more user inputs (e.g., 718) via one or more input devices corresponding to the selection of a first media item (e.g., 710F) (e.g., one or more gesture inputs (e.g., one or more air gestures) and / or one or more non-gesture inputs (e.g., one-handed pinch gesture and / or two-handed pinch-out gesture while the user's gaze is directed at and / or maintained over the first media item) (e.g., selecting the first media item without selecting any other media items (e.g., selecting only the first media item)). In response to detecting one or more user inputs corresponding to the selection of a first media item, the computer system, via a display generation component, displays the first media to a selected media user interface (e.g., 719) that is different from the media library user interface. The system displays a media item (e.g., 710F). In some embodiments, the selected media user interface overlays the media library user interface. In some embodiments, in response to detecting one or more user inputs corresponding to the selection of a first media item, the computer system discontinues displaying the media library user interface or discontinues displaying at least a portion of the media library user interface. While the first media item is displayed within the selected media user interface, the computer system detects a fifth set (e.g., 732) (e.g., one or more air gestures) of one or more user gestures via one or more input devices, including pinch gestures (e.g., placing two fingers next to each other and / or moving two fingers closer to each other) and drag gestures (e.g., moving a hand in a certain direction (e.g., moving a pinched hand)) (e.g., pinch-and-drag gesture in a first direction).In response to detecting a fifth set of one or more user gestures, the computer system discontinues displaying the selected media user interface and displays the media library user interface (e.g., 708) via a display generation component (as described with reference to Figure 7F, for example) (e.g., in a location previously occupied by the selected media user interface). In some embodiments, discontinuing the display of the selected media user interface and displaying the media library user interface is performed according to a determination that a drag gesture (e.g., a pinch-and-drag gesture) corresponds to a particular direction (e.g., left, right, up, and / or down). In some embodiments, a pinch-and-drag gesture in a direction different from a particular direction results in a different action (e.g., replacing the display of a first media item in the selected media user interface with a different media item in the selected media user interface). Discontinuing the display of the selected media user interface and displaying the media library user interface in response to detecting a fifth set of one or more user gestures provides the user with visual feedback regarding the state of the system (e.g., that the system has detected a fifth set of one or more user gestures), which provides improved visual feedback.
[0212] In some embodiments, while displaying a media library user interface (e.g., 708), the computer system detects one or more user inputs (e.g., 718) via one or more input devices corresponding to the selection of a first media item (e.g., 710F) (e.g., one or more gesture inputs (e.g., one or more air gestures) and / or one or more non-gesture inputs (e.g., one-handed pinch gesture and / or two-handed pinch-out gesture while the user's gaze is directed at and / or maintained over the first media item) (e.g., selecting the first media item without selecting any other media items (e.g., selecting only the first media item)). In response to detecting one or more user inputs corresponding to the selection of a first media item, the computer system, via a display generation component, displays the first media to a selected media user interface (e.g., 719) that is different from the media library user interface. The system displays a media item (e.g., 710F). In some embodiments, the selected media user interface overlays the media library user interface. In some embodiments, in response to detecting one or more user inputs corresponding to the selection of a first media item, the computer system discontinues displaying the media library user interface or discontinues displaying at least a portion of the media library user interface. While the first media item is displayed within the selected media user interface, the computer system detects a sixth set of one or more user gestures (e.g., one or more air gestures) (e.g., 732) via one or more input devices, including pinch gestures (e.g., placing two fingers adjacent to each other and / or moving two fingers closer to each other) and drag gestures (e.g., moving a hand in a certain direction (e.g., moving a pinched hand)) (e.g., a pinch-and-drag gesture in a first direction).In response to detecting a sixth set of one or more user gestures, and in accordance with the determination that a drag gesture (e.g., a pinch-and-drag gesture) corresponds to a first direction (e.g., left, right, up, and / or down), the computer system discontinues displaying the first media item in the selected media user interface and displays a second media item different from the first media item in the selected media user interface via a display generation component. In some embodiments, displaying the second media item includes replacing the display of the first media item in the selected media user interface with the display of the second media item in the selected media user interface. In some embodiments, in response to detecting a sixth set of one or more user gestures, and in accordance with the determination that a drag gesture (e.g., a pinch-and-drag gesture) corresponds to a second direction different from the first direction (e.g., left, right, up, and / or down), the computer system discontinues displaying the selected media user interface and displays the media library user interface (e.g., as described with reference to Figure 7F) via a display generation component.
[0213] In some embodiments, upon detecting a sixth set of one or more user gestures, and determining that the drag gesture is in a third direction different from the first and second directions (and optionally opposite to the first direction), the computer system stops displaying the first media item in the selected media user interface and displays the third media item, which is different from the first and second media items in the selected media user interface, via the display generation component. In some embodiments, the computer system replaces the display of the first media item in the selected media user interface with the display of the third media item in the selected media user interface.
[0214] In response to the detection of a sixth set of one or more user gestures and the determination that the drag gesture corresponds to a first direction, discontinuing the display of a first media item in the selected media user interface and displaying a second media item different from the first media item in the selected media user interface provides the user with visual feedback regarding the system state (e.g., that the system has detected a sixth set of one or more user gestures and determined that the sixth set of gestures corresponds to a first direction), which provides improved visual feedback.
[0215] In response to the detection of a sixth set of one or more user gestures, discontinuing the display of the selected media user interface and displaying the media library user interface, in accordance with the determination that the drag gesture corresponds to a second direction, provides the user with visual feedback regarding the system state (for example, that the system has detected a sixth set of one or more user gestures and determined that the sixth set of one or more user gestures corresponds to a second direction), which provides improved visual feedback.
[0216] In some embodiments, while displaying the media library user interface (e.g., 708), the computer system detects one or more user inputs (e.g., 718) via one or more input devices (e.g., one or more gesture inputs (e.g., one or more air gestures) and / or one or more non-gesture inputs) corresponding to the selection of a first media item (e.g., 710F) (e.g., one-handed pinch gesture and / or two-handed pinch-out gesture while the user's gaze is directed towards and / or maintained over the first media item) (e.g., selecting the first media item without selecting any other media items (e.g., selecting only the first media item)). In response to detecting one or more user inputs corresponding to the selection of the first media item, the computer system displays the first media item (e.g., 710F) in a selected media user interface (e.g., 719) different from the media library user interface via a display generation component. In some embodiments, the selected media user interface overlays the media library user interface. In some embodiments, upon detecting one or more user inputs corresponding to the selection of a first media item, the computer system ceases displaying the media library user interface or ceases displaying at least a portion of the media library user interface. While the first media item is being displayed within the selected media user interface, in accordance with a determination that the user's hands (e.g., one or more of the user's hands) are in a first state (e.g., a raised state and / or a first pose), the computer system displays a first set of user interface controls (e.g., 728A, 728B, 728C, and / or 740) (e.g., selectable controls, close options, share options, time information, date information, location information, and / or playback controls (e.g., play options, pause options, fast forward options, and / or rewind options)) via display generation components.In some embodiments, while displaying a first set of user interface controls, the computer system detects one or more selection inputs corresponding to selections of the first user interface controls in the first set of user interface controls, and modifies the display of the first media item in response to the detection of one or more selection inputs (e.g., closing the first media item (e.g., stopping the display), starting and / or pausing playback of the first media item, skipping playback of the first media item forward and / or backward, slowing down and / or accelerating playback of the first media item). In some embodiments, the first set of user interface controls includes a first user interface control (e.g., a close option) that can be selected to close the first media item (e.g., stopping its display). In some embodiments, the first set of one or more user interface controls includes a second user interface control that can be selected to initiate a process for sharing the first media item with one or more external electronic devices (e.g., a share option). In some embodiments, a first set of one or more user interface controls includes a third user interface control (e.g., a play option) that can be selected to resume and / or start playback of a first media item. In some embodiments, a first set of one or more user interface controls includes a fourth user interface control (e.g., a pause option) that can be selected to pause playback of a first media item. In some embodiments, a first set of one or more user interface controls includes a fifth user interface control (e.g., a fast-forward option) that can be selected to skip and / or accelerate playback of a first media item. In some embodiments, a first set of one or more user interface controls includes a sixth user interface control (e.g., a rewind option) that can be selected to skip backward, slow down, and / or reverse playback of a first media item.Upon determining that the user's hands (e.g., one or more of the user's hands) are in a second state different from the first state (e.g., a lowered state and / or a second pose), the computer system stops displaying the first set of user interface controls (as illustrated, for example, with reference to Figure 7N).
[0217] In some embodiments, a first set of user interface controls is displayed without being overlaid on a first media item (e.g., above and / or below the first media item). In some embodiments, while the first set of user interface controls is being displayed, the computer system detects, via one or more input devices, that the user's hand has moved from a first state to a second state, and in response to detecting that the user's hand has moved from a first state to a second state, it stops displaying the first set of user interface controls.
[0218] Displaying a first set of user interface controls in accordance with the determination that the user's hand is in a first state provides the user with visual feedback about the system state (e.g., that the system has detected that the user's hand is in a first state), which provides improved visual feedback.
[0219] Disabling the display of the first set of user interface controls when the user's hands are in the second state, and displaying the first set of user interface controls when the user's hands are in the first state, provides additional control options without disrupting the user interface.
[0220] In some embodiments, while displaying a media library user interface (e.g., 708), the computer system detects via one or more input devices a seventh set (e.g., 718) of one or more user gestures corresponding to the selection of a first media item (e.g., 710F) (e.g., one or more air gestures) (e.g., one-handed pinch gesture and / or two-handed pinch-out gesture while the user's gaze is directed at and / or maintained over the first media item) (e.g., selecting the first media item without selecting any other media items (e.g., selecting only the first media item)). In response to detecting a seventh set of one or more user gestures, the computer system transitions from displaying a first set of lighting effects (e.g., a first set of visual lighting characteristics) to displaying a second set of lighting effects (e.g., a second set of visual lighting characteristics) that is different from the first set of lighting effects, and then, following the display of the first set of lighting effects, displays an intermediate set of lighting effects (e.g., an intermediate set of visual lighting characteristics different from the first set and the second set of visual lighting characteristics) via a display generation component, wherein the intermediate set of lighting effects includes the first set of lighting effects and lighting The method includes displaying a second set of lighting effects that is different from the second set of effects (in some embodiments, the first set of lighting effects has a first value (e.g., a first numerical value) for a first lighting characteristic (e.g., luminance, contrast, saturation), the second set of lighting effects has a second value different from the first value for the first lighting characteristic, and the intermediate set of lighting effects has a third value different from the first and second values for the first lighting characteristic, the third value being between the first and second values), and, following the display of the intermediate set of lighting effects, displaying a second set of lighting effects via a display generation component (e.g., as described with reference to Figures 7D to 7F).In some embodiments, if the computer system is a head-mounted device, the appearance of a first set and / or second set of lighting effects changes in response to the computer system detecting that the user has repositioned itself in the physical environment and / or that the user has rotated its head (for example, while the user is wearing the computer system).
[0221] In some embodiments, the transition from displaying a first set of lighting effects to displaying a second set of lighting effects includes a gradual transition from the first set of lighting effects to the second set of lighting effects (e.g., multiple intermediate lighting effects applied between the first set of lighting effects and the second set of lighting effects). In some embodiments, the transition from displaying a first set of lighting effects to displaying a second set of lighting effects includes gradually applying a light spill effect where multiple rays (e.g., multiple rays of various colors, lengths, and / or intensities) extend from the media window (e.g., gradually increasing the brightness and / or intensity of the light spill effect). In some embodiments, the light spill effect is determined based on the selected media item (e.g., different depending on which media item is selected). In some embodiments, the transition from displaying a first set of lighting effects to displaying a second set of lighting effects includes gradually darkening background content displayed simultaneously with the media library user interface. In some embodiments, in response to the detection of one or more sets of user gestures, the selected first media item is displayed in the selected media item user interface. In some embodiments, background content is displayed simultaneously with the media library user interface at a first time point and simultaneously with the selected media item user interface at a second time point. In some embodiments, background content is displayed in a bright state while displayed simultaneously with the media library user interface and in a dark state while displayed simultaneously with the selected media item user interface. In some embodiments, if the computer system is a head-mounted device, the appearance of a first set and / or second set of lighting effects changes in response to the computer system detecting that the user has repositioned itself in the physical environment and / or that the user has rotated its head (e.g., while the user is wearing the computer system).
[0222] Transitioning from displaying a first set of lighting effects to displaying a second set of lighting effects in response to the detection of a seventh set of one or more user gestures provides the user with visual feedback regarding the system state (e.g., that the system has detected a seventh set of one or more user gestures), which provides improved visual feedback.
[0223] In some embodiments, the computer system displays a first media item (e.g., 710F) in a selected media user interface (e.g., 719) different from the media library user interface (e.g., 708) via a display generation component (e.g., a selected media user interface that indicates the selection of a first media item (e.g., the selection of only the first media item) (e.g., a selected media user interface where the first media item is visually highlighted (e.g., displayed in a larger size than other media items, and / or is the only media item in the displayed media library)). While displaying the first media item in the selected media user interface, the computer system detects an eighth set (e.g., 732) (e.g., one or more air gestures) via one or more input devices that correspond to a user request to close the selected media user interface (e.g., a pinch-and-drag gesture (e.g., a pinch-and-drag gesture in a given direction)).In response to detecting an eighth set of one or more user gestures, the computer system transitions from displaying a third set of lighting effects (e.g., a third set of visual lighting characteristics) (e.g., a third set of lighting effects displayed simultaneously with a first media item in a selected media user interface) to displaying a fourth set of lighting effects different from the third set of lighting effects (e.g., a fourth set of visual lighting characteristics), and then, following the display of the third set of lighting effects, displays a second intermediate set of lighting effects (e.g., a second intermediate set of visual lighting characteristics different from the third and fourth sets of visual lighting characteristics) via a display generation component. The present invention includes, wherein the second intermediate set of lighting effects is different from the third and fourth sets of lighting effects (in some embodiments, the third set of lighting effects has a first value (e.g., a first numerical value) for a first lighting characteristic (e.g., luminance, contrast, saturation), the fourth set of lighting effects has a second value different from the first value for the first lighting characteristic, and the second intermediate set of lighting effects has a third value different from the first and second values for the first lighting characteristic, the third value being between the first and second values), and, following the display of the second intermediate set of lighting effects, the display of the fourth set of lighting effects via a display generation component (as described with reference to, for example, Figure 7F).
[0224] In some embodiments, the transition from displaying a third set of lighting effects to displaying a fourth set of lighting effects includes a gradual transition from the third set of lighting effects to the fourth set of lighting effects (e.g., multiple intermediate lighting effects applied between the third set and the fourth set of lighting effects). In some embodiments, the transition from displaying a third set of lighting effects to displaying a fourth set of lighting effects includes gradually reducing the light spill effect (e.g., gradually reducing the brightness and / or intensity of the light spill effect) of multiple rays (e.g., multiple rays of various colors, lengths, and / or intensities) extending from the media window (e.g., extending from the selected media user interface). In some embodiments, the light spill effect is determined based on the selected media item (e.g., it differs depending on which media item is selected). In some embodiments, the transition from displaying a third set of lighting effects to displaying a fourth set of lighting effects includes gradually brightening background content displayed simultaneously with the selected media user interface. In some embodiments, in response to detecting an eighth set of one or more user gestures, the computer system discontinues displaying the selected media user interface and displays the media library user interface. In some embodiments, background content is displayed simultaneously with the selected media user interface at a first time point and simultaneously with the media library user interface at a second time point. In some embodiments, background content is displayed in a dimmed state while displayed simultaneously with the selected media user interface and in a brightened state while displayed simultaneously with the media library user interface.
[0225] Transitioning from displaying a third set of lighting effects to displaying a fourth set of lighting effects in response to the detection of an eighth set of one or more user gestures provides the user with visual feedback regarding the system state (e.g., the system has detected an eighth set of one or more user gestures), which provides improved visual feedback.
[0226] In some embodiments, the computer system displays a first media item (e.g., 741 or 743) in a selected media user interface (e.g., 719) different from the media library user interface (e.g., 708) (e.g., a selected media user interface that indicates the selection of a first media item (e.g., the selection of only the first media item) (e.g., a selected media user interface where the first media item is visually highlighted (e.g., displayed in a larger size than other media items, and / or is the only media item in the displayed media library)) via a display generation component, and simultaneously displays a navigation user interface element (e.g., 740) (e.g., a scrubber bar) for navigating visual content (e.g., navigating multiple frames (e.g., images from a video), and / or navigating multiple media items). Displaying the navigation user interface element simultaneously with the first media item in the selected media user interface allows the user to navigate the content with less input, thereby reducing the number of inputs required to perform an action.
[0227] In some embodiments, a navigation user interface element (e.g., 740) includes one or more selectable controls that, when selected, perform a respective function (e.g., a play option, a pause option, a fast-forward option, a rewind option, and / or a skip option) associated with one or more media items (e.g., a pause option shown in the scrubber 740 in Figures 7M-7N). In some embodiments, while the navigation user interface element is displayed, the computer system detects one or more selection inputs (e.g., one or more air gestures) corresponding to a selection of a first selectable control among the one or more selectable controls, and in response to the detection of one or more selection inputs, the computer system modifies the display of the first media item (e.g., start and / or pause playback of the first media item, skip playback of the first media item forward and / or backward, slow down and / or accelerate playback of the first media item). In some embodiments, one or more selectable controls include a first selectable control (e.g., a play option) for resuming and / or starting playback of the first media item. In some embodiments, one or more selectable controls include a second control (e.g., a pause option) that can be selected to pause playback of a first media item. In some embodiments, one or more selectable controls include a third control (e.g., a fast-forward option) that can be selected to skip forward and / or accelerate playback of the first media item. In some embodiments, one or more selectable controls include a fourth control (e.g., a rewind option) that can be selected to skip backward, slow down, and / or reverse playback of the first media item. Displaying one or more selectable controls allows the user to perform various respective functions with fewer inputs, thereby reducing the number of inputs required to perform an action.
[0228] In some embodiments, while displaying a media library user interface (e.g., 708), the computer system detects one or more user inputs (e.g., 718) via one or more input devices corresponding to the selection of a first media item (e.g., 710F) (e.g., one or more touch inputs, one or more non-touch inputs, and / or one or more gestures (e.g., one or more air gestures)) (e.g., one-handed pinch gesture and / or two-handed pinch-out gesture while the user's gaze is directed at and / or maintained over the first media item). In response to detecting one or more user inputs corresponding to the selection of a first media item, the computer system displays the first media item at a first angular size via a display generation component, according to a determination that a first set of criteria is met (for example, according to a determination that a first user setting (e.g., immersive viewing setting) is enabled and / or disabled, according to a determination that the first media item is of a specific type, and / or according to a determination that one or more user inputs of a specific type have been detected). In response to a determination that the first set of criteria is not met, the computer system displays the first media item at a second angular size different from the first angular size via a display generation component (for example, as described with reference to Figure 7F). In some embodiments, the first angular size corresponds to an immersive viewing experience (e.g., having a larger angular size relative to the media item being viewed), and the second angular size corresponds to a non-immersive viewing experience (e.g., the second angular size is smaller than the first angular size). In some embodiments, when the computer system is a head-mounted device, while the first media item is displayed at a first angular size, the computer system changes the perspective in which the first media item is displayed in response to the computer system detecting that the user has repositioned themselves in the physical environment and / or that the user has rotated their head (for example, while wearing the computer system).Displaying a first media item at a first angular size based on the determination that a first set of criteria has been met provides the user with visual feedback about the system's state (e.g., that the first set of criteria has been met), which provides improved visual feedback.
[0229] In some embodiments, the computer system displays a first media item (e.g., 710F) via a display generation component (e.g., in a selected media user interface distinct from the media library user interface) (e.g., a selected media user interface that indicates the selection of the first media item (e.g., the selection of only the first media item)) (e.g., a selected media user interface in which the first media item is visually highlighted relative to other media items (e.g., displayed in a larger size than other media items, and / or being the only media item in the displayed media library)), and the first media item is displayed (e.g., in the selected media user interface) with vignetting applied to the first media item (e.g., darkening, thinning, obscuring, and / or blurring the edges and / or corners of the first media item). Displaying a representation of the first media item with vignetting applied to the first media item indicates to the user that the representation of the first media item is displayed in a selected state, for example, which provides improved visual feedback.
[0230] In some embodiments, the computer system detects a user gaze (e.g., 714) corresponding to a first position in the media library user interface (e.g., 708) via one or more input devices (e.g., at a sixth time point after a second time point) (e.g., after or while displaying a representation of a first media item in a third manner) (e.g., detect and / or determine that the user is gazing at a first position in the media library user interface (e.g., a first position in the media library user interface corresponding to a representation of a first media item)). While continuously detecting the user's gaze corresponding to a first position in the media library user interface (for example, at a seventh time point after a sixth time point), the computer system detects selection gestures (e.g., 718) (e.g., air gestures) (e.g., movement of the user's hands, head, and / or body) (in some embodiments, one or more non-gesture inputs) (e.g., one-handed pinch gestures and / or two-handed pinch-out gestures while the user's gaze is directed at and / or maintained over the first media item) via one or more input devices. CheckDepending on what is performed, and according to the determination that the selection gesture corresponds to a first type of selection gesture (e.g., air gesture) (e.g., one-handed pinch gesture, one-handed double pinch gesture, two-handed pinch gesture, two-handed pinch-out gesture, partially completed one-handed pinch gesture, and / or completed one-handed pinch gesture), the computer system displays the first media item in a ninth manner (e.g., using a ninth set of visual characteristics) via the display generation component (e.g., displaying the first media item in an intermediate selection state (e.g., a transition state that transitions from the media library user interface to the selected media user interface)) (e.g., displaying the first media item in a size smaller than the size of the first media item displayed in the selected media user interface, and larger than the size of the representation of the first media item in the media library user interface), the selection gesture However, in accordance with the determination that a second type of selection gesture (e.g., a one-handed pinch gesture, a one-handed double pinch gesture, a two-handed pinch gesture, a two-handed pinch-out gesture, a partially completed one-handed pinch gesture, and / or a completed one-handed pinch gesture) corresponds to a first type of selection gesture, the computer system displays the first media item via a display generation component in a tenth method different from the ninth method (e.g., using a tenth set of visual characteristics different from the ninth set of visual characteristics) (e.g., in a selected media user interface different from a media library user interface indicating user selection of the first media item). In some embodiments, displaying the first media item in the ninth method includes displaying the first media item in a transition state, and displaying the first media item in the tenth method includes displaying the first media item in a selected state.In some embodiments, displaying a first media item in the ninth manner includes displaying the first media item in a first size larger than the size of its representation in the media library user interface, and displaying a first media item in the tenth manner includes displaying the first media item in a second size larger than the first size (as illustrated, for example, with reference to Figure 7F). Displaying the first media item in the ninth and / or tenth manner in response to the detection of a selection gesture corresponding to the selection of the first media item provides the user with visual feedback regarding the state of the system (e.g., that the system has detected a selection gesture corresponding to the selection of the first media item), which provides improved visual feedback.
[0231] In some embodiments, aspects / operations of methods 800, 900, 1000, and 1100 may be interchangeable, substituted, and / or added to among these methods. For example, the media library user interface displayed in method 800 is optionally the user interface displayed in method 900, and / or the first media item displayed in method 800 is optionally the first media item displayed in methods 900 and / or 1000. For brevity, those details will not be repeated here.
[0232] Figure 9 is a flowchart of exemplary method 900 for interacting with media items and user interfaces according to several embodiments. In some embodiments, method 1000 is performed with a computer system (e.g., computer system 101 in Figure 1) (e.g., smartphone, smartwatch, tablet, and / or wearable device), mouse, keyboard, remote control, visual input device (e.g., camera), audio input device (e.g., microphone), and / or biometric sensor (e.g., fingerprint sensor, face recognition sensor, and / or iris recognition sensor)) that communicates with display generating components (e.g., 702) (e.g., display controller, touch-sensitive display system, display (e.g., integrated and / or connected), 3D display, transparent display, projector, and / or head-up display) and one or more input devices (e.g., 715) (e.g., touch-sensitive surface (e.g., touch-sensitive display)). In some embodiments, Method 1000 is stored in a non-temporary (or temporary) computer-readable storage medium and on one or more processors 202 of the computer system 101 (for example, Figure 1 It is controlled by instructions executed by one or more processors of a computer system, such as control 110. Some operations of method 900 are optionally combined, and / or the order of some operations is optionally changed.
[0233] In some embodiments, a computer system (e.g., 700) displays a first zoom level user interface (e.g., user interface 708 in Figure 7D, user interface 719 in Figure 7F or Figure 7J) (e.g., a user interface including one or more content items (e.g., one or more selectable content items and / or media items (e.g., photographs and / or videos)) and / or representations of one or more content items (e.g., selectable representations of one or more media items)) (e.g., 708, 719) (902). While the user interface is being displayed (904), the computer system detects one or more user inputs (e.g., 718, 732, 733, 734) (e.g., one or more tap inputs, one or more gestures (e.g., one or more air gestures), and / or one or more other inputs) via one or more input devices (906).In response to detecting one or more user inputs corresponding to a zoom-in user command (908), and in accordance with the determination that the user's gaze (e.g., 714) (e.g., the user's gaze detected while and / or after detecting one or more user inputs) corresponds to a first position in the user interface (910) (e.g., in accordance with the determination that the user is gazing at the first position in the user interface), the computer system, via the display generation component, displays the user interface at a second zoom level greater than the first zoom level (e.g., user interface 708 in Figure 7E, user interface 7G in Figure 7G). Displaying interface 719 (user interface 719 in Figure 7H, user interface 719 in Figure 7K) (912), and displaying the user interface at a second zoom level includes zooming the user interface using a first zoom center selected based on the first position (e.g., maintaining the first position of the user interface at its current display position on the display generating component while enlarging and / or zooming the user interface (e.g., maintaining the first position of the user interface at its current display position on the display generating component for at least part of the zoom operation)).In accordance with the determination that the user's gaze (for example, the user's gaze detected while detecting one or more user inputs and / or after detecting one or more user inputs) corresponds to a second position in the user interface that is different from the first position (914) (for example, in accordance with the determination that the user is fixated on the second position in the user interface), the computer system, via the display generation component, displays the user interface (for example, user interface 708 in Figure 7E, user interface 719 in Figure 7G, user interface 7H in Figure 7H) at a third zoom level greater than the first zoom level (for example, a third zoom level equal to or different from the second zoom level). Displaying the interface 719 (user interface 719 in Figure 7K) (916) and displaying the user interface at a third zoom level includes zooming the user interface using a second zoom center selected based on a second position, the second zoom center being in a different location from the first zoom center (e.g., maintaining the second position of the user interface at its current display position on the display-generating component while enlarging and / or zooming the user interface (e.g., maintaining the second position of the user interface at its current display position on the display-generating component for at least part of the zoom operation)). Displaying the user interface at a second zoom level using a first zoom center selected based on the user's gaze position in response to detecting one or more user inputs corresponding to a zoom-in user command provides the user with visual feedback regarding the state of the system (e.g., that the system has detected one or more user inputs and has detected the user's gaze position), which provides improved visual feedback.
[0234] In some embodiments, one or more user inputs corresponding to a zoom-in user command (e.g., 718, 732, 733, or 734) include a first pinch gesture (e.g., an air gesture) (e.g., a gesture in which two fingers move closer together (e.g., a gesture in which the index finger and thumb of one hand move closer together)) (e.g., a one-handed pinch gesture (e.g., two fingers of one hand move closer together)) and a second pinch gesture (e.g., an air gesture) (e.g., the first and second pinch gestures occurring and / or detected within a threshold duration of each other). Displaying the user interface at a second zoom level using a first zoom center selected based on the user's gaze position in response to a first and second pinch gesture provides the user with visual feedback regarding the system's state (e.g., that the system has detected the first and second pinch gestures and the user's gaze position), which provides improved visual feedback.
[0235] In some embodiments, one or more user inputs corresponding to a zoom-in user command (e.g., 718, 732, 733, or 734) include a two-handed pinch-out gesture (e.g., an air gesture) (e.g., a gesture where a first hand moves away from another hand) (e.g., a gesture where a first hand making a pinch shape (e.g., a shape where the index finger and thumb of the hand are in contact) moves away from a second hand making a pinch shape). Displaying the user interface at a second zoom level using a first zoom center selected based on the user's gaze position in response to a two-handed pinch-out gesture provides the user with visual feedback regarding the state of the system (e.g., that the system has detected a two-handed pinch-out gesture and detected the user's gaze position), which provides improved visual feedback.
[0236] In some embodiments, upon detecting one or more user inputs corresponding to a zoom-in user command (e.g., 718, 732, 733, or 734), and in accordance with the determination that the user interface is displaying a first media item of a first type (e.g., a non-panoramic image), the computer system displays the first media item at a first size (e.g., a first coverage area and / or a first set of dimensions) (e.g., a default maximum size for a media item of the first type). In accordance with the determination that the user interface is displaying a second media item of a second type distinct from the first type (e.g., a panoramic image (e.g., an image produced by stitching together multiple image captures in a particular direction) (e.g., an image having a set of dimensions identified as panoramic dimensions) (e.g., an image having an aspect ratio greater than a threshold aspect ratio (e.g., an image having an aspect ratio greater than and / or greater than 16:9)), the computer system displays the second media item in a second size (e.g., a second coverage area and / or second set of dimensions) that is larger than the first size (e.g., a size greater than the default maximum size of the first type of media item) (e.g., media window 704 in Figures 7K to 7L). In some embodiments, a panoramic image can be enlarged to a larger size than a non-panoramic image. By automatically displaying a second media item in a second size larger than the first size, based on the user interface's determination that it is displaying a second type of media item, the user can view a second type of media item (e.g., a panoramic image) in a larger size without requiring additional user input, which executes the action when a set of conditions is met without requiring further user input.
[0237] In some embodiments, upon detection of one or more user inputs corresponding to a zoom-in user command (e.g., 718, 732, 733, or 734), the computer system displays the first media item as a flat object (e.g., a two-dimensional object, a non-curved object, a flat planar object, and / or an object having a flat, non-curved surface) in accordance with the determination that the user interface is displaying a first media item of a first type (e.g., a panoramic image (e.g., an image produced by stitching multiple image captures together in a particular direction) (e.g., an image having a set of dimensions (e.g., width and / or height) identified as panoramic dimensions)) in accordance with the determination that the user interface is displaying a second media item of a second type (e.g., a panoramic image (e.g., an image produced by stitching multiple image captures together in a particular direction) (e.g., an image having a set of dimensions (e.g., width and / or height) identified as panoramic dimensions)) in accordance with the determination that the user interface is displaying a second media item as a curved object (e.g., a three-dimensional object, a curved surface object, and / or an object having one or more curved surfaces) (e.g., media window 704 in Figures 7K to 7L) (as described with reference to Figures 7J to 7L, for example). In some embodiments, panoramic images are curved as they are zoomed in, while non-panoramic images are not. In some embodiments, when the computer system is a head-mounted device, a second media item occupies more of the user's field of view when it is displayed as a curved object, as opposed to when the first media item is displayed as a flat object. By automatically displaying the second media item as a curved object according to the user interface's determination that it is displaying a second type of media item, the user can display the second type of media item (e.g., a panoramic image) as a curved object without requiring additional user input, which performs the action when a set of conditions is met without requiring further user input.
[0238] In some embodiments, displaying the user interface at a first zoom level includes displaying a representation of a first media item (e.g., a thumbnail representation of the first media item) at a first size (e.g., 710F in Figure 7D). A three-dimensional environment (e.g., 706) at least partially surrounds the user interface (e.g., 708) and includes background content behind the user interface (e.g., a representation of a physical or virtual environment). In some embodiments, the background content at least partially surrounds the user interface. In some embodiments, the three-dimensional environment and / or background content are displayed by a display-generating component (e.g., behind the user interface). In some embodiments, the three-dimensional environment and / or background content are visible to the user (e.g., behind the user interface) but not displayed by the display-generating component (e.g., the three-dimensional environment and / or background content are tangible physical objects visible to the user behind the user interface without being displayed by the display-generating component). In some embodiments, upon detecting one or more user inputs corresponding to a zoom-in user command (e.g., 718), the computer system transitions the representation of the first media item from being displayed at a first size (e.g., 710F in Figure 7D) to being displayed at a second size larger than the first size (e.g., 710F in Figure 7F) (e.g., enlarges the display of the first media item). It then reduces the visual emphasis of the background content on the first media item (e.g., the three-dimensional environment 706 in Figure 7F) (e.g., darkens the background content). In some embodiments, the background content transitions from being displayed at a first brightness level to being displayed at a second brightness level at the same time that the representation of the first media item transitions from being displayed at a first size to being displayed at a second size.Displaying background content in a darkened state when the user zooms in on the first media item provides the user with feedback about the system's state (e.g., that the system is displaying the first media item in a zoomed-in state), which provides improved visual feedback.
[0239] In some embodiments, displaying the user interface at a first zoom level includes displaying a representation of a first media item (e.g., a thumbnail representation of the first media item) at a first size (e.g., 710F in Figure 7D). In some embodiments, upon detecting one or more user inputs corresponding to a zoom-in user command (e.g., 718), the computer system transitions the representation of the first media item from being displayed at a first size (e.g., 710F in Figure 7D) to being displayed at a second size larger than the first size (e.g., 710F in Figure 7F) (e.g., enlarging the display of the first media item). The system then displays a light spill effect (e.g., 726-1) extending from the user interface (e.g., 704) (e.g., extending from the first media item). In some embodiments, the light spill effect includes one or more visual properties (e.g., brightness, intensity, size and / or length, color, saturation, contrast) determined based on the visual content (e.g., visual properties) of the first media item (e.g., different media items produce different light spill effects). In some embodiments, the light spill effect includes one or more of the following: glow around the edges of the item, the appearance of light surrounding the item, the appearance of rays of light around the item, and / or the appearance of a light source behind the middle or center of the item. In some embodiments, the size of the representation of the first media item gradually increases in response to one or more user inputs corresponding to a zoom-in user command. In some embodiments, the light spill effect extending from the user interface to the background content is gradually modified (e.g., gradually enhanced) concurrently with the gradual increase in the size of the representation of the first media item. Displaying the light spill effect when the user zooms in on the first media item provides the user with feedback about the state of the system (e.g., that the system is displaying the first media item in a zoomed-in state), which provides improved visual feedback.
[0240] In some embodiments, displaying the user interface at a first zoom level includes displaying a first media item at a first size (e.g., 731 in Figures 7J and 7K) (e.g., displaying a representation of the first media item within a selected media user interface indicating user selection of the first media item). In some embodiments, displaying the user interface at a second zoom level includes displaying a first media item at a second size larger than the first size (e.g., 731 in Figure 7L), displaying the first media item with a blur effect applied to at least a first edge of the first media item (e.g., 736A or 736B) according to the determination that the second size is larger than a predetermined threshold size, and displaying the first media item without applying a blur effect to any edge of the first media item (e.g., Figure 7J) according to the determination that the second size is not larger than a predetermined threshold size. In some embodiments, a blurring effect is applied to the first edge of a first media item to indicate that additional content of the first media item extending beyond the first edge is not displayed, and / or that the user can scroll in the direction of the first edge to view the additional content of the first media item. In some embodiments, the blurring effect includes blurring the visual content at the first edge of the first media item, and / or otherwise visually obscuring the visual content at the first edge of the first media item. In some embodiments, displaying the first media item without any blurring effect applied to any edge of the first media item indicates that the entire first media item is displayed. Displaying the first media item with a blurring effect applied to the first edge of the first media item provides the user with feedback about the state of the system (e.g., that there is additional content of the first media item extending beyond the undisplayed first edge), which provides improved visual feedback.
[0241] In some embodiments, the user interface is a media library user interface (e.g., 708) that includes representations of multiple media items (e.g., 710A-710F) in a media library (e.g., a collection of media items associated with a device (e.g., stored in a device) and / or associated with a user), including representations of a first media item and a second media item. In some embodiments, displaying the user interface at a first zoom level includes simultaneously displaying a representation of a first media item at a first size (e.g., having a first set of dimensions (e.g., height and / or width)) and a representation of a second media item at a second size (in some embodiments, the second size is different from or the same as the first size) (e.g., 708 in Figure 7D). In some embodiments, displaying the user interface at a second zoom level includes simultaneously displaying a representation of a first media item at a third size larger than the first size, and a representation of a second media item at a fourth size larger than the second size (e.g., Figure 7E). In some embodiments, the fourth size is different from or the same as the third size. In some embodiments, displaying the user interface at a first zoom level includes displaying a media library grid having representations of multiple media items, and displaying the user interface at a second zoom level includes zooming in on the media library grid (e.g., displaying representations of fewer media items at a larger size). Displaying representations of the first media items at a third size and representations of the second media items at a fourth size in response to detecting one or more user inputs corresponding to a zoom-in user command provides the user with visual feedback about the state of the system (e.g., that the system has detected one or more user inputs corresponding to a zoom-in user command), which provides improved visual feedback.
[0242] In some embodiments, displaying the user interface at a first zoom level includes displaying the user interface at a first size (for example, having a first set of dimensions (e.g., height and / or width)). In some embodiments, upon detecting one or more user inputs corresponding to a zoom-in user command, and in accordance with the determination that the user interface is a first user interface (e.g., 719) (e.g., a selected media user interface (e.g., a user interface that displays media items selected by the user, and / or a user interface that indicates and / or responds to the user selection of media items)), the computer system displays the user interface (e.g., the selected media user interface 719 in Figures 7J to 7L) at a second size larger than the first size. Upon determining that the user interface is a second user interface (e.g., a media library user interface (e.g., a user interface that displays representations of multiple media items in the media library)) that is different from a first user interface (e.g., 708), the computer system maintains the user interface at the first size (e.g., the media library user interface 708 in Figures 7D-7E) (in some embodiments, the selected media user interface (e.g., 719) can be enlarged to a larger size in response to a zoom-in command, but the media library user interface (e.g., 708) is not enlarged to a larger size in response to a zoom-in command). Displaying the user interface at a second size larger than the first size in response to the detection of one or more user inputs corresponding to a zoom-in user command provides the user with visual feedback regarding the state of the system (e.g., that the system has detected one or more user inputs corresponding to a zoom-in user command), which provides improved visual feedback.
[0243] In some embodiments, aspects / operations of methods 800, 900, 1000, and 1100 may be interchangeable, substituted, and / or added to among these methods. For example, the media library user interface displayed in method 800 is optionally the user interface displayed in method 900, and / or the first media item displayed in method 800 is optionally the first media item displayed in methods 900 and / or 1000. For brevity, those details will not be repeated here.
[0244] Figure 10 is a flowchart of exemplary method 1000 for interacting with media items and user interfaces according to several embodiments. In some embodiments, method 1000 is performed on a computer system (e.g., 700) (e.g., computer system 101 in Figure 1) that communicates with display generating components (e.g., 702) (e.g., display controller, touch-sensitive display system, display (e.g., integrated and / or connected), 3D display, transparent display, projector, and / or head-up display) and one or more input devices (e.g., 715) (e.g., touch-sensitive surface (e.g., touch-sensitive display)), a mouse, keyboard, remote control, visual input device (e.g., camera), audio input device (e.g., microphone), and / or biometric sensor (e.g., fingerprint sensor, face recognition sensor, and / or iris recognition sensor)). In some embodiments, method 1000 is stored on a non-temporary (or temporary) computer-readable storage medium and on one or more processors 202 of computer system 101 (e.g., Figure 1 It is controlled by instructions executed by one or more processors of a computer system, such as control 110. Some operations of method 1000 are optionally combined, and / or the order of some operations is optionally changed.
[0245] In some embodiments, a computer system (e.g., 700) detects one or more user inputs (e.g., 718) (e.g., one or more tap inputs, one or more gestures (e.g., one or more air gestures), and / or one or more other inputs) (e.g., one-handed pinch gesture and / or two-handed pinch-out gesture while the user's gaze is directed at and / or maintained over the first media item) (e.g., selection of the first media item in a media library, selection of the first media item among multiple media items in a media library, selection of the first media item among multiple media items, and / or selection of the first media item among multiple displayed media items (e.g., selection of a representation of the first media item among multiple displayed representations of a media item)) via one or more input devices (1002). In response to detecting one or more user inputs corresponding to the selection of a first media item (1004), and in accordance with the determination that the first media item is a media item containing a distinct type of depth information (e.g., a stereoscopic media item having media simultaneously captured from two different cameras (or sets of cameras) displayed by displaying images from a first set of one or more cameras for the user's first eye and images from a second set of one or more cameras for the user's second eye) (1006), the computer system displays the first media item in a first manner (1008) (e.g., Figure 7F) (e.g., using a first set of visual characteristics), and in accordance with the determination that the first media item is a media item that does not contain a distinct type of depth information (1010), the computer system displays the first media item in a second manner (e.g., using a second set of visual characteristics) different from the first manner (e.g., Figure 7J) (e.g., as described with reference to Figures 7F and 7J) (1012).Displaying the first media item in a first manner based on the determination that the first media item contains depth information of a distinct type provides the user with visual feedback about the system's state (e.g., that the system has determined the first media item contains depth information of a distinct type), which provides improved visual feedback.
[0246] In some embodiments, displaying a first media item in a first manner includes displaying a first media item having a first type of boundary surrounding the first media item (e.g., a peripheral edge and / or boundary) (e.g., a first boundary having a first shape, a first boundary having a first set of visual characteristics) (e.g., a first boundary having a refractive edge and / or a first boundary having rounded corners) (e.g., Figure 7F). In some embodiments, displaying a first media item in a second manner includes displaying a first media item having a second type of boundary (e.g., a peripheral edge and / or boundary) different from the first boundary surrounding the first media item (e.g., as described with reference to Figure 7F) (e.g., Figure 7J) (e.g., a second boundary having a second shape, a second boundary having a second set of visual characteristics) (e.g., a second boundary having a non-refracting edge and / or a second boundary having non-circular (e.g., rectangular and / or pointed) corners). In some embodiments, a media library (e.g., presented in a media library user interface) includes a set of media items of a first type (e.g., multiple media items) and a set of media items of a second type (e.g., multiple media items). In response to detecting one or more user inputs corresponding to the selection of individual media items of the first type, the computer system displays the individual media items of the first type in a first manner, which includes displaying the individual media items of the first type having a boundary of the first type surrounding them. In response to detecting one or more user inputs corresponding to the selection of individual media items of a second type, the computer system displays the individual media items of the second type in a second manner, which includes displaying the individual media items of the second type having a boundary of the second type surrounding them. In some embodiments, the media items of the first type are displayed in a first manner, which includes being displayed with a boundary of the first type, and the media items of the second type are displayed in a second manner, which includes being displayed with a boundary of the second type.Displaying the boundaries of a first media item differently based on whether or not it contains depth information of a particular type provides the user with visual feedback regarding the system's state (e.g., whether or not the first media item contains depth information of a particular type), which provides improved visual feedback.
[0247] In some embodiments, when a first media item is displayed in a first manner, content (e.g., background content) (e.g., background content displayed behind the first media item and partially surrounding the first media item) that at least partially surrounds (e.g., completely surrounds) the first media item (e.g., 706) has a first appearance (e.g., Figure 7F) (e.g., having a third set of visual characteristics) (e.g., having selectable controls displayed outside the boundaries of the first media item and / or having a third set of lighting effects applied to the background content) (e.g., modified by a display generating component to have). When the first media item is displayed in the second manner, the content (e.g., background content) that at least partially surrounds (e.g., completely surrounds) the first media item has a second appearance different from the first appearance (e.g., Figure 7J) (e.g., without displaying selectable controls outside the boundaries of the first media item and / or without having a fourth set of lighting effects applied to the background content) (e.g., having a fourth set of visual characteristics) (e.g., modified to have a second appearance by a display generating component). In some embodiments, the content that at least partially surrounds the first media item is displayed by a display generating component. In some embodiments, the content that at least partially surrounds the first media item is not displayed by a display generating component (e.g., the content that at least partially surrounds the first media item includes one or more tangible physical objects that are visible to the user behind the user interface without the content being displayed by a display generating component). Displaying the content surrounding a first media item differently based on whether the first media item contains a particular type of depth information provides the user with visual feedback about the system state (e.g., whether the first media item contains a particular type of depth information), which provides improved visual feedback.
[0248] In some embodiments, the computer system displays a set of rays extending from a first media item (e.g., 726-1, 726-2, 726-3, 726-4, 730-1, 730-2, 730-3, or 730-4) (e.g., a light spill illumination effect) (e.g., rays extending from the outer boundary of the first media item into background content that at least partially surrounds the first media item). In some embodiments, the rays are a visual effect in which one or more visual properties of the rays (e.g., brightness, intensity, size, length, color, saturation, and / or contrast) are determined based on the visual content (e.g., visual properties) of the first media item (e.g., different media items result in different light spill rays). In some embodiments, displaying the first media item in a second manner includes discontinuing the display of a set of rays extending from the first media item. In some embodiments, the set of rays extending from the first media item changes over time (for example, the set of rays extending from the first media item changes over time as the visual content of the first media item changes over time (for example, as the visual content of the first media item (e.g., video content) is played)). Displaying the set of rays extending from the first media item in response to the detection of one or more user inputs corresponding to the selection of the first media item provides the user with visual feedback about the state of the system (e.g., that the system has detected one or more user inputs corresponding to the selection of the first media item), which provides improved visual feedback.
[0249] In some embodiments, one or more visual characteristics of a set of rays (e.g., luminance, intensity, size, length, color, saturation, and / or contrast) (e.g., 726-1, 726-2, 726-3, 726-4, 730-1, 730-2, 730-3, or 730-4) are determined based on one or more colors at the edge of the first media item (e.g., the outer boundary or within a predetermined distance from the edge) (for example, as described with reference to Figure 7F) (for example, in some embodiments, the visual characteristics of rays extending from a first edge of the first media item are determined based on the color displayed at the first edge of the first media item, and / or the visual characteristics of rays extending from a second edge of the first media item are determined based on the color displayed at the second edge of the first media item). Displaying a set of rays extending from a first media item in response to the detection of one or more user inputs corresponding to the selection of a first media item provides the user with visual feedback about the system's state (e.g., that the system has detected one or more user inputs corresponding to the selection of a first media item), which provides improved visual feedback.
[0250] In some embodiments, a set of rays (e.g., 726-1, 726-2, 726-3, 726-4, 730-1, 730-2, 730-3, or 730-4) includes a first ray having a first length and a second ray having a second length different from the first length (e.g., the set of rays includes rays having different lengths or variable lengths). In some embodiments, the set of rays includes a third ray having a third length different from the first and second lengths. In some embodiments, a first ray extends from a first side of a first media item (e.g., top, bottom, left, and / or right); a second ray extends from a second side of the first media item that is different in size from the first; a third ray extends from a third side of the first media item that is different from the first and second sides; and a fourth ray extends from a fourth side of the first media item that is different from the first, second, and third sides. Displaying a set of rays extending from a first media item in response to the detection of one or more user inputs corresponding to the selection of a first media item provides the user with visual feedback regarding the state of the system (e.g., that the system has detected one or more user inputs corresponding to the selection of a first media item), which provides improved visual feedback.
[0251] In some embodiments, the first media item includes a distinct type of depth information (e.g., a stereoscopic media item with media simultaneously captured from two different cameras (or sets of cameras), displayed by showing images from a first set of one or more cameras for the user's first eye and images from a second set of one or more cameras for the user's second eye), and the first media item is captured by a computer system (e.g., 700) (e.g., using multiple cameras connected to and / or integrated with the computer system). In some embodiments, the computer system captures the first media item via one or more input devices (e.g., using multiple cameras connected to and / or integrated with the computer system) before detecting one or more user inputs. In some embodiments, the media item includes a distinct type of depth information if the first media item is a stereoscopic capture (e.g., a stereoscopic media item). In some embodiments, stereoscopic capture includes two images simultaneously captured from two different cameras spaced apart (e.g., spaced approximately the same distance as human eyes), and the two images are displayed simultaneously to the user (a first image for the user's first eye and a second image for the user's second eye) to reproduce the depth of the captured scene. In some embodiments, when the computer system is a head-mounted device and the first media item contains depth information, the computer system displays a first perspective view of the physical environment contained in the media item on a first display of the computer system so that the user perceives stereoscopic depth between the contents contained in the media item, and the computer system displays a second perspective view of the physical environment contained in the media item on a second display of the electronic device.Displaying the first media item in a first manner based on the determination that the first media item contains depth information of a distinct type provides the user with visual feedback about the system's state (e.g., that the system has determined the first media item contains depth information of a distinct type), which provides improved visual feedback.
[0252] In some embodiments, displaying a first media item in a first manner includes displaying a first media item having a first set of lighting effects (e.g., 726-1, 726-2, 726-3, or 726-4) (e.g., a first set of light spill lighting effects extending from the first media item (e.g., a first set of light spill lighting effects having one or more visual characteristics (e.g., brightness, intensity, size, length, color, saturation, and / or contrast) determined based on the visual content of the first media item). In some embodiments, displaying a first media item in a second manner includes displaying a second set of lighting effects (e.g., 730-1, 730-2, 730-3, or 730-4) (e.g., a second set of lighting effects different from the first set of lighting effects, or a second set of lighting effects the same as the first set of lighting effects) (e.g., a second set of light spill lighting effects extending from the first media item (e.g., based on the visual content of the first media item) The method includes displaying a first media item having a second set of light spill lighting effects having one or more visual characteristics (e.g., brightness, intensity, size, length, color, saturation, and / or contrast) determined by the first media item. For example, in some embodiments, the first media item is displayed with lighting effects applied to the first media item, regardless of whether the first media item is a media item containing a distinct type of depth information (e.g., a stereoscopic media item) or a media item not containing a distinct type of depth information (e.g., not a stereoscopic media item). In some embodiments, the first media item is displayed with light spill rays extending from the first media item, regardless of whether the first media item is a media item containing a distinct type of depth information (e.g., a stereoscopic media item) or a media item not containing a distinct type of depth information (e.g., not a stereoscopic media item).In some embodiments, the light spill rays extending from the first media item differ based on whether the first media item contains or does not contain a distinct type of depth information (for example, one or more algorithms for determining the light spill rays extending from the first media item differ based on whether the first media item contains or does not contain a distinct type of depth information). Displaying the first media item using a first or second lighting effect set in response to the detection of one or more user inputs corresponding to the selection of the first media item provides the user with visual feedback regarding the state of the system (e.g., that the system has detected one or more user inputs corresponding to the selection of the first media item...
Claims
1. It is a method, In a computer system that communicates with display generation components, It is a user interface, A first representation of a stereoscopic media item, wherein the first representation of the stereoscopic media item includes at least a first edge, Displaying a user interface via the display generation component, the user interface includes a visual effect which obscures at least a first portion of the stereoscopic media item, and which extends inward from at least the first edge of the first representation of the stereoscopic media item toward the interior of the first representation of the stereoscopic media item, and which is displayed on the first edge and the second edge of the first representation of the stereoscopic media item. While the first representation of the stereoscopic media item is being displayed, a first change in the user's viewpoint is detected, which includes a change in the angle of the user's viewpoint relative to the display of the first representation of the stereoscopic media item. In response to detecting the first change in the user's viewpoint, In accordance with the determination that the first change in the user's viewpoint is in a first direction, the size of the visual effect on at least the second edge of the first representation of the stereoscopic media item is increased. A method comprising increasing the size of the visual effect on at least the first edge of the first representation of the stereoscopic media item, in accordance with the determination that the first change in the user's viewpoint is in a second direction different from the first direction.
2. The stereoscopic media item is captured from a set of cameras, the first camera of the set of cameras captures a first perspective view of the physical environment, the second camera of the set of cameras captures a second perspective view of the physical environment, the second perspective view differs from the first perspective view in that it displays the first representation of the stereoscopic media item. The first perspective view of the physical environment is displayed to the user's first eye, not the user's second eye, via the display generation component. The method according to claim 1, further comprising displaying the second perspective view to the user's second eye rather than the user's first eye via the display generation component.
3. While the first representation of the stereoscopic media item is being displayed, a second change in the user's viewpoint is detected. In response to detecting the second change in the user's viewpoint, the appearance of the first representation of the stereoscopic media item is changed based on the second change in the user's viewpoint. The method according to claim 1, further comprising:
4. Before altering the appearance of the first representation of the stereoscopic media item, the visual effect obscures the first portion of the first representation of the stereoscopic media item. Modifying the appearance of the first representation of the stereoscopic media item means modifying the visual effect such that the visual effect obscures a second part of the first representation of the stereoscopic media item that is different from the first part. The method according to claim 3, including the method described in claim 3.
5. Changing the appearance of the first representation of the stereoscopic media item is: The first method involves changing the appearance of the first representation of the stereoscopic media item, which is displayed to the user's left eye rather than the user's right eye, The method according to claim 3, comprising modifying the appearance of the first representation of the stereoscopic media item, which is displayed to the user's right eye rather than the user's left eye, in a second manner, the second manner being different from the first manner.
6. The method according to claim 1, wherein the visual effect has a visual characteristic that decreases over a plurality of values as the visual effect extends inward from at least the first edge of the first representation of the stereoscopic media item toward the interior of the first representation of the stereoscopic media item.
7. The method according to claim 6, wherein the first representation of the stereoscopic media item includes content at the first edge of the first representation of the stereoscopic media item, and the visual effect is blurring of the content at the first edge of the first representation of the stereoscopic media item.
8. Displaying the user interface includes displaying a blurred area, and the stereoscopic media item includes a first plurality of edges. In accordance with the determination that the first plurality of edges of the stereoscopic media item include the first content, the blurred region includes a blurred representation of the first content, and the blurred representation of the first content is based on the extrapolation of the first content. The method according to claim 1, wherein, in accordance with the determination that the first plurality of edges of the stereoscopic media item include second content, the blurred region includes a blurred representation of the second content, and the blurred representation of the second content is based on extrapolation of the second content.
9. The method according to claim 1, wherein the first representation of the stereoscopic media item is displayed with a vignette effect.
10. The method according to claim 1, wherein displaying the user interface includes displaying a virtual portal, and the first representation of the stereoscopic media item is displayed within the virtual portal.
11. The first representation of the stereoscopic media item includes a foreground portion and a background portion, and the method is While the first representation of the stereoscopic media item is being displayed, a third change in the user's viewpoint is detected. The method according to claim 1, further comprising detecting the third change in the user's viewpoint, and, based on the third change in the user's viewpoint, moving the foreground portion of the first representation of the stereoscopic media item relative to the background portion of the first representation of the stereoscopic media item via the display generation component.
12. The method according to claim 10, wherein displaying the virtual portal includes displaying a first portion of the virtual portal at a first location in a first representation of a physical environment located at a first distance from a first viewpoint of a user, and displaying the first representation of the stereoscopic media item includes displaying the first representation of the stereoscopic media item at a second location in the first representation of the physical environment located at a second distance from the first viewpoint of the user, wherein the second distance is greater than the first distance.
13. The method according to claim 10, wherein the virtual portal is a three-dimensional virtual object having a certain amount of thickness, and the amount of thickness displayed to the user is directly correlated to the angle of the user's viewpoint relative to the display of the virtual portal.
14. The method according to claim 1, wherein displaying the first representation of the stereoscopic media item includes displaying a specular reflection effect at the first edge of the first representation of the stereoscopic media item.
15. The first representation of the stereoscopic media item is displayed in a first location within a second representation of the physical environment, the first location being a first distance from the user's first viewpoint, and the method is as follows: While the first representation of the stereoscopic media item is displayed in the first location within the second representation of the physical environment, a request is detected to display the first representation of the stereoscopic media item in an immersive view, wherein displaying the first representation of the stereoscopic media item in an immersive view means A second location within the second representation of the physical environment, the second location being a second distance from the user's first viewpoint, wherein the second distance is shorter than the first distance, from which the first representation of the stereoscopic media item is displayed. This includes increasing the size of the first representation of the stereoscopic media item along a plane of the first representation of the stereoscopic media item that is parallel to the path between the first location in the second representation of the physical environment and the second location in the second representation of the physical environment, In response to detecting the request to display the first representation of the stereoscopic media item in the immersive appearance, the first representation of the stereoscopic media item is displayed in the immersive appearance, The method according to claim 2, further comprising:
16. When the first representation of the stereoscopic media item is displayed in the immersive view, The method according to claim 15, wherein the content included in the first representation of the stereoscopic media item is displayed at a separate scale having a predetermined relationship with the scale of the captured object when the stereoscopic media item was captured.
17. Displaying the first representation of the stereoscopic media item having an immersive appearance before detecting the request to display the first representation of the stereoscopic media item having an immersive appearance includes displaying the first representation of the stereoscopic media item having a non-immersive appearance. The first representation of the stereoscopic media item, when displayed in the immersive view, occupies a first angular range of the user's second viewpoint. The method according to claim 15, wherein the first representation of the stereoscopic media item, when displayed in the non-immersive view, occupies a second angular range of the user's second viewpoint, the second angular range being smaller than the first angular range.
18. The stereoscopic media item includes a second plurality of edges, and the method is While the first representation of the stereoscopic media item is displayed in the immersive appearance, In accordance with the determination that the second plurality of edges of the stereoscopic media item include the third content, the first edge of the first representation of the stereoscopic media item displays a representation of the third content, wherein the representation of the third content is based on the extrapolation of the third content. The method according to claim 15, further comprising, in accordance with the determination that the second plurality of edges of the stereoscopic media item include the fourth content, displaying an expression of the fourth content at the first edge of the first expression of the stereoscopic media item, wherein the expression of the fourth content is based on an extrapolation of the fourth content.
19. A non-stereoscopic media item, wherein the non-stereoscopic media item displays a representation of the non-stereoscopic media item that is displayed without the aforementioned visual effects. The method according to claim 1, further comprising:
20. A second representation of a second stereoscopic media item, distinct from the first representation of the stereoscopic media item, wherein the second representation of the second stereoscopic media item is displayed together with a second visual effect, the second visual effect obscuring at least a first portion of the second representation of the second stereoscopic media item, and displaying the second representation of the second stereoscopic media item extending inward from at least a first edge toward the interior of the second representation of the second stereoscopic media item. The method according to claim 1, further comprising:
21. A computer program that causes a computer system to perform the method described in any one of claims 1 to 20.
22. A computer system, A memory storing the computer program described in claim 21, One or more processors capable of executing the computer program stored in the memory Equipped with, The computer system is configured to communicate with a display generation component.
23. A computer system configured to communicate with display generation components, A computer system comprising means for performing the method described in any one of claims 1 to 20.