Device for, method of, and graphical user interface for capturing and displaying medium

The computer system addresses inefficiencies in media capture and display by providing an intuitive, real-time media capture preview and power-efficient interfaces for virtual/extended reality environments, enhancing user interaction and conserving energy.

JP2025172725APending Publication Date: 2025-11-26APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025118467
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-22
Filing Date
2025-07-14
Publication Date
2025-11-26

Smart Images

  • Figure 2025172725000001_ABST
    Figure 2025172725000001_ABST
Patent Text Reader

Abstract

To provide a method, a device and a user interface which are further efficient and intuitive for a user and is directed to capturing and display of a medium.SOLUTION: A user interface display method according to the present invention has a step of displaying an extended reality camera user interface through a display generation element in a computer system communicating with the display generation element and at least one or more cameras. The extended reality camera interface has a recording indicator for indicating expression of a physical environment and a recording region in a visual field of a camera. The recording indicator has a first edge region having a visual parameter. A value of the visual parameter is gradually decreased as a distance from the first edge region of the recording indicator is increased.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is related to U.S. Patent Application No. 17 / 992,789, filed November 22, 2022, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING AND DISPLAYING MEDIA," U.S. Provisional Patent Application No. 63 / 409,690, filed September 23, 2022, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING AND DISPLAYING MEDIA," U.S. Provisional Patent Application No. 63 / 338,864, filed May 5, 2022, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING AND DISPLAYING MEDIA," and U.S. Provisional Patent Application No. 63 / 338,864, filed December 3, 2021, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR This application claims priority to U.S. Provisional Patent Application No. 63 / 285,897, entitled "PATENT CAPTURING AND DISPLAYING MEDIA," the contents of each of which are incorporated herein by reference in their entirety.

[0002] Technical Field The present disclosure generally relates to computer systems in communication with display generation components, optionally one or more input devices, and one or more cameras, that provide computer-generated experiences, including, but not limited to, electronic devices that provide virtual reality and mixed reality experiences via a display. [Background technology]

[0003] The development of computer systems for capturing and / or displaying media in various environments, such as extended reality environments, has increased significantly in recent years. Exemplary extended reality environments include at least some virtual elements that replace or augment the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with the virtual / extended reality environment. Exemplary virtual elements include virtual objects such as digital images, video, text, icons, and control elements such as buttons and other graphics. Summary of the Invention

[0004] Some methods and interfaces for capturing and / or displaying media in various environments are cumbersome, inefficient, and limited. For example, systems that provide insufficient visual feedback for capturing media, systems that require a complex series of inputs to execute the media capture process, and systems that display media in a complex, tedious, and error-prone manner impose a significant cognitive burden on users and detract from the experience of the virtual / extended reality environment. In addition, these methods are unnecessarily time-consuming, thereby wasting computer system energy. This latter consideration is particularly important in battery-operated devices.

[0005] Thus, there is a need for computer systems with improved methods and interfaces for providing users with computer-generated experiences that make capturing and / or displaying media in a variety of environments more efficient and intuitive for users. Such methods and interfaces, optionally complement or replace conventional methods for capturing and / or displaying media in a variety of environments. Such methods and interfaces reduce the number, extent, and / or type of inputs from users by helping users understand the connection between provided inputs and device responses to those inputs, thereby creating a more efficient human-machine interface.

[0006] The above-mentioned drawbacks and other problems associated with user interfaces of computer systems are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to the display generating components, the output devices including one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or instruction sets stored in the memory for performing a plurality of functions. In some embodiments, a user interacts with the GUI through stylus and / or finger contacts and gestures on the touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body as captured by cameras and other movement sensors, and voice input as captured by one or more audio input devices.In some embodiments, the functions performed through the interactions optionally include image editing, drawing, presenting, word processing, creating spreadsheets, playing games, making phone calls, video conferencing, emailing, instant messaging, training support, digital photography, digital videography, web browsing, playing digital music, note taking, and / or playing digital videos, and executable instructions to perform those functions are optionally contained in a transient and / or non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0007] There is a need for electronic devices with improved methods and interfaces for capturing and / or displaying media in a variety of environments. Such methods and interfaces can complement or replace conventional methods for capturing and / or displaying media. Such methods and interfaces reduce the number, extent, and / or type of input from a user, creating a more efficient human-machine interface. For battery-operated computing devices, such methods and interfaces conserve power, increasing the time between battery charges and reducing the amount of processing power.

[0008] According to some embodiments, a method is described that is executed on a computer system in communication with a display generation component and one or more cameras. The method includes detecting, via a display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on the representation of the physical environment, the representation of the physical environment changing as a portion of the physical environment corresponding to the representation of the physical environment changes and / or a user's viewpoint changes; and in response to detecting the request to display the media capture user interface, displaying, via the display generation component, a media capture preview including a representation of a portion of a field of view of one or more cameras, with content that updates as a portion of the physical environment within the portion of the field of view of the one or more cameras changes, the media capture preview indicating a boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed, the media capture preview being displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before the request to display the media capture user interface is detected, and the media capture preview being displayed in place of a second portion of the representation of the physical environment, the first portion of the representation of the physical environment being updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes.

[0009] According to some embodiments, a non-transitory computer-readable storage medium is described, the non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs detecting, via the display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on the representation of the physical environment, the representation of the physical environment changing as portions of the physical environment corresponding to the representation of the physical environment change and / or as a user's viewpoint changes, and in response to detecting the request to display the media capture user interface, displaying, via the display generation component, a media capture preview including a representation of a portion of the field of view of the one or more cameras. and displaying a media capture preview with content that updates as portions of the physical environment within portions of the field of view of the one or more cameras change, wherein the media capture preview indicates a boundary of the media to be captured in response to detecting a media capture input while the media capture user interface is displayed, the media capture preview being displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before a request to display the media capture user interface is detected, the media capture preview being displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment being updated as portions of the physical environment corresponding to the first portion of the representation of the physical environment change and / or the user's viewpoint changes.

[0010] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs detecting, via the display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on the representation of the physical environment, the representation of the physical environment changing as portions of the physical environment corresponding to the representation of the physical environment change and / or as a user's viewpoint changes, and in response to detecting the request to display the media capture user interface, generating, via the display generation component, a media capture preview including a representation of a portion of the field of view of the one or more cameras. and displaying a media capture preview with content that updates as portions of the physical environment within the portions of the field of view of the camera change, wherein the media capture preview indicates boundaries of media to be captured in response to detecting a media capture input while the media capture user interface is displayed, the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before a request to display the media capture user interface is detected, the media capture preview is displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as portions of the physical environment corresponding to the first portion of the representation of the physical environment change and / or as the user's viewpoint changes.

[0011] According to some embodiments, a computer system in communication with a display generation component and one or more cameras is described, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs detecting, via the display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on the representation of the physical environment, the representation of the physical environment changing as portions of the physical environment corresponding to the representation of the physical environment change and / or as a user's viewpoint changes; and, in response to detecting the request to display the media capture user interface, displaying, via the display generation component, a media capture preview including representations of portions of the field of view of the one or more cameras. and displaying a media capture preview with content that updates as the portion of the physical environment within the portion changes, wherein the media capture preview indicates a boundary of the media to be captured in response to detecting a media capture input while the media capture user interface is displayed, the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before a request to display the media capture user interface is detected, the media capture preview is displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as the portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes.

[0012] According to some embodiments, a computer system in communication with a display generation component and one or more cameras is described. The computer system comprises means for detecting, via a display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on the representation of the physical environment, the representation of the physical environment changing as a portion of the physical environment corresponding to the representation of the physical environment changes and / or a user's viewpoint changes; and means for displaying, via the display generation component, a media capture preview including a representation of a portion of the field of view of one or more cameras, with content that updates as a portion of the physical environment within the portion of the field of view of the one or more cameras changes in response to detecting the request to display the media capture user interface, the media capture preview indicating a boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed, the media capture preview being displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before the request to display the media capture user interface is detected, and the media capture preview being displayed in place of a second portion of the representation of the physical environment, the first portion of the representation of the physical environment being updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes.

[0013] According to some embodiments, a computer program product is described, comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras. The one or more programs include instructions to detect, via a display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on the representation of the physical environment, the representation of the physical environment changing as a portion of the physical environment corresponding to the representation of the physical environment changes and / or a user's viewpoint changes; and in response to detecting the request to display the media capture user interface, display, via the display generation component, a media capture preview including a representation of a portion of a field of view of one or more cameras, with content that updates as a portion of the physical environment within the portion of a field of view of the one or more cameras changes, the media capture preview indicating a boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed, the media capture preview being displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before the request to display the media capture user interface was detected, and the media capture preview being displayed in place of a second portion of the representation of the physical environment, the first portion of the representation of the physical environment being updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes.

[0014] According to some embodiments, a method is described that is executed on a computer system in communication with a display generation component and one or more cameras. The method includes displaying an extended reality user interface via a display generation component including a preview of one or more camera field of view overlaid on a first portion of the three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed together with an individual spatial configuration relative to the user's viewpoint; detecting a change in the user's viewpoint pose from the first pose to a second pose different from the first pose; and in response to detecting the change in the user's viewpoint pose from the first pose to the second pose, shifting the preview of the one or more camera field of view away from the individual spatial configuration relative to the user's viewpoint in a direction determined based on the change in the user's viewpoint pose from the first pose to the second pose, wherein the shifting of the preview of the one or more camera field of view is performed at a first speed, and while the preview of the one or more camera field of view shifts based on the change in the user's viewpoint pose, the representation of the three-dimensional environment changes based on the change in the user's viewpoint pose at a second speed different from the first speed.

[0015] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs displaying, via the display generating component, an extended reality user interface including a preview of the field of view of the one or more cameras overlaid on a first portion of the three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration relative to the user's viewpoint; and and detecting a change in the user's viewpoint pose to a second pose different from the first pose; and in response to detecting the change in the user's viewpoint pose from the first pose to the second pose, shifting a preview of the field of view of one or more cameras away from the distinct spatial configuration relative to the user's viewpoint in a direction determined based on the change in the user's viewpoint pose from the first pose to the second pose, wherein the shifting of the preview of the field of view of the one or more cameras is performed at a first rate, and while the preview of the field of view of the one or more cameras is shifting based on the change in the user's viewpoint pose, the representation of the three-dimensional environment changes based on the change in the user's viewpoint pose at a second rate different from the first rate.

[0016] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs displaying, via the display generating component, an extended reality user interface including a preview of the field of view of the one or more cameras overlaid on a first portion of the three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration relative to the user's viewpoint, and and detecting a change in the user's viewpoint pose to a second pose different from the first pose; and in response to detecting the change in the user's viewpoint pose from the first pose to the second pose, shifting a preview of a field of view of one or more cameras away from the distinct spatial configuration relative to the user's viewpoint in a direction determined based on the change in the user's viewpoint pose from the first pose to the second pose, wherein the shifting of the preview of the field of view of the one or more cameras is performed at a first rate, and while the preview of the field of view of the one or more cameras is shifting based on the change in the user's viewpoint pose, the representation of the three-dimensional environment changes based on the change in the user's viewpoint pose at a second rate different from the first rate.

[0017] According to some embodiments, a computer system in communication with a display generation component and one or more cameras is described. The computer system includes one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs displaying, via the display generation component, an extended reality user interface including a preview of the field of view of the one or more cameras overlaid on a first portion of the three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration relative to the user's viewpoint; and a display generating component that displays ... and, in response to detecting the change in the user's viewpoint pose from the first pose to the second pose, shifting a preview of a field of view of one or more cameras away from the distinct spatial configuration relative to the user's viewpoint in a direction determined based on the change in the user's viewpoint pose from the first pose to the second pose, wherein the shifting of the preview of the field of view of the one or more cameras is performed at a first rate, and while the preview of the field of view of the one or more cameras is shifting based on the change in the user's viewpoint pose, the representation of the three-dimensional environment changes based on the change in the user's viewpoint pose at a second rate different from the first rate.

[0018] According to some embodiments, a computer system in communication with a display generation component and one or more cameras is described. The computer system comprises: means for displaying, via a display generation component, an extended reality user interface including a preview of one or more camera field of view overlaid on a first portion of the three-dimensional environment visible from the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed together with an individual spatial configuration relative to the user's viewpoint; means for detecting a change in the user's viewpoint pose from the first pose to a second pose different from the first pose; and means for shifting the preview of the one or more camera field of view away from the individual spatial configuration relative to the user's viewpoint in a direction determined based on the change in the user's viewpoint pose from the first pose to the second pose in response to detecting the change in the user's viewpoint pose from the first pose to the second pose, wherein the shifting of the preview of the one or more camera field of view is performed at a first speed, and while the preview of the one or more camera field of view shifts based on the change in the user's viewpoint pose, the representation of the three-dimensional environment changes based on the change in the user's viewpoint pose at a second speed different from the first speed.

[0019] According to some embodiments, a computer program product is described, the computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs displaying, via the display generation component, an extended reality user interface including a preview of one or more camera fields of view overlaid on a first portion of a three-dimensional environment visible at a user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration relative to the user's viewpoint, and wherein the preview includes a representation of the first portion of the three-dimensional environment and a distinct spatial configuration relative to the user's viewpoint, the preview including a preview of one or more camera fields of view overlaid on a first portion of a three-dimensional environment visible at a user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration relative to the user's viewpoint, the preview including a preview ... and, in response to detecting the change in the user's viewpoint pose from the first pose to the second pose, shifting a preview of a field of view of one or more cameras away from the distinct spatial configuration relative to the user's viewpoint in a direction determined based on the change in the user's viewpoint pose from the first pose to the second pose, wherein the shifting of the preview of the field of view of the one or more cameras is performed at a first rate, and while the preview of the field of view of the one or more cameras is shifting based on the change in the user's viewpoint pose, the representation of the three-dimensional environment changes based on the change in the user's viewpoint pose at a second rate different from the first rate.

[0020] According to some embodiments, the method is performed on a computer system in communication with a display generation component, the method including: detecting, while displaying an extended reality environment user interface, a request to display captured media including immersive content that provides a first set of visual cues that a user is at least partially surrounded by the immersive content when viewed from one or more discrete ranges of viewpoints; and, in response to detecting the request to display the captured media, displaying the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that the user's first perspective is outside the discrete ranges of the one or more viewpoints.

[0021] According to some embodiments, a non-transitory computer-readable storage medium is described, the non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions to: while displaying an extended reality environment user interface, detect a request to display captured media including immersive content that provides a first set of visual cues that a user is at least partially surrounded by the immersive content when viewed from one or more discrete ranges of viewpoints; and, in response to detecting the request to display the captured media, display the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that the user's first perspective is outside the discrete ranges of the one or more viewpoints.

[0022] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions to: while displaying an extended reality environment user interface, detect a request to display captured media including immersive content that provides a first set of visual cues that a user is at least partially surrounded by the immersive content when viewed from one or more discrete ranges of viewpoints; and, in response to detecting the request to display the captured media, display the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that the user's first viewpoint is outside the discrete range of the one or more viewpoints.

[0023] According to some embodiments, a computer system in communication with a display generation component includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to: while displaying an extended reality environment user interface, detect a request to display captured media including immersive content that provides a first set of visual cues that a user is at least partially surrounded by the immersive content when viewed from one or more discrete ranges of viewpoints; and, in response to detecting the request to display the captured media, display the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that the user's first perspective is outside the discrete range of the one or more viewpoints.

[0024] According to some embodiments, a computer system in communication with a display generation component is described, comprising: means for detecting, while displaying an extended reality environment user interface, a request to display captured media including immersive content that provides a first set of visual cues that a user is at least partially surrounded by the immersive content when viewed from one or more discrete ranges of viewpoints; and means for displaying the captured media in response to detecting the request to display the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that the user's first perspective is outside the discrete ranges of the one or more viewpoints.

[0025] According to some embodiments, a computer program product is described comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: while displaying an extended reality environment user interface, detecting a request to display captured media including immersive content that provides a first set of visual cues that a user is at least partially surrounded by the immersive content when viewed from one or more discrete ranges of viewpoints; and, in response to detecting the request to display the captured media, displaying the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that the user's first perspective is outside the discrete range of the one or more perspectives.

[0026] According to some embodiments, a method is described that is executed on a computer system in communication with a display generation component and one or more cameras. The method includes displaying, via the display generation component, an extended reality camera user interface that includes a representation of a physical environment and a recording indicator that indicates a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance from the first edge region of the recording indicator increases.

[0027] According to some embodiments, a non-transitory computer-readable storage medium is described, the non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for displaying, via the display generation component, an extended reality camera user interface including a representation of a physical environment and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance from the first edge region of the recording indicator increases.

[0028] According to some embodiments, a temporary computer-readable storage medium is described, the temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for displaying, via the display generation component, an extended reality camera user interface including a representation of a physical environment and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance from the first edge region of the recording indicator increases.

[0029] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to, via the display generation component, display an extended reality camera user interface including a representation of a physical environment and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance from the first edge region of the recording indicator increases.

[0030] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described, the computer system comprising means for displaying, via the display generation component, an extended reality camera user interface including a representation of a physical environment and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance from the first edge region of the recording indicator increases.

[0031] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs include instructions for displaying, via the display generation component, an extended reality camera user interface, the extended reality camera user interface including a representation of a physical environment and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance from the first edge region of the recording indicator increases.

[0032] According to some embodiments, a method is described that is executed in a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The method includes detecting a request to display a camera user interface via the one or more input devices, and in response to detecting the request to display the camera user interface, displaying a camera user interface, the camera user interface including a reticle virtual object that indicates a capture area of ​​the one or more cameras. Displaying the camera user interface includes displaying the camera user interface with a tutorial within the camera user interface, the tutorial providing information on how to capture media using the computer system while the camera user interface is displayed, in accordance with a determination that the set of one or more criteria is met, and displaying the camera user interface without displaying the tutorial in accordance with a determination that the set of one or more criteria is not met.

[0033] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more input devices, and one or more cameras, the one or more programs including instructions for detecting a request to display a camera user interface via the one or more input devices, and in response to detecting the request to display the camera user interface, displaying the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras, wherein displaying the camera user interface includes, in accordance with a determination that a set of one or more criteria is met, displaying the camera user interface with a tutorial within the camera user interface, the tutorial providing information on how to capture media using the computer system while the camera user interface is displayed, and in accordance with a determination that the set of one or more criteria is not met, displaying the camera user interface without displaying the tutorial.

[0034] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more input devices, and one or more cameras, the one or more programs including instructions for detecting a request to display a camera user interface via the one or more input devices, and in response to detecting the request to display the camera user interface, displaying the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras, wherein displaying the camera user interface includes, in accordance with a determination that a set of one or more criteria is met, displaying the camera user interface with a tutorial within the camera user interface, the tutorial providing information on how to capture media using the computer system while the camera user interface is displayed, and in accordance with a determination that the set of one or more criteria is not met, displaying the camera user interface without displaying the tutorial.

[0035] According to some embodiments, a computer system is described, the computer system including one or more processors, the computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for detecting a request to display a camera user interface via the one or more input devices, and in response to detecting the request to display the camera user interface, displaying the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras, wherein displaying the camera user interface includes, in accordance with a determination that a set of one or more criteria is met, displaying the camera user interface with a tutorial within the camera user interface, the tutorial providing information on how to capture media using the computer system while the camera user interface is displayed, and in accordance with a determination that the set of one or more criteria is not met, displaying the camera user interface without displaying the tutorial.

[0036] According to some embodiments, a computer system is described that is configured to communicate with a display generation component, one or more input devices, and one or more cameras, the computer system comprising: means for detecting a request to display a camera user interface via the one or more input devices; and means for displaying a camera user interface in response to detecting the request to display the camera user interface, the camera user interface including a reticle virtual object that indicates a capture area of ​​the one or more cameras, wherein displaying the camera user interface includes: displaying the camera user interface with a tutorial within the camera user interface in accordance with a determination that a set of one or more criteria is met, the tutorial providing information on how to capture media using the computer system while the camera user interface is displayed; and displaying the camera user interface without displaying the tutorial in accordance with a determination that the set of one or more criteria is not met.

[0037] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs including instructions for detecting a request to display a camera user interface via the one or more input devices, and in response to detecting the request to display the camera user interface, displaying the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras, wherein displaying the camera user interface includes: in accordance with a determination that a set of one or more criteria is met, displaying the camera user interface with a tutorial within the camera user interface, the tutorial providing information on how to capture media using the computer system while the camera user interface is displayed; and in accordance with a determination that the set of one or more criteria is not met, displaying the camera user interface without displaying the tutorial.

[0038] According to some embodiments, a method is described that is executed on a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The method includes displaying, via a display generation component, a user interface including a representation of a physical environment, where a first portion of the representation of the physical environment is inside a capture area of ​​one or more cameras and a second portion of the representation of the physical environment is outside the capture area of ​​the one or more cameras; and a viewfinder, where the viewfinder includes a boundary; detecting, while displaying the user interface, a first request to capture media via one or more input devices; and in response to detecting the first request to capture media, capturing a first media item including at least the first portion of the representation of the physical environment using the one or more cameras and altering an appearance of the viewfinder, where altering the appearance of the viewfinder includes altering an appearance of a first portion of content within a threshold distance on a first side of a boundary of the viewfinder and altering an appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder, different from the first side of the boundary of the viewfinder.

[0039] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs displaying, via the display generation component, a user interface including a representation of a physical environment, where a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside a capture area of ​​the one or more cameras, and a viewfinder, where the viewfinder includes a boundary. and instructions for detecting, via one or more input devices, a first request to capture media while displaying the user interface; and, in response to detecting the first request to capture media, capturing a first media item using one or more cameras, the first media item comprising at least a first portion of a representation of the physical environment; and altering an appearance of the viewfinder, wherein altering the appearance of the viewfinder includes altering an appearance of a first portion of the content within a threshold distance on a first side of a boundary of the viewfinder and altering an appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder, different from the first side of the boundary of the viewfinder.

[0040] According to some embodiments, a temporary computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs displaying, via the display generation component, a user interface including a representation of a physical environment, where a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside a capture area of ​​the one or more cameras, and a viewfinder, where the viewfinder includes a boundary, and the user can view the user interface. The system includes instructions for detecting, via one or more input devices, a first request to capture media while displaying the interface; and, in response to detecting the first request to capture media, capturing a first media item using one or more cameras, the first media item including at least a first portion of a representation of the physical environment; and altering the appearance of the viewfinder, wherein altering the appearance of the viewfinder includes altering the appearance of a first portion of the content within a threshold distance on a first side of a boundary of the viewfinder; and altering the appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder, different from the first side of the boundary of the viewfinder.

[0041] According to some embodiments, a computer system is described, the computer system comprising: one or more processors configured to communicate with a display generation component, one or more input devices, and one or more cameras; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs displaying, via the display generation component, a user interface including a representation of a physical environment, where a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside a capture area of ​​the one or more cameras; and a viewfinder, where the viewfinder includes a boundary. and, while displaying the user interface, detecting a first request to capture media via one or more input devices; and, in response to detecting the first request to capture media, capturing a first media item using one or more cameras, the first media item including at least a first portion of a representation of the physical environment; and altering the appearance of the viewfinder, wherein altering the appearance of the viewfinder includes altering the appearance of a first portion of the content within a threshold distance on a first side of a boundary of the viewfinder and altering the appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder, different from the first side of the boundary of the viewfinder.

[0042] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component, one or more input devices, and one or more cameras, and the computer system comprises: means for displaying, via the display generation component, a user interface including a representation of a physical environment, where a first portion of the representation of the physical environment is inside a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside the capture area of ​​the one or more cameras; and a viewfinder, where the viewfinder includes a boundary; means for detecting, while displaying the user interface, a first request to capture media via the one or more input devices; and means for capturing, using the one or more cameras, a first media item including at least the first portion of the representation of the physical environment in response to detecting the first request to capture media, and altering the appearance of the viewfinder, wherein altering the appearance of the viewfinder includes altering the appearance of a first portion of content within a threshold distance on a first side of a boundary of the viewfinder and altering the appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder, different from the first side of the boundary of the viewfinder.

[0043] According to some embodiments, a computer program product is described, the computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs displaying, via the display generation component, a user interface including a representation of a physical environment, where a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside a capture area of ​​the one or more cameras, and a viewfinder, where the viewfinder includes a boundary. The method includes instructions for detecting, via one or more input devices, a first request to capture media while displaying the interface; and, in response to detecting the first request to capture media, capturing a first media item using one or more cameras, the first media item comprising at least a first portion of a representation of the physical environment; and altering the appearance of the viewfinder, wherein altering the appearance of the viewfinder includes altering the appearance of a first portion of the content within a threshold distance on a first side of a boundary of the viewfinder and altering the appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.

[0044] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification, and claims. Furthermore, it should be noted that the language used in this specification has been selected solely for the purposes of readability and explanation, and not to define or limit the subject matter of the present invention. [Brief explanation of the drawings]

[0045] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout:

[0046] [Figure 1] FIG. 1 is a block diagram illustrating an operating environment for a computer system for providing an XR experience, according to some embodiments.

[0047] [Figure 2] FIG. 1 is a block diagram illustrating a controller of a computer system configured to manage and coordinate an XR experience for a user, according to some embodiments.

[0048] [Figure 3] FIG. 1 is a block diagram illustrating display generation components of a computer system configured to provide a user with visual components of an XR experience, according to some embodiments.

[0049] [Figure 4] FIG. 1 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input, according to some embodiments.

[0050] [Figure 5] FIG. 1 is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input, according to some embodiments.

[0051] [Figure 6] FIG. 1 is a flow diagram illustrating a glint-assisted gaze tracking pipeline according to some embodiments.

[0052] [Figure 7A] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7B]1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7C] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7D] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7E] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7F] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7G] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7H] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7I] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7J] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7K] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7L] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7M] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7N]1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7O] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7P] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments. [Figure 7Q] 1 illustrates exemplary techniques for capturing and / or displaying media in some environments, according to some embodiments.

[0053] [Figure 8] FIG. 1 is a flow diagram of a method for capturing media, according to some embodiments.

[0054] [Figure 9] 1 is a flow diagram of a method for displaying a preview of media, according to some embodiments.

[0055] [Figure 10] FIG. 1 is a flow diagram of a method for displaying previously captured media, according to some embodiments.

[0056] [Figure 11A] 1 illustrates an exemplary technique for displaying a representation of a physical environment with a recording indicator, according to some embodiments. [Figure 11B] 1 illustrates an exemplary technique for displaying a representation of a physical environment with a recording indicator, according to some embodiments. [Figure 11C] 1 illustrates an exemplary technique for displaying a representation of a physical environment with a recording indicator, according to some embodiments. [Figure 11D] 1 illustrates an exemplary technique for displaying a representation of a physical environment with a recording indicator, according to some embodiments.

[0057] [Figure 12] FIG. 1 is a flow diagram of a method for displaying a representation of a physical environment using a recording indicator, according to some embodiments.

[0058] [Figure 13A] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13B] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13C] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13D] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13E1] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13E2] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13E3] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13E4] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13E5] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13F] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13G] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13H] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13I] 1 illustrates an exemplary technique for displaying a camera user interface. [Figure 13J] 1 illustrates an exemplary technique for displaying a camera user interface.

[0059] [Figure 14]FIG. 1 is a flow diagram of a method for displaying information related to captured media, according to some embodiments.

[0060] [Figure 15A] FIG. 1 is a flow diagram of a method for modifying the appearance of a viewfinder, according to some embodiments. [Figure 15B] FIG. 1 is a flow diagram of a method for modifying the appearance of a viewfinder, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0061] The present disclosure relates to a user interface that provides an extended reality (XR) experience to a user, according to some embodiments.

[0062] FIGS. 1-6 provide an illustration of an exemplary computer system for providing an XR experience to a user. FIGS. 7A-7Q illustrate an exemplary technique for capturing and / or displaying media in various environments, according to some embodiments. FIG. 8 is a flow diagram of a method for capturing and viewing media, according to various embodiments. FIG. 9 is a flow diagram of a method for displaying a preview of media, according to various embodiments. FIG. 10 is a flow diagram of a method for displaying previously captured media, according to various embodiments. The user interfaces of FIGS. 7A-7Q illustrate the processes of FIGS. 8, 9, and 10. FIGS. 11A-11D illustrate an exemplary technique for displaying a representation of a physical environment with a recording indicator, according to some embodiments. FIG. 12 is a flow diagram of a method for displaying a representation of a physical environment with a recording indicator, according to some embodiments. The user interfaces of FIGS. 11A-11D illustrate the process of FIG. 12. FIGS. 13A-13J illustrate an exemplary technique for displaying a camera user interface, according to some embodiments. FIG. 14 is a flow diagram of a method for displaying information related to captured media, according to some embodiments. Figures 15A-15B are flow diagrams of methods for changing the appearance of a viewfinder, according to some embodiments. The user interfaces of Figures 13A-13J illustrate the processes of Figures 14, 15A, and 15B.

[0063] The processes described below enhance device usability and make user-device interfaces more efficient (e.g., by helping users provide appropriate inputs and reducing user errors when operating / interacting with the device) through various techniques, including providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional controls, performing an operation without requiring further user input when a set of conditions is met, improving privacy and / or security, providing a more diverse, detailed, and / or realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and improve device battery life by allowing users to use the device more quickly and efficiently. Saving battery power, and therefore weight, improves device ergonomics. These techniques also enable real-time communication and the use of fewer and / or less accurate sensors, resulting in more compact, lighter, and less expensive devices, and allowing devices to be used in a variety of lighting conditions. These techniques reduce energy use and thereby reduce the heat given off by the device, which is particularly important for wearable devices where a device that is well within the operating parameters for the device components may become uncomfortable for the user to wear if it is generating too much heat.

[0064] Furthermore, for methods described herein in which one or more steps are conditioned on one or more conditions being satisfied, it should be understood that the described method can be repeated in multiple iterations, such that over the course of the iterations, all of the conditions on which the method steps are conditioned are satisfied in different iterations of the method. For example, if a method requires performing a first step if a condition is satisfied and a second step if the condition is not satisfied, one skilled in the art will understand that the steps recited in the claim are repeated in a particular order until the conditions are satisfied and then no longer satisfied. Thus, a method described with one or more steps that depend on one or more conditions being satisfied can be rewritten as a method that is repeated until each condition recited in the method is satisfied. However, this is not required for system or computer-readable medium claims in which the system or computer-readable medium includes instructions for performing a conditional action based on the satisfaction of the corresponding one or more conditions, and thus can determine whether a contingency is met without explicitly repeating the method steps until all conditions on which the method steps are conditioned are satisfied. Those skilled in the art will also understand that, as with methods having conditional steps, the system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all of the conditional steps have been performed.

[0065] 1, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a velocity sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a consumer electronics device, a wearable device, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., within a head-mounted or handheld device).

[0066] When describing an XR experience, various terms are used to individually refer to several related, but distinct, environments that a user senses and / or can interact with (e.g., using inputs detected by computer system 101 that cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to computer system 101 generating the XR experience). The following is a subset of these terms:

[0067] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through their senses, such as sight, touch, hearing, taste, and smell.

[0068] Extended reality: In contrast, an extended reality (XR) environment refers to a wholly or partially mimicked environment that people sense and / or interact with through electronic systems. In XR, a subset of a person's body movements or representations thereof are tracked, and one or more properties of one or more virtual objects simulated within the XR environment are adjusted accordingly to behave according to at least one law of physics. For example, an XR system may detect a person's head rotation and adjust the graphical content and sound field presented to the person accordingly, in a manner similar to how such views and sounds change in a physical environment. In some circumstances (e.g., for accessibility reasons), adjustments to the property(ies) of virtual object(s) in the XR environment may be made in response to representations of body movements (e.g., voice commands). A person may sense and / or interact with an XR object using any one of these senses, including sight, hearing, touch, taste, and smell. For example, a person may sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. In another example, audio objects may enable audio transparency that selectively incorporates ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, a person may sense and / or interact with only audio objects.

[0069] Examples of XR include virtual reality and mixed reality.

[0070] Virtual Reality: A virtual reality (VR) environment refers to an emulated environment designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with virtual objects in the VR environment through a simulation of the person's presence in the computer-generated environment and / or through a simulation of a subset of the person's physical movement within the computer-generated environment.

[0071] Mixed Reality: A mixed reality (MR) environment refers to a mimicked environment designed to incorporate sensory input from or representations of a physical environment in addition to including computer-generated sensory input (e.g., virtual objects), as opposed to a VR environment designed to be based entirely on computer-generated sensory input. On a virtuality continuum, a mixed reality environment is anywhere between, but not including, a complete physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input may respond to changes in sensory input from the physical environment. Some electronic systems for presenting MR environments may also track location and / or orientation relative to the physical environment to allow virtual objects to interact with real objects (i.e., physical items or representations thereof from the physical environment). For example, the system may take into account movement so that a virtual tree appears stationary relative to the physical ground.

[0072] Examples of mixed reality include extended reality and augmented virtuality.

[0073] Extended reality: An extended reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person uses the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system composites the images or videos with the virtual objects and presents the composite on the opaque display. The person uses the system to indirectly view the physical environment through the images or videos of the physical environment and perceive the virtual objects superimposed on the physical environment. As used herein, video of a physical environment shown on an opaque display is referred to as "pass-through video," meaning that the system captures images of the physical environment using one or more image sensors and uses those images in presenting the AR environment on the opaque display. Alternatively, the system may include a projection system that projects virtual objects, e.g., as holograms, into the physical environment or onto a physical surface, such that a person using the system perceives the virtual objects superimposed on the physical environment. An extended reality environment also refers to an imitative environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, when providing pass-through video, the system may distort one or more sensor images to impose a selected perspective (e.g., viewpoint) other than the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be distorted by graphically modifying (e.g., enlarging) a portion thereof, such that the modified portion becomes a non-photorealistic, altered version that represents the originally captured image.As a further example, the representation of the physical environment may be altered by graphically removing or obscuring portions of it.

[0074] Augmented Virtuality: An augmented virtuality (AV) environment refers to a mimicking environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs may be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, while people with faces are realistically recreated from images of physical people. As another example, virtual objects may adopt the shape or color of physical items imaged by one or more imaging sensors. As a further example, virtual objects may adopt shadows that match the position of the sun in the physical environment.

[0075] Perspective-Locked Virtual Object: A virtual object is perspective-locked when the computer system displays the virtual object in the same location and / or position within the user's perspective, even as the user's perspective shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user's perspective is locked to the forward-facing orientation of the user's head (e.g., the user's perspective is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's perspective remains fixed even as the user's line of sight moves without moving the user's head. In embodiments in which the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's perspective is the extended reality view being presented to the user on the display generating component of the computer system. For example, a perspective-locked virtual object displayed in the upper left corner of the user's perspective when the user's perspective is in a first orientation (e.g., the user's head is facing north) will continue to be displayed in the upper left corner of the user's perspective even if the user's perspective changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position at which a viewpoint-locked virtual object is displayed in a user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."

[0076] Environment-Locked Virtual Object: A virtual object is environment-locked (or "world-locked") when a computer system displays the virtual object at a location and / or position within a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) locations and / or objects within a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint shifts, the locations and / or objects within the environment relative to the user's viewpoint change, resulting in the environment-locked virtual object appearing at a different location and / or position within the user's viewpoint. For example, an environment-locked virtual object locked to a tree directly in front of the user will appear centered within the user's viewpoint. If the user's viewpoint shifts to the right (e.g., the user's head is turned to the right) and the tree becomes more left-leaning within the user's viewpoint (e.g., the position of the tree within the user's viewpoint shifts), the environment-locked virtual object locked to the tree will appear more left-leaning within the user's viewpoint. In other words, the location and / or position at which the environment-locked virtual object appears within the user's viewpoint depends on the position and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system fixed to a fixed location and / or object in the physical environment) to determine a position at which to display an environment-locked virtual object in the user's viewpoint. The environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object) or can be locked to a moving portion of the environment (e.g., a vehicle, an animal, a person, or a representation of a part of the user's body that moves independent of the user's viewpoint, such as the user's hand, wrist, arm, or leg), so that the virtual object moves as the viewpoint or part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.

[0077] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits delayed-following behavior, which reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting delayed-following behavior, the computer system intentionally delays the movement of the virtual object when it detects movement of the reference point that the virtual object is following (e.g., a part of the environment, the viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 and 300 cm from the viewpoint). For example, when the reference point (e.g., a part of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but at a second speed that is slower than the first speed (e.g., until the reference point stops or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits delayed-following behavior, the device ignores small amounts of movement of the reference point (e.g., ignores movement of the reference point that is less than a threshold amount, such as movement between 0 and 5 degrees or movement between 0 and 50 cm). For example, when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment or a viewpoint to which the virtual object is locked) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object initially increases (e.g., because the virtual object is displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases as the amount of movement of the reference point increases beyond a threshold (e.g., a “delayed following” threshold) as the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point.In some embodiments, a virtual object maintaining a substantially fixed position relative to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., above / below, left / right, and / or forward / backward relative to the position of the reference point).

[0078] Hardware: There are many different types of electronic systems that allow a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed over a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system to provide audio output. A head-mounted system may have one or more speaker(s) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. The head-mounted system may have a transparent or translucent display rather than an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser-scanned light source, or any combination of these technologies. The medium may be a light guide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. A projection-based system may employ retinal projection technology that projects a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface.In some embodiments, controller 110 is configured to manage and coordinate the XR experience for the user. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Controller 110 is described in more detail below with respect to FIG. 2. In some embodiments, controller 110 is a computing device that is local or remote to scene 105 (e.g., the physical environment). For example, controller 110 is a local server located within scene 105. In another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is contained within the housing (e.g., physical housing) of one or more of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the foregoing.

[0079] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least a visual component of an XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with respect to FIG. 3. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.

[0080] According to some embodiments, the display generation component 120 provides an XR experience to the user while the user is virtually and / or physically present in the scene 105.

[0081] In some embodiments, the display generating component is worn on a part of the user's body (e.g., on their head, their hand, etc.). Thus, display generating component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, display generating component 120 surrounds the user's field of view. In some embodiments, display generating component 120 is a handheld device (e.g., a smartphone or tablet) configured to present XR content, where the user holds the device with a display pointed toward the user's field of view and a camera pointed toward scene 105. In some embodiments, the handheld device is optionally located within a housing worn on the user's head. In some embodiments, the handheld device is optionally located on a support (e.g., a tripod) in front of the user. In some embodiments, display generating component 120 is an XR chamber, housing, or room configured to present XR content without the user wearing or holding display generating component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface illustrating interactions with XR content that are triggered based on interactions occurring in the space in front of a handheld or tripod-mounted device may be implemented similarly to an HMD in which the interactions occur in the space in front of the HMD and the XR content responses are displayed via the HMD. Similarly, a user interface illustrating interactions with XR content that are triggered based on movement of a handheld or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)) may be implemented similarly to an HMD in which the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hands)).

[0082] While relevant features of operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the exemplary embodiments disclosed herein.

[0083] 2 is a block diagram of an example controller 110 according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more pertinent aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0084] In some embodiments, one or more communication buses 204 include circuitry that interconnects and controls communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0085] Memory 220 includes high-speed random-access memory, such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from the one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240:

[0086] Operating system 230 includes instructions for handling various basic system services and performing hardware-dependent tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for respective groups of one or more users). To that end, in various embodiments, XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.

[0087] 1 , and optionally one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data acquisition unit 241 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0088] In some embodiments, tracking unit 242 is configured to map scene 105 and track the position / location of at least display generating component 120 relative to scene 105 of FIG. 1 , and optionally relative to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, tracking unit 242 includes instructions and / or logic therefor, as well as heuristics and metadata therefor. In some embodiments, tracking unit 242 includes hand tracking unit 244 and / or eye tracking unit 243. In some embodiments, hand tracking unit 244 is configured to track the position / location of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 , relative to display generating component 120, and / or relative to a coordinate system defined relative to the user's hand. Hand tracking unit 244 is described in more detail below with respect to FIG. 4. In some embodiments, eye tracking unit 243 is configured to track the position and movement of the user's gaze (or, more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or the user (e.g., the user's hands)), or relative to XR content displayed via display generation component 120. Eye tracking unit 243 is described in more detail below with respect to FIG. 5.

[0089] In some embodiments, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120 and, optionally, by one or more of output devices 155 and / or peripheral devices 195. To that end, in various embodiments, coordination unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0090] In some embodiments, data dissemination unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least display generation component 120, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data dissemination unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0091] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 can be located within separate computing devices.

[0092] Furthermore, Figure 2 is intended more to illustrate the functionality of various features that may be present in particular implementations, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 2 can be implemented within a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functionality and how functions are allocated among them, will vary from implementation to implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0093] 3 is a block diagram of an example of a display generation component 120, according to some embodiments. While certain features are shown, those skilled in the art will understand from this disclosure that, for the sake of brevity, various other features are not shown so as to not obscure more pertinent aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward-facing and / or outward-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0094] In some embodiments, the one or more communication buses 304 include circuitry that interconnects and controls communications between system components. In some embodiments, the one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.

[0095] In some embodiments, the one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emissive element display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to a waveguide display, such as a diffractive, reflective, polarized, holographic, etc. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content. In some embodiments, the one or more XR displays 312 are capable of presenting mixed reality (MR) or virtual reality (VR) content.

[0096] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as eye-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as hand-tracking cameras). In some embodiments, the one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene as the user would view it if the display generating component 120 (e.g., an HMD) were not present (and may be referred to as a scene camera). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., with a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, and / or the like.

[0097] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from the one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or its non-transitory computer-readable storage medium, stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and an XR presentation module 340:

[0098] The operating system 330 includes instructions for handling various basic system services and for performing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. To that end, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.

[0099] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of Figure 1. To that end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0100] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. To that end, in various embodiments, the XR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0101] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an extended reality) based on the media content data. To that end, in various embodiments, the XR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0102] In some embodiments, data dissemination unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least controller 110, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To that end, in various embodiments, data dissemination unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.

[0103] Although the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data dissemination unit 348 are shown as residing on a single device (e.g., the display generation component 120 of FIG. 1), it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data dissemination unit 348 can be located in separate computing devices.

[0104] Furthermore, Figure 3 is intended more to illustrate the functionality of various features that may be present in a particular implementation, as opposed to a structural overview of the embodiments described herein. As will be recognized by those skilled in the art, items shown separately can be combined and some items can be separated. For example, some functional modules shown separately in Figure 3 can be implemented within a single module, and various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the division of specific functionality and how functions are allocated among them, will vary from implementation to implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0105] 4 is a schematic diagram of an example embodiment of a hand tracking device 140. In some embodiments, hand tracking device 140 (FIG. 1) is controlled by hand tracking unit 244 (FIG. 2) to track the location / position of one or more parts of a user's hand and / or the movement of one or more parts of a user's hand relative to scene 105 of FIG. 1 (e.g., relative to a portion of the physical environment surrounding the user, relative to display generating components 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to the user's hand). In some embodiments, hand tracking device 140 is part of display generating components 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, hand tracking device 140 is separate from display generating components 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0106] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images with sufficient resolution to allow for differentiation of the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body, or all of the body, and can have either zoom capabilities or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor, or a portion thereof, is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.

[0107] In some embodiments, image sensor 404 outputs a sequence of frames containing 3D map data (and possibly color image data) to controller 110, which extracts high-level information from the map data. This high-level information is provided, typically via an application program interface (API), to an application running on the controller, which drives display generation component 120 accordingly. For example, a user can interact with software running on controller 110 by moving their hand 406 and changing the posture of their hand.

[0108] In some embodiments, the image sensor 404 projects a spot pattern onto a scene including the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the pattern's spots. This approach is advantageous in that it does not require the user to hold or wear any type of beacon, sensor, or other marker. This provides depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In this disclosure, the image sensor 404 is assumed to define a set of orthogonal x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) can use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurement, based on single or multiple cameras or other types of sensors.

[0109] In some embodiments, the hand tracking device 140 captures and processes a time sequence of depth maps containing the user's hand while the user moves the hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or a processor in the controller 110 processes the 3D map data to extract patch descriptors of the hand in these depth maps. The software matches these descriptors with patch descriptors stored in the database 408, based on a previous learning process, to estimate the pose of the hand in each frame. The pose typically includes the 3D locations of the user's wrist joints and fingertips.

[0110] The software can also analyze hand and / or finger trajectories across multiple frames in a sequence to identify gestures. The pose estimation functionality described herein may be interleaved with motion tracking functionality, whereby patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to discover pose changes that occur across the remaining frames. The pose, motion, and gesture information is provided to an application program running on controller 110 via the API described above. This program can, for example, move and modify an image presented on display generation component 120 or perform other functions in response to the pose and / or gesture information.

[0111] In some embodiments, the gesture includes an air gesture, which is detected without (or independent of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to another of the user's hands, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of the user's body part (e.g., a tap gesture involving movement of a hand in a predetermined posture by a predetermined amount and / or speed, or a shake gesture involving a predetermined speed or amount of rotation of the user's body part).

[0112] In some embodiments, input gestures used in various examples and embodiments described herein include air gestures performed by movement of a user's finger(s) relative to other finger(s) or part(s) of the user's hand to interact with an XR environment (e.g., a virtual or mixed reality environment), according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body, including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture that includes movement of the hand in a predetermined pose by a predetermined amount and / or speed, or a shake gesture that includes rotation of a part of the user's body at a predetermined speed or amount).

[0113] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides a computer system with information about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touchscreen or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is detected attention (e.g., gaze) to a user interface element in combination with (e.g., simultaneous with) movement of the user's finger(s) and / or hand to perform pinch and / or tap input, as described in more detail below.

[0114] In some embodiments, an input gesture directed at a user interface object is performed directly or indirectly with reference to the user interface object. For example, user input is performed directly at a user interface object in response to performing an input gesture with the user's hand at a position corresponding to the user interface object's position in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, an input gesture is performed indirectly at a user interface object in response to detecting the user's attention (e.g., gaze) to the user interface object while performing the input gesture while the user's hand position is not at a position corresponding to the user interface object's position in the three-dimensional environment. For example, for a direct input gesture, a user can direct the user's input at a user interface object by initiating the gesture at or near a position corresponding to the user interface object's displayed position (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0-5 cm, measured from an outer edge of the option or a central portion of the option). For indirect input gestures, a user can direct their input to a user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user initiates an input gesture (e.g., at any position detectable by the computer system) (e.g., at a position that does not correspond to the displayed position of the user interface object).

[0115] In some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment, according to some embodiments. For example, pinch inputs and tap inputs, as described below, are performed as air gestures.

[0116] In some embodiments, the pinch input is part of an air gesture, including one or more of a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double pinch gesture. For example, a pinch gesture that is an air gesture includes moving two or more fingers of a hand to contact each other, i.e., optionally with a short break (e.g., within 0-1 second) after contact with each other. A long pinch gesture that is an air gesture includes moving two or more fingers of a hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before detecting a break in contact with each other. For example, a long pinch gesture includes a user holding a pinch gesture (e.g., when two or more fingers are in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture that is an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) that are detected immediately in succession (e.g., within a predetermined period of time) after each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaking contact between two or more fingers), and performs a second pinch input within a predetermined period of time (e.g., within 1 second or 2 seconds) after releasing the first pinch input.

[0117] In some embodiments, a pinch-and-drag gesture that is an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., followed by) a drag input that changes the position of a user's hand from a first position (e.g., a start position of the drag) to a second position (e.g., an end position of the drag). In some embodiments, a user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., spreading two or more fingers apart) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., a user pinches two or more fingers together and moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by a user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves from a first position to a second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture that is an air gesture includes an input (e.g., a pinch input and / or a tap input) performed using both of a user's hands. For example, the input gesture includes two (e.g., or more) pinch inputs performed in conjunction with each other (e.g., simultaneously or within a predetermined period of time). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) performed using a first hand of the user and a second pinch input performed using the other hand (e.g., a second of the user's hands) in conjunction with performing the pinch input using the first hand. In some embodiments, a movement between a user's hands (e.g., to increase and / or decrease the distance or relative orientation between the user's hands).

[0118] In some embodiments, a tap input (e.g., directed toward a user interface element) performed as an air gesture includes movement(s) of a user's finger(s) toward the user interface element, movement of a user's hand toward a user interface element, optionally with the user's finger(s) extended toward the user interface element, a downward movement of a user's finger (e.g., mimicking a mouse click action or a tap on a touchscreen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture, moving the finger or hand away from the user's viewpoint and / or toward the object that is the target of the tap input followed by an end of the movement. In some embodiments, an end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's viewpoint and / or toward the object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the direction of acceleration of the movement of the finger or hand).

[0119] In some embodiments, the user's attention is determined to be directed to a portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, the device determines that the user's attention is directed to the portion of the three-dimensional environment based on detecting a gaze directed to the portion of the three-dimensional environment with one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell time) while the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, and / or requiring the gaze to be directed to the portion of the three-dimensional environment, and if one of the additional conditions is not met, the device determines that the user's attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until one or more additional conditions are met).

[0120] In some embodiments, detection of a ready configuration of a user or a portion of a user is detected by a computer system, and detection of a ready configuration of the hands is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs performed with the hands (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand geometry (e.g., a pre-pinch geometry with the thumb and one or more fingers extended and spaced apart, ready to perform a pinch or grab gesture, or a pre-tap geometry with one or more fingers extended and the palm facing away from the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular manner (e.g., above the user's waist, moved toward an area in front of the user below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface is responsive to attentional (e.g., gaze) input.

[0121] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or alternatively may be provided on a tangible, non-transitory medium, such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in memory associated with the controller 110. Alternatively, or additionally, some or all of the described functionality of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). While the controller 110 is shown in FIG. 4 as, by way of example, a separate unit from the image sensor 404, some or all of the processing functionality of the controller may be implemented by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device), or otherwise associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, handheld device, or head-mounted device) or using any other suitable computerized device, such as a game console or media player. The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device that is controlled by the sensor output.

[0122] FIG. 4 also includes a schematic diagram of a depth map 410 captured by the image sensor 404, according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. A pixel 412 corresponding to the hand 406 is segmented from the background and wrist in this map. The intensity of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z-distance from the image sensor 404, with increasing gray levels as depth increases. The controller 110 processes these depth values ​​to identify and segment components of the image (i.e., groups of adjacent pixels) that have characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and frame-to-frame motion of the depth map sequence.

[0123] 4 also schematically illustrates a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406, according to some embodiments. In FIG. 4, the hand skeleton 414 is overlaid on a hand background 416 that was segmented from the original depth map. In some embodiments, key feature points on the hand (e.g., knuckles, fingertips, center of the palm, end of the hand where it connects to the wrist, etc.), and optionally the wrist or arm connected to the hand, are identified and positioned on the hand skeleton 414. In some embodiments, the location and movement of these key feature points over multiple image frames are used by the controller 110 to determine hand gestures performed by the hand or the current state of the hand, according to some embodiments.

[0124] FIG. 5 shows an exemplary embodiment of eye tracking device 130 ( FIG. 1 ). In some embodiments, eye tracking device 130 is controlled by eye tracking unit 243 ( FIG. 2 ) to track the position and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, eye tracking device 130 is integrated with display generation component 120. For example, in some embodiments, if display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed in a wearable frame, the head-mounted device includes both components for generating XR content for viewing by the user and components for tracking the user's gaze relative to the XR content. In some embodiments, eye tracking device 130 is separate from display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, eye tracking device 130 is optionally a device separate from the handheld device or the XR chamber. In some embodiments, eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, head-mounted eye tracking device 130 is optionally used in conjunction with head-mounted or non-head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally used in combination with head-mounted display generating components. In some embodiments, eye tracking device 130 is not a head-mounted device, and is optionally part of non-head-mounted display generating components.

[0125] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that displays frames including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display that allows the user to view the physical environment directly and display virtual objects on the transparent or translucent display. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, allowing an individual using the system to observe the virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.

[0126] As shown in FIG. 5 , in some embodiments, eye tracking device 130 (e.g., gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) camera or near-IR (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be aimed at the user's eyes to receive reflected IR or NIR light from the light source directly from the eyes, or alternatively, may be aimed at a “hot” mirror positioned between the user's eyes and a display panel that reflects IR or NIR light from the eyes to the eye tracking camera while allowing visual light to pass through. Eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes the images to generate eye tracking information, and communicates the eye tracking information to controller 110. In some embodiments, the user's eyes are tracked separately by their respective eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by a separate eye-tracking camera and lighting source.

[0127] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the eye tracking device's parameters for the particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility before delivery of the AR / VR equipment to the end user. The device-specific calibration process may be an automatic or manual calibration process. The user-specific calibration process may include estimation of a particular user's eye parameters, such as pupil location, central visual location, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 have been determined, images captured by the eye tracking camera can be processed using glint-assisted methods to determine the user's current visual axis and viewpoint relative to the display.

[0128] As shown in FIG. 5, eye tracking device 130 (e.g., 130A or 130B) includes an eyepiece(s) 520 and a gaze tracking system including at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking occurs and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 may be positioned between the user's eye(s) 592 and the display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, a projector, etc.) and may be directed at a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (e.g., as shown at the top of FIG. 5), or may be directed at the user's eye(s) 592 to receive reflected IR or NIR light from the eye(s) 592 (e.g., as shown at the bottom of FIG. 5).

[0129] In some embodiments, controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye tracking camera 540 for various purposes, such as in processing frames 562 for display. Controller 110 optionally estimates the user's viewpoint on display 510 based on gaze tracking input 542 obtained from eye tracking camera 540, using a glint-assisted method or other suitable method. The viewpoint estimated from gaze tracking input 542 is optionally used to determine the direction the user is currently looking.

[0130] Some possible use cases of the user's current gaze direction are described below, but are not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined user's gaze direction. For example, the controller 110 may generate virtual content with higher resolution in a central visual area determined from the user's current gaze direction than in a peripheral area. As another example, the controller may position or move virtual content within a view based at least in part on the user's current gaze direction. As another example, the controller may display particular virtual content within a view based at least in part on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 can orient an external camera to capture the physical environment of the XR experience and focus in the determined direction. The external camera's autofocus mechanism can then focus on an object or surface within the environment the user is currently viewing on the display 510. As another exemplary use case, eyepiece 520 may be a focusable lens, and eye-tracking information is used by the controller to adjust the focus of eyepiece 520 so that the virtual object the user is currently looking at has the proper binocular coordination to match the convergence of the user's eyes 592. Controller 110 can utilize the eye-tracking information to orient and focus eyepiece 520 so that close objects the user is looking at appear at the correct distance.

[0131] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye tracking camera (e.g., eye tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)) attached to the wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eye(s) 592. In some embodiments, the light sources may be arranged in a ring or circle around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520, as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be used.

[0132] In some embodiments, the display 510 emits light in the visible light range and not in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. Note that the location and angle of the eye tracking camera(s) 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0133] Embodiments of an eye tracking system such as that shown in FIG. 5 may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, extended reality, and / or augmented virtual experience.

[0134] FIG. 6 illustrates a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the glint-assisted gaze tracking system tracks the pupil contour and glint in the current frame using prior information from the previous frame when analyzing the current frame. When not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.

[0135] As shown in FIG. 6, an eye-tracking camera can capture left and right images of a user's left and right eyes. The captured images are then input into an eye-tracking pipeline for processing beginning at 610. As indicated by the arrow returning to element 600, the eye-tracking system can continue to capture images of the user's eyes at a rate of, for example, 60-120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.

[0136] At 610, if the tracking status is yes for the currently captured image, the method proceeds to element 640. If the tracking status is no at 610, the image is analyzed to detect the user's pupil and glint in the image, as shown at 620. If the pupil and glint are successfully detected at 630, the method proceeds to element 640. If not, the method returns to element 610 to process the next image of the user's eyes.

[0137] At 640, proceeding from element 610, the current frame is analyzed to track pupils and glints based in part on previous information from the previous frame. At 640, proceeding from element 630, a tracking state is initialized based on the detected pupils and glints in the current frame. The results of the processing at element 640 are checked to ensure that the tracking or detection results are reliable. For example, the results can be checked to determine whether a sufficient number of glints are successfully tracked or detected in the current frame to perform pupil and gaze estimation. At 650, if the results are not reliable, the tracking state is set to no at element 660 and the method returns to element 610 to process the next image of the user's eyes. At 650, if the results are reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes) and the pupil and glint information is passed to element 680 to estimate the user's gaze point.

[0138] 6 is intended to serve as an example of eye-tracking technology that may be used in particular implementations. As will be recognized by those skilled in the art, other eye-tracking technologies, now existing or developed in the future, may be used in place of or in combination with the glint-assisted eye-tracking technology described herein in computer system 101 to provide a user with an XR experience according to various embodiments.

[0139] In this disclosure, various input methods are described with respect to interaction with a computer system. Where one example is provided using one input device or input method and another example is provided using a different input device or input method, it should be understood that each example may be compatible with, and optionally utilize, the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. Where one example is provided using one output device or output method and another example is provided using a different output device or output method, it should be understood that each example may be compatible with, and optionally utilize, the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. Where one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with, and optionally utilize, the method described with respect to the other example. Thus, this disclosure discloses embodiments that are combinations of features of multiple examples, without exhaustively listing all features of the embodiments in the description of each exemplary embodiment. User Interface and Related Processes

[0140] Attention is now directed to embodiments of a user interface ("UI") and associated processes that may be implemented on a computer system, such as a portable multifunction device or a head-mounted device, in communication with display generation components and (optionally) one or more cameras and one or more input devices.

[0141] Figures 7A-7Q illustrate exemplary techniques for capturing and / or displaying media in various environments, according to some embodiments. Figure 8 is a flow diagram of a method for capturing media, according to various embodiments. Figure 9 is a flow diagram of a method for displaying a preview of media, according to various embodiments. Figure 10 is a flow diagram of a method for displaying previously captured media. The user interfaces of Figures 7A-7Q are used to illustrate processes described below, including the processes of Figures 8, 9, and 10.

[0142] Figures 7A-7Q illustrate exemplary techniques for capturing and viewing media, according to some embodiments. The schematic diagrams and user interfaces in Figures 7A-7Q are used to illustrate the processes described below, including the processes in Figures 8, 9, and 10.

[0143] 7A shows a user 712 holding a computer system 700 including a display 702 in a physical environment (e.g., a room in a house). The physical environment includes a couch 709a, a photo 709b, a first individual 709c1, a second individual 709c2, a television 709d, and a table 709e. The display 702 presents a representation of the physical environment 704 (e.g., using “pass-through video” as described above). The user 712 is holding the computer system 700 such that the couch 709a, the photo 709b, the first individual 709c1, and the second individual 709c2 are visible from the user's perspective, which is determined based on the location of a portion of the computer system 700, including one or more cameras used to acquire visual information about the physical environment and to generate a virtual environment based on the visual information about the visual environment, for virtual pass-through. In the embodiment of FIGS. 7A-7Q, the user's perspective corresponds to the field of view of one or more cameras (e.g., a camera on the back of computer system 700) in communication with computer system 700. Thus, in the case of virtual pass-through, as computer system 700 is moved throughout the physical environment, the field of view of one or more cameras changes, thereby changing the user's perspective. Because couch 709a, photo 709b, individual 709c1, and second individual 709c2 are visible from the user's perspective in FIG. 7A, display 702 includes depictions of couch 709a, photo 709b, first individual 709c1, and second individual 709c2. When user 712 views display 702, user 712 can see a representation of physical environment 704 along with one or more virtual objects that computer system 700 can display (e.g., as shown in FIGS. 7B-7Q). Thus, computer system 700 presents an extended reality environment via display 702.

[0144] Although computer system 700 is a tablet in FIG. 7A , in some embodiments, computer system 700 may be one or more other devices, such as a handheld device (e.g., a smartphone) and / or a head-mounted device. In some embodiments, when computer system 700 is a head-mounted device, the representation of physical environment 704 is an extended reality environment. In some embodiments, while the representation of physical environment 704 is an extended reality environment, the representation of physical environment 704 includes immersive visual characteristics, including the display of depth data (e.g., the foreground and background of the representation of physical environment 704 are displayed differently to present the visual effect of depth when viewed by a user of computer system 700). In some embodiments, computer system 700 includes one or more components of computer system 101, and / or display 702 includes components of display generation component 120. In some embodiments, display 702 presents a representation of a virtual environment (e.g., instead of the physical environment in FIG. 7A ).

[0145] 7B-7E illustrate a method for capturing spatial (e.g., immersive) media. In FIGS. 7B-7E, the computer system 700 remains in the physical environment shown in FIG. 7A, as shown in schematic diagram 701, which will be described in more detail below. In FIGS. 7B-7E, the computer system 700 is shown enlarged to better illustrate the content visible on the display 702. As shown in FIG. 7B, the computer system 700 displays a control center virtual object 707 (e.g., in response to a swipe gesture on the display 702 performed by a user 712). The control center virtual object 707 includes multiple virtual objects. Each virtual object included in the control center virtual object 707 is selectable. When selected, each virtual object included in the control center virtual object 707 causes the computer system 700 to perform a respective action (e.g., change the playback state of the computer system 700, change the volume at which the computer system 700 outputs audio, display applications currently installed on the computer system 700, and / or any other suitable action).

[0146] 7B , computer system 700 presents a representation of physical environment 704. The representation of physical environment 704 corresponds to a user's viewpoint (e.g., the representation of physical environment 704 includes content that is visible from the user's viewpoint). That is, as the user's viewpoint changes, the representation of physical environment 704 changes based on the change in the user's viewpoint of computer system 700. In some embodiments, the representation of physical environment 704 is a pass-through representation of at least a portion of the physical environment surrounding computer system 700.

[0147] 7B , the representation of the physical environment 704 visually contrasts with the display of the control center virtual object 707. The representation of the physical environment 704 includes a first amount of shading / blurring, and the display of the control center virtual object 707 is displayed with a second amount of shading / blurring that is different from the first amount of shading / blurring (e.g., no shading / blurring). In some embodiments, the representation of the physical environment 704 does not contrast with the display of the control center virtual object 707. In some embodiments, the representation of the physical environment 704 does not have any amount of blurring / blurring.

[0148] 7B-7Q include a schematic diagram 701 of a physical environment. Computer system 700 is represented by an indication 703 within schematic diagram 701. That is, the location and orientation of indication 703 in schematic diagram 701 represents the location and orientation of computer system 700 within the physical environment. While schematic diagram 701 depicts the physical environment shown in FIG. 7A, it should be appreciated that this is merely an example and that the techniques described herein can work with other types of physical environments. Schematic diagram 701 is merely a visual aid. Computer system 700 does not display schematic diagram 701.

[0149] FIG. 7B illustrates computer system 700 as having hardware button 711a (e.g., a hardware input device / mechanism) (e.g., a physical input device) and hardware button 711b. Additionally, FIG. 7B illustrates body part 712a of user 712. Body part 712a represents one of the fingers of user 704 (e.g., the user's index finger, ring finger, pinky finger, middle finger, or thumb). In some embodiments, the representation of body part 712a is any other part of the user's 704's body (e.g., a wrist, arm, hand, and / or any other suitable body part) capable of activating hardware button 711a or hardware button 711b. In FIG. 7B, computer system 700 detects activation of hardware button 711a by body part 712a, or computer system 700 detects input 750b directed toward camera virtual object 707a. In some embodiments, input 750b is a tap input on camera virtual object 707a (e.g., an air tap in space corresponding to the location of the display of camera virtual object 707a). In some embodiments, input 750b is a gaze input (e.g., a sustained gaze) directed toward the display of camera virtual object 707a. In some embodiments, input 750b is an air tap input combined with detection of a gaze directed toward the display of camera virtual object 707a. In some embodiments, input 750b is a gaze and blink directed toward the display of camera virtual object 707a.

[0150] 7C , in response to activating a hardware button 711 a with a body part 712 a or detecting input 750 b directed at a camera virtual object 707 a, the computer system 700 displays a media capture preview 708, a timer virtual object 713, a camera shutter virtual object 714, a reposition virtual object 716, an erase virtual object 719, and a photowell virtual object 715. The computer system 700 displays the media capture preview 708 as overlaid on top of a representation of the physical environment 704. As shown in FIG. 7C , the display of the media capture preview 708 is smaller (e.g., occupies less space on the display 702) than the representation of the physical environment 704. In some embodiments, the computer system 700 is a head-mounted device that presents a representation of the physical environment 704 along with one or more virtual objects that the computer system 700 displays via display generating components that surround (or substantially surround) the user's field of view. In embodiments in which computer system 700 is an HMD, the user's viewpoint is locked to the forward direction of the user's head, such that the representation of one or more virtual objects, such as physical environment 704 and media capture preview 708, shifts as the user's head moves (e.g., because computer system 700 moves as the user's head moves).

[0151] The timer virtual object 713, the camera shutter virtual object 714, the reposition virtual object 716, the erase virtual object 719, and the photo well virtual object 715 are all anchored to the media capture preview 708. That is, the locations of the display of the timer virtual object 713, the camera shutter virtual object 714, the reposition virtual object 716, and the photo well virtual object 715 are associated with the location of the display of the media capture preview 708. In some embodiments, when the location of the display of the media capture preview 708 changes, the locations of the display of the timer virtual object 713, the camera shutter virtual object 714, the reposition virtual object 716, the erase virtual object 719, and the photo well virtual object 715 change (e.g., see Figures 7F-7G). As shown in Figure 7C, the computer system 700 displays the media capture preview 708 above / on top of the camera shutter virtual object 714 and in the center of the display 702. As shown in FIG. 7C , media capture preview 708 includes a portion of a representation of physical environment 704 that was visible before computer system 700 displayed media capture preview 708. For example, in FIG. 7B (e.g., before computer system 700 displayed media capture preview 708), the representation of the physical environment includes couch 709a, picture 709b, first individual 709c1, and second individual 709c2. Thus, as shown in FIG. 7C , media capture preview 708 includes depictions of couch 709a1, picture 709b1, first individual 709c3, and second individual 709c4. Media capture preview 708 provides a preview of content that will be captured in response to computer system 700 detecting a request to capture media. The content displayed within media capture preview 708 is based on the field of view of one or more cameras in communication with computer system 700 (e.g., the content displayed within media capture preview 708 is within the field of view of one or more cameras in communication with computer system 700).The content displayed within media capture preview 708 changes based on changes in the field of view of one or more cameras. In some embodiments, computer system 700 includes two cameras, and the content displayed within media capture preview 708 is content that falls within the field of view of both cameras, which allows for the capture of immersive content.

[0152] 7C , the user's viewpoint corresponding to the representation of physical environment 704 has a wider range of viewing angles than the angular range of the field of view corresponding to media capture preview 708. This causes the representation of physical environment 704 to show a greater amount of the physical environment than the amount of the physical environment shown in media capture preview 708 (e.g., only a portion of the sofa is visible in media capture preview 708, whereas the entire sofa is visible in the representation of physical environment 704). Thus, media captured while media capture preview 708 appears as shown in FIG. 7C will include only a portion of the sofa, rather than the entire sofa. In some embodiments, the representation of the physical environment included in media capture preview 708 is displayed at a first scale, and the representation of physical environment 704 is presented at a second scale that is larger than the first scale. In some embodiments, computer system 700 includes two cameras with different but overlapping fields of view, media capture preview 708 represents a portion of the physical environment common to the fields of view of both cameras (e.g., the portion where the FOVs overlap), and the representation of physical environment 704 includes content within the FOV of a first camera and / or a second camera of the two cameras (e.g., both overlapping and non-overlapping). In some embodiments, computer system 700's display of media capture preview 708 includes content included in the representation of physical environment 704. In some embodiments, the representation of physical environment 704 is displayed from an immersive perspective, and the content included in the display of media capture preview 708 is displayed from a non-immersive perspective.

[0153] As shown in FIG. 7C , the media capture preview 708 is displayed with a visual appearance that does not include dimming and / or blurring, and the representation of the physical environment 704 is displayed with dimming and / or blurring. This provides contrast between the display of the media capture preview 708 and the representation of the physical environment 704. In some embodiments, the representation of the physical environment 704 is not dimmed and / or blurred before the computer system 700 detects the input 750b or before the computer system 700 detects the activation of the hardware button 711a with the body portion 712a of FIG. 7B . In some embodiments, the representation of the physical environment 704 is dimmed and / or blurred (e.g., faded out) in response to the computer system 700 detecting the input 750b or in response to the computer system 700 detecting the activation of the hardware button 711a with the body portion 712a.

[0154] With respect to virtual objects anchored in the display of the media capture preview 708, a timer virtual object 713 provides an indication of the amount of time (e.g., minutes, seconds, hours) that has elapsed since the computer system 700 began the media capture process. The photo well virtual object 715 includes a representation of the most recently captured media item (e.g., a still photo or a video). In some embodiments, the photo well virtual object 715 includes a representation of the most recently captured media item captured by the computer system 700. In some embodiments, the photo well virtual object 715 includes a representation of the most recently captured media item captured by an external device in communication with the computer system 700. As shown in FIG. 7C , the photo well virtual object 715 includes a representation of a fountain. Thus, the most recently captured media item includes a depiction of a fountain.

[0155] Selection of the camera shutter virtual object 714 initiates a process on the computer system 700 to capture media including the content shown in the media capture preview 708. Repositioning the virtual object 716 allows the user 712 to reposition the location of the display of the media capture preview 708. For example, moving the display location of the repositioned virtual object 716 to the left causes the display of the media capture preview 708 to be moved to the left. Selecting the erase virtual object 719 causes the computer system 700 to cease displaying the media capture preview 708. In some embodiments, the representation of the physical environment 704 becomes less blurred and / or less shaded when the media capture preview 708 is no longer displayed.

[0156] 7C , computer system 700 detects activation of hardware button 711a by body part 712a, or computer system 700 detects input 750c directed toward camera shutter virtual object 714. In some embodiments, input 750c is a tap on camera shutter virtual object 714 (e.g., an air tap in space corresponding to the location of the display of camera shutter virtual object 714). In some embodiments, input 750c is a gaze (e.g., sustained gaze) input directed toward the display of camera shutter virtual object 714. In some embodiments, input 750c is an air tap input combined with gaze detection in the direction of the display of camera shutter virtual object 714. In some embodiments, input 750c is a gaze and blink directed toward the display of camera shutter virtual object 714. In some embodiments, the activation of the hardware button 711 a or input 750 c by the body part 712 a is a long press (e.g., press and hold) (e.g., the duration of the activation of the hardware button or input 750 c by the body part 712 a spans several seconds). In some embodiments, the activation of the hardware button 711 a or input 750 c by the body part 712 a is a short press (e.g., press and release) (e.g., the duration of the activation of the hardware button 711 a or input 750 c by the body part 712 a is less than one second). In some embodiments, a specific air gesture is detected that is recognized as a request to capture media (e.g., as described above with respect to selecting a virtual object in an XR environment).

[0157] 7D , in response to detecting activation of body portion 712a of hardware button 711a or in response to detecting input 750c, computer system 700 initiates a media capture process. In FIG. 7D , a determination is made that the activation of hardware button 711 or input 750c by body portion 712a is a long press. Because a determination is made that the activation of body portion 712a of hardware button 711 or input 750c is a long press, video media is captured (e.g., not still media). The media capture process records content displayed in media capture preview 708 while the media capture process is in progress. In some embodiments, the field of view of one or more cameras in communication with computer system 700 is changed during the media capture process, changing what is displayed in media capture preview 708 and changing the content being captured by the media capture process. In some embodiments, in accordance with a determination that the activation of hardware button 711 or input 750c by body portion 712a is a short press, still media (e.g., a photograph) is captured via the media capture process.

[0158] As shown in FIG. 7D , the display of the timer virtual object 713 indicates “00:05” (e.g., 5 seconds). The timer virtual object 713 in FIG. 7D indicates that 5 seconds have elapsed since the computer system 700 started the media capture process. Further, as shown in FIG. 7D , the display of the camera shutter virtual object 714 includes a square. The display of the camera shutter virtual object 714 with a square indicates that the computer system 700 is currently recording video media. In some embodiments, the shape, size, and / or color of the camera shutter virtual object 714 is updated to indicate that the computer system 700 is currently recording video media. In FIG. 7D , the computer system 700 detects activation of the hardware button 711 a by the body part 712 a, or the computer system 700 detects an input 750 d directed toward the camera shutter virtual object 714. In some embodiments, the input 750 d is a tap input on the camera shutter virtual object 714 (e.g., an air tap in space corresponding to the location of the display of the camera shutter virtual object 714). In some embodiments, input 750b is a gaze input directed toward the display of camera shutter virtual object 714. In some embodiments, input 750d is an air tap combined with gaze detection in the direction of the display of camera shutter virtual object 714. In some embodiments, input 750d is a gaze and blink directed toward the display of camera shutter virtual object 714.

[0159] In Figure 7E, the computer system 700 ceases the media capture process in response to activation of the hardware button 711a by the body part 712a or detecting input 750d directed at the camera shutter virtual object 714. Because the computer system 700 is no longer performing the media capture process, the display of the camera shutter virtual object 714 in Figure 7E does not include a square. As shown in Figure 7E, the computer system 700 displays a photo well virtual object 715 along with a representation of the video captured in Figures 7C-7D (e.g., a video of the physical environment). Additionally, because the computer system 700 is no longer performing the media capture process, the display of the timer virtual object 713 shows 0:00 as shown in Figure 7E.

[0160] As shown in Figure 7E, schematic diagram 701 includes movement indicator 721. Movement indicator 721 indicates that computer system 700 is beginning to move within the physical environment. In Figure 7E, computer system 700 begins to move laterally to the right within the physical environment.

[0161] 7F-7H illustrate a method of computer system 700 displaying a media capture preview 708 that lags behind (e.g., tracks) movement of computer system 700. 7E-7G illustrate one continuous lateral movement of computer system 700 to the right within the physical environment. In some embodiments, computer system 700 is moved laterally to the left, up, and / or down within the physical environment. In some embodiments, computer system 700 is moved in a combination of different directions (e.g., up and left and / or down and right).

[0162] In FIG. 7F, computer system 700 is positioned laterally to the right of computer system 700's previous position (e.g., computer system 700's position in FIG. 7E). Movement of computer system 700 changes the user's viewpoint. As described above, the representation of physical environment 704 corresponds to the portion of the physical environment that is visible from the user's viewpoint. Thus, as the user's viewpoint changes, the representation of physical environment 704 changes accordingly. Further, as described above, media capture preview 708 corresponds to the field of view of one or more cameras in communication with computer system 700. Movement of computer system 700 changes the field of view of the one or more cameras. Thus, as the field of view of the one or more cameras changes, the content displayed within media capture preview 708 changes accordingly.

[0163] As shown in Figure 7F, the computer system 700 displays the media capture preview 708 off-center (e.g., to the right, left, above, or below the center of the display 702). In Figure 7F, the computer system 700 is moving within the physical environment at a first speed (e.g., 1 ft / sec, 3 ft / sec, 5 ft / sec), and the computer system 700 displays the media capture preview 708 as moving at a second speed (e.g., the second speed relative to the physical environment) that is slower than the first speed (e.g., if the computer system is moving at 5 ft / sec, the media capture preview is moving at 3 ft / sec). The difference between the speed at which the computer system 700 moves within the physical environment and the speed at which the computer system 700 displays the media capture preview 708 as moving relative to the physical environment creates a lag visual effect, showing the media capture preview 708 lagging behind (e.g., following) the movement of the computer system 700. From strictly the perspective of the display 702, the media capture preview 708 is moved on the screen in a direction opposite to the direction of movement of the computer system 700 to create the visual effect of the media capture preview 708 moving slower relative to the physical environment than the computer system 700. In some embodiments, while the computer system 700 is moving, the representation of the physical environment 704 has a first set of parallax characteristics, and the computer system 700 displays the media capture preview 708 with a second set of parallax characteristics that differ from the first set of parallax characteristics. For example, in FIG. 7F , as the computer system 700 moves within the physical environment, there is a first amount of shift between the depiction of the first individual 709c3 and the second individual 709c4 and the depiction of picture 709b1 within the media capture preview 708, and there is a second amount of shift between the depiction of the first individual 709c1 and the second individual 709c2 and the depiction of picture 709b within the representation of the physical environment 704.Because the representation of physical environment 704 has a different set of parallax characteristics than media capture preview 708, a first shift amount between the depiction of first individual 709c3 and second individual 709c4 in media capture preview 708 and the depiction in picture 709b1 is different from the shift amount between the depiction of first individual 709c1 and second individual 709c2 in the representation of physical environment 704 and the depiction in picture 709b. In some embodiments, media capture preview 708 is displayed without any parallax effect (e.g., the first shift amount is zero), while the representation of physical environment 704 includes a non-zero degree of shift, such that the representation of physical environment 704 is perceived as having depth while media capture preview 708 does not (e.g., appears flat). In some embodiments, while computer system 700 is moving, computer system 700 applies a first image stabilization technique to the representation of physical environment 704, and computer system 700 applies a second, different stabilization technique to the display of media capture preview 708. In some embodiments, while computer system 700 is moving, computer system 700 applies a lesser amount of digital image stabilization to the representation of physical environment 704 than the amount of digital stabilization that computer system 700 applies to media capture preview 708.

[0164] 7F , computer system 700 is moved laterally to the right. As computer system 700 is moved laterally to the right, computer system 700 displays media capture preview 708 to the left of center of display 702. In some embodiments, computer system 700 is moved laterally to the left in the physical environment, causing computer system 700 to display media capture preview 708 to the right of center of display 702. In some embodiments, computer system 700 is moved vertically upward in the physical environment, causing computer system 700 to display media capture preview 708 below center of display 702. In some embodiments, computer system 700 is moved vertically downward in the physical environment, causing computer system 700 to display media capture preview 708 above center of display 702. In some embodiments, the computer system 700 is moved forward (e.g., in the z-direction) within the physical environment (e.g., into a page) such that the computer system 700 displays the media capture preview 708 as larger relative to the representation of the physical environment 704 for a period of time before transitioning the size of the media capture preview 708 to have the same relative size relative to the representation of the physical environment 704 before the start of the movement, in such embodiments, the rate at which the media capture preview 708 transitions between the two sizes lags behind the rate of the forward movement of the computer system 700. In some embodiments, the computer system 700 is moved backward (e.g., out of a page) within the physical environment such that the computer system displays the media capture preview 708 as smaller relative to the representation of the physical environment 704 for a period of time before transitioning the size of the media capture preview 708 to have the same relative size relative to the representation of the physical environment 704 before the start of the movement, in such embodiments, the rate at which the media capture preview 708 transitions between the two sizes lags behind the rate of the forward movement of the computer system 700.In some embodiments, as computer system 700 moves within a physical environment, computer system 700 detects movement of computer system 700 with a first amount of tracking lag (e.g., measured as a function of distance over time). For example, while computer system 700 is located at the position indicated by indication 703 of schematic diagram 701 in FIG. 7F, computer system 700 detects that it is positioned at the position indicated by previous position indicator 717 in FIG. 7F for an amount of time corresponding to the first tracking lag. In some embodiments, the second speed at which computer system 700 displays media capture preview 708 as moving is configured to be less than the first amount of tracking lag to present a lag visual effect (e.g., to reduce visual artifacts).

[0165] As shown in Figure 7F, schematic diagram 701 includes a previous position indicator 717. Previous position indicator 717 indicates a previous position of computer system 700 (e.g., the position of computer system 700 in Figure 7E). In Figure 7F, computer system 700 continues to move laterally to the right within the physical environment, as indicated by schematic diagram 701 including movement indication 721.

[0166] In Figure 7G, computer system 700 is positioned further to the right in the physical environment than the position of computer system 700 in Figure 7F and has ceased movement. As shown in Figure 7G, schematic diagram 701 includes a previous position indication 717 that indicates the previous position of computer system 700 (e.g., the position of computer system 700 in Figure 7F). Schematic diagram 701 also includes an indication 703 that indicates the current location of computer system 700. In Figure 7G, schematic diagram 701 does not include movement indicator 721 because computer system 700 is no longer moving in Figure 7F.

[0167] 7G, the computer system 700 displays the media capture preview 708 to the left of center on the display 702. That is, although the computer system 700 is no longer moving in FIG. 7G, the computer system 700 continues to display the media capture preview 708 as if it were moving at a second speed (e.g., 3 feet / second when the computer system is moving at 5 feet / second) so that the media capture preview 708 can "catch up" to the computer system 700. In some embodiments, when the computer system 700 stops moving, the computer system 700 displays the media capture preview 708 as if it were moving at a third speed (e.g., 5 feet / second) that is faster than the second speed so that the media capture preview 708 is recentered more quickly than if it were moving at the second speed.

[0168] 7G, the representation of physical environment 704 is updated relative to the representation of physical environment 704 in FIG. 7H (e.g., the representation of physical environment 704 in FIG. 7G includes table 709e in the background and a portion of couch 709a in the foreground). As described above, the user's viewpoint changes based on changes in the field of view of one or more cameras in communication with computer system 700. Thus, as computer system 700 is moved throughout the physical environment (e.g., changing the field of view of one or more cameras in communication with computer system 700), the user's viewpoint changes and the representation of physical environment 704 changes accordingly. Further, in 7G, the display of media capture preview 708 is updated. As computer system 700 moves within the physical environment, the field of view of one or more cameras in communication with the computer system changes, thereby changing what is displayed in media capture preview 708.

[0169] In FIG. 7H, the media capture preview 708 has caught up with the previous movement of the computer system 700 (e.g., movement of the computer system 700 as described above in connection with FIGS. 7E-7G). As shown in FIG. 7H, because the media capture preview 708 has caught up with the previous movement of the computer system 700, the computer system 700 displays the media capture preview 708 in the center of the display 702. The computer system's 700 display of the media capture preview 708 in the center of the display 702 makes a first portion of the representation of the physical environment 704 (e.g., a portion of the torso of the second individual 709c2) that was previously visible (e.g., visible in FIG. 7G) invisible (e.g., because the media capture preview 708 is now displayed overlaid on top of the first portion) and makes a second portion of the representation of the physical environment 704 (e.g., the torso of the first individual 709c1) that was previously invisible visible (e.g., because the media capture preview 708 is no longer displayed overlaid on top of the second portion). 7H , the computer system 700 detects input 750h directed toward the photowell virtual object 715. In some embodiments, the input 750h is a gaze input (e.g., a sustained gaze) directed toward the display of the photowell virtual object 715. In some embodiments, the input 750h is a tap on the photowell virtual object 715 (e.g., an air tap in space corresponding to the location of the display of the photowell virtual object). In some embodiments, the input 750h is an air tap combined with detection of a gaze toward the display of the photowell virtual object 715. In some embodiments, the input 750h is a gaze and blink directed toward the display of the photowell virtual object 715.

[0170] As shown in FIG. 7I, in response to detecting input 750h, computer system 700 displays a previously captured media item 730. The previously captured media item 730 is the media item most recently captured by computer system 700 (e.g., the media item captured in FIGS. 7C and 7D). In some embodiments, the previously captured media item 730 is the media item most recently captured by an external device (e.g., a device separate from computer system 700) in communication with computer system 700. The previously captured media item 730 is displayed as a square. In some embodiments, the previously captured media item 730 is displayed as a rectangle, a triangle, or any other suitable shape.

[0171] The computer system 700 displays the previously captured media item 730 along with a library virtual object 731, a reposition virtual object 732, a erase virtual object 733, an identifier virtual object 734, a shared virtual object 735, and a projected shape virtual object 736. In some embodiments, the above-listed virtual objects are anchored to the display of the previously captured media item 730 (e.g., as described above in the description of FIG. 7C ). In some embodiments, the library virtual object 731, the reposition virtual object 732, the erase virtual object 733, the identifier virtual object 734, the shared virtual object 735, and the projected shape virtual object 736 are displayed in a different spatial configuration than the spatial configuration shown in FIG. 7I . In some embodiments, in response to detecting input 750h, the computer system 700 displays a user interface that includes a subset of the virtual objects displayed in FIG. 7I .

[0172] Selection of the library virtual object 731 causes multiple representations of previously captured media items to be displayed. In some embodiments, the library virtual object 731 is displayed simultaneously with the media capture preview 708. In some embodiments, the display of the multiple representations of previously captured media items replaces the display of the previously captured media item 730. Repositioning the virtual object 732 allows the user to reposition the location of the display of the representation of the previously captured media item 730 in the same way that the media capture preview 708 can be repositioned using the reposition virtual object 716. Selection of the erase virtual object 733 causes the computer system 700 to cease displaying the previously captured media item 730. In some embodiments, selection of the erase virtual object 733 causes the computer system 700 to cease displaying the library virtual object 731, the reposition virtual object 732, the erase virtual object 733, the identifier virtual object 734, the shared virtual object 735, and the projected shape virtual object 736. In some embodiments, selection of the erase virtual object 733 causes the computer system 700 to display the media capture preview 708. The identifier virtual object 734 provides an indication of where and when the previously captured media item 730 was captured. In some embodiments, the identifier virtual object 734 provides different information for the previously captured media item 730 (e.g., the resolution of the previously captured media item, the date the previously captured media item was captured). Selection of the shared virtual object 735 initiates a process on the computer system 700 to share the previously captured media item 730 with an external device (e.g., a device separate from the computer system 700). In FIG. 7I , the computer system 700 detects an input 750i directed at the projected shape virtual object 736.In some embodiments, input 750i is a gaze input directed in the direction of the display of projected shape virtual object 736. In some embodiments, input 750i is a tap input on projected shape virtual object 736 (e.g., an air tap in space corresponding to the location of the display of projected shape virtual object 736). In some embodiments, input 750i is an air tap input combined with gaze detection in the direction of the display of projected shape virtual object 736. In some embodiments, input 750i is a gaze and blink directed in the direction of the display of projected shape virtual object 736.

[0173] As shown in Figure 7J, in response to detecting input 750i, computer system 700 displays a representation of the previously captured media item 730 as a circle. While computer system 700 displays the previously captured media item 730 as a circle as shown in Figure 7J, computer system 700 also displays the library virtual object 731, reposition virtual object 732, erase virtual object 733, identifier virtual object 734, shared virtual object 735, and projected shape virtual object 736 as described above in the description of Figure 7I. In some embodiments, in response to detecting input 750i, computer system 700 displays the previously captured media item 730 as a three-dimensional sphere. As shown in Figure 7I, the schematic diagram includes a movement indicator 721. The movement indicator 721 indicates that computer system 700 is beginning to move (e.g., laterally to the left) within the physical environment. In FIG. 7J, as indicated by the schematic diagram 701 including the movement indication 721, the computer system 700 begins to move laterally to the left within the physical environment, returning to the initial position of the computer system 700 (e.g., the position of the computer system 700 in FIG. 7A).

[0174] 7K-7Q illustrate how a user may interact with a three-dimensional representation of a previously captured media item. FIGS. 7K-7Q include an elapsed time indication 744. The elapsed time indication 744 is a visual aid that indicates the amount of time that has passed between each view. The elapsed time indication 744 is not displayed by the computer system 700. In FIG. 7K, the computer system 700 is located at an initial location within the physical environment (e.g., the location of the computer system 700 in FIGS. 7A-7E). Between FIGS. 7J and 7K, the computer system 700 receives a request to display a spatial capture virtual object 740. As shown in FIG. 7K, in response to receiving the request to display the spatial capture virtual object 740, the computer system 700 displays the spatial capture virtual object 740 overlaid on top of the representation of the physical environment 704. The display of the spatial capture virtual object 740 obscures portions of the representation of the physical environment 704. The spatial capture virtual object 740 is a representation of a previously captured video media item. In some embodiments, computer system 700 is a head-mounted device that presents spatial capture virtual object 740 as part of an extended reality environment, and while computer system 700 is moving, spatial capture virtual object 740 is displayed with a parallax effect (e.g., as discussed above with respect to FIG. 7F ) that causes the media item represented by spatial capture virtual object 740 to have an amount of depth between the foreground and background of content within the media item. In some embodiments, spatial capture virtual object 740 is a representation of previously captured still media. In some embodiments, the previously captured video media item was captured using one or more cameras in communication with computer system 700. In some embodiments, the request to display spatial capture virtual object 740 includes one or more inputs directed at one or more hardware buttons in communication with computer system 700.In some embodiments, the request to display the spatial capture virtual object 740 includes one or more inputs directed at one or more virtual objects displayed by the computer system 700. In some embodiments, the request to display the spatial capture virtual object 740 includes detection of a gaze (e.g., made by a user) in a direction corresponding to one of the previously captured sub-representations of the media item. In embodiments in which the computer system 700 is an HMD, the user's viewpoint is locked to the forward direction of the user's head, and therefore the representations of the physical environment 704 and one or more virtual objects, such as the spatial capture virtual object 740, shift as the user's head moves (e.g., because the computer system 700 also moves as the user's head moves).

[0175] The computer system 700 displays a spatial capture virtual object 740 at a first location in the representation of the physical environment 704 that corresponds to a first location in the physical environment (e.g., the location indicated by the spatial capture indicator 740a in the schematic diagram 701). The computer system 700 displays the spatial capture virtual object 740 when the first location in the physical environment is within the field of view of one or more cameras in communication with the computer system 700. The display of the spatial capture virtual object 740 is environment-locked to the first location in the representation of the physical environment 704. That is, the location of the display of the spatial capture virtual object 740 in the representation of the physical environment 704 does not change even if the real-world position of the computer system 700 changes. However, the display of the spatial capture virtual object 740 is updated as the user's viewpoint changes in the physical environment. For example, as the user's viewpoint approaches the first location in the physical environment, the computer system 700 displays the spatial capture virtual object 740 larger (e.g., to provide the visual effect that the user is getting closer to the spatial capture virtual object 740). In contrast, as the user's viewpoint moves away from the first location in the physical environment, the computer system 700 displays the spatial capture virtual object 740 smaller (e.g., to provide the visual effect that the user has moved further away from the spatial capture virtual object 740).

[0176] 7K, in response to receiving a request to display the spatial capture virtual object 740, the computer system 700 simultaneously displays multiple sub-representations of previously captured media items 743 along with the display of the spatial capture virtual object 740. As shown in FIG. 7K, the computer system 700 displays multiple sub-representations of the previously captured media items 743 below the display of the spatial capture virtual object 740. Each sub-representation of the previously captured media represents a respective media item (e.g., video media or still media) previously captured by the computer system 700 or an external device in communication with the computer system 700. The sub-representation of the previously captured media item 743a displayed at the center of the multiple sub-representations of the previously captured media items 743 corresponds to the media item displayed in focus (e.g., represented by the spatial capture virtual object 740). In some embodiments, a user can navigate through the multiple sub-representations of the previously captured media items 743 to select the sub-representation that is displayed in focus at the location of the spatial capture virtual object 740. In some embodiments, a user can switch which sub-representation is focused and displayed by performing a motion input (e.g., a pinch-and-drag gesture). In some embodiments, while the spatial capture virtual object 740 is displayed, the computer system 700 increases the size of the display spatial capture virtual object 740 in response to the computer system 700 detecting (e.g., via one or more cameras in communication with the computer system 700) that the user has performed a pinch gesture. In some embodiments, the computer system 700 decreases the size of the display of the spatial capture virtual object 740 in response to the computer system 700 detecting (e.g., via one or more cameras in communication with the computer system 700) that the user has performed a de-pinch air gesture.

[0177] 7K, the schematic diagram 701 includes a spatial capture indicator 740a. The spatial capture virtual object 740 does not physically exist in the physical environment. Rather, the spatial capture virtual object 740 is a virtual object that is displayed only within the representation of the physical environment 704. The spatial capture indicator 740a indicates the spatial orientation / position of the display of the spatial capture virtual object 740 within the representation of the physical environment 704 and also indicates the location to which the spatial capture virtual object 740 is environment-locked.

[0178] 7K, the computer system 700 plays the video media item represented by the spatial capture virtual object 740 while the spatial capture virtual object 740 is displayed. The computer system 700 automatically (e.g., without intervening user input) begins playing the video media item represented by the spatial capture virtual object 740 when the computer system 700 displays the spatial capture virtual object 740. In some embodiments, playback of the video media item represented by the spatial capture virtual object 740 is not automatic. In some embodiments, the computer system 700 outputs spatial audio as part of playing the video media item represented by the spatial capture virtual object 740. In some embodiments, the video media item represented by the spatial capture virtual object 740 is a stereoscopic video media item. A stereoscopic video media item presents a user with two different images of the same scene. In some embodiments, a first image is from a first camera and a second image is from a second camera, where the first and second cameras have slightly different perspectives of the scene. As a result, the first image varies slightly from the second image. When viewed by a user, the two images are superimposed on one another, creating the illusion of depth in the resulting image. In some embodiments, using the techniques described above in connection with FIG. 7J , a user can change the shape in which the computer system 700 displays the spatial capture virtual object 740. In some embodiments, a user can change the shape of the spatial capture virtual object 740 from a flattened stereoscopic projection to a spherical stereoscopic projection, and vice versa. In some embodiments, the video media item represented by the spatial capture virtual object 740 is played on an external electronic device that is not capable of playing stereoscopic video.In some embodiments, spatial audio is used when the video media item represented by the spatial capture virtual object 740 is played on an external electronic device that is not capable of playing stereoscopic video. In some embodiments, pursuant to a determination that a user of the computer system 700 has an interpupillary distance that is different from the default interpupillary distance setting of the computer system 700, the computer system 700 plays the video media item with a shift (e.g., the computer system 700 plays the video media item represented by the spatial capture virtual object 740 at a scale that is different from the scale at which the video media item was captured). The interpupillary distance is the distance (e.g., in millimeters) between the centers of the pupils of an individual's eyes. In some embodiments, following a determination that a user of computer system 700 has an interpupillary distance that differs from the default interpupillary distance setting of computer system 700, computer system 700 modifies (e.g., shifts) the first image and / or the second image of the stereoscopic video image (e.g., to the left or right) based on the difference between the user's interpupillary distance and the interpupillary distance setting of computer system 700, without changing the scale of playback of the stereoscopic video media item.

[0179] The video media item represented by the spatial capture virtual object 740 is a video of an individual progressing along a trajectory. In FIG. 7K, the computer system 700 plays the video media item represented by the spatial capture virtual object 740 from a non-immersive perspective. Playback of content from a non-immersive perspective cannot be presented from multiple perspectives depending on detected changes in the orientation / location of the computer system 700. When presented from a non-immersive perspective, video media is presented from only one perspective, regardless of whether the orientation / location of the computer system 700 changes. In FIG. 7K, the computer system 700 is located a first distance (e.g., a virtual distance) away from a first location in the physical environment where the computer system 700 displays the spatial capture virtual object 740 in a representation of the physical environment 704. In FIG. 7K, the computer system 700 begins to move toward the first location in the physical environment, as indicated by the schematic diagram 701 including the movement indication 721.

[0180] As shown in FIG. 7L , the elapsed time indication 744 is shown as “00:01.” Thus, one second has elapsed between FIG. 7K and FIG. 7L . In FIG. 7L , the computer system 700 continues to display the spatial capture virtual object 740 within the representation of the physical environment 704. Further, in FIG. 7L , the computer system 700 is positioned closer to the first location in the physical environment than the position of the computer system 700 in FIG. 7K . Thus, in FIG. 7L , because the computer system 700 is located closer to the first location in the physical environment, the computer system 700 increases the size of the display of the spatial capture virtual object 740 (e.g., compared to the size of the display of the spatial capture virtual object 740 in FIG. 7K ). As described above, as the user's viewpoint changes, the display of the spatial capture virtual object 740 changes based on the user's changing viewpoint. Because the computer system 700 displays the spatial capture virtual object 740 larger, less of the representation of the physical environment 704 is visible. In some embodiments, computer system 700 moves backward within the physical environment, which reduces the size of the display of spatial capture virtual object 740. In some embodiments, while computer system 700 moves laterally sideways within the physical environment, a representation of physical environment 704 is presented with respective visual parallax effects, and computer system 700 displays spatial capture virtual object 740 with respective visual parallax effects.

[0181] In Figure 7L, the computer system 700 continues playing the video media represented by the spatial capture virtual object 740. Accordingly, in Figure 7L, the computer system 700 updates the representation of the video media represented by the spatial capture virtual object 740 (e.g., the representation of the video media shows the user as having progressed along the trajectory) to indicate that the playback of the video media has progressed by one second (e.g., the amount of time that has elapsed between Figures 7K and 7L). In Figure 7L, the computer system 700 continues moving toward a first location in the physical environment, as indicated by the movement indication 721 in the schematic diagram 701.

[0182] As shown in Figure 7M, the elapsed time indication 744 shows "00:05." Thus, four seconds have elapsed between Figures 7L and 7M. In Figure 7M, the computer system 700 is positioned closer to the first location in the physical environment than the position of the computer system 700 in Figure 7L. Because the computer system 700 is positioned closer to the first location in the physical environment in Figure 7M, the computer system 700 increases the size of the display of the spatial capture virtual object 740 (e.g., compared to the size of the display of the spatial capture virtual object 740 in Figure 7L). Because the computer system 700 displays the spatial capture virtual object 740 larger, less of the representation of the physical environment 704 is visible.

[0183] 7M, the computer system 700 continues playing the video media item represented by the spatial capture virtual object 740. Thus, in FIG. 7M, the computer system 700 displays a frame of a representation of the video media item four seconds after the frame of video media shown in FIG. 7L, indicating that the playback of the video media has advanced four seconds. In some embodiments, the computer system 700 ceases displaying the spatial capture virtual object 740 as the computer system 700 moves past a first location in the physical environment. In some embodiments, the computer system 700 is a head-mounted device that presents the spatial capture virtual object 740 as part of an extended reality environment, and the closer the computer system 700 is to a first location in the physical environment, the more immersive the video media item represented by the spatial capture virtual object 740 becomes (e.g., has greater depth between the foreground and background and / or greater responsiveness to shifts in the orientation of the computer system 700). In some embodiments, in FIG. 7M, the spatial capture virtual object 740 is displayed as a full-screen configuration. The representation of the physical environment 704 is visually obscured (e.g., the representation of the physical environment 704 is blurred, dimmed, and / or blacked out) while the spatial capture virtual object 740 is displayed in a full-screen configuration. In some embodiments, the spatial capture virtual object 740 is displayed in a full-screen configuration, but the computer displays a different virtual environment (e.g., a virtual movie theater environment and / or a virtual drive-in environment) in place of the representation of the physical environment 704. In Figure 7M, the computer system 700 begins to rotate clockwise (e.g., 90 degrees clockwise) within the physical environment, as indicated by the movement indication 721 in the schematic diagram 701.

[0184] In Figure 7N, the user's perspective has been rotated 90 degrees clockwise relative to the user's perspective in Figure 7M. As described above, moving computer system 700 changes the user's perspective, which in turn changes the representation of physical environment 704. Thus, in Figure 7N, the representation of physical environment 704 corresponds to the change in the user's perspective. In Figure 7N, first individual 709c1, second individual 709c2, photo 709b, and couch 709a, which are part of the physical environment, are not visible from the user's perspective in Figure 7N. Thus, the representation of physical environment 704 in Figure 7N does not include depictions of couch 709a, picture 709b, first individual 709c1, and second individual 709c2. Rather, the representation of physical environment 704 in 7N includes a depiction of the right side of the physical environment.

[0185] 7N , a first location of the physical environment (e.g., the location where computer system 700 displays spatial capture virtual object 740 within the physical representation of physical environment 704) is not within the field of view of one or more cameras in communication with computer system 700. Thus, in FIG. 7N , computer system 700 does not display spatial capture virtual object 740. In some embodiments, when spatial capture virtual object 740 is displayed as a virtual environment configuration within a full-screen display that does not represent physical environment 704 (e.g., as described above in connection with FIG. 7M ), the computer system displays a perspective view of the virtual environment corresponding to the user's viewpoint that is rotated 90 degrees clockwise from the initial perspective view of the virtual environment. In FIG. 7N , computer system 700 begins to rotate counterclockwise (e.g., 90 degrees to the left), as indicated by movement indication 721 in schematic diagram 701.

[0186] In Figure 7O, the user's perspective has rotated 90 degrees counterclockwise relative to the user's perspective in Figure 7N. As shown in Figure 7O, the elapsed time indication 744 shows "00:06." Thus, one second has elapsed between Figure 7M and Figure 7O. In Figure 7O, a first location of the physical environment (e.g., the location where the computer system 700 displays the spatial capture virtual object 740 within the representation of the physical environment 704) is within the field of view of one or more cameras in communication with the computer system 700. Thus, as shown in Figure 7O, the computer system 700 displays the spatial capture virtual object 740 within the representation of the physical environment 704. In Figure 7O, the computer system 700 continues playing the video media represented by the spatial capture virtual object 740. Thus, in Figure 7O, the computer system 700 displays a frame of the video media item one second after the frame of the video media item shown in Figure 7M, indicating that one second has advanced in the playback of the video media item.

[0187] In FIG. 7O, computer system 700 begins to move toward a first location in the physical environment, as indicated by diagram 701 including movement indication 721 .

[0188] In FIG. 7P , as indicated by the positioning of indication 703 in schematic diagram 701, computer system 700 is at a first location within the physical environment (e.g., a location where computer system 700 displays spatial capture virtual object 740 in a representation of physical environment 704). Because computer system 700 is positioned at the first location within the physical environment, computer system 700 plays video media represented by spatial capture virtual object 740 from an immersive perspective. Immersive visual content is visual content that includes content from multiple perspectives captured from the same first point (e.g., location) within the physical environment at a given time. Playback (e.g., playback) of content from an immersive (e.g., first-person) perspective includes playing back the content from a perspective that matches the first viewpoint and can provide multiple different perspectives (e.g., fields of view) depending on user input. In some embodiments, while computer system 700 is at a first location in the physical environment, computer system 700 is moved forward in the physical environment (e.g., beyond the first location in the physical environment) to an updated position, causing the computer system to display spatial capture virtual object 740 from a non-immersive perspective view at a location in the representation of physical environment 704 that corresponds to a location in front of the updated position of computer system 700 in the physical environment. In some embodiments, while computer system 700 is at the first location in the physical environment, computer system 700 is moved backward in the physical environment (e.g., moved so that computer system 700 is positioned in front of the first location in the physical environment), causing computer system 700 to display spatial capture virtual object 740 from a non-immersive perspective view at a location in the representation of physical environment 704 that corresponds to the first location. Thus, by changing the position of computer system 700 within the physical environment, a user has the ability to control when computer system 700 displays media items from an immersive perspective view or a non-immersive perspective view.

[0189] 7P, computer system 700's playback of video media represented by spatial capture virtual object 740 from an immersive perspective view occupies the entirety of display 702. While computer system 700 plays video media represented by spatial capture virtual object 740 from an immersive perspective view, a representation of physical environment 704 is not visible. In some embodiments, a representation of physical environment 704 is visible while computer system 700 plays video media represented by spatial capture virtual object 740 from an immersive perspective view.

[0190] As shown in FIG. 7P , the time indication is shown as 00:07. Thus, one second has elapsed between FIG. 7O and FIG. 7P . In FIG. 7P , the computer system 700 continues playing the video media represented by the spatial capture virtual object 740. Thus, in FIG. 7P , the computer system 700 displays a frame of video media that is one second before the frame of video media shown in FIG. 7O to indicate that one second has advanced in the playback of the video media (e.g., the spatial capture virtual object 740 shows the user as having advanced along a trajectory). In some embodiments, in response to detecting that the computer system is at a first location within the physical environment, the computer system resumes playing the video media represented by the spatial capture virtual object 740 from the beginning. In FIG. 7P , the computer system 700 begins to rotate clockwise (e.g., 90 degrees to the right), as indicated by the movement indication 721 in the schematic diagram 701.

[0191] In Figure 7Q, the user's viewpoint is rotated 90 degrees clockwise in the physical environment relative to the user's viewpoint in Figure 7P. In Figure 7Q, as indicated by indication 703 in schematic diagram 701, computer system 700 is located at a first location in the physical environment. In response to the user's viewpoint being rotated 90 degrees clockwise in the physical environment as shown in Figure 7Q, computer system 700 displays playback of video media represented by spatial capture virtual object 740 from a perspective different from the perspective from which computer system 700 displays the video media in Figure 7P. More specifically, computer system 700 displays the video media from a perspective individually facing to the right of the trajectory shown on the video media. As described above, while computer system 700 is positioned at a first location in the physical environment, computer system 700 displays video media represented by spatial capture virtual object 740 from an immersive perspective. That is, the perspective of the playback of the media item represented by the spatial capture virtual object 740 changes based on changes in the user's viewpoint in the physical environment.

[0192] Additional description regarding FIGS. 7A-7Q is provided below with reference to methods 800, 900, and 1000 described with respect to FIGS. 7A-7Q.

[0193] 8 is a flow diagram of an exemplary method 800 for capturing media, according to some embodiments. In some embodiments, method 800 is performed on a computer system (e.g., computer system 101 and / or computer system 700 of FIG. 1 ) (e.g., a smartphone, tablet, and / or head-mounted device) that includes display generation components (e.g., display generation components 120 of FIGS. 1 , 3, and 4 ) (e.g., a head-up display, a display, a display controller, a touch-sensitive display system, a display (e.g., integrated and / or connected), a 3D display, a transparent display, a projector, a touchscreen, and / or a projector), and one or more cameras (e.g., a camera pointing down the user's hand (e.g., color sensors, infrared sensors, and other depth-sensing cameras) or a camera pointing forward from the user's head), and optionally a physical input mechanism. In some embodiments, the computer system is in communication with one or more eye-tracking sensors (e.g., optical and / or IR cameras configured to track the direction of gaze of a user of the computer system and / or the user's attention). In some embodiments, the first camera has an FOV that is outside at least a portion of the FOV of the second camera. In some embodiments, the second camera has an FOV that is outside at least a portion of the FOV of the first camera. In some embodiments, the first camera is positioned on a side of the computer system opposite to a side of the computer system on which the second camera is positioned. In some embodiments, method 800 is governed by instructions stored on a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 of FIG. 1 ). Some operations of method 800 are, optionally, combined and / or the order of some operations is, optionally, changed.

[0194] The computer system, via a display generation component (e.g., 702), displays a first user interface overlaid on a representation of a physical environment (e.g., 704), while the representation of the physical environment changes in response to a portion of the physical environment corresponding to the representation of the physical environment changing and / or a user's (e.g., 712) viewpoint changing (e.g., captured by one or more cameras in communication with the computer system) (e.g., a mixed reality and / or extended reality environment that is a representation of the physical environment), and detects (802) a request to display a media capture user interface (e.g., selection of 711a of 712a and / or 750b in FIG. 7B ). In some embodiments, as part of receiving the request to display the media capture user interface, the computer system detects input (e.g., press, swipe, and / or tap) on a hardware button. In some embodiments, as part of receiving the request to display the media capture user interface, the computer system detects the user's attention at one or more locations of the first user interface and / or one or more movement voice command inputs (e.g., via one or more microphones in communication with the computer system).

[0195] In response to detecting a request to display a media capture user interface (e.g., to capture immersive or semi-immersive visual media, capture three-dimensional media, capture three-dimensional stereo media, and / or capture spatial media) (e.g., to capture video content that can present content from multiple perspectives in response to detected changes in the orientation of the user and / or the computer system), the computer system (e.g., after initiating media capture) displays, via a display generation component (e.g., after initiating media capture), a media capture preview (e.g., 708) that includes representations (e.g., virtual representations and / or virtual objects) of portions of the field of view of one or more cameras, along with content that updates as portions of the physical environment within the portions of the field of view of one or more cameras change (e.g., a live preview and / or changes in the appearance of the physical environment that change the field of view of one or more cameras) (e.g., immersive media content as described above in connection with FIG. 7K) (804).

[0196] The media capture preview indicates the boundaries of the media to be captured in response to detecting a media capture input (e.g., activation of 711a and / or 750d of 712a in FIG. 7D) while the media capture user interface is displayed (806).

[0197] The media capture preview is displayed while a first portion of the representation of the physical environment (e.g., portion 709b including 704) is visible (e.g., as described above in connection with FIG. 7C ), and the first portion of the representation of the physical environment was visible (e.g., unobstructed) before a request to display the media capture user interface was detected (808).

[0198] The media capture preview is displayed in place of (e.g., overlaid on) a second portion of the representation of the physical environment (e.g., portion 709c1, including portion 704) (e.g., as described above in connection with FIG. 7C ), and the first portion of the representation of the physical environment is updated (810) as the portion of the physical environment corresponding to the first portion of the representation of the physical environment changes (e.g., as described above in connection with FIG. 7C ) and / or the user's viewpoint changes (e.g., one or more other portions of the representation of the physical environment remain visible while the media capture preview is displayed and / or visible). In some embodiments, the computer system 700 displays the media capture preview with a black border. In some embodiments, in response to detecting a request to capture media, the computer system begins capturing media (e.g., three-dimensional media, stereo media, and / or spatial media). In some embodiments, the first content captured by the first camera is different from the second content captured by the second camera. In some embodiments, the media capture preview is displayed at the center (e.g., middle) of the computer system (e.g., at the center of the display generation components), near the nose of the user wearing the computer system, and at the center of the three-dimensional representation of the physical environment. In some embodiments, the media capture preview is displayed between one or more virtual objects (e.g., a time lapse virtual object, a capture control virtual object, a camera roll virtual object, a close button virtual object). In some embodiments, the computer system displays the media capture preview in the center of the display above the shutter button virtual object (e.g., as described above in connection with FIG. 7C). In some embodiments, the captured media is played back using one or more techniques, such as those described above in connection with FIGS. 7K-7Q. In some embodiments, the media capture preview includes content of the physical environment, and the content is visible in the representation of the physical environment before the media capture preview is displayed.In some embodiments, during playback in a non-immersive perspective view, multiple perspective views from points in the physical environment other than the first point can be displayed in response to user input. Displaying a media capture preview indicating the boundaries of the media to be captured in response to detecting a media capture input provides the user with improved feedback regarding what content from the user's point of view will be captured. Displaying a media capture preview while portions of the representation of the physical environment are visible provides the user with the ability to better compose and capture desired media while also maintaining awareness of the physical environment, which improves media capture operations and reduces the risk of failing to capture transient events that may be missed if the capture operation is inefficient or difficult to use. Improving media capture operations improves system usability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).

[0199] In some embodiments, the representation of the physical environment (e.g., 704 in FIG. 7B ) is a pass-through representation (e.g., the pass-through representation is virtual (e.g., a representation of camera image data captured by one or more cameras integrated into the computer system) and / or optical pass-through (e.g., light passing directly through a portion of the system (e.g., a transparent portion) relative to the user)) of the real-world environment (e.g., as seen in FIG. 7A ) of the computer system (e.g., a portion of the real-world environment surrounding the computer system). In some embodiments, as the position and / or orientation of the computer system changes, the pass-through representation is updated to reflect the change in position and / or orientation of the computer system. Providing a pass-through representation of the computer system's real-world environment provides the user with visual feedback regarding the location (e.g., position and / or orientation) of the computer system within the real-world environment, which provides improved visual feedback, particularly when the pass-through representation is visible while a media capture preview is displayed.

[0200] In some embodiments, a representation (e.g., content within the representation) of a physical environment (e.g., 704) is visible (e.g., displayed or visible via an optical pass-through) at a first scale (e.g., 1:1, 1:2, 1:4, and / or any other suitable scale), and a representation (e.g., content within the representation) of a portion of the field of view of one or more cameras (e.g., a second portion of the representation of the physical environment) included in the media capture preview (e.g., 708) is displayed at a second scale, the first scale being larger than the second scale (e.g., the same object appearing in both representations appears larger in the representation of the physical environment) (e.g., as described above in connection with FIG. 7C ). Displaying the media capture preview at a smaller scale than the scale at which the representation of the physical environment is visible provides the user with improved visual feedback as to which representation is which and reduces user confusion. Doing so also saves display space and allows both to be visible together, which provides the user with the ability to better organize and capture desired media while maintaining awareness of the physical environment, which improves media capture operations and reduces the risk of failing to capture transient events that may be missed if the capture operations are inefficient or difficult to use. Improving media capture operations improves system usability and makes the user-system interface more efficient (e.g., by helping users provide appropriate inputs and reducing user errors when operating / interacting with the device).

[0201] In some embodiments, the representation of the physical environment (e.g., 704 in FIG. 7G ) (e.g., a separate portion) includes first content (e.g., the right half of 709E in FIG. 7G ) that is not included in (e.g., outside of) (e.g., not displayed within) the representation of the field of view (e.g., 708 in FIG. 7G ) included in the media capture preview. Having a representation of the physical environment that includes content not included in the media capture preview provides the user with the ability to view additional content that may be included in (but is not currently included in) the media capture if the user shifts their viewpoint, as well as providing the user with greater awareness of their current physical environment while configuring the media capture. Doing so improves the media capture operation and reduces the risk of failing to capture transient events and / or content that may be missed if the capture operation is inefficient or difficult to use. Improving the media capture operation improves system usability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).

[0202] In some embodiments, the representation of the physical environment (e.g., 704 of FIG. 7C ) includes a portion (e.g., a first portion, a particular portion) of the physical environment (e.g., different from the first portion of the representation of the physical environment (e.g., a second portion of the physical environment has a narrower angular range than the first portion of the physical environment)) (e.g., 709c1 of FIG. 7C ), and the representation of the field of view of one or more cameras included in the media capture preview (e.g., 708 of FIG. 7C ) includes a portion (e.g., 709c3 of FIG. 7C ) of the physical environment (e.g., the media capture preview includes objects that are also included in the representation of the physical environment). In some embodiments, the second portion of the physical environment included in the representation of the physical environment has a different visual appearance (e.g., blurred and / or dimmed) than the second portion of the physical environment included in the media capture preview.

[0203] In some embodiments, the computer system communicates (e.g., direct communication (e.g., wired communication) and / or wireless communication) with a physical input mechanism (e.g., 711a or 711b) (e.g., a hardware button) (e.g., a hardware input device / mechanism) (e.g., a physical input device), and the request to display the media capture user interface includes activation (e.g., actuation and / or selection (e.g., button press)) of the physical input mechanism (e.g., activation of 711a or 712a in FIG. 7B). Displaying a media capture preview including representations (e.g., virtual representations and / or virtual objects) of portions of the field of view of one or more cameras in response to detecting activation of the physical input mechanism allows the computer system to perform display operations that provide the user with greater control over the computer system without displaying additional controls, which provides additional control options without cluttering the user interface.

[0204] In some embodiments, while displaying the media capture preview, the computer system detects (e.g., via one or more input devices in communication with the computer system) input corresponding to a request to capture media (e.g., activation of 750c and / or 711a of 712a of FIG. 7C ). In response to detecting the input corresponding to the request to capture media, following a determination that the input corresponding to the request to capture media is of a first type (e.g., a single, rapid press of a hardware button (e.g., a virtual shutter button), a tap gesture on a touch-sensitive surface, or a rapid air gesture (e.g., an air gesture of shorter duration than a second type of input described below) (e.g., a pinch-and-release or any other suitable air gesture as described above with respect to selecting a virtual object in an XR environment), the computer system initiates a process to capture (e.g., using one or more cameras of the computer system) media content of the first type (e.g., still media (e.g., a photo)) and displays the media (e.g., as described above with respect to FIG. 7C ). 7C ), initiate a process for capturing (e.g., using one or more cameras of the computer system) the second type of media content (e.g., video) (e.g., different from the first type of media content) (e.g., as described above in connection with FIG. 7C ) pursuant to a determination that the input corresponding to the request to capture is of a second type (e.g., different from the first type) (e.g., a long press of a button, a sustained gaze directed toward the display of the virtual shutter object, a touch-and-hold gesture at a location in space corresponding to the display of the virtual shutter object, a touch-and-hold gesture on the virtual shutter button, a sustained air gesture (e.g., a held pinch), or any other suitable air gesture such as those described above with respect to selecting a virtual object in an XR environment).Initiating either a first type of media capture operation or a second type of media capture operation based on whether a type of input (e.g., a short press or a long press) is received allows the computer system to provide the user with additional control options regarding the type of media capture operation the computer system performs without displaying additional controls, which provides additional control options without cluttering the user interface.

[0205] In some embodiments, the representations of the portions of the fields of view of one or more cameras included in the media capture preview (e.g., 708) have a first set of visual parallax characteristics, and the representation of the physical environment has a second set of visual parallax characteristics that are different from the first set of parallax characteristics (e.g., as described above in connection with FIG. 7F ). In some embodiments, the media capture preview includes a first foreground and a first background. In some embodiments, as the portions of the physical environment within the portions of the fields of view of the one or more cameras change, there is a first amount of shift between the first foreground and the first background. In some embodiments, the representation of the physical environment includes a second foreground and a second background. In some embodiments, as the portions of the physical environment within the portions of the fields of view of the one or more cameras change, there is a second amount of shift (e.g., different from the first shift amount) between the second foreground and the second background (e.g., the first foreground moves at a first velocity relative to the first background and the second foreground moves at a second velocity (different from the first velocity) relative to the second background) as the portions of the physical environment within the portions of the fields of view of the one or more cameras change). Providing a first set of visual parallax characteristics to a representation of a portion of the field of view of one or more cameras included in the media capture preview and a second set of visual parallax characteristics to a representation of the physical environment provides visual feedback to a user as to whether the computer system is moving, and also provides feedback as to which representations are previews of the media capture and which representations are previews of the physical environment, providing improved visual feedback.

[0206] In some embodiments, the representation of the physical environment (e.g., 704) is an immersive perspective (e.g., first-person perspective) (e.g., a representation of the environment is presented from multiple perspectives in response to detecting a change in the orientation of the user and / or computer system), and the representation of the portion of the field of view of one or more cameras included in the media capture preview is a non-immersive perspective (e.g., third-person perspective) (e.g., as described above with respect to FIG. 7C ). In some embodiments, the immersive perspective includes content of the representation of the environment that includes content from multiple perspectives captured from the same point (e.g., location) in the environment. Providing a representation of the physical environment from an immersive perspective and providing the portion of the field of view of one or more cameras included in the media capture preview from a non-immersive perspective provides improved visual feedback as to which representation is which and also provides feedback as to how the captured media may be previewed when viewed non-immersively (e.g., in a non-immersive photo album).

[0207] In some embodiments, prior to detecting a request to display a media capture user interface, a third portion (e.g., a portion the same as or different from the first portion) of the representation of the physical environment (e.g., the entire representation of the environment) (e.g., the first portion of the representation of the environment less than the entire representation of the environment) has a first visual appearance including visual characteristics (e.g., dimming, blurring, and / or superimposed filters) having a first magnitude (e.g., no amount, a low non-zero amount (e.g., 0%, 10%, or 20% of the maximum amount)). In some embodiments, in response to detecting a request to display a media capture user interface (e.g., selection of 750b and / or 711a of 712a of FIG. 7B ), the computer system modifies the third portion of the representation of the physical environment to have a second visual appearance (e.g., different from the first visual appearance) including visual characteristics having a second magnitude, e.g., as described above in connection with FIG. 7C . In some embodiments, the second magnitude of the visual characteristics is different (e.g., larger or smaller than) the first magnitude of the visual characteristics (e.g., as described above in connection with FIG. 7C ). In some embodiments, the magnitude of the visual characteristic is an amount of dimming to be applied to the third portion of the representation, where, prior to receiving the request, a first magnitude is no dimming applied and a second amount is a non-zero level of dimming to be applied. Altering the visual appearance of the third portion of the representation of the physical environment in response to detecting a request to display a media capture user interface provides visual feedback to the user regarding the state of the computer system (e.g., that the computer system has detected a request to display a media capture user interface) and also enhances the media capture preview and provides improved visual feedback.

[0208] In some embodiments, displaying the media capture preview includes displaying one or more virtual objects (e.g., 713, 714, 715, and 716) (e.g., that, when selected, cause the computer system to perform an action (e.g., a media-related action (e.g., capturing a photo, capturing a video)) having a spatial relationship (e.g., a predetermined distance and relative orientation at which the one or more virtual objects are displayed with respect to the display of the media capture preview and / or a position (e.g., above, below, and / or to the side) at which the one or more virtual objects are displayed with respect to the display of the media capture preview) with the display of the media capture preview (e.g., 708) (e.g., each virtual object of the one or more virtual objects has a distinct spatial relationship with the media capture preview). 7C-7E ), and while displaying the media capture preview and the one or more virtual objects at the first display location (e.g., locations 713, 714, 715, and 716 in FIGS. 7C-7E ), the computer system detects a change in the pose of the user's viewpoint (e.g., via one or more sensors in communication with the computer system) (e.g., as described above with respect to FIG. 7E ). In some embodiments, in response to detecting a change in the user's viewpoint pose, the computer system displays the media capture preview and one or more virtual objects at a second display location (e.g., locations 708, 713, 714, 715, and 716 in FIG. 7F) that is different from the first location (e.g., at a different location on the display generation component), and the computer system maintains a spatial relationship between the display of the media capture preview and the one or more virtual objects.Maintaining a spatial relationship between the representation of the media capture preview and the one or more virtual objects in response to detecting a change in viewpoint pose provides visual feedback to the user that allows the user to easily locate the representation of the one or more virtual objects, which provides improved visual feedback. Displaying the one or more virtual objects based on the computer system's representation of the media capture preview causes the computer system to perform display operations that provide the user with additional control options related to media capture without requiring further user input, which reduces the number of inputs required to perform the operations.

[0209] In some embodiments, the one or more virtual objects include a time lapse virtual object (e.g., 713) that provides an indication of the amount of time (e.g., seconds, minutes, and / or hours) that has elapsed since the start of a process to capture media (e.g., video recording). Displaying a time lapse virtual object that indicates the amount of time that has elapsed since a computer system began a process to capture media provides visual feedback regarding the video media recording status of a computer system and provides improved visual feedback.

[0210] In some embodiments, the one or more virtual objects include a shutter button virtual object (e.g., 714) (e.g., a software shutter button), which, when selected (e.g., selected via detection of a user's gaze (e.g., gaze and dwell) directed at the shutter button virtual object, and in some embodiments, in combination with detection of a user performing one or more gestures (e.g., an air pinch gesture, a de-pinch air gesture, an air tap, and / or an air swipe) (e.g., as described above with reference to selecting a virtual object in an XR environment) and / or selected via detection of a tap on the shutter button virtual object), causes initiation of a process for capturing media (e.g., causing a computer system to initiate a process for capturing media (e.g., still media or video media) (e.g., capturing media using one or more cameras in communication with the computer system). In some embodiments, the display of the shutter button virtual object is updated to indicate that the computer system is recording video media. Displaying a shutter button virtual object anchored to the display of the media capture preview provides visual feedback regarding the display state of the computer system (e.g., that the computer system is currently displaying the media capture preview) and also provides actions that can be performed by interacting with the object, which provides improved visual feedback.

[0211] In some embodiments, one or more virtual objects, including a camera roll virtual (e.g., 715) object, when selected (e.g., via detection of a user's gaze (e.g., gaze and dwell) directed at the camera roll virtual object), in combination with detection of a user performing one or more gestures (e.g., air pinch gesture, de-pinch air gesture, air tap, and / or air swipe) (e.g., as described above with reference to selecting a virtual object in an XR environment) and / or via detection of a tap on the camera roll virtual object, cause (e.g., via a display generation component) a display (e.g., causing a computer system to display previously captured media items) (e.g., still media or video media) (e.g., media items previously captured using one or more cameras in communication with the computer system) (e.g., media items previously captured using an external device (e.g., a smartphone) (e.g., a device separate from the computer system)). In some embodiments, in response to detecting selection of the camera roll virtual object, the computer system displays a user interface including a subset of the one or more virtual objects. Displaying a camera roll virtual object that is anchored to the display of the media capture preview provides visual feedback regarding the display state of the computer system (e.g., that the computer system is currently displaying the media capture preview) and also provides actions that can be performed by interacting with the object, which provides improved visual feedback.

[0212] In some embodiments, while displaying the media capture preview, the computer system detects selection (e.g., 750h) of a camera roll virtual object. In some embodiments, in response to detecting the camera roll virtual object (e.g., 715 of FIG. 7H ), the computer system, via the display generation component, displays a representation of a previously captured media item (e.g., 730) (e.g., ceases displaying the media capture preview). In some embodiments, displaying the representation of the previously captured media is a first erase virtual object (e.g., 733) that, when selected (e.g., selected via detection of a user's gaze (e.g., gaze and dwell) directed at the location of the display of the first erase virtual object, in some embodiments selected in combination with detection of the user performing one or more gestures (e.g., air pinch gesture, de-pinch air gesture, air tap, and / or air swipe), and / or selected via the user tapping the first erase virtual object (e.g., as described above with reference to selecting a virtual object in an XR environment)), causes the computer system to cease displaying the representation of the previously captured media item (e.g., causes the computer system to cease displaying the previously captured media item). a first erased virtual object, a shared virtual object (e.g., 735), which, when selected (e.g., selected via detection of a user's gaze (e.g., gaze and dwell) directed at the location of the representation of the shared virtual object, and in some embodiments selected in combination with detection of the user performing one or more gestures (e.g., air pinch gesture, de-pinch air gesture, air tap, and / or air swipe) (e.g., as described above with reference to selecting a virtual object in an XR environment), and / or selected via a user tapping on the shared virtual object), causes initiation of a process (e.g.,a shared virtual object, a media library virtual object (e.g., 731), which, when selected (e.g., via detection of a user's gaze (e.g., gaze and dwell) directed at a location of a representation of the media library virtual object, in some embodiments in combination with detection of a user performing one or more gestures (e.g., air pinch gesture, de-pinch air gesture, air tap, and / or air swipe) (e.g., as described above with reference to selecting virtual objects in an XR environment), and / or via a user tap on the media library virtual object), pulls together representations of the plurality of previously captured media items. The methods include: causing the computer system to display a plurality of previously captured media items (e.g., non-immersive and / or immersive media items), a media library virtual object, and / or displaying (e.g., simultaneously displaying) a resize virtual object (e.g., 736) indicating a change in size (e.g., an increase in size or a decrease in size) of a representation of a previously captured media item based on detection (e.g., detected by one or more cameras in communication with the computer system) of one or more gestures (e.g., gestures on a touch-sensitive display and / or air gestures (e.g., air pinch-and-drag gestures, air swipe, and / or air taps)) (e.g., as described above with respect to selecting a virtual object in an XR environment). In some embodiments, the magnitude of the change in size is based on characteristics of the gesture (e.g., the amount of resizing is based on the distance of the dragging motion of a pinch-and-drag gesture). In some embodiments, selecting the first erased virtual object causes the media capture preview to be redisplayed. Displaying the plurality of virtual objects based on the previously captured representation of the media item provides visual feedback regarding the display state of the computer system (e.g., the computer system is currently displaying a representation of the previously captured media item);It also provides actions that can be performed by interacting with the object, which provides improved visual feedback.

[0213] In some embodiments, the computer system receives one or more sets of inputs including an input corresponding to a camera roll virtual object (e.g., 715). In some embodiments, the one or more sets of inputs are a selection of a camera roll virtual object followed by a selection of a virtual object corresponding to a subset of previously captured media items (e.g., all media items captured from an immersive perspective view). In some embodiments, in response to receiving the set of one or more inputs, the computer system displays a representation (e.g., a still photograph and / or video) of a first previously captured media item (e.g., 730 in FIGS. 7I and 7J ) of the plurality of previously captured media items at a first location (e.g., a central location on a display generation component) (e.g., the first representation of the first previously captured media item is selectable) (e.g., the first previously captured media item is captured by the computer system), and while displaying the representation of the first previously captured media item at the first location, the computer system receives a request (e.g., an air pinch gesture, a de-pinch air gesture, an air tap, and / or an air swipe) to navigate to a different previously captured media item of the plurality of previously captured media items (e.g., as described above with respect to selecting a virtual object in an XR environment). In some embodiments, in response to receiving a request to navigate to a different previously captured media item among the plurality of previously captured media items, the computer system replaces the display of a representation of the first previously captured media item at the first location with a display of a representation of a second previously captured media item (e.g., different from the first previously captured media item) among the plurality of previously captured media items (e.g., as described above in connection with FIG. 7K).In some embodiments, replacing the display of the representation of the first previously captured media item includes ceasing to display the first previously captured media item. In some embodiments, the first previously captured media item is displayed (e.g., at a second location different from the first location) while the second previously captured media item is displayed at the first location. In some embodiments, the representation is displayed in place of a portion of the representation of the physical environment. In some embodiments, while the second previously captured media item is displayed, the computer system receives a second request to navigate to a different (e.g., different from the first previously captured media item and the second previously captured media item), where the second request is a repeat of the initial request (e.g., the initial request and the second request include the same type of gesture) and / or the second request includes a gesture that is the opposite of (e.g., performed in the opposite direction) the gesture included in the initial request. In some embodiments, the request to navigate to a previously captured media item of the plurality of previously captured media items is an air gesture (e.g., as described above with reference to selecting a virtual object in an XR environment), and the second previously captured media item is selected by the computer system based on the magnitude of the air gesture (air gesture direction, air gesture speed, and / or air gesture intensity). Replacing the display of the representation of the first previously captured media item at the first location with the display of the representation of the second previously captured media item provides visual feedback regarding the state of the computer system (e.g., that the computer system received a request to navigate to a different previously captured media item while the first previously captured media item was being displayed), which provides improved visual feedback.

[0214] In some embodiments, the one or more virtual objects include a second, erased virtual object (e.g., 719), which, when selected (e.g., selected via detection of a user's gaze (e.g., gaze and dwell) directed at a location of the display of the second, erased virtual object, and in some embodiments, in combination with detection of a user performing one or more gestures (e.g., air pinch gesture, de-pinch air gesture, air tap, and / or air swipe) (e.g., as described above with reference to selecting virtual objects in an XR environment) and / or when selected via a user tap on the second, erased virtual object), causes the display of the media capture preview (e.g., 708) to cease (e.g., cause the computer system to cease displaying the media capture preview). In some embodiments, ceasing display of the media capture preview results in the display of a portion (e.g., a second portion) of the representation of the environment that was not displayed while the media capture preview was displayed. Displaying a second, erased virtual object anchored to the display of the media capture preview provides visual feedback regarding the display state of the computer system (e.g., that the computer system is currently displaying the media capture preview) and also provides actions that may be performed by interacting with the object, which provides improved visual feedback.

[0215] In some embodiments, the one or more virtual objects include a repositioned virtual object (e.g., 716), and while the media capture preview (e.g., 708) is displayed at a first location within the media capture user interface, the computer system detects a set of one or more inputs (e.g., gesture(s) and / or air gesture(s) on the touch-sensitive surface) that include inputs corresponding to the repositioned virtual object. In some embodiments, in response to detecting the set of one or more inputs that include inputs corresponding to the repositioned virtual object, the computer system moves the media capture preview from the first location to a second location (e.g., as described above in connection with FIG. 7C ) (e.g., the media capture preview is moved (e.g., along the x-axis, y-axis, and / or z-axis of the media capture user interface) based on the direction, magnitude, and / or speed of the request to move the representation of the repositioned virtual object). In some embodiments, the set of one or more inputs includes an input that selects the repositioned virtual object and one or more inputs that specify a subsequent target location and / or direction of movement. Displaying a fixed repositioning object in the display of the media capture preview provides visual feedback regarding the display state of the computer system (e.g., that the computer system is currently displaying the media capture preview) and also provides actions that can be performed by interacting with the object, which provides improved visual feedback.

[0216] In some embodiments, while the media capture preview (e.g., 708) is displayed, the computer system, via a display generation component, simultaneously displays a media library virtual object (e.g., 731), which, when selected (e.g., via detection of a user's gaze (e.g., gaze and dwell) directed at the location of the display of the media library virtual object, in some embodiments in combination with detection of a user performing one or more gestures (e.g., air pinch gesture, de-pinch air gesture, air tap, and / or air swipe) (e.g., as described above with reference to selecting a virtual object within an XR environment) and / or via a user tap on the media library virtual object, causes the display of multiple previously captured media items (e.g., as described above with reference to FIG. 7C ), and causes the computer system to display multiple previously captured media items. In some embodiments, selection of a media library virtual object causes the display of the media library virtual object to cease. In some embodiments, selection of a media library virtual object causes the display of the media capture preview to cease. In some embodiments, selection of a media library virtual object causes multiple previously captured media items to be displayed below, above, and / or beside the display of the media capture preview. Simultaneous display of media library virtual objects (e.g., displayed while the media capture preview is displayed) provides visual feedback regarding the display state of the computer system (e.g., that the computer system is currently displaying the media capture preview) and also provides actions that can be performed by interacting with the objects, which provides improved visual feedback.

[0217] In some embodiments, the portion of the field of view of one or more cameras included in the media capture preview (e.g., 708) (e.g., the horizontal and / or vertical field of view of one or more cameras) has a first viewing angle range (e.g., 0-45°, 0-90°, 40-180°, or any other suitable angle range) (e.g., the portion of the field of view of one or more cameras included in the media capture preview is based on an overlap of the field of view of one or more cameras) (e.g., the media capture preview includes only data from the overlapping portion of the field of view of one or more cameras) (e.g., the media capture preview The view does not include data from non-overlapping portions of the field of view of one or more cameras), the representation of the physical environment (e.g., 704) represents a second field of view (e.g., horizontal and / or vertical field of view of one or more cameras) of one or more cameras having a second viewing angle range (e.g., 0-45°, 0-90°, 40-180°, or any other suitable angle range) (e.g., the second field of view includes non-overlapping data from one or more cameras), and the first angular range is narrower (e.g., smaller) than the second angular range (e.g., the second angle is wider than the first angular range) (e.g., as described above in connection with FIG. 7C ). In some embodiments, the first angular range is a subset of the second angular range. Displaying a representation of the physical environment that includes a wider field of view than the media capture preview provides the user with the ability to view additional content that may be included (but is not currently included) in the media capture preview if the user shifts their viewpoint, as well as providing the user with a greater awareness of the user's current physical environment while configuring media capture. Doing so improves media capture operations and reduces the risk of failing to capture transient events and / or content that may be missed if the capture operations are inefficient or difficult to use. Improving media capture operations improves usability of the system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device).

[0218] In some embodiments, the representation of the portion of the field of view of the one or more cameras included in the media capture preview (e.g., 708) includes first content (e.g., 709c3 and 709c4 of FIG. 7F) (e.g., three-dimensional content) within the field of view of a first camera of the one or more cameras (e.g., content captured by the first camera in response to detecting a request to capture media and / or content saved and / or stored by a computer system), and (e.g., as described above in FIG. 7C) a field of view of a second camera of the one or more cameras (e.g., content captured by the second camera in response to detecting a request to capture media and / or content saved and / or stored by a computer system). In some embodiments, portions of the physical environment within the field of view of the first camera but not the second camera are not included in the media capture preview. In some embodiments, portions of the physical environment within the field of view of the first camera but not the second camera are not included in the media capture preview, but are part of the representation of the physical environment.

[0219] In some embodiments, aspects / operations of methods 800, 900, 1000, 1200, 1400, and 1500 may be interchanged, substituted, and / or added between these methods. For example, media items captured in method 800 may be displayed as part of method 1000. For the sake of brevity, those details will not be repeated here.

[0220] 9 is a flow diagram of an example method 900 for displaying a preview of media, according to some embodiments. In some embodiments, method 900 is performed on a computer system (e.g., computer system 101 and / or computer system 700 of FIG. 1 ) (e.g., a smartphone, tablet, and / or head-mounted device) in communication with a display generation component (e.g., display generation component 120 and / or 702 of FIGS. 1 , 3, and 4 ) (e.g., a display controller, a touch-sensitive display system, a display (e.g., integrated and / or connected), a 3D display, a see-through display, a projector, a head-up display, and / or a head-mounted display) and one or more cameras (e.g., a camera pointing down the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing camera) or a camera pointing forward from the user's head). In some embodiments, the computer system is in communication with one or more eye-tracking sensors (e.g., optical and / or IR cameras configured to track the direction of a user's gaze and / or the attention of a user of the computer system). In some embodiments, the computer system includes a first camera. In some embodiments, method 900 is governed by instructions stored on a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 202 (e.g., control 110 in FIG. 1 ) of computer system 101. Some operations of method 900 are, optionally, combined, and / or the order of some operations is, optionally, changed.

[0221] The computer system displays, via a display generating component, an extended reality user interface including a preview (e.g., a real-time preview) of a field of view of one or more cameras (e.g., 708 of 7E) (e.g., virtual cameras or physical cameras configured to communicate with the computer system) overlaid on a first portion (e.g., the first portion of the environment is visible via the display generating component) of a three-dimensional environment (e.g., 704) visible from the user's (e.g., 712) perspective (e.g., the environment visible on the display generating component) while the user's (e.g., 712) (e.g., and / or the computer system) perspective is in a first pose (e.g., an orientation and / or position within the physical environment and / or the virtual environment) (e.g., and / or while the user's head is detected as being present), the preview including a representation (e.g., a three-dimensional representation, a spatial representation, and / or a two-dimensional representation) of the first portion of the three-dimensional environment displayed in a distinct spatial configuration relative to the user's perspective (e.g., 708 of FIGS. 7C-7E) (902). In some embodiments, the field of view of the first camera is a portion of the field of view of the first camera, but not the entire field of view of the first camera. In some embodiments, the preview includes a representation of at least a portion of the field of view of a second camera (e.g., different from the first camera) that includes a representation of a portion of the physical environment. In some embodiments, the field of view of the second camera is a portion of the field of view of the second camera. In some embodiments, the preview is displayed at the center (e.g., middle) of the computer system (e.g., at the center of the display generation component), near the nose of a user wearing the computer system, and at the center of the three-dimensional representation of the physical environment.

[0222] The computer system detects (904) a change in the pose of the user's viewpoint (and / or the user's head) from a first pose to a second pose different from the first pose (e.g., as described above in connection with FIGS. 7E-7F) (e.g., while displaying an extended reality environment user interface including a preview of the first camera's field of view overlaid on the first location on the extended reality user interface). In some embodiments, while the computer system and / or user's viewpoint is in the second pose, the first camera's field of view is directed toward a second portion of the physical environment and / or the second portion of the physical environment is visible from the user's viewpoint while the first camera is in the second pose.

[0223] In response to detecting a change in the user's viewpoint pose from the first pose to the second pose, the computer system shifts (e.g., changes and / or transitions the display of) one or more camera field of view previews (e.g., to pan, change, and / or update) away from a distinct spatial configuration (e.g., 708 in Figures 7F and 7G) relative to the user's viewpoint in a direction determined based on the change in the user's viewpoint pose from the first pose to the second pose (e.g., based on the direction and / or speed of the change in the viewpoint pose), the shifting of the one or more camera field of view previews occurs at a first rate (e.g., 708 in Figures 7F and 7G), and while the one or more camera field of view previews are shifting based on the change in the user's viewpoint pose, the representation of the three-dimensional environment changes (906) based on the change in the user's viewpoint pose at a second rate (e.g., while continuing to update the preview at the same rate at which the viewpoint movement is updated) that is different from the first rate (e.g., as described above with respect to Figure 7F). In some embodiments, the preview is updated to include a representation of a second portion of the environment at the first velocity while being moved at the second velocity. In some embodiments, displaying one or more objects in the representation of the field of view is updated at the same rate as one or more objects in the viewpoint of the physical environment. In some embodiments, in response to detecting a change in orientation of the user's viewpoint from a first orientation to a second orientation, the first portion of the environment becomes invisible. In some embodiments, the first portion of the environment is not encompassed by the second portion of the environment, surrounds the second portion of the environment, is separate from the second portion of the environment, is distinct from the second portion of the environment, includes the second portion of the environment, and / or vice versa. In some embodiments, locations included in the first portion of the environment are distinct from locations included in the first portion of the environment.Shifting the preview of one or more cameras' fields of view based on changes in the user's viewpoint pose at a first rate, while the representation of the three-dimensional environment changes based on changes in the user's viewpoint pose at a second rate different from the first rate (e.g., in response to detecting a change in the user's viewpoint pose from the first pose to the second pose), allows the computer system to automatically perform an action that reduces the amount of movement between elements shown to the user and reduces the likelihood of motion sickness, performing the action when a set of conditions is met without requiring further user input. Doing so may also prompt the user to reduce changes to the viewpoint while in media capture mode, as the user may tend to maintain focus on the preview. Reducing viewpoint changes while capturing media can improve media capture operations, reducing the risk of failing to capture transient events and / or content that could be missed if the capture operation is inefficient or difficult to use. Improving media capture operations improves system usability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).

[0224] In some embodiments, the computer system (e.g., 700) tracks changes in the pose of the user's (e.g., 712) viewpoint with a first amount of tracking lag (e.g., a delay between the actual change in viewpoint and the computer system configured to track the amount / degree of that movement (e.g., a delay of 0.1, 0.2, 0.3, 0.4, or 0.5 meters / second)) (e.g., as described above in connection with FIGS. 7E-7F ), and the first rate introduces an amount of visual delay in updating the location of the preview of the field of view of one or more cameras (e.g., 708) that is greater than the amount of visual delay in updating the location of the preview of the field of view of one or more cameras that is introduced based on (e.g., based solely on) the first amount of detected tracking lag (e.g., as described above in connection with FIG. 7F ) (e.g., the detected tracking lag is 0.1 meters / second and the first rate is 0.09 meters / second). In some embodiments, the maximum rate of shifts in the preview of the field of view of one or more cameras is limited (e.g., capped) to a value less than the amount (e.g., current amount) of tracking lag. (In some embodiments, the amount of tracking lag is proportional to the rate of pose of the eye point.) Maintaining the first velocity less than the amount of tracking lag reduces the occurrence of disruptions in the display of the preview caused by tracking lag and automatically adjusts the speed of the preview movement based on system parameters without requiring further user input.

[0225] In some embodiments, while the preview of the field of view of one or more cameras is shifting at a first rate (e.g., 708 in FIGS. 7F and 7G ), the computer system detects that the user's viewpoint is changing pose less than a first threshold amount (e.g., changing at a rate below a threshold, or not changing (e.g., a static pose) (e.g., the user is currently positioned in a second pose)). In some embodiments, in response to detecting that the user's viewpoint is changing pose less than the threshold amount (e.g., as described above in connection with FIG. 7G ), the computer system shifts the preview of the field of view of one or more cameras toward the distinct spatial configuration at a third rate faster than the first rate (e.g., as described above in connection with FIG. 7H ) (e.g., in response to detecting that the user has stopped moving, the preview of the one or more cameras “snap” back to its original position with the distinct spatial configuration relative to the user's viewpoint). In some embodiments, the computer system shifts the preview of the field of view of one or more cameras in the opposite direction (e.g., opposite the direction of the shift away from the distinct spatial configuration) as the computer system shifts the preview of the field of view of one or more cameras toward the distinct spatial configuration. Shifting a preview of the field of view of one or more cameras toward a distinct spatial configuration in response to detecting that the user's viewpoint is changing pose by less than a threshold amount causes the computer system to display the preview of the field of view at a location on the display that is centered on the user's field of view, which performs the operation without requiring further user input.

[0226] In some embodiments, while the previews of the one or more cameras' fields of view are shifting at a first rate (e.g., 708 in FIGS. 7F and 7G ), the computer system detects that the user's viewpoint is changing pose less than a second threshold amount (e.g., changing at a rate below a threshold or not changing (e.g., a static pose)) (e.g., a non-zero amount) (e.g., as described above in connection with FIG. 7G ). In some embodiments, in response to detecting that the user's viewpoint is changing pose less than the second threshold amount, the computer system ceases shifting (e.g., automatically (e.g., without user input)) the previews of the one or more cameras' fields of view away from the distinct spatial configuration (e.g., as described above in FIG. 7H ), and the computer system displays previews of the one or more cameras' fields of view having the distinct spatial configuration relative to the user's viewpoint (e.g., 708 in FIG. 7H ) (e.g., the previews of the one or more cameras' fields of view are displayed in a position where the previews of the fields of view were displayed before the computer system detected the movement of the user's viewpoint changing pose) (e.g., the previews of the one or more cameras' fields of view are displayed overlaid on a first portion of the three-dimensional environment). In some embodiments, after ceasing shifting the preview of the field of view of the one or more cameras, the preview of the field of view of the one or more cameras shifts a second time away from the distinct spatial configuration at a rate different from the rate at which the representation of the three-dimensional environment is changing in response to the computer system detecting that the user's viewpoint is changing pose a second time at a rate greater than a second threshold amount. Ceasing shifting the preview of the field of view of the one or more cameras in response to detecting that the user's viewpoint is changing pose less than the second threshold amount causes the computer system to display the preview of the field of view at a location on its display that is the center of the user's field of view, which performs the operation without requiring further user input.

[0227] In some embodiments, the change in pose (e.g., as described above in connection with FIG. 7F ) includes (e.g., corresponds to) a lateral movement along the plane of the user's viewpoint (e.g., a change in the user's viewpoint pose corresponds to the user's viewpoint moving left and / or right from the user's viewpoint). Shifting the preview of the field of view of one or more cameras based on the change in the user's viewpoint pose in the lateral direction provides visual feedback regarding the direction the computer system is moving, which provides improved visual feedback.

[0228] In some embodiments, the change in pose (e.g., as described above in connection with FIG. 7F ) includes (e.g., corresponds to) longitudinal movement along the plane of the user's viewpoint (e.g., a change in the user's viewpoint pose corresponds to the user's viewpoint moving upward and / or downward). In some embodiments, while the user is in a first pose, the user's head is located at a first distance from the ground, and while the user is in a second pose, the user's head is located at a second distance from the ground that is greater / less than the first distance). Shifting the preview of the field of view of one or more cameras based on the change in the user's viewpoint pose in the longitudinal direction provides visual feedback regarding the direction the computer system is moving, which provides improved visual feedback.

[0229] In some embodiments, the change in pose (e.g., as described above in connection with FIG. 7F ) includes (e.g., corresponds to) forward and backward movement perpendicular to the plane of the user's viewpoint (e.g., detecting a change in the pose of the user's viewpoint corresponds to the user's viewpoint moving forward and / or backward). In some embodiments, while the user's viewpoint is in a first pose, the computer system moves backward within the physical environment, which causes the computer system to cease displaying previews of the field of view of the one or more cameras. Shifting the previews of the field of view of the one or more cameras forward and backward based on the change in the pose of the user's viewpoint provides visual feedback regarding the direction the computer system is moving, which provides improved visual feedback.

[0230] In some embodiments, the preview of the field of view of one or more cameras (e.g., 708) does not overlay a second portion of the three-dimensional environment that is visible from the user's perspective (e.g., the portion of 704 that includes the second individual 709c2 in FIG. 7F). In some embodiments, the three-dimensional physical environment is the real-world environment in which the user is currently located. In some embodiments, the camera preview covers at least a portion of the view of the physical world, while at least a portion of the view of the physical world is displayed near or adjacent to the edge of the camera preview. Displaying a preview of the field of view of one or more cameras while at least a portion of the three-dimensional environment is visible provides the user with the ability to better compose and capture desired media while also maintaining awareness of the physical environment, which improves media capture operations and reduces the risk of failing to capture transient events that may be missed if the capture operation is inefficient or difficult to use. Improving media capture operations improves system usability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).

[0231] In some embodiments, the representation of the three-dimensional environment included in the media capture preview (e.g., 708) changes based on a change in the user's viewpoint pose (e.g., as described above in connection with FIG. 7G), and while detecting a change in the user's viewpoint pose from a first pose to a second pose (e.g., as described above in connection with FIG. 7F), the computer system performs (e.g., automatically (e.g., without user input) (e.g., the first visual stabilization corresponds to digital image stabilization)) first visual stabilization on the representation of the three-dimensional environment (e.g., to average out (e.g., attenuate) involuntary head movements (e.g., head moments performed by the user that do not correspond to a change in the user's viewpoint pose from the first pose to the second pose)) included in the preview of one or more cameras' fields of view (e.g., as described above in connection with FIG. 7F). In some embodiments, the first visual stabilization does not occur based on a change in the user's viewpoint pose. The first visual stabilization is an optical image stabilization technique that adjusts one or more glass elements in one or more cameras communicating with a computer system based on changes in the user's gaze. In some embodiments, the first visual stabilization includes a digital stabilization technique that includes zooming and / or cropping a representation of the three-dimensional environment included in the media capture preview based on changes in the user's gaze pose. Applying the first visual stabilization to the representation of the three-dimensional environment included in the preview of the field of view of one or more cameras causes the computer system to automatically perform display operations, without requiring user input, that increase the clarity with which the user views the representation of the three-dimensional environment included in the field of view of one or more cameras, thereby reducing the number of inputs required to perform the operations. Applying the first visual stabilization to the representation of the three-dimensional environment improves the media capture operation and reduces the risk of failing to capture a transient event that could be missed if the capture operation is inefficient or difficult to use.Improving media capture operations improves the usability of the system and makes the user-system interface more efficient (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device).

[0232] In some embodiments, performing the first stabilization includes applying a first amount of visual stabilization to a representation of the three-dimensional environment included in the preview of the field of view of one or more cameras (e.g., 708), and while detecting a change in the pose of the user's viewpoint from the first pose to a second pose (e.g., as described above in connection with FIGS. 7E-7G ), the computer system performs (e.g., via the computer system) (e.g., automatically (e.g., without user input)) second visual stabilization (e.g., the second visual stabilization corresponds to digital visual stabilization) on a second portion of the representation of the three-dimensional environment (e.g., the representation of the three-dimensional environment not included in the preview of the field of view of one or more cameras) (e.g., not performing the second visual stabilization on the preview of the field of view of one or more cameras). In some embodiments, the second visual stabilization applies a second amount of visual stabilization to the second portion of the representation of the three-dimensional environment that is less than the first amount of visual stabilization (e.g., as described above in connection with FIG. 7F ) (e.g., the representation of the three-dimensional environment included in the preview of the field of view of one or more cameras is sharper (e.g., clearer) than the representation of the three-dimensional environment). In some embodiments, the second visual stabilization is the same process (e.g., method) as the first visual stabilization. In some embodiments, the second visual stabilization is performed while the first visual stabilization is being performed. Applying the second visual stabilization to the second portion of the representation of the three-dimensional environment while detecting a change in the user's viewpoint pose causes the computer system to automatically perform display operations that increase the clarity with which the user views the second portion of the representation of the three-dimensional environment without requiring user input, which reduces the number of inputs required to perform the operations. Applying less stabilization to the representation of the three-dimensional environment provides more accurate visual feedback regarding the current physical environment, which improves visual feedback.

[0233] In some embodiments, aspects / operations of methods 800, 900, 1000, 1200, 1400, and 1500 may be interchanged, substituted, and / or added between these methods. For example, the media capture preview displayed in method 800 is optionally shifted using the methods described in method 900. For the sake of brevity, those details will not be repeated here.

[0234] 10 is a flow diagram of an exemplary method 1000 for displaying previously captured media, according to some embodiments. In some embodiments, method 1000 is performed on a computer system (e.g., computer system 101 and / or computer system 700 of FIG. 1 ) (e.g., a smartphone, a tablet, and / or a head-mounted device) in communication with a display generation component (e.g., display generation component 120 of FIGS. 1 , 3, and 4 and / or FIG. 702) (e.g., a display controller, a touch-sensitive display system, a display (e.g., integrated and / or connected), a 3D display, a see-through display, a projector, a head-up display, and / or a head-mounted display). In some embodiments, the computer system is in communication with one or more eye-tracking sensors (e.g., an optical camera and / or an IR camera configured to track the direction of gaze of a user of the computer system). In some embodiments, method 1000 is governed by instructions stored on a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 202 (e.g., control 110 of FIG. 1 ) of computer system 101. Some operations of method 1000 are, optionally, combined, and / or the order of some operations is, optionally, changed.

[0235] While displaying the extended reality environment user interface, the computer system detects (1002) a request to display captured media (e.g., as described above in connection with FIG. 7K) including immersive content (e.g., 730) that provides a first set of visual cues (e.g., visual indication of depth, visual indication of perspective, visual response to shifts in the user's viewpoint orientation) that the user is at least partially surrounded by content (e.g., immersive or semi-immersive visual media, three-dimensional media, three-dimensional stereo media, and / or spatial media) (e.g., video content that can be presented from multiple perspectives in response to detected changes in the orientation of the user and / or the computer system) when viewed from one or more distinct ranges of viewpoints. In some embodiments, the media content is immersive or semi-immersive visual media. In some embodiments, the immersive or semi-immersive visual media is visual media that includes content for multiple perspectives captured from the same first point (e.g., location) in the physical environment at a given time. In some embodiments, playing (e.g., playing back) visual media from an immersive (e.g., first-person) perspective includes playing the media content from a perspective that corresponds to a first point in the physical environment at a given time, and can provide multiple different perspectives (e.g., fields of view), all from the first point in the physical environment, in response to user input. In some embodiments, playing visual media from a non-immersive (e.g., third-person) perspective includes playing the media content from a perspective other than the first point in the physical environment (e.g., a shift in perspective (e.g., in response to user input)). In some embodiments, during playback in a non-immersive perspective, multiple perspectives corresponding to the first point in the physical environment are not displayed in response to user input.In some embodiments, as part of detecting a request to play the captured media, the computer system detects a user's attention (e.g., the user's gaze and / or the user's viewpoint) to one or more locations on the computer system (e.g., one or more locations corresponding to one or more virtual objects), detects one or more inputs to one or more hardware input mechanisms positioned on and / or coupled to the computer system, and / or detects one or more voice commands directed to playing the captured media. In some embodiments, the media was captured using one or more techniques such as those described above in connection with FIGS. 7C-7D . In some embodiments, the captured media is not non-immersive (e.g., is not immersive visual media and / or is not semi-immersive visual media). In some embodiments, non-immersive video media is video content that cannot be presented from multiple perspectives in response to detected changes in the orientation of the user and / or the computer system. In some embodiments, non-immersive video media may be presented in only one perspective (e.g., first-person perspective or third-person perspective), regardless of whether the computer system detects a change in the orientation of the computer system and / or a user of the computer system.

[0236] In response to detecting a request to display the captured media, the computer system displays (1004) the captured media as a three-dimensional representation (e.g., 740) (e.g., a non-static representation) of the captured media displayed at a location within the three-dimensional environment (e.g., location 740a within 701) selected by the computer system such that the user's first viewpoint (e.g., 712) is outside the individual range of one or more viewpoints (e.g., a non-immersive perspective) (e.g., a third-person perspective) (displaying the captured media being played). In some embodiments, the three-dimensional representation of the captured media replaces one or more portions of a representation of the environment (e.g., the virtual environment and / or the physical environment) that was previously displayed (e.g., displayed before the request to play the captured media was received) (e.g., using one or more techniques such as those described above in connection with FIG. 7C ). In some embodiments, the computer system displays the three-dimensional representation of the captured media over a shutter button virtual object and / or centered on a display generation component (e.g., as described above in connection with FIG. 7C ). In some embodiments, in response to detecting a request to display the captured media, the computer system displays multiple representations of previously captured media items simultaneously with a three-dimensional representation of the captured media. In some embodiments, in response to receiving a request to play the captured media, the computer system ceases displaying a representation (e.g., a static representation) of the captured media. Displaying the captured media in a location that meets a set of conditions specified for the three-dimensional representation (e.g., the location is outside the individual ranges of one or more viewpoints) automatically enables the computer system to present the three-dimensional representation from a non-immersive perspective without user input, which reduces the number of inputs required to perform an action. Doing so also provides improved visual feedback that the initially displayed viewpoint is not within the individual ranges of one or more viewpoints.

[0237] In some embodiments, the location selected by the computer system is a location within the physical environment (e.g., location 740a in 701), the three-dimensional representation of the captured media is an environment-locked virtual object, and while displaying the three-dimensional representation of the captured media, the computer system detects that the user's viewpoint has changed (e.g., as described above in connection with FIGS. 7K-7P) (e.g., the user moves left / right, up / down, and / or towards / away from the real-world environment) (e.g., the user looks up, down, right, or left) (e.g., the field of view of one or more cameras integrated into the computer system has changed) (e.g., detected via one or more cameras integrated into the computer system) (e.g., detected via an external device (e.g., a smartwatch, one or more cameras, and / or an external computer system) in communication (e.g., wireless communication) with the computer system). In some embodiments, in response to detecting that the user's viewpoint has changed, the computer system maintains the display of the three-dimensional representation of the captured media at the location within the three-dimensional environment selected by the computer system (e.g., as described above in connection with FIG. 7K). In some embodiments, the three-dimensional representation is environment-locked. In some embodiments, in response to detecting a change in the user's viewpoint, the size of the three-dimensional representation of the captured media changes (e.g., the three-dimensional representation of the captured media appears larger and / or smaller). In some embodiments, in response to detecting a change in the user's viewpoint, the three-dimensional representation of the captured media is no longer displayed. Maintaining the display of the three-dimensional representation of the captured media at a location within the three-dimensional environment selected by the computer system in response to detecting a change in the user's viewpoint causes the computer system to automatically perform display operations that allow the user to move and maintain a view of the three-dimensional representation of the captured media without requiring user input, which reduces the number of inputs required to perform operations.

[0238] In some embodiments, the user's first viewpoint corresponds to a first viewpoint location that is a first distance (e.g., 10 meters, 5 meters, or 3 meters) from a location selected by the computer system, and while displaying the three-dimensional representation of the media captured at the location selected by the computer system (e.g., location 740 in FIG. 7K), and while the user is at the first viewpoint location, the computer system detects a change in the pose of the user's viewpoint (e.g., repositioning the user's entire body, repositioning a part of the user's body (e.g., head, hands, arms, and / or legs), repositioning the user's head) relative to a second viewpoint of the user that corresponds to a second viewpoint location that is a second distance (e.g., 5 meters, 3 meters, or 1 meter) from the location selected by the computer system (e.g., the change in the user's positioning is detected by one or more cameras and / or external devices of the user, as described above), and the second distance is less than the first distance (e.g., as described above in FIG. 7L). In some embodiments, in response to detecting a change in the pose of the user's viewpoint relative to the user's second viewpoint, the computer system displays the three-dimensional representation having a second set of visual cues (e.g., 740 in FIG. 7L) that the user is at least partially surrounded by content. In some embodiments, the second set of visual cues includes at least a second visual cue that, when viewed from the user's first viewpoint (e.g., 740 in FIG. 7L), the user is at least partially surrounded by content that is not provided while displaying the three-dimensional representation. In some embodiments, as the user's viewpoint approaches a location selected by the computer system, the three-dimensional representation of the captured media becomes more immersive (albeit less immersive than when the user's viewpoint is within the range of one or more viewpoints providing the first set of visual cues).In some embodiments, the second set of visual cues does not include at least a first visual cue that the user is at least partially surrounded by content included in the first set of visual cues (e.g., a visual indication of depth, a visual indication of perspective, a visual response to a shift in the orientation of the user's viewpoint). In some embodiments, the second set of visual cues does not include at least a first visual cue that the user is at least partially surrounded by content included in the first set of visual cues (e.g., a visual indication of depth, a visual indication of perspective, a visual response to a shift in the orientation of the user's viewpoint). By displaying the three-dimensional representation with different sets of visual cues as the user approaches a location selected by the computer system, the computer system automatically, without requiring user input, performs display operations that enable the user to perceive the three-dimensional representation of the captured media differently based on the location of the user's viewpoint relative to the location selected by the computer, which performs the operations without requiring further user input. Displaying the three-dimensional representation with a different set of visual cues as the user approaches a location selected by the computer system provides the user with visual feedback as to what action the user needs to take, as the three-dimensional representation is displayed from an immersive perspective, providing improved visual feedback.

[0239] In some embodiments, while the user is at the second viewpoint location and while the three-dimensional representation of the captured media is displayed, the computer system detects a change in the user's viewpoint pose relative to the user's third viewpoint corresponding to the third viewpoint location (e.g., as described above in connection with FIG. 7M). In some embodiments, in response to detecting a change in the user's viewpoint pose relative to the user's third viewpoint, the computer system ceases displaying the three-dimensional representation of the captured media (e.g., as described above in connection with FIG. 7M) in response to determining that a first set of display criteria are met, the first criterion being met when a distance between the third viewpoint location and a location selected by the computer system is greater than a first threshold distance (e.g., 0.1 meters, 0.5 meters, 1 meter, 2 meters, 3 meters) (e.g., the user passes the location selected by the computer system). In some embodiments, the change in pose from the user's second viewpoint to the user's third viewpoint is in the same direction (e.g., along the same plane) (along the same path) as the change in pose from the user's first viewpoint to the user's second viewpoint (e.g., the directional component of the vector for the change in pose from the user's first viewpoint to the user's second viewpoint is the same as the directional component of the vector for the change in pose from the user's second viewpoint to the user's third viewpoint). In some embodiments, in response to detecting a change in pose of the user's viewpoint relative to the user's third viewpoint, the computer system maintains display of the three-dimensional representation of the captured media following a determination that the first set of display criteria is not met. In some embodiments, the computer system ceases displaying the three-dimensional representation of the captured media when the distance from the third viewpoint location to the location selected by the computer system is greater than a first threshold distance.In some embodiments, the first set of display criteria includes criteria that are met when the location selected by the computer system is not visible from the user's third viewpoint (e.g., the location selected by the computer system is no longer in front of the user). Ceasing to display the three-dimensional representation in accordance with a determination that the distance between the third viewpoint location and the location selected by the computer system is greater than a first threshold distance provides visual feedback regarding the location (e.g., the position of the computer system relative to the position selected by the computer system), which provides improved visual feedback.

[0240] In some embodiments, while the user is at the second viewpoint location and while the three-dimensional representation of the captured media is displayed, the computer system detects a change in the user's viewpoint pose relative to the user's fourth viewpoint corresponding to the fourth viewpoint location (e.g., as described above in connection with FIG. 7P). In some embodiments, in response to detecting a change in the user's viewpoint pose relative to the user's fourth viewpoint, and in accordance with determining that a second set of display criteria are met, the second set of display criteria including a second criterion that is met when the distance between the fourth viewpoint location and a location selected by the computer system is less than a second threshold distance (e.g., 0.1 meters, 0.5 meters, 1 meter, 2 meters, or 3 meters) (e.g., the user moves toward the location selected by the computer system), the computer system displays the three-dimensional representation of the captured media (e.g., as described above in connection with FIG. 7P) at a separate location selected by the computer system that is farther from the fourth viewpoint location than the location selected by the computer system (e.g., the three-dimensional representation of the captured media is moved away from the user's current location). In some embodiments, the separate location is a location within the physical environment that is visible from the user's fourth viewpoint. In some embodiments, the second set of display criteria includes criteria that are met when the location selected by the computer system is not visible from the user's fourth viewpoint (e.g., the location selected by the computer system is no longer in front of the user). In some embodiments, the distinct location is a location that is visible from the user's fourth viewpoint. In some embodiments, the computer system displays the three-dimensional representation of the media captured at the distinct location selected by the computer system when the distance from the fourth viewpoint location to the location selected by the computer system is greater than a second threshold distance.In some embodiments, pursuant to a determination that the viewpoint location has moved past the location of the three-dimensional representation (e.g., when the distance between the fourth viewpoint location and the location selected by the computer system is greater (or less) than a third threshold distance), the computer system displays the three-dimensional representation of the captured media at a separate location selected by the computer system that is farther from the fourth viewpoint location than the location selected by the computer system. In some embodiments, the change in pose from the user's second viewpoint to the user's fourth viewpoint is in the same direction as the change in pose from the user's first viewpoint to the user's second viewpoint (e.g., the directional component of the vector of the change in pose from the user's first viewpoint to the user's second viewpoint is the same as the directional component of the vector of the change in pose from the user's second viewpoint to the user's fourth viewpoint). Displaying the three-dimensional representation of the captured media at a separate location selected by the computer system that is farther from the fourth viewpoint location than the location selected by the computer system when certain specified conditions are met automatically changes the display of the three-dimensional representation so that the three-dimensional representation of the captured media is easily viewable by the user, which performs an action when a set of conditions is met without requiring further user input. Displaying a three-dimensional representation of media captured at a separate location that is farther from the fourth viewpoint location than the location selected by the computer system provides visual feedback to the user regarding the location of the computer system (e.g., the position of the computer system relative to the position selected by the computing system), which provides improved visual feedback.

[0241] In some embodiments, the user's first viewpoint corresponds to a fifth viewpoint location that is a fifth distance from the location selected by the computer system, and the computer displays a three-dimensional representation of the captured media at the selected location, and while the user is at the fifth viewpoint location, the computer system detects a change in the user's viewpoint pose (e.g., repositioning the user's entire body, repositioning a part of the user's body (e.g., head, hand, arm, leg), repositioning the user's head) to a sixth viewpoint of the user that corresponds to a sixth viewpoint location that is a sixth distance from the location selected by the computer system (e.g., the change in the user's positioning is detected by one or more cameras and / or external devices of the user as described above) (e.g., as described above in connection with FIG. 7P). In some embodiments, in response to detecting a change in the user's viewpoint pose to the user's sixth viewpoint, the computer system displays a three-dimensional representation having a third set of visual cues that the user is at least partially surrounded by content. In some embodiments, as the user's viewpoint moves further from the location selected by the computer system, the three-dimensional representation of the captured media becomes less immersive (although less immersive than when the user's viewpoint is within the range of one or more viewpoints providing the first set of visual cues). In some embodiments, the third set of visual cues does not include at least a third visual cue that the user is at least partially surrounded by content included in the first set of visual cues. In some embodiments, the third set of visual cues does not include at least a fourth visual cue that the user is at least partially surrounded by content provided while displaying the three-dimensional representation when viewed from the user's first viewpoint.Displaying the three-dimensional representation with a different set of visual cues as the user moves further from the location selected by the computer system causes the computer system to perform display operations that allow the user to perceive the three-dimensional representation of the captured media differently based on the user's viewpoint without displaying additional controls, which provides additional control options without cluttering the user interface.

[0242] In some embodiments, the three-dimensional representation of the captured media includes a plurality of virtual objects, including a first virtual object and a second virtual object, and the computer system detects a change in the pose of the user's viewpoint (e.g., a change in positioning of the user's entire body, a change in positioning of a first part of the user's body) (e.g., a change in the user's viewpoint) relative to a seventh viewpoint of the user (e.g., a lateral movement of the user, a left-right movement of the user, and / or a movement of the user along a horizontal plane). In some embodiments, in response to detecting a change in the user's viewpoint pose relative to the user's seventh viewpoint, the computer system, via the display generation component, displays (e.g., as described above in connection with FIG. 7K ) a first virtual object (e.g., the foreground of the content of 740) moving relative to a second virtual object (e.g., the background of the content of 740) based on the change in the user's viewpoint pose (e.g., displaying a parallax effect in which the first and second virtual objects shift differently as the user's viewpoint pose changes) (e.g., the first virtual object is displayed in the foreground of the three-dimensional representation of the captured media and moves at a first variable speed based on the change in the user's viewpoint pose, and the second virtual object is displayed in the background of the three-dimensional representation of the captured media and moves at a second variable speed based on the change in the user's viewpoint pose). In some embodiments, the first variable speed is greater than the second variable speed at any given time. Displaying the first virtual object moving relative to the second virtual object in response to detecting a change in the user's viewpoint pose provides visual feedback to the user regarding the depth data associated with the captured media, which provides improved visual feedback.

[0243] In some embodiments, the three-dimensional representation of the captured media (e.g., 740) is displayed as a first type of projection (e.g., a first type of shape), and while displaying the three-dimensional representation of the captured media as the first type of projection, the computer system detects a request (e.g., selection of 736) (e.g., selection of a virtual arrow object) (e.g., one or more gestures corresponding to selection of a virtual object) to display the three-dimensional representation as a second type of projection that is different from (e.g., displayed as a different shape, a different size, and / or a different location) than the first type of projection. In some embodiments, in response to detecting the request to display the three-dimensional representation as a second type of projection, the computer system displays the three-dimensional representation as a second type of projection (e.g., as described above in connection with FIG. 7K). In some embodiments, displaying the three-dimensional representation as a second type of projection includes displaying less of the captured media than when the three-dimensional representation is displayed as a first type of projection. In some embodiments, displaying the three-dimensional representation as a first type of projection (e.g., a spherical projection) includes distorting the three-dimensional representation of the captured media along edges of the projection, and displaying the three-dimensional representation as a second type of projection (e.g., a flattened projection) does not ...

Claims

1. 1. A method comprising:

1. A computer system in communication with a display generation component and one or more cameras, comprising: Detecting, via the display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on a representation of a physical environment, the representation of the physical environment changing as portions of the physical environment corresponding to the representation of the physical environment change and / or as a viewpoint of the user changes; in response to detecting the request to display the media capture user interface, displaying, via the display generation component, a media capture preview including a representation of a portion of a field of view of the one or more cameras, with content that updates as the portion of the physical environment within the portion of the field of view of the one or more cameras changes; the media capture preview indicates a boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed; the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before the request to display the media capture user interface is detected; The method, wherein the media capture preview is displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as the portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes.

2. The method of claim 1 , wherein the representation of the physical environment is a pass-through representation of the real-world environment of the computer system.

3. 3. The method of claim 1 or 2, wherein the representation of the physical environment is visible at a first scale and the representation of the portion of the field of view of the one or more cameras included in the media capture preview is displayed at a second scale, the first scale being larger than the second scale.

4. The method of claim 1 , wherein the representation of the physical environment includes first content that is not included in the representation of the field of view included in the media capture preview.

5. 5. The method of claim 1, wherein the representation of the physical environment includes a portion of the physical environment, and the representation of the field of view of the one or more cameras included in the media capture preview includes the portion of the physical environment.

6. The method of claim 1 , wherein the computer system is in communication with a physical input mechanism, and the request to display the media capture user interface includes activating the physical input mechanism.

7. detecting an input corresponding to a request to capture media while displaying the media capture preview; In response to detecting the input corresponding to the request to capture media, In response to determining that the input corresponding to the request to capture media is of a first type, initiate a process for capturing media content of a first type; In accordance with determining that the input corresponding to the request to capture media is of a second type, initiating a process for capturing media content of a second type; The method of claim 1 , further comprising:

8. 8. The method of claim 1, wherein the representation of the portion of the field of view of the one or more cameras included in the media capture preview has a first set of visual parallax characteristics, and the representation of the physical environment has a second set of visual parallax characteristics that is different from the first set of parallax characteristics.

9. 9. The method of claim 1, wherein the representation of the physical environment is an immersive perspective view and the representation of the portion of the field of view of the one or more cameras included in the media capture preview is a non-immersive perspective view.

10. Prior to detecting the request to display the media capture user interface, a third portion of the representation of the physical environment has a first visual appearance including visual features having a first magnitude, and the method further comprises:

10. The method of claim 1, further comprising, in response to detecting the request to display the media capture user interface, modifying the third portion of the representation of the physical environment to have a second visual appearance including the visual characteristic having a second magnitude, the second magnitude of the visual characteristic being different from the first magnitude of the visual characteristic.

11. displaying the media capture preview includes displaying one or more virtual objects having a spatial relationship with the representation of the media capture preview, the media capture preview and the one or more virtual objects being displayed at a first display location, the method comprising: detecting a change in the pose of the user's viewpoint while displaying the media capture preview and the one or more virtual objects at the first display location; In response to detecting the change in the pose of the gaze point of the user, displaying the media capture preview and the one or more virtual objects at a second display location different from the first location; The method of claim 1 , further comprising: maintaining the spatial relationship between the representation of the media capture preview and the one or more virtual objects.

12. The method of claim 11 , wherein the one or more virtual objects include a time lapse virtual object that provides an indication of an amount of time that has elapsed since the start of a process for capturing media.

13. The method of claim 11 or 12, wherein the one or more virtual objects include a shutter button virtual object that, when selected, causes the initiation of a process for capturing media.

14. The method of claim 11 , wherein the one or more virtual objects include a camera roll virtual object that, when selected, causes the display of previously captured media items.

15. Detecting a selection of the camera roll virtual object while displaying the media capture preview; and displaying, via the display generation component, a representation of the previously captured media item in response to detecting a selection of the camera roll virtual object, wherein displaying the representation of the previously captured media includes: a first erase virtual object that, when selected, causes the representation of the previously captured media item to cease display; a shared virtual object that, when selected, causes the initiation of a process for sharing the representation of the previously captured media item; a media library virtual object that, when selected, causes said display of a plurality of previously captured media items; and / or displaying a resize virtual object indicating a change in size of the representation of the previously captured media item based on the detection of one or more gestures; 15. The method of claim 14, comprising:

16. receiving a set of one or more inputs including an input corresponding to the camera roll virtual object; displaying, at a first location, a representation of a first previously captured media item of a plurality of previously captured media items in response to receiving the set of one or more inputs; receiving a request to navigate to a different previously captured media item of the plurality of previously captured media items while displaying the representation of the first previously captured media item at the first location; In response to receiving the request to navigate to a different previously captured media item from the plurality of previously captured media items, replacing the display of the representation of the first previously captured media item at the first location with a display of a representation of a second previously captured media item from the plurality of previously captured media items; 16. The method of claim 14 or 15, further comprising:

17. The method of claim 11 , wherein the one or more virtual objects include a second, erasing virtual object that, when selected, causes the display of the media capture preview to cease.

18. The one or more virtual objects include a repositioning virtual object, and the method further comprises: detecting a set of one or more inputs, including an input corresponding to the repositioned virtual object, while the media capture preview is displayed at a first location within the media capture user interface; 18. The method of claim 11, further comprising: in response to detecting the set of one or more inputs including the input corresponding to the repositioned virtual object, moving the media capture preview from the first location to a second location.

19. 19. The method of claim 1, further comprising simultaneously displaying, via the display generation component, a media library virtual object that, when selected, causes the display of multiple previously captured media items while the media capture preview is displayed.

20. the portion of the field of view of the one or more cameras included in the media capture preview has a first viewing angle range; the representation of the physical environment represents a second field of view of the one or more cameras having a second range of viewing angles; The first angular range is narrower than the second angular range.

20. The method of any one of claims 1 to 19.

21. 21. The method of claim 1, wherein the representation of the portion of the field of view of the one or more cameras included in the media capture preview includes first content that is within the field of view of a first camera of the one or more cameras and that is within the field of view of a second camera of the one or more cameras.

22. 22. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 1 to 21.

23. A computer system in communication with a display generation component and one or more cameras, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 1 to 21.

24. a computer system in communication with a display generation component and one or more cameras, A computer system comprising means for carrying out the method of any one of claims 1 to 21.

25. 22. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 1 to 21.

26. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs comprising: detect, via the display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on a representation of a physical environment, the representation of the physical environment changing as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as a viewpoint of the user changes; a non-transitory computer-readable storage medium comprising instructions for, in response to detecting the request to display the media capture user interface, displaying, via the display generation component, a media capture preview including a representation of a portion of a field of view of the one or more cameras, with content that updates as the portion of the physical environment within the portion of the field of view of the one or more cameras changes; the media capture preview indicates a boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed; the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before the request to display the media capture user interface is detected; the media capture preview is displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as the portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes.

27. a computer system in communication with a display generation component and one or more cameras, one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: detect, via the display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on a representation of a physical environment, the representation of the physical environment changing as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as a viewpoint of the user changes; responsive to detecting the request to display the media capture user interface, display, via the display generation component, a media capture preview including a representation of a portion of a field of view of the one or more cameras, with content that updates as the portion of the physical environment within the portion of the field of view of the one or more cameras changes; the media capture preview indicates a boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed; the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before the request to display the media capture user interface is detected; The media capture preview is displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as the portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes.

28. a computer system in communication with a display generation component and one or more cameras, means for detecting, via the display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on a representation of a physical environment, the representation of the physical environment changing as portions of the physical environment corresponding to the representation of the physical environment change and / or as the user's viewpoint changes; means for displaying, via the display generation component, a media capture preview including a representation of a portion of a field of view of the one or more cameras, with content that updates as the portion of the physical environment within the portion of the field of view of the one or more cameras changes, in response to detecting the request to display the media capture user interface; the media capture preview indicates a boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed; the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before the request to display the media capture user interface is detected; The media capture preview is displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as the portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes.

29. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs comprising: detect, via the display generation component, a request to display a media capture user interface while displaying a first user interface overlaid on a representation of a physical environment, the representation of the physical environment changing as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as a viewpoint of the user changes; responsive to detecting the request to display the media capture user interface, display, via the display generation component, a media capture preview including a representation of a portion of a field of view of the one or more cameras, with content that updates as the portion of the physical environment within the portion of the field of view of the one or more cameras changes; the media capture preview indicates a boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed; the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible before the request to display the media capture user interface is detected; the media capture preview is displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as the portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or as the user's viewpoint changes.

30. 1. A method comprising:

1. A computer system in communication with a display generation component and one or more cameras, comprising: displaying, via the display generation component, an extended reality user interface including a preview of the field of view of the one or more cameras overlaid on a first portion of a three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration for the user's viewpoint; detecting a change in the pose of the viewpoint of the user from the first pose to a second pose different from the first pose; in response to detecting the change in the pose of the user's viewpoint from the first pose to the second pose, shifting the previews of the field of view of the one or more cameras away from the individual spatial configuration relative to the user's viewpoint in a direction determined based on the change in the pose of the user's viewpoint from the first pose to the second pose, wherein the shifting of the previews of the field of view of the one or more cameras is performed at a first rate, and while the previews of the field of view of the one or more cameras are shifting based on the change in the pose of the user's viewpoint, the representation of the three-dimensional environment changes based on the change in the pose of the user's viewpoint at a second rate different from the first rate.

31. the computer system tracks changes in the pose of the user's viewpoint with a first amount of tracking lag; the first rate introduces a visual delay in updating the location of the preview of the field of view of the one or more cameras that is greater than a visual delay in updating the location of the preview of the field of view of the one or more cameras that is introduced based on the first amount of detection tracking lag.

31. The method of claim 30.

32. detecting that the user's viewpoint is changing pose by less than a first threshold amount while the preview of the field of view of the one or more cameras is shifting at the first rate; shifting the preview of the field of view of the one or more cameras toward the respective spatial configuration at a third rate faster than the first rate in response to detecting that the viewpoint of the user is changing pose by less than the threshold amount; 32. The method of claim 30 or 31, further comprising:

33. detecting that the user's viewpoint is changing pose by less than a second threshold amount while the preview of the field of view of the one or more cameras is shifting at the first rate; in response to detecting that the gaze point of the user is changing pose less than the second threshold amount; ceasing to shift the preview of the field of view of the one or more cameras away from the individual spatial configuration; displaying the preview of the field of view of the one or more cameras with the respective spatial configuration relative to the viewpoint of the user; 33. The method of any one of claims 30 to 32, further comprising:

34. 34. The method of any one of claims 30 to 33, wherein the change in pose comprises a lateral movement along a plane of the user's viewpoint.

35. 35. The method of any one of claims 30 to 34, wherein the change in pose comprises a longitudinal movement along the plane of the user's viewpoint.

36. 36. The method of any one of claims 30 to 35, wherein the change in pose comprises a forward or backward movement perpendicular to the plane of the user's viewpoint.

37. 37. The method of any one of claims 30 to 36, wherein the preview of the field of view of the one or more cameras does not overlay a second portion of the three-dimensional environment that is visible from the viewpoint of the user.

38. The representation of the three-dimensional environment included in the media capture preview changes based on the change in pose of the viewpoint of the user, the method comprising:

38. The method of any one of claims 30 to 37, further comprising performing a first visual stabilization on the representation of the three-dimensional environment included in the preview of the field of view of the one or more cameras while detecting the change in pose of the viewpoint of the user from the first pose to the second pose.

39. performing the first stabilization includes applying a first amount of visual stabilization to the representation of the three-dimensional environment included in the preview of the field of view of the one or more cameras, the method comprising:

39. The method of claim 38, further comprising: while detecting the change in pose of the viewpoint of the user from the first pose to the second pose, performing second visual stabilization on a second portion of the representation of the three-dimensional environment, the second visual stabilization applying a second amount of visual stabilization to the second portion of the representation of the three-dimensional environment, the second amount of visual stabilization being less than the first amount of visual stabilization.

40. 40. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 30 to 39.

41. A computer system in communication with a display generation component and one or more cameras, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 30 to 39.

42. a computer system in communication with a display generation component and one or more cameras, A computer system comprising means for carrying out the method of any one of claims 30 to 39.

43. 40. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 30 to 39.

44. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs comprising: displaying, via the display generation component, an extended reality user interface including a preview of the field of view of the one or more cameras overlaid on a first portion of the three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration for the user's viewpoint; detecting a change in the pose of the viewpoint of the user from the first pose to a second pose different from the first pose; a first pose to the first camera and a second pose to the second camera; a second pose to the first camera and a third pose to the second camera; a second pose to the first camera and a third pose to the second camera; a second pose to the first camera and a third pose to the second camera; a second pose to the first camera and a third pose to the second camera; a second pose to the first camera and a third pose to the second camera; a second pose to the first camera and a third pose to the second camera;

45. a computer system in communication with a display generation component and one or more cameras, one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: displaying, via the display generation component, an extended reality user interface including a preview of the field of view of the one or more cameras overlaid on a first portion of the three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration for the user's viewpoint; detecting a change in the pose of the viewpoint of the user from the first pose to a second pose different from the first pose; 1. A computer system comprising: instructions for, in response to detecting a change in the pose of the user's viewpoint from the first pose to the second pose, shifting the previews of the field of view of the one or more cameras away from the individual spatial configuration relative to the user's viewpoint in a direction determined based on the change in the pose of the user's viewpoint from the first pose to the second pose; wherein the shifting of the previews of the field of view of the one or more cameras is performed at a first rate, and while the previews of the field of view of the one or more cameras are shifting based on the change in the pose of the user's viewpoint, the representation of the three-dimensional environment changes based on the change in the pose of the user's viewpoint at a second rate different from the first rate.

46. a computer system in communication with a display generation component and one or more cameras, means for displaying, via the display generation component, an extended reality user interface including a preview of the field of view of the one or more cameras overlaid on a first portion of a three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a particular spatial configuration for the user's viewpoint; means for detecting a change in the pose of the viewpoint of the user from the first pose to a second pose different from the first pose; and means for shifting the preview of the field of view of the one or more cameras away from the individual spatial configuration relative to the user's viewpoint in a direction determined based on the change in the pose of the user's viewpoint from the first pose to the second pose in response to detecting the change in the pose of the user's viewpoint from the first pose to the second pose, wherein the shifting of the preview of the field of view of the one or more cameras is performed at a first rate, and while the preview of the field of view of the one or more cameras is shifting based on the change in the pose of the user's viewpoint, the representation of the three-dimensional environment changes based on the change in the pose of the user's viewpoint at a second rate different from the first rate.

47. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs comprising: displaying, via the display generation component, an extended reality user interface including a preview of the field of view of the one or more cameras overlaid on a first portion of the three-dimensional environment visible at the user's viewpoint while the user's viewpoint is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and displayed with a distinct spatial configuration for the user's viewpoint; detecting a change in the pose of the viewpoint of the user from the first pose to a second pose different from the first pose; responsive to detecting a change in the pose of the user's viewpoint from the first pose to the second pose, shift the previews of the field of view of the one or more cameras away from the individual spatial configuration relative to the user's viewpoint in a direction determined based on the change in the pose of the user's viewpoint from the first pose to the second pose, wherein the shifting of the previews of the field of view of the one or more cameras is performed at a first rate, and while the previews of the field of view of the one or more cameras are shifting based on the change in the pose of the user's viewpoint, the representation of the three-dimensional environment changes based on the change in the pose of the user's viewpoint at a second rate different from the first rate.

48. 1. A method comprising:

1. A computer system in communication with a display generation component, comprising: While displaying an extended reality environment user interface, detecting a request to display captured media including the immersive content that, when viewed from one or more distinct ranges of viewpoints, provides a first set of visual cues that the user is at least partially surrounded by the immersive content; In response to detecting the request to display the captured media, displaying the captured media as a three-dimensional representation of the captured media displayed at a location within a three-dimensional environment selected by the computer system such that a first viewpoint of the user is outside the individual ranges of one or more viewpoints.

49. the location selected by the computer system is a location within a physical environment, and the three-dimensional representation of the captured media is an environment-locked virtual object, the method comprising: detecting a change in the user's viewpoint while displaying the three-dimensional representation of the captured media; 49. The method of claim 48, further comprising: in response to detecting that the viewpoint of the user has changed, maintaining the display of the three-dimensional representation of the captured media at the location within the three-dimensional environment selected by the computer system.

50. The first viewpoint of the user corresponds to a first viewpoint location that is a first distance from the location selected by the computer system, and the method includes: detecting, while displaying the three-dimensional representation of the captured media at the location selected by the computer system and while the user is at the first viewpoint location, a change in pose of the user's viewpoint relative to a second viewpoint of the user corresponding to a second viewpoint location at a second distance from the location selected by the computer system, the second distance being less than the first distance; 50. The method of claim 48 or 49, further comprising: in response to detecting the change in the pose of the user's viewpoint relative to the user's second viewpoint, displaying the three-dimensional representation along with a second set of visual cues that the user is at least partially surrounded by the content, the second set of visual cues including at least a second visual cue that the user is at least partially surrounded by the content that was not provided while displaying the three-dimensional representation when viewed from the user's first viewpoint.

51. detecting a change in a pose of the user's viewpoint relative to a third viewpoint of the user corresponding to a third viewpoint location while the user is at the second viewpoint location and while the three-dimensional representation of the captured media is being displayed; In response to detecting the change in pose of the viewpoint of the user relative to the third viewpoint of the user, ceasing to display the three-dimensional representation of the captured media in accordance with a determination that a first set of display criteria is satisfied, the first set of display criteria including a first criterion that is satisfied when a distance between the third viewpoint location and the location selected by the computer system is greater than a first threshold distance; 51. The method of claim 50, further comprising:

52. detecting a change in a pose of the user's viewpoint relative to a fourth viewpoint of the user corresponding to a fourth viewpoint location while the user is at the second viewpoint location and while the three-dimensional representation of the captured media is being displayed; In response to detecting the change in pose of the user's viewpoint relative to the fourth viewpoint of the user, displaying the three-dimensional representation of the captured media at a separate location selected by the computer system that is farther from the fourth viewpoint location than the location selected by the computer system in accordance with a determination that a second set of display criteria is satisfied, the second set of display criteria including a second criterion that is satisfied when a distance between the fourth viewpoint location and the location selected by the computer system is less than a second threshold distance; 51. The method of claim 50, further comprising:

53. The first viewpoint of the user corresponds to a fifth viewpoint location that is a fifth distance from the location selected by the computer system, and the method further comprises: displaying the three-dimensional representation of the captured media at the location selected by the computer system, and detecting, while the user is at the fifth viewpoint location, a change in the pose of the user's viewpoint relative to a sixth viewpoint of the user corresponding to a sixth viewpoint location that is a sixth distance from the location selected by the computer system; 53. The method of any one of claims 48 to 52, further comprising: in response to detecting the change in the user's viewpoint pose to the sixth viewpoint of the user, displaying the three-dimensional representation along with a third set of visual cues that the user is at least partially surrounded by the content.

54. The three-dimensional representation of the captured media includes a plurality of virtual objects including a first virtual object and a second virtual object, and the method further comprises: detecting a change in pose of the user's viewpoint relative to a seventh viewpoint of the user; 54. The method of any one of claims 48 to 53, further comprising: in response to detecting the change in pose of the user's viewpoint relative to the seventh viewpoint of the user, displaying, via the display generation component, the first virtual object moving relative to the second virtual object based on the change in pose of the user.

55. The three-dimensional representation of the captured media is displayed as a first type of projection, and the method comprises: While displaying the three-dimensional representation of the captured media as a first projection, detecting a request to display the three-dimensional representation as a second type of projection different from the first type of projection; 55. The method of any one of claims 48 to 54, further comprising, in response to detecting the request to display the three-dimensional representation as a projection of a second type, displaying the three-dimensional representation as a projection of the second type.

56. The first type of projection and the second type of projection independently comprise: Spherical stereographic projection and 56. The method of claim 55, wherein the method is selected from the group consisting of: a stereographic projection of the flattened shape; and

57. The three-dimensional representation of the captured media is displayed at a first size, and the method includes: detecting a set of one or more gestures while displaying the three-dimensional representation at the first size; 57. The method of any one of claims 48 to 56, further comprising: in response to detecting the set of one or more gestures, enlarging the display of the three-dimensional representation of the captured media to a second size larger than the first size.

58. prior to detecting the set of one or more gestures, the extended reality environment user interface includes a first portion of a representation of a physical environment; enlarging the display of the three-dimensional representation of the captured media to the second size, which is larger than the first size, includes displaying the three-dimensional representation of the captured media at the second size in place of the first portion of the representation of the physical environment; 58. The method of claim 57, comprising:

59. detecting a second set of one or more gestures that include a movement component while the three-dimensional representation of the captured media is displayed; In response to detecting the second set of one or more gestures, ceasing to display the three-dimensional representation of the captured media; and displaying a second three-dimensional representation of the second captured media at the location selected by the computer system; and 59. The method of any one of claims 48 to 58, further comprising:

60. The method comprises: receiving a request to play the captured media while displaying the captured media as a three-dimensional representation of the captured media; In response to receiving the request to play the captured media, and in accordance with a determination that the captured media includes audio data, playing the captured media, wherein playing the captured media item outputs spatial audio corresponding to the audio data.

60. The method of any one of claims 48 to 59, comprising:

61. The computer system is in communication with an external device, and the captured media includes depth data, and the method includes: receiving a request to play the captured media on the external device while displaying the captured media as a three-dimensional representation of the captured media; 61. The method of claim 48, further comprising: in response to receiving the request to play the captured media on the external device and in accordance with a determination that the external device is not capable of displaying the depth data contained in the captured media, commencing playback of the captured media on the external device without stereoscopic depth effects.

62. 62. The method of claim 61, wherein playing the captured media on the external device includes outputting spatial audio corresponding to the captured media.

63. The computer system has a default interpupillary distance value setting, and the method includes: detecting a request to play the captured media; 63. The method of any one of claims 48 to 62, further comprising: in response to detecting the request to play the captured media and in accordance with a determination that the user's eyes have an interpupillary distance value different from the default interpupillary distance value setting, commencing playback of the captured media item with a first amount of visual shift based on the user's interpupillary distance.

64. 64. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for implementing the method of any one of claims 48 to 63.

65. 1. A computer system in communication with a display generation component, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 48 to 63.

66. a computer system in communication with a display generation component, 64. A computer system comprising means for carrying out the method of any one of claims 48 to 63.

67. 64. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for performing the method of any one of claims 48 to 63.

68. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising: While displaying an extended reality environment user interface, detect a request to display captured media including the immersive content that, when viewed from one or more distinct ranges of viewpoints, provides a first set of visual cues that the user is at least partially surrounded by the immersive content; a non-transitory computer-readable storage medium comprising instructions for, in response to detecting the request to display the captured media, displaying the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that a first viewpoint of the user is outside the individual ranges of one or more viewpoints.

69. a computer system in communication with a display generation component, one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: While displaying an extended reality environment user interface, detect a request to display captured media including the immersive content that, when viewed from one or more distinct ranges of viewpoints, provides a first set of visual cues that the user is at least partially surrounded by the immersive content; a computer system including instructions for, in response to detecting the request to display the captured media, displaying the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that a first viewpoint of the user is outside the individual ranges of one or more viewpoints.

70. a computer system in communication with a display generation component, means for detecting, while displaying an extended reality environment user interface, a request to display captured media including the immersive content that, when viewed from one or more distinct ranges of viewpoints, provides a first set of visual cues that the user is at least partially surrounded by the immersive content; means for, in response to detecting the request to display the captured media, displaying the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that a first viewpoint of the user is outside the individual ranges of one or more viewpoints.

71. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising: While displaying an extended reality environment user interface, detect a request to display captured media including the immersive content that, when viewed from one or more distinct ranges of viewpoints, provides a first set of visual cues that the user is at least partially surrounded by the immersive content; a computer program product, comprising instructions for, in response to detecting the request to display the captured media, displaying the captured media as a three-dimensional representation of the captured media displayed at a location selected by the computer system such that a first viewpoint of the user is outside the individual ranges of one or more viewpoints.

72. 1. A method comprising:

1. A computer system in communication with a display generation component and one or more cameras, comprising: via the display generation component, an extended reality camera user interface, the extended reality camera user interface comprising: Representation of the physical environment and and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance of the recording indicator from the first edge region increases.

73. Detecting an input corresponding to a request to capture media while the extended reality camera user interface is displayed; capturing media comprising a representation of at least a portion of the physical environment within the recording area in response to detecting the input corresponding to a request to capture media; 73. The method of claim 72, further comprising:

74. 74. The method of claim 73, wherein the captured media is still media.

75. 74. The method of claim 73, wherein the captured media is video media.

76. 76. The method of any one of claims 73 to 75, wherein the captured media includes a representation of the field of view of the one or more cameras that is different from a representation of the physical environment within the recording area.

77. 77. The method of any one of claims 72 to 76, wherein the visual parameter is a color gradient.

78. 78. The method of claim 77, wherein the recording indicator includes a second edge region that is farther from the center of the recording area than the first edge region, the second edge region being larger than the first edge region.

79. 79. A method according to any one of claims 72 to 78, wherein the one or more cameras in communication with the computer system have an optimum capture distance for capturing depth data, and the size and / or shape of the recording indicator assists in positioning the computer system at the optimum capture distance of the one or more cameras relative to one or more objects within the field of view of the one or more cameras.

80. 80. A method according to any one of claims 72 to 79, wherein the recording indicator is displayed at a fixed simulated depth within the representation of the physical environment.

81. the recording indicator includes one or more corners, and the method further comprises:

81. The method of any one of claims 72 to 80, further comprising, in accordance with a determination that a set of display criteria is satisfied while the recording indicator is displayed, displaying, via the display generation component, a secondary recording indicator at the one or more corners of the recording indicator.

82. 82. The method of claim 81, wherein the recording indicator is displayed in a first plane and the secondary recording indicator is displayed in the first plane.

83. displaying the secondary recording indicator displaying the secondary recording indicator with a first amount of visual emphasis relative to the recording indicator in accordance with a determination that the physical environment has a first amount of luminance; 83. The method of claim 81 or 82, comprising, in response to a determination that the physical environment has a second amount of luminance that is less than the first amount of luminance, displaying the secondary recording indicator with a second amount of visual emphasis relative to the recording indicator, the second amount of visual emphasis being greater than the first amount of visual emphasis.

84. 84. A method according to any one of claims 81 to 83, wherein the set of display criteria includes criteria that are met when one or more subject conditions are suitable for depth capture.

85. displaying the secondary recording indicator around a representation of a second subject while the secondary recording indicator is displayed with a first visual appearance indicating that the current state is not suitable for depth capture; changing the visual appearance of the secondary recording indicator from the first visual appearance to a second visual appearance indicating that the current state is appropriate for depth capture according to a determination that a set of one or more depth capture criteria is satisfied; 85. The method of any one of claims 81 to 84, further comprising:

86. detecting an input corresponding to a request to display the extended reality camera user interface prior to displaying the extended reality camera user interface; and displaying the extended reality camera user interface in response to detecting the input corresponding to the request to display the extended reality camera user interface, wherein displaying the extended reality camera user interface includes animating the recording indicator to fade in; 86. The method of any one of claims 72 to 85, comprising:

87. 87. The method of any one of claims 72 to 86, wherein the representation of the physical environment includes a third portion surrounding the recording indicator, and the third portion of the representation of the physical environment and the recording area have substantially the same amount of brightness modification due to a displayed user interface element.

88. 88. The method of claim 87, wherein the recording indicator has a third edge region, the first edge region being darker than the third edge region.

89. 88. The method of claim 87, wherein the recording indicator has a fourth edge region, the fourth edge region being darker than the first edge region.

90. displaying within the record indicator a capture virtual object that, when selected, causes initiation of a process for capturing media; 90. The method of any one of claims 72 to 89, further comprising:

91. displaying within the recording indicator a camera roll virtual object that, when selected, causes the initiation of a process for displaying previously captured media; 91. The method of any one of claims 72 to 90, further comprising:

92. 92. The method of any one of claims 72 to 91, wherein the recording indicator includes one or more rounded corners.

93. detecting a change in pose of the one or more cameras while the recording indicator is displayed surrounding a fourth portion of the representation of the physical environment; In response to detecting the change in the pose of the one or more cameras, displaying the record indicator around a fifth portion of the representation of the physical environment without displaying the record indicator around the fourth portion of the representation of the physical environment; and 93. The method of any one of claims 72 to 92, further comprising:

94. 94. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 72 to 93.

95. 1. A computer system configured to communicate with a display generation component and one or more cameras, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 72 to 93.

96. 1. A computer system configured to communicate with a display generation component and one or more cameras, comprising:

94. A computer system comprising means for carrying out the method of any one of claims 72 to 93.

97. 94. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs comprising instructions for performing the method of any one of claims 72 to 93.

98. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs comprising: via the display generation component, an extended reality camera user interface, the extended reality camera user interface comprising: Representation of the physical environment and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance of the recording indicator from the first edge region increases.

99. 1. A computer system configured to communicate with a display generation component and one or more cameras, comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: via the display generation component, an extended reality camera user interface, the extended reality camera user interface comprising: Representation of the physical environment and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance of the recording indicator from the first edge region increases.

100. 1. A computer system configured to communicate with a display generation component and one or more cameras, comprising: via the display generation component, an extended reality camera user interface, the extended reality camera user interface comprising: Representation of the physical environment and a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance of the recording indicator from the first edge region increases.

101. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs comprising: via the display generation component, an extended reality camera user interface, the extended reality camera user interface comprising: Representation of the physical environment and 11. A computer program product comprising: instructions for displaying an extended reality camera user interface comprising: a recording indicator indicating a recording area within a field of view of the one or more cameras, the recording indicator including at least a first edge region having a visual parameter that decreases through a plurality of different values ​​of the visual parameter in a visible portion of the recording indicator, the value of the parameter gradually decreasing as a distance of the recording indicator from the first edge region increases.

102. 1. A method comprising:

1. A computer system in communication with a display generating component, one or more input devices, and one or more cameras, comprising: detecting a request to display a camera user interface via the one or more input devices; and in response to detecting the request to display the camera user interface, displaying the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras, wherein displaying the camera user interface comprises: displaying the camera user interface along with a tutorial within the camera user interface in accordance with a determination that a set of one or more criteria is satisfied, the tutorial providing information on how to capture media using the computer system while the camera user interface is being displayed; and displaying the camera user interface without displaying the tutorial in accordance with a determination that the set of one or more criteria is not met.

103. 103. The method of claim 102, wherein the set of one or more criteria includes criteria that are met when the camera user interface is first displayed.

104. 104. The method of claim 102 or 103, wherein displaying the tutorial includes displaying instructions for capturing a first media item using the one or more cameras.

105. 105. The method of any one of claims 102 to 104, wherein the tutorial comprises a video.

106. 106. The method of any one of claims 102 to 105, wherein displaying the camera user interface includes displaying a viewfinder virtual object, and wherein displaying the tutorial includes displaying the tutorial overlaid on at least a portion of the viewfinder virtual object.

107. 107. The method of any one of claims 102 to 106, wherein the computer system communicates with a hardware input mechanism that, when activated, causes initiation of a media capture process, the tutorial includes a representation of the hardware input, and displaying the tutorial includes displaying a representation of the input corresponding to activation of the hardware input mechanism.

108. 108. The method of claim 107, wherein the hardware input mechanism is not visible to the user while the user is operating the computer system.

109. Detecting a first activation of the hardware input mechanism, the first activation of the hardware input mechanism being a first type of input; capturing a second media item using the one or more cameras in response to detecting the first activation of the hardware input mechanism; 109. The method of claim 107 or 108, further comprising:

110. detecting a second activation of the hardware input mechanism, the second activation of the hardware input mechanism corresponding to a second type of input including maintaining the input for a predetermined period of time; capturing a third media item in response to detecting the second activation of the hardware input mechanism; and 110. The method of any one of claims 107 to 109, further comprising:

111. detecting a third activation of the hardware input mechanism while a camera user interface is displayed with the tutorial; ceasing the display of the tutorial in response to detecting the third activation of the hardware input mechanism; and 111. The method of any one of claims 107 to 110, further comprising:

112. pursuant to a determination that a set of criteria is met, the camera user interface includes a camera shutter virtual object that, when selected, initiates a process for capturing a media item; In accordance with a determination that the set of criteria is not met, the camera user interface does not include the camera shutter virtual object for initiating a process of capturing a media item.

112. The method of any one of claims 102 to 111.

113. 113. The method of claim 112, wherein the set of criteria includes criteria that are met when a setting of the computer system is enabled.

114. 114. The method of any one of claims 102 to 113, wherein the camera user interface includes a proximity virtual object that, when selected, causes the camera user interface to cease display.

115. 115. The method of any one of claims 102 to 114, wherein the camera user interface is displayed within an extended reality environment and a first portion of the extended reality environment is displayed within the reticle virtual object.

116. Detecting a request to capture a fourth media item; capturing the fourth media item in response to detecting the request to capture the fourth media item, wherein the fourth media item is a stereoscopic media item; and 116. The method of any one of claims 102 to 115, further comprising:

117. 117. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 102 to 116.

118. 1. A computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 102 to 116.

119. 1. A computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, comprising:

117. A computer system comprising means for carrying out the method of any one of claims 102 to 116.

120. 117. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs comprising instructions for performing the method of any one of claims 102 to 116.

121. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more input devices, and one or more cameras, the one or more programs comprising: detecting a request to display a camera user interface via the one or more input devices; 10. A non-transitory computer-readable storage medium comprising instructions for, in response to detecting the request to display the camera user interface, displaying the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras, wherein displaying the camera user interface comprises: displaying the camera user interface along with a tutorial within the camera user interface in accordance with a determination that a set of one or more criteria is satisfied, the tutorial providing information on how to capture media using the computer system while the camera user interface is being displayed; and displaying the camera user interface without displaying the tutorial in accordance with a determination that the set of one or more criteria is not met.

122. 1. A computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: detecting a request to display a camera user interface via the one or more input devices; responsive to detecting the request to display the camera user interface, displaying the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras, the displaying the camera user interface comprising: displaying the camera user interface along with a tutorial within the camera user interface in accordance with a determination that a set of one or more criteria is satisfied, the tutorial providing information on how to capture media using the computer system while the camera user interface is being displayed; and displaying the camera user interface without displaying the tutorial according to a determination that the set of one or more criteria is not met.

123. 1. A computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, comprising: means for detecting, via said one or more input devices, a request to display a camera user interface; and means for displaying the camera user interface in response to detecting the request to display the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras. displaying the camera user interface along with a tutorial within the camera user interface in accordance with a determination that a set of one or more criteria is satisfied, the tutorial providing information on how to capture media using the computer system while the camera user interface is being displayed; and displaying the camera user interface without displaying the tutorial according to a determination that the set of one or more criteria is not met.

124. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs comprising: detecting a request to display a camera user interface via the one or more input devices; responsive to detecting the request to display the camera user interface, displaying the camera user interface, the camera user interface including a reticle virtual object indicating a capture area of ​​the one or more cameras, wherein displaying the camera user interface comprises: displaying the camera user interface along with a tutorial within the camera user interface in accordance with a determination that a set of one or more criteria is satisfied, the tutorial providing information on how to capture media using the computer system while the camera user interface is being displayed; and displaying the camera user interface without displaying the tutorial according to a determination that the set of one or more criteria is not met.

125. 1. A method comprising: A computer system in communication with a display generating component, one or more cameras, and one or more input devices, comprising: displaying a user interface via the display generation component, the user interface comprising: a representation of a physical environment, wherein a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside the capture area of ​​the one or more cameras; a viewfinder, the viewfinder including a boundary; and Detecting a first request to capture media via the one or more input devices while displaying the user interface; In response to detecting the first request to capture media, using the one or more cameras to capture a first media item including at least the first portion of the representation of the physical environment; Modifying an appearance of the viewfinder, wherein modifying the appearance of the viewfinder comprises: modifying the appearance of a first portion of content within a threshold distance on a first side of the boundary of the viewfinder; and modifying the appearance of a second portion of content within the threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder. And, A method comprising:

126. 126. The method of claim 125, wherein the first media item is a stereoscopic media item.

127. Prior to detecting a first request to capture the media, the viewfinder is displayed in a first appearance, and changing the appearance of the viewfinder includes displaying the viewfinder in a second appearance different from the first appearance, the method comprising: changing the appearance of the viewfinder from the second appearance to the first appearance after the viewfinder has been displayed in the second appearance for a period of time; 127. The method of claim 125 or 126, further comprising:

128. 128. The method of any one of claims 125 to 127, wherein the appearance of the first portion of content and the appearance of the second portion of content are modified in the same way.

129. 129. The method of claim 128, wherein modifying the appearance of the viewfinder includes modifying the appearance of a third portion of content within the threshold distance on a third side of the boundary of the viewfinder, and wherein the appearance of the first portion of content, the second portion of content, and the third portion of content are modified in the same manner.

130. The boundary of the viewfinder is a reticle virtual object, and before the first request to capture media is detected, the reticle virtual object is displayed with a first appearance, and the method further comprises:

130. The method of any one of claims 125 to 129, further comprising, in response to detecting the first request to capture media, changing the appearance of the reticle virtual object from the first appearance to a second appearance different from the first appearance.

131. 131. The method of claim 130, wherein displaying the viewfinder includes displaying one or more elements within the viewfinder, and modifying the appearance of the reticle virtual object includes modifying the appearance of the reticle virtual object relative to the one or more elements displayed within the viewfinder.

132. Displaying the user interface includes displaying one or more corners, and modifying the appearance of the viewfinder includes: modifying the appearance of the one or more corners in a first manner; and modifying the appearance of at least the first side of the boundary of the viewfinder in a second manner different from the first manner.

133. 133. The method of any one of claims 125 to 132, wherein altering the appearance of the viewfinder comprises altering a first set of one or more optical properties of a first portion of the content within the viewfinder.

134. 134. The method of claim 133, wherein the first set of one or more optical characteristics includes contrast of the content in the viewfinder.

135. 135. The method of claim 133 or 134, wherein the first set of one or more optical properties includes brightness of the content in the viewfinder.

136. 136. A method according to any one of claims 133 to 135, wherein the first set of one or more optical properties comprises translucency of the content within the viewfinder.

137. 137. The method of any one of claims 133 to 136, wherein the first set of one or more optical properties includes a size of the content in the viewfinder.

138. 138. The method of any one of claims 125 to 137, wherein displaying the user interface comprises displaying a first set of virtual control objects within the boundary of the viewfinder.

139. the first set of virtual objects includes a first proximity virtual object, and the method further comprises: Detecting an input corresponding to a selection of the first proximity virtual object; 139. The method of claim 138, further comprising ceasing the display of the user interface in response to detecting the input corresponding to a selection of the first proximate virtual object.

140. The first set of virtual objects includes a first media review virtual object, and the method further comprises: Detecting an input corresponding to a selection of the first media review virtual object; 140. The method of claim 138 or 139, further comprising: displaying one or more representations of previously captured media items in response to detecting the input corresponding to selection of the first media review virtual object.

141. 141. The method of any one of claims 138 to 140, wherein the first set of virtual objects includes a record time virtual object that indicates an amount of time that has elapsed since the computer system initiated a first video capture operation.

142. 142. The method of any one of claims 125 to 141, wherein displaying the user interface includes displaying a second set of virtual control objects outside the boundary of the viewfinder.

143. the second set of virtual control objects includes a set of camera mode virtual objects respectively corresponding to respective operational modes of the one or more cameras, the set of camera mode virtual objects including a first camera mode virtual object and a second camera mode virtual object, and the method further comprises: Detecting an input corresponding to a selection of an individual camera mode virtual object within the set of camera mode virtual objects; in response to detecting the input corresponding to a selection of an individual camera mode virtual object within the set of camera mode virtual objects; configuring the one or more cameras to operate in a first mode in accordance with a determination that the input corresponds to a selection of the first camera mode virtual object; configuring the one or more cameras to operate in a second mode in accordance with a determination that the input corresponds to a selection of the second camera mode virtual object; 143. The method of claim 142, further comprising:

144. the second set of virtual control objects includes a second proximity virtual object, and the method further comprises: detecting an input corresponding to a selection of the second proximity virtual object; ceasing to display the user interface in response to detecting the input corresponding to selection of the second virtual object; and 144. The method of claim 142 or 143, further comprising:

145. displaying a representation of the first media item fading into the display of the user interface after the first media item is captured; 145. The method of any one of claims 125 to 144, further comprising:

146. Displaying the representation of the first media item fading into the display of the user interface includes: displaying the representation of the first media item transitioning from a first size to a second size, the first size being larger than the second size; displaying the representation of the first media item moving from a first location within the user interface to a second location within the user interface, the second location corresponding to a corner of the viewfinder; 146. The method of claim 145, comprising:

147. 147. The method of claim 145 or 146, wherein before displaying the representation of the first media item, the viewfinder includes a second reticle virtual object, and the display of the representation of the first media item replaces the display of the second reticle virtual object.

148. Displaying the user interface includes displaying a second media review virtual object at a third location within the user interface, the second media review virtual object being displayed at a third size, and displaying the representation of the first media item fading in the user interface includes: displaying the representation of the first media item transitioning from a fourth size to a fifth size, the fourth size being larger than the fifth size and the fourth size being larger than the third size; and displaying the representation of the first media item as moving from a fourth location within the user interface to the third location within the user interface.

149. 149. The method of any one of claims 125 to 148, wherein the first media item is a still photograph or a video.

150. detecting a second request to capture media after changing the appearance of the viewfinder; In response to detecting the second request to capture media, displaying a first type of feedback in accordance with a determination that the second request to capture media corresponds to a request to capture a still photograph; displaying a second type of feedback different from the first type of feedback in accordance with a determination that the second request to capture media corresponds to a request to capture video; and 150. The method of claim 149, further comprising:

151. the first request to capture media corresponds to a first type of input, the first media item is a still photograph, and the first type of input corresponds to a short press of a first hardware input mechanism; 151. The method of any one of claims 125 to 150.

152. detecting a third request to capture media after changing the appearance of the viewfinder; capturing a video media item in response to detecting the third request to capture media and in accordance with a determination that the third request to capture media corresponds to a second type of input, the second type of input corresponding to a long press of a second hardware input mechanism; 152. The method of any one of claims 125 to 151, further comprising:

153. the third request to capture media corresponds to the second type of input, and the method further comprises: In response to detecting the third request to capture media, displaying an indication that capturing the video media item corresponds to a second video capture operation, wherein displaying the indication includes: displaying the indication in a first appearance in accordance with a determination that a set of criteria is not met; displaying the indication in a second appearance in accordance with a determination that the set of criteria is met; and and ceasing detection of the second type of input while the indication is displayed; and In response to ceasing to detect the second type of input, ceasing performance of the second video capture operation in accordance with a determination that the set of criteria was not met prior to ceasing detection of the second type of input; continuing to perform the second video capture operation in accordance with a determination that the set of criteria has been met before ceasing to detect the second type of input; and 153. The method of claim 152, further comprising:

154. the indication indicating an amount of time that has elapsed since the computer system initiated the second video capture operation, and displaying the indication includes: displaying the indication in the first appearance; Detecting that the set of criteria is met while the indication is displayed in the first appearance; and and changing the appearance of the indication from the first appearance to the second appearance in response to detecting that the set of criteria is met.

155. detecting, while the computer system is capturing the video media item, an input corresponding to activation of a third hardware input mechanism; halting the capture of the video media item in response to detecting the input corresponding to activation of the third hardware input mechanism; 155. The method of any one of claims 152 to 154, further comprising:

156. 156. The method of any one of claims 125 to 155, wherein the first side of the boundary and the second side of the boundary are on opposite sides of the boundary.

157. 157. The method of any one of claims 125 to 156, wherein a first portion of the content is at the threshold distance from the first side of the boundary of the viewfinder and a second portion of the content is at the threshold distance from the second side of the boundary of the viewfinder.

158. 158. The method of any one of claims 125 to 157, wherein displaying the user interface includes displaying a third reticle virtual object, the third reticle virtual object indicating the capture area of ​​the one or more cameras.

159. 159. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more input devices, and one or more cameras, the one or more programs including instructions for performing the method of any one of claims 125 to 158.

160. 1. A computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, the computer system comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of claims 125 to 158.

161. 1. A computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, comprising:

159. A computer system comprising means for carrying out the method of any one of claims 125 to 158.

162. 159. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs comprising instructions for performing the method of any one of claims 125 to 158.

163. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component, one or more input devices, and one or more cameras, the one or more programs comprising: a user interface via the display generation component, a representation of a physical environment, wherein a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside the capture area of ​​the one or more cameras; displaying a user interface including a viewfinder, the viewfinder including a boundary; Detecting a first request to capture media via the one or more input devices while displaying the user interface; In response to detecting the first request to capture media, using the one or more cameras to capture a first media item comprising at least the first portion of the representation of the physical environment; 1. A non-transitory computer-readable storage medium comprising instructions for altering an appearance of the viewfinder, the altering the appearance of the viewfinder comprising: modifying the appearance of a first portion of content within a threshold distance on a first side of the boundary of the viewfinder; and modifying the appearance of a second portion of content within the threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.

164. 1. A computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, comprising: one or more processors; a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising: a user interface via the display generation component, a representation of a physical environment, wherein a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside the capture area of ​​the one or more cameras; displaying a user interface including a viewfinder, the viewfinder including a boundary; Detecting a first request to capture media via the one or more input devices while displaying the user interface; In response to detecting the first request to capture media, using the one or more cameras to capture a first media item comprising at least the first portion of the representation of the physical environment; 10. A computer system comprising instructions for modifying an appearance of the viewfinder, the modifying the appearance of the viewfinder comprising: modifying the appearance of a first portion of content within a threshold distance on a first side of the boundary of the viewfinder; and modifying the appearance of a second portion of content within the threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.

165. 1. A computer system configured to communicate with a display generation component, one or more input devices, and one or more cameras, comprising: a user interface via the display generation component, a representation of a physical environment, wherein a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside the capture area of ​​the one or more cameras; a viewfinder, the viewfinder including a boundary; and means for displaying a user interface, the viewfinder including a boundary; means for detecting a first request to capture media via the one or more input devices while displaying the user interface; In response to detecting the first request to capture media, using the one or more cameras to capture a first media item comprising at least the first portion of the representation of the physical environment; and means for modifying the appearance of the viewfinder, the modifying the appearance of the viewfinder comprising: modifying the appearance of a first portion of content within a threshold distance on a first side of the boundary of the viewfinder; and modifying the appearance of a second portion of content within the threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.

166. 1. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras, the one or more programs comprising: a user interface via the display generation component, a representation of a physical environment, wherein a first portion of the representation of the physical environment is within a capture area of ​​the one or more cameras and a second portion of the representation of the physical environment is outside the capture area of ​​the one or more cameras; displaying a user interface including a viewfinder, the viewfinder including a boundary; Detecting a first request to capture media via the one or more input devices while displaying the user interface; In response to detecting the first request to capture media, using the one or more cameras to capture a first media item comprising at least the first portion of the representation of the physical environment; 10. A computer program product comprising instructions for modifying an appearance of the viewfinder, the modifying the appearance of the viewfinder comprising: modifying the appearance of a first portion of content within a threshold distance on a first side of the boundary of the viewfinder; and modifying the appearance of a second portion of content within the threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.