Device, method, and graphical user interface for capturing and displaying media
The computer system addresses inefficiencies in media capture and display by reducing user inputs and conserving power through intuitive interfaces and adaptive media previews, enhancing user interaction in virtual/extended reality environments.
Patent Information
- Application Number
- JP2024532566
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2022-11-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-11-23
AI Technical Summary
Existing methods for capturing and displaying media in virtual/extended reality environments are cumbersome, inefficient, and impose a significant cognitive burden on users, often requiring complex inputs and wasting energy, particularly in battery-operated devices.
A computer system with improved methods and interfaces that reduce the number, degree, and type of user inputs by providing intuitive media capture and display through a display generation component and cameras, including a media capture preview that indicates the boundary of media to be captured, and adjusting user interfaces based on user pose changes.
Enhances user interaction efficiency, reduces processing power, and conserves battery life by simplifying media capture and display processes in virtual/extended reality environments.
Smart Images

Figure 0007713597000001 
Figure 0007713597000002 
Figure 0007713597000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. patent application Ser. No. 17 / 992,789, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING AND DISPLAYING MEDIA", filed Nov. 22, 2022; U.S. provisional patent application Ser. No. 63 / 409,690, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING AND DISPLAYING MEDIA", filed Sep. 23, 2022; U.S. provisional patent application Ser. No. 63 / 338,864, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING AND DISPLAYING MEDIA", filed May 5, 2022; and U.S. provisional patent application Ser. No. 63 / 285,897, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING AND DISPLAYING MEDIA", filed Dec. 3, 2021. The entire content of each of these applications is hereby incorporated by reference into this specification.
[0002] Technical Field The present disclosure generally relates to computer systems that provide computer - generated experiences, including, but not limited to, electronic devices that provide virtual reality and mixed - reality experiences via a display, a display - generating component, optionally one or more input devices, and one or more cameras that communicate with each other.
Background Art
[0003] The development of computer systems for capturing and / or displaying media in various environments, such as extended reality environments, has increased significantly in recent years. Exemplary extended reality environments include at least some virtual elements that replace or enhance the physical world. Input devices such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays for computer systems and other electronic computing devices are used to interact with virtual / extended reality environments. Exemplary virtual elements include virtual objects such as digital images, videos, text, icons, and control elements such as buttons and other graphics.
Summary of the Invention
[0004] Some methods and interfaces for capturing and / or displaying media in various environments are cumbersome, inefficient, and limited. For example, systems that provide inadequate visual feedback for capturing media, systems that require a series of complex inputs to execute a media capture process, and systems where the display of media is complex, boring, and error-prone impose a significant cognitive burden on the user and detract from the experience in virtual / extended reality environments. In addition, those methods take more time than necessary, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-operated devices.
[0005] Accordingly, there is a need for a computer system having an improved method and interface for providing a computer-generated experience to a user that makes it more efficient and intuitive for the user to capture and / or display media in various environments. Such a method and interface optionally complements or replaces conventional methods for capturing and / or displaying media in various environments. Such a method and interface reduces the number, degree, and / or type of inputs from the user by assisting the user in understanding the connection between the provided input and the device response to that input, thereby creating a more efficient human-machine interface.
[0006] The above-mentioned drawbacks and other problems associated with the user interface of a computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also known as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, the computer system has one or more output devices in addition to a display generation component, and the output devices include one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, programs, or sets of instructions stored in the memory for performing a plurality of functions. In some embodiments, the user interacts with the GUI through a stylus and / or finger contact and gestures on a touch-sensitive surface, the movement of the user's eyes and hands in space relative to the GUI (and / or the computer system) when captured by a camera and other motion sensors, and voice input when captured by one or more audio input devices.In some embodiments, the functions executed through the interaction optionally include image editing, drawing, presenting, word processing, spreadsheet creation, game play, making a phone call, video conferencing, sending an email, instant messaging, training support, digital photography, digital video shooting, web browsing, playing digital music, taking notes, and / or playing digital video. The executable instructions for performing those functions are optionally included in a primary computer-readable storage medium and / or a non-transitory computer-readable storage medium, or other computer program products configured to be executed by one or more processors.
[0007] There is a need for electronic devices having improved methods and interfaces for capturing and / or displaying media in various environments. Such methods and interfaces can complement or replace conventional methods for capturing and / or displaying media. Such methods and interfaces reduce the number, degree, and / or type of user inputs and create a more efficient human-machine interface. In the case of battery-operated computing devices, such methods and interfaces conserve power, increase the time between battery charges, and reduce the amount of processing power.
[0008] According to some embodiments, a method is described that is executed in a computer system communicating with a display generation component and one or more cameras. The method includes, while displaying, via the display generation component, a first user interface overlaid on a representation of a physical environment, where the representation of the physical environment changes as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as the user's viewpoint changes, detecting a request to display a media capture user interface; and in response to detecting the request to display a media capture user interface, displaying, via the display generation component, a media capture preview that includes a representation of a portion of the field of view of the one or more cameras, with content that is updated as a portion of the physical environment within the portion of the field of view of the one or more cameras changes, wherein the media capture preview indicates the boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed, the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment is visible prior to detecting the request to display the media capture user interface, the media capture preview is displayed instead of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or as the user's viewpoint changes.
[0009] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs, via the display generation component, display a first user interface overlaid on a representation of a physical environment, where the representation of the physical environment changes as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as the user's viewpoint changes. While displaying the first user interface, detect a request to display a media capture user interface, and in response to detecting the request to display the media capture user interface, via the display generation component, display a media capture preview including a representation of a portion of the field of view of the one or more cameras, along with content that is updated as a portion of the physical environment within the portion of the field of view of the one or more cameras changes. The media capture preview indicates the boundary of media that is captured in response to detecting a media capture input while the media capture user interface is being displayed. The media capture preview is displayed while a first portion of the representation of the physical environment is visible, where the first portion of the representation of the physical environment is visible before the request to display the media capture user interface is detected. The media capture preview is displayed instead of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or as the user's viewpoint changes.
[0010] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs, via the display generation component, display a first user interface overlaid on a representation of a physical environment, where the representation of the physical environment changes as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as the user's viewpoint changes, and detect a request to display a media capture user interface while the first user interface is being displayed. In response to detecting the request to display the media capture user interface, display, via the display generation component, a media capture preview including a representation of a portion of the field of view of the one or more cameras, along with content that is updated as a portion of the physical environment within the portion of the field of view of the one or more cameras changes. The media capture preview indicates the boundary of media to be captured in response to detecting a media capture input while the media capture user interface is being displayed. The media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible prior to detecting the request to display the media capture user interface. The media capture preview is displayed instead of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or as the user's viewpoint changes.
[0011] According to some embodiments, a computer system that communicates with a display generation component and one or more cameras is described. The computer system includes one or more processors and a memory that stores one or more programs configured to be executed by the one or more processors. The one or more programs, via the display generation component, display a first user interface overlaid on a representation of a physical environment, where the representation of the physical environment changes as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as the user's viewpoint changes. While displaying the first user interface, detect a request to display a media capture user interface, and in response to detecting the request to display a media capture user interface, via the display generation component, display a media capture preview that includes a representation of a portion of the field of view of the one or more cameras, along with content that is updated as a portion of the physical environment within the portion of the field of view of the one or more cameras changes. The media capture preview indicates the boundary of media that is captured in response to detecting a media capture input while the media capture user interface is displayed. The media capture preview is displayed while a first portion of the representation of the physical environment is visible, where the first portion of the representation of the physical environment is visible before the request to display the media capture user interface is detected. The media capture preview is displayed instead of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or as the user's viewpoint changes.
[0012] According to some embodiments, a computer system communicating with a display generation component and one or more cameras is described. The computer system, via the display generation component, while displaying a first user interface overlaid on a representation of a physical environment, where the representation of the physical environment changes as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as the user's viewpoint changes, means for detecting a request to display a media capture user interface; and in response to detecting a request to display a media capture user interface, via the display generation component, means for displaying a media capture preview including a representation of a portion of the field of view of the one or more cameras, with content that is updated as a portion of the physical environment within the portion of the field of view of the one or more cameras changes, wherein the media capture preview indicates the boundary of media to be captured in response to detecting a media capture input while the media capture user interface is being displayed, the media capture preview is displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible prior to detecting a request to display the media capture user interface, the media capture preview is displayed in place of a second portion of the representation of the physical environment, and the first portion of the representation of the physical environment is updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or as the user's viewpoint changes.
[0013] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs, via the display generation component, detect a request to display a media capture user interface while displaying a first user interface overlaid on a representation of a physical environment, the representation of the physical environment changing as a portion of the physical environment corresponding to the representation of the physical environment changes and / or as the user's viewpoint changes, and in response to detecting the request to display the media capture user interface, via the display generation component, display a media capture preview including a representation of a portion of the field of view of the one or more cameras, with content that is updated as a portion of the physical environment within the portion of the field of view of the one or more cameras changes, the media capture preview indicating the boundary of media to be captured in response to detecting a media capture input while the media capture user interface is displayed, the media capture preview being displayed while a first portion of the representation of the physical environment is visible, the first portion of the representation of the physical environment being visible prior to detecting the request to display the media capture user interface, and the media capture preview being displayed in place of a second portion of the representation of the physical environment, the first portion of the representation of the physical environment being updated as a portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or as the user's viewpoint changes.
[0014] According to some embodiments, a method is described that is executed in a computer system communicating with a display generation component and one or more cameras. The method includes, while the user's perspective is in a first pose, via the display generation component, a preview of the field of view of one or more cameras overlaid on a first portion of a three-dimensional environment visible from the user's perspective, the preview including a representation of the first portion of the three-dimensional environment and being displayed with an individual spatial configuration relative to the user's perspective, displaying an extended reality user interface including the preview; detecting a change in the pose of the user's perspective from the first pose to a second pose different from the first pose; and in response to detecting a change in the pose of the user's perspective from the first pose to the second pose, shifting the preview of the field of view of one or more cameras away from the individual spatial configuration relative to the user's perspective in a direction determined based on the change in the pose of the user's perspective from the first pose to the second pose, wherein the shift of the preview of the field of view of one or more cameras is performed at a first speed, and while the preview of the field of view of one or more cameras is shifting based on the change in the pose of the user's perspective, the representation of the three-dimensional environment changes based on the change in the pose of the user's perspective at a second speed different from the first speed.
[0015] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more cameras. The one or more programs display an extended reality user interface including a preview of the field of view of one or more cameras overlaid on a first portion of a three-dimensional environment visible from the user's perspective via the display generation component while the user's perspective is in a first pose. The preview includes a representation of the first portion of the three-dimensional environment and is displayed with an individual spatial configuration relative to the user's perspective. The one or more programs detect a change in the pose of the user's perspective from the first pose to a second pose different from the first pose, and in response to detecting the change in the pose of the user's perspective from the first pose to the second pose, shift the preview of the field of view of the one or more cameras away from the individual spatial configuration relative to the user's perspective in a direction determined based on the change in the pose of the user's perspective from the first pose to the second pose. The shift of the preview of the field of view of the one or more cameras is performed at a first speed, and while the preview of the field of view of the one or more cameras is shifting based on the change in the pose of the user's perspective, the representation of the three-dimensional environment changes based on the change in the pose of the user's perspective at a second speed different from the first speed.
[0016] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more cameras. The one or more programs, while the user's perspective is in a first pose, via the display generation component, are previews of the fields of view of one or more cameras overlaid on a first portion of a three-dimensional environment visible in the user's perspective, the preview including a representation of the first portion of the three-dimensional environment and being displayed with an individual spatial configuration relative to the user's perspective, display an extended reality user interface including the preview, detect a change in the pose of the user's perspective from the first pose to a second pose different from the first pose, and in a direction determined based on the change in the pose of the user's perspective from the first pose to the second pose, shift the preview of the fields of view of one or more cameras away from the individual spatial configuration relative to the user's perspective. The shift of the preview of the fields of view of one or more cameras is performed at a first speed, and while the preview of the fields of view of one or more cameras is shifting based on the change in the pose of the user's perspective, the representation of the three-dimensional environment changes based on the change in the pose of the user's perspective at a second speed different from the first speed.
[0017] According to some embodiments, a computer system communicating with a display generation component and one or more cameras is described. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs display an extended reality user interface including a preview of the fields of view of one or more cameras overlaid on a first portion of a three-dimensional environment visible from the user's perspective while the user's perspective is in a first pose, the preview including a representation of the first portion of the three-dimensional environment and being displayed with an individual spatial configuration relative to the user's perspective. The one or more programs detect a change in the pose of the user's perspective from the first pose to a second pose different from the first pose, and in response to detecting the change in the pose of the user's perspective from the first pose to the second pose, shift the preview of the fields of view of the one or more cameras away from the individual spatial configuration relative to the user's perspective in a direction determined based on the change in the pose of the user's perspective from the first pose to the second pose. The shift of the preview of the fields of view of the one or more cameras is performed at a first speed, and while the preview of the fields of view of the one or more cameras is shifting based on the change in the pose of the user's perspective, the representation of the three-dimensional environment changes based on the change in the pose of the user's perspective at a second speed different from the first speed.
[0018] According to some embodiments, a computer system communicating with a display generation component and one or more cameras is described. The computer system, while the user's perspective is in a first pose, via the display generation component, a preview of the field of view of one or more cameras overlaid on a first portion of a three-dimensional environment visible in the user's perspective, the preview including a representation of the first portion of the three-dimensional environment and being displayed with an individual spatial configuration relative to the user's perspective, means for displaying an extended reality user interface including the preview; means for detecting a change in the pose of the user's perspective from the first pose to a second pose different from the first pose; and means for shifting the preview of the field of view of one or more cameras away from the individual spatial configuration relative to the user's perspective in a direction determined based on the change in the pose of the user's perspective from the first pose to the second pose in response to detecting the change in the pose of the user's perspective from the first pose to the second pose, wherein the shift of the preview of the field of view of one or more cameras is performed at a first speed, and while the preview of the field of view of one or more cameras is shifting based on the change in the pose of the user's perspective, the representation of the three-dimensional environment changes based on the change in the pose of the user's perspective at a second speed different from the first speed.
[0019] According to some embodiments, a computer program product is described. The computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component. The one or more programs include a preview of the field of view of one or more cameras overlaid on a first portion of a three-dimensional environment visible from a user's perspective via the display generation component while the user's perspective is in a first pose. The preview includes a representation of the first portion of the three-dimensional environment and is displayed with an individual spatial configuration relative to the user's perspective. The one or more programs display an extended reality user interface including the preview, detect a change in the pose of the user's perspective from the first pose to a second pose different from the first pose, and, in response to detecting the change in the pose of the user's perspective from the first pose to the second pose, shift the preview of the field of view of the one or more cameras away from the individual spatial configuration relative to the user's perspective in a direction determined based on the change in the pose of the user's perspective from the first pose to the second pose. The shift of the preview of the field of view of the one or more cameras is performed at a first speed, and while the preview of the field of view of the one or more cameras is shifting based on the change in the pose of the user's perspective, the representation of the three-dimensional environment changes based on the change in the pose of the user's perspective at a second speed different from the first speed.
[0020] According to some embodiments, the method is executed in a computer system that communicates with a display generation component. The method, while displaying an extended reality environment user interface, detects a request to display captured media that includes immersive content providing a first set of visual cues that the user is at least partially surrounded by the immersive content when viewed from an individual range of one or more viewpoints, and in response to detecting the request to display the captured media, displays the captured media as a three-dimensional representation of the captured media at a location selected by the computer system such that a first viewpoint of the user is outside an individual range of the one or more viewpoints.
[0021] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs detecting, while displaying an extended reality environment user interface, a request to display captured media that includes immersive content providing a first set of visual cues that the user is at least partially surrounded by the immersive content when viewed from an individual range of one or more viewpoints, and in response to detecting the request to display the captured media, displaying the captured media as a three-dimensional representation of the captured media at a location selected by the computer system such that a first viewpoint of the user is outside an individual range of the one or more viewpoints, including instructions.
[0022] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, the one or more programs including instructions to detect a request to display captured media including immersive content that provides a first set of visual cues that the user is at least partially surrounded by the immersive content when viewed from an individual range of one or more viewpoints while an extended reality environment user interface is being displayed, and in response to detecting the request to display the captured media, to display the captured media as a three-dimensional representation of the captured media at a location selected by the computer system such that a first viewpoint of the user is outside an individual range of the one or more viewpoints.
[0023] According to some embodiments, a computer system communicates with a display generation component. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions to detect a request to display captured media including immersive content that provides a first set of visual cues that the user is at least partially surrounded by the immersive content when viewed from an individual range of one or more viewpoints while an extended reality environment user interface is being displayed, and in response to detecting the request to display the captured media, to display the captured media as a three-dimensional representation of the captured media at a location selected by the computer system such that a first viewpoint of the user is outside an individual range of the one or more viewpoints.
[0024] According to some embodiments, a computer system that communicates with a display generation component is described. The computer system, while displaying an extended reality environment user interface, provides a first set of visual cues that the user is at least partially surrounded by immersive content when viewed from the individual ranges of one or more viewpoints, and means for detecting a request to display captured media including immersive content, and in response to detecting a request to display captured media, means for displaying the captured media as a three-dimensional representation of the captured media to be displayed at a location selected by the computer system such that the user's first viewpoint is outside the individual ranges of the one or more viewpoints.
[0025] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, the one or more programs detecting a request to display captured media including immersive content that provides a first set of visual cues that the user is at least partially surrounded by immersive content when viewed from the individual ranges of one or more viewpoints while the extended reality environment user interface is being displayed, and in response to detecting a request to display captured media, displaying the captured media as a three-dimensional representation of the captured media to be displayed at a location selected by the computer system such that the user's first viewpoint is outside the individual ranges of the one or more viewpoints.
[0026] According to some embodiments, a method is described that is executed in a computer system communicating with a display generation component and one or more cameras. The method includes displaying, via the display generation component, an extended reality camera user interface, where the extended reality camera user interface includes a representation of a physical environment and a recording indicator indicating a recording area within the field of view of one or more cameras, and the recording indicator includes at least a first edge region having a visual parameter that decreases through a plurality of different values of the visual parameter in a visible portion of the recording indicator, and the value of the parameter gradually decreases as the distance from the first edge region of the recording indicator increases.
[0027] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more cameras, and the one or more programs include instructions to display, via the display generation component, an extended reality camera user interface, where the extended reality camera user interface includes a representation of a physical environment and a recording indicator indicating a recording area within the field of view of one or more cameras, and the recording indicator includes at least a first edge region having a visual parameter that decreases through a plurality of different values of the visual parameter in a visible portion of the recording indicator, and the value of the parameter gradually decreases as the distance from the first edge region of the recording indicator increases.
[0028] According to some embodiments, a non - transitory computer - readable storage medium is described. The non - transitory computer - readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras. The one or more programs include instructions to display, via the display generation component, an extended reality camera user interface, where the extended reality camera user interface includes a representation of a physical environment and a recording indicator indicating a recording area within the field of view of the one or more cameras. The recording indicator includes at least a first edge region having a visual parameter that decreases through a plurality of different values of the visual parameter in the visible portion of the recording indicator, and the value of the parameter gradually decreases as the distance from the first edge region of the recording indicator increases.
[0029] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system includes one or more processors and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions to display, via the display generation component, an extended reality camera user interface, where the extended reality camera user interface includes a representation of a physical environment and a recording indicator indicating a recording area within the field of view of the one or more cameras. The recording indicator includes at least a first edge region having a visual parameter that decreases through a plurality of different values of the visual parameter in the visible portion of the recording indicator, and the value of the parameter gradually decreases as the distance from the first edge region of the recording indicator increases.
[0030] According to some embodiments, a computer system configured to communicate with a display generation component and one or more cameras is described. The computer system, via the display generation component, provides an extended reality camera user interface, where the extended reality camera user interface includes a representation of the physical environment and a recording indicator indicating a recording area within the field of view of one or more cameras. The recording indicator includes at least a first edge region having a visual parameter that decreases through a plurality of different values of the visual parameter in the visible portion of the recording indicator, and the value of the parameter gradually decreases as the distance from the first edge region of the recording indicator increases. The computer system comprises means for displaying the extended reality camera user interface including the recording indicator.
[0031] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more cameras. The one or more programs, via the display generation component, provide an extended reality camera user interface, where the extended reality camera user interface includes a representation of the physical environment and a recording indicator indicating a recording area within the field of view of one or more cameras. The recording indicator includes at least a first edge region having a visual parameter that decreases through a plurality of different values of the visual parameter in the visible portion of the recording indicator, and the value of the parameter gradually decreases as the distance from the first edge region of the recording indicator increases. The one or more programs include instructions for displaying the extended reality camera user interface including the recording indicator.
[0032] According to some embodiments, a method is described that is executed in a computer system communicating with a display generation component, one or more input devices, and one or more cameras. The method includes detecting, via one or more input devices, a request to display a camera user interface; and in response to detecting the request to display a camera user interface, displaying a camera user interface that includes a reticle virtual object indicating a capture area of one or more cameras. Displaying the camera user interface includes, in accordance with a determination that one or more sets of criteria are met, displaying, within the camera user interface, a tutorial that provides information on how to capture media using the computer system while the camera user interface is being displayed, together with the camera user interface; and in accordance with a determination that one or more sets of criteria are not met, displaying the camera user interface without displaying the tutorial.
[0033] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more input devices, and one or more cameras. The one or more programs detect a request to display a camera user interface via the one or more input devices, and in response to detecting the request to display the camera user interface, include instructions to display a camera user interface, the camera user interface including a reticle virtual object indicating a capture area of the one or more cameras. Displaying the camera user interface includes, in accordance with a determination that a set of one or more criteria is satisfied, displaying, within the camera user interface, a tutorial, the tutorial providing information on how to capture media using the computer system while the camera user interface is being displayed, and, in accordance with a determination that the set of one or more criteria is not satisfied, displaying the camera user interface without displaying the tutorial.
[0034] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The one or more programs detect a request to display a camera user interface via the one or more input devices, and in response to detecting the request to display the camera user interface, display a camera user interface that includes a reticle virtual object indicating a capture area of the one or more cameras. Displaying the camera user interface includes, in accordance with a determination that a set of one or more criteria is satisfied, displaying, within the camera user interface, a tutorial that provides information on how to capture media using the computer system while the camera user interface is being displayed, and, in accordance with a determination that the set of one or more criteria is not satisfied, displaying the camera user interface without displaying the tutorial.
[0035] According to some embodiments, a computer system is described. The computer system includes one or more processors, and the computer system is configured to communicate with a display generation component, one or more input devices, and one or more cameras. The memory stores one or more programs configured to be executed by the one or more processors. The one or more programs detect a request to display a camera user interface via one or more input devices, and in response to detecting the request to display the camera user interface, include instructions to display a camera user interface that includes a reticle virtual object indicating a capture region of one or more cameras. Displaying the camera user interface includes, in accordance with a determination that a set of one or more criteria is satisfied, displaying, within the camera user interface, a tutorial that provides information on how to capture media using the computer system while the camera user interface is being displayed, and, in accordance with a determination that the set of one or more criteria is not satisfied, displaying the camera user interface without displaying the tutorial.
[0036] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component, one or more input devices, and one or more cameras. The computer system includes means for detecting a request to display a camera user interface via one or more input devices, and means for displaying a camera user interface in response to detecting the request to display the camera user interface. The camera user interface includes a reticle virtual object indicating a capture area of one or more cameras. Displaying the camera user interface includes displaying the camera user interface together with a tutorial within the camera user interface according to a determination that a set of one or more criteria is met. The tutorial provides information on how to capture media using the computer system while the camera user interface is being displayed. Displaying the camera user interface also includes displaying the camera user interface without displaying the tutorial according to a determination that the set of one or more criteria is not met.
[0037] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The one or more programs detect a request to display a camera user interface via the one or more input devices, and in response to detecting the request to display the camera user interface, display a camera user interface that includes a reticle virtual object indicating a capture area of the one or more cameras. Displaying the camera user interface includes, in accordance with a determination that one or more sets of criteria are met, displaying, within the camera user interface, a tutorial that provides information on how to capture media using the computer system while the camera user interface is being displayed, and, in accordance with a determination that the one or more sets of criteria are not met, displaying the camera user interface without displaying the tutorial.
[0038] According to some embodiments, a method is described that is executed in a computer system communicating with a display generation component, one or more input devices, and one or more cameras. The method includes, via the display generation component, displaying a user interface that includes a representation of a physical environment, where a first portion of the representation of the physical environment is inside the capture region of the one or more cameras and a second portion of the representation of the physical environment is outside the capture region of the one or more cameras, and a viewfinder that includes a boundary; detecting, via the one or more input devices while the user interface is being displayed, a first request to capture media; in response to detecting the first request to capture media, capturing a first media item that includes at least the first portion of the representation of the physical environment using the one or more cameras and changing the appearance of the viewfinder; and changing the appearance of the viewfinder includes changing the appearance of a first portion of the content within a threshold distance on a first side of the boundary of the viewfinder and changing the appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.
[0039] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, one or more input devices, and one or more cameras. The one or more programs, via the display generation component, display a user interface that includes a representation of a physical environment, where a first portion of the representation of the physical environment is inside the capture region of the one or more cameras and a second portion of the representation of the physical environment is outside the capture region of the one or more cameras, and a viewfinder that includes a boundary. While displaying the user interface, detect a first request to capture media via the one or more input devices, and in response to detecting the first request to capture media, capture a first media item that includes at least the first portion of the representation of the physical environment using the one or more cameras and change the appearance of the viewfinder. Changing the appearance of the viewfinder includes changing the appearance of a first portion of the content within a threshold distance on a first side of the boundary of the viewfinder and changing the appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.
[0040] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generation component, one or more input devices, and one or more cameras. The one or more programs, via the display generation component, display a user interface that includes a representation of a physical environment, where a first portion of the representation of the physical environment is inside the capture region of the one or more cameras and a second portion of the representation of the physical environment is outside the capture region of the one or more cameras, and a viewfinder that includes a boundary. While displaying the user interface, detect a first request to capture media via the one or more input devices, and in response to detecting the first request to capture media, capture a first media item that includes at least the first portion of the representation of the physical environment using the one or more cameras, and change the appearance of the viewfinder. Changing the appearance of the viewfinder includes changing the appearance of a first portion of the content within a threshold distance on a first side of the boundary of the viewfinder and changing the appearance of a second portion of the content within a threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.
[0041] According to some embodiments, a computer system is described. The computer system includes a display generation component, one or more input devices, and one or more processors configured to communicate with one or more cameras, and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs, via the display generation component, display a user interface that includes a representation of a physical environment, where a first part of the representation of the physical environment is inside the capture region of the one or more cameras and a second part of the representation of the physical environment is outside the capture region of the one or more cameras, and a viewfinder that includes a boundary. While displaying the user interface, detect a first request to capture media via the one or more input devices, and in response to detecting the first request to capture media, capture a first media item that includes at least the first part of the representation of the physical environment using the one or more cameras, and change the appearance of the viewfinder. Changing the appearance of the viewfinder includes changing the appearance of a first part of the content within a threshold distance of a first side of the boundary of the viewfinder and changing the appearance of a second part of the content within a threshold distance of a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.
[0042] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component, one or more input devices, and one or more cameras. The computer system, via the display generation component, provides a user interface that is a representation of a physical environment. A first portion of the representation of the physical environment is inside the capture region of the one or more cameras, and a second portion of the representation of the physical environment is outside the capture region of the one or more cameras. The user interface includes a viewfinder that includes a boundary. The computer system has means for displaying the user interface, means for detecting, via the one or more input devices while the user interface is being displayed, a first request to capture media, means for capturing, in response to detecting the first request to capture media, a first media item that includes at least the first portion of the representation of the physical environment using the one or more cameras, and means for changing the appearance of the viewfinder. Changing the appearance of the viewfinder includes changing the appearance of a first portion of the content that is within a threshold distance on a first side of the boundary of the viewfinder and changing the appearance of a second portion of the content that is within a threshold distance on a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.
[0043] According to some embodiments, a computer program product is described. The computer program product comprises one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component, one or more input devices, and one or more cameras. The one or more programs, via the display generation component, display a user interface that includes a representation of a physical environment, wherein a first portion of the representation of the physical environment is inside a capture region of the one or more cameras and a second portion of the representation of the physical environment is outside the capture region of the one or more cameras, and a viewfinder that includes a boundary. While displaying the user interface, detect a first request to capture media via the one or more input devices, and in response to detecting the first request to capture media, capture a first media item that includes at least the first portion of the representation of the physical environment using the one or more cameras, and change an appearance of the viewfinder, wherein changing the appearance of the viewfinder includes changing an appearance of a first portion of the content within a threshold distance of a first side of the boundary of the viewfinder and changing an appearance of a second portion of the content within a threshold distance of a second side of the boundary of the viewfinder that is different from the first side of the boundary of the viewfinder.
[0044] Note that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will be apparent to those skilled in the art, particularly in view of the drawings, specification, and claims. Further, note that the language used herein has been selected solely for readability and for the purpose of explanation and not for the purpose of defining or limiting the subject matter of the invention.
Brief Description of the Drawings
[0045] To better understand the various embodiments described, the following "Modes for Carrying Out the Invention" should be referred to in conjunction with the following drawings, and like reference numerals refer to corresponding parts throughout the following figures.
[0046]
Figure 1
[0047]
Figure 2
[0048]
Figure 3
[0049]
Figure 4
[0050]
Figure 5
[0051]
Figure 6
[0052]
Figure 7A
Figure 7B
Figure 7C
Figure 7D
Figure 7E
Figure 7F
Figure 7G
Figure 7H
Figure 7I
Figure 7J
Figure 7K
Figure 7L
Figure 7M
Figure 7N
Figure 7O
Figure 7P
Figure 7Q
[0053]
Figure 8
[0054]
Figure 9
[0055]
Figure 10
[0056]
Figure 11A
Figure 11B
Figure 11C
Figure 11D
[0057]
Figure 12
[0058]
Figure 13A
Figure 13B
Figure 13C
Figure 13D
Figure 13E1
Figure 13E2
Figure 13E3
Figure 13E4
Figure 13E5
Figure 13F
Figure 13G
Figure 13H
Figure 13I
Figure 13J
[0059]
Figure 14
[0060]
Figure 15A
Figure 15B
Mode for Carrying Out the Invention
[0061] The present disclosure relates to a user interface for providing a user with an extended reality (XR) experience according to some embodiments.
[0062] Figures 1-6 provide an illustration of an exemplary computer system for providing XR experiences to a user. Figures 7A-7Q show exemplary techniques for capturing and / or displaying media in various environments according to some embodiments. Figure 8 is a flowchart of a method for capturing and viewing media according to various embodiments. Figure 9 is a flowchart of a method for displaying a preview of media according to various embodiments. Figure 10 is a flowchart of a method for displaying previously captured media according to various embodiments. The user interfaces of Figures 7A-7Q illustrate the processes of Figures 8, 9, and 10. Figures 11A-11D show exemplary techniques for displaying a representation of a physical environment using a recording indicator according to some embodiments. Figure 12 is a flowchart of a method for displaying a representation of a physical environment using a recording indicator according to some embodiments. The user interfaces of Figures 11A-11D illustrate the process of Figure 12. Figures 13A-13J show exemplary techniques for displaying a camera user interface according to some embodiments. Figure 14 is a flowchart of a method for displaying information related to captured media according to some embodiments. Figures 15A-15B are flowcharts of methods for changing the appearance of a viewfinder according to some embodiments. The user interfaces of Figures 13A-13J illustrate the processes of Figures 14, 15A, and 15B.
[0063] The processes described below enhance the operability of a device and make the user device interface more efficient through various techniques, including, for example, assisting the user in providing appropriate input and reducing user errors when operating / interacting with the device, that provide improved visual feedback to the user, reduce the number of inputs required to perform an operation, provide additional control options without confusing the user interface with additional displayed controls, perform an operation without requiring further user input when a set of conditions is met, improve privacy and / or security, provide a more diverse, detailed, and / or realistic user experience while saving memory space, and / or include additional technologies. These techniques also reduce power usage and improve the battery life of the device by enabling the user to use the device more quickly and efficiently. Saving battery power, and thus weight, improves the ergonomics of the device. These techniques also enable real-time communication, enable the use of fewer and / or less accurate sensors, result in more compact, lighter, and less expensive devices, and enable the device to be used under various lighting conditions. These techniques reduce energy usage, thereby reducing the heat emitted by the device, which is particularly important for wearable devices that can become uncomfortable for the user to wear if the device generates too much heat, as is adequately within the operating parameters for the device components.
[0064] Furthermore, in the methods described herein that are conditional on one or more conditions being met by one or more steps, it should be understood that the described methods can be repeated in multiple iterations such that all of the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires performing a first step when a condition is met and a second step when the condition is not met, one of ordinary skill in the art will understand that the steps recited in the claims will be repeated in a particular order until the condition is met and then ceases to be met. Thus, a method described in terms of one or more steps that are dependent on one or more conditions being met can be rewritten as a method that is repeated until each condition described in the method is met. However, this is not required in claims for a system or computer-readable medium that includes instructions for performing conditional operations based on the fulfillment of the corresponding one or more conditions, and thus can determine whether an eventuality is met without explicitly repeating the steps of the method until all of the conditions for which the steps of the method are conditional are met. One of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as necessary to ensure that all of the conditional steps are executed.
[0065] In some embodiments, as shown in FIG. 1, the XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., speakers 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., within a head-mounted device or a handheld device).
[0066] When describing the XR experience, various related but distinct environments that a user perceives and / or with which a user can interact (e.g., using inputs detected by the computer system 101 to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101 that generates the XR experience) are individually referred to using various terms. The following is a subset of these terms.
[0067] Physical Environment: The physical environment refers to the physical world that people can perceive and / or interact with without the aid of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through senses such as vision, touch, hearing, taste, and smell.
[0068] Extended Reality: In contrast, an extended reality (XR) environment refers to an environment that is wholly or partially simulated in which people perceive and / or interact through an electronic system. In XR, a subset of a person's body movements or their representation is tracked, and in response, one or more characteristics of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics. For example, an XR system can detect the rotation of a person's head and, in response, adjust the graphic content and sound field presented to the person in a manner similar to how such views and sounds would change in the physical environment. Depending on the situation (e.g., for accessibility reasons), the adjustment of the characteristic(s) of the virtual object(s) in the XR environment may be performed in response to a representation of a body movement (e.g., a voice command). A person may use any one of these senses including vision, hearing, touch, taste, and smell to perceive and / or interact with the XR object. For example, a person can perceive and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of a point audio source within a 3D space. In another example, an audio object can enable audio transparency that selectively incorporates ambient sound from the physical environment, with or without including computer-generated audio. In some XR environments, a person may only perceive and / or interact with the audio object.
[0069] Examples of XR include virtual reality and mixed reality.
[0070] Virtual Reality: A virtual reality (VR) environment refers to an imitation environment designed to be based entirely on computer-generated sensory inputs for one or more senses. A VR environment includes a plurality of virtual objects that a person can perceive and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can perceive and / or interact with the virtual objects in a VR environment through a simulation of the person's presence within the computer-generated environment and / or through a simulation of a subset of the person's physical movements within the computer-generated environment.
[0071] Mixed Reality: In contrast to a VR environment designed to be based entirely on computer-generated sensory inputs, a mixed reality (MR) environment refers to an imitation environment designed to incorporate sensory inputs or representations thereof from the physical environment in addition to including computer-generated sensory inputs (e.g., virtual objects). On the virtual continuum, an MR environment can be anywhere between, but not including, a complete physical environment at one end and a virtual reality environment at the other end. In some MR environments, the computer-generated sensory inputs can be responsive to changes in the sensory inputs from the physical environment. Also, some electronic systems for presenting an MR environment may track the location and / or orientation with respect to the physical environment in order to enable virtual objects to interact with real objects (i.e., physical articles or representations thereof from the physical environment). For example, the system may take movement into account so that a virtual tree appears stationary with respect to the physical ground.
[0072] Examples of mixed reality include extended reality and augmented virtuality.
[0073] Extended Reality: An extended reality (AR) environment refers to an emulated environment in which one or more virtual objects are superimposed on a physical environment or its representation. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on the transparent or translucent display, whereby a person can use the system to perceive virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture an image or video of the physical environment, which is a representation of the physical environment. The system synthesizes the image or video with virtual objects and presents the composite on the opaque display. A person uses this system to indirectly view the physical environment via the image or video of the physical environment and to perceive virtual objects superimposed on the physical environment. As used herein, the video of the physical environment shown on the opaque display is referred to as a "pass-through video," meaning that the system uses one or more image sensors to capture an image of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system may have a projection system that projects virtual objects, for example, as holograms, into the physical environment or onto a physical surface, whereby a person can use the system to perceive virtual objects superimposed on the physical environment. An extended reality environment also refers to an emulated environment in which the representation of the physical environment is transformed by computer-generated sensory information. For example, when providing a pass-through video, the system may transform one or more sensor images to map to a selected perspective (e.g., viewpoint) different from the perspective captured by the imaging sensor. As another example, the representation of the physical environment may be transformed by graphically modifying a portion thereof (e.g., magnifying), thereby creating a modified version that represents the original captured image but is non-photorealistic.As a further example, the representation of the physical environment may be transformed by graphically removing or obscuring a part thereof.
[0074] Augmented Virtuality: An augmented virtuality (AV) environment refers to an imitative environment in which a virtual environment or a computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but people with faces are realistically reproduced from images of physical people. As another example, virtual objects may adopt the shape or color of physical articles imaged by one or more imaging sensors. As a further example, virtual objects can adopt shadows that coincide with the position of the sun in the physical environment.
[0075] Viewpoint-locked virtual object: A virtual object is viewpoint-locked when the computer system displays the virtual object at the same location and / or position within the user's viewpoint even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked in the forward direction of the user's head (e.g., the user's viewpoint is at least a portion of the user's field of view when the user is looking straight ahead). Thus, the user's viewpoint remains fixed even when the user's line of sight moves without the user moving their head. In embodiments where the computer system has a display generation component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the extended reality view presented to the user on the computer system's display generation component. For example, a viewpoint-locked virtual object displayed at the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed at the upper left corner of the user's viewpoint even if the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the location and / or position at which the viewpoint-locked virtual object is displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head such that the virtual object is also referred to as a "head-locked virtual object".
[0076] Environment-Locked Virtual Object: A virtual object is environment-locked (or "world-locked") when a computer system displays the virtual object at a location and / or position within the user's view that is based on (e.g., selected with reference to and / or fixed to) a location and / or object within a three-dimensional environment (e.g., a physical environment or a virtual environment). When the user's view shifts, the location and / or object within the environment relative to the user's view changes, and as a result, the environment-locked virtual object is displayed at a different location and / or position within the user's view. For example, an environment-locked virtual object locked to a tree directly in front of the user is displayed at the center of the user's view. If the user's view shifts to the right (e.g., the user's head is turned to the right) such that the tree becomes leftward in the user's view (e.g., the position of the tree in the user's view shifts), the environment-locked virtual object locked to the tree is displayed leftward in the user's view. In other words, the location and / or position at which the environment-locked virtual object is displayed in the user's view depends on the location and / or position and / or orientation of the location and / or object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system fixed to a fixed location and / or object in the physical environment) to determine the position at which to display the environment-locked virtual object in the user's view. The environment-locked virtual object can be locked to a stationary part of the environment (e.g., the floor, a wall, a table, or other stationary object), or to a movable part of the environment (e.g., a vehicle, an animal, a person, or a representation of a part of the user's body such as the user's hand, wrist, arm, foot, etc. that moves independently of the user's view), such that the virtual object moves as the view or the part of the environment moves in order to maintain a fixed relationship between the virtual object and the part of the environment.
[0077] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits a delayed following behavior that reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object is following. In some embodiments, when exhibiting the delayed following behavior, the computer system intentionally delays the movement of the virtual object when detecting the movement of a reference point (e.g., a part of the environment, a viewpoint, or a point fixed relative to the viewpoint such as a point between 5 and 300 cm from the viewpoint) that the virtual object is following. For example, when the reference point (e.g., a part of the environment or a viewpoint) moves at a first speed, the virtual object is moved by the device so as to remain locked to the reference point but moves at a second speed that is slower than the first speed (e.g., until the reference point stops or decelerates its movement, at which point the virtual object begins to catch up to the reference point). In some embodiments, when the virtual object exhibits the delayed following behavior, the device ignores small movements of the reference point (e.g., movements of the reference point that are less than a threshold movement amount, such as movements of 0 to 5 degrees or 0 to 50 cm). For example, when the reference point (e.g., the part of the environment or the viewpoint to which the virtual object is locked) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is displayed so as to maintain a position fixed or substantially fixed relative to a viewpoint or a part of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., the part of the environment or the viewpoint to which the virtual object is locked) moves by a second amount that is greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is displayed so as to maintain a position fixed or substantially fixed relative to a viewpoint or a part of the environment different from the reference point to which the virtual object is locked), and then decreases as the virtual object is moved by the computer system so as to maintain a position fixed or substantially fixed relative to the reference point, as the movement amount of the reference point increases beyond a threshold (e.g., a "delayed following" threshold).In some embodiments, a virtual object that maintains a position substantially fixed relative to a reference point includes the virtual object being displayed within a threshold distance (e.g., 1, 2, 3, 5, 15, 20, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the position of the reference point).
[0078] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on top of a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable controllers or handheld controllers with or without tactile feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may include speakers integrated into the head-mounted system and / or other audio output devices to provide audio output. A head-mounted system may have one or more speakers (singular or plural) and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed towards a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scan light sources, or any combination of these technologies. The medium may be an optical waveguide, hologram medium, optical coupler, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface.In some embodiments, the controller 110 is configured to manage and adjust the XR experience for the user. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. The controller 110 will be described in more detail below with respect to FIG. 2. In some embodiments, the controller 110 is a computing device that is local or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the scene 105. In some embodiments, the controller 110 is communicatively coupled to a display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., BLUETOOTH, IEEE802.11x, IEEE802.16x, IEEE802.3x, etc.). In another example, the controller 110 is included within a housing (e.g., a physical housing) of one or more of the display generation component 120 (e.g., a portable electronic device including an HMD, or a display and one or more processors, etc.), one or more of the input devices 125, one or more of the output devices 155, one or more of the sensors 190, and / or one or more of the peripheral devices 195, or shares the same physical housing or support structure as one or more of the above.
[0079] In some embodiments, the display generation component 120 is configured to provide the user with an XR experience (e.g., at least a visual component of the XR experience). In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The display generation component 120 will be described in more detail below with respect to FIG. 3. In some embodiments, the functions of the controller 110 are provided by and / or combined with the display generation component 120.
[0080] According to some embodiments, the display generation component 120 provides an XR experience to a user while the user is virtually and / or physically present within the scene 105.
[0081] In some embodiments, the display generation component is worn on a part of the user's body (e.g., the user's own head or hand). Accordingly, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smartphone or a tablet) configured to present XR content, and the user holds a device having a display directed towards the user's field of view and a camera directed towards the scene 105. In some embodiments, the handheld device is optionally disposed within a housing worn on the user's head. In some embodiments, the handheld device is optionally disposed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR chamber, housing, or room configured to present XR content while the user is not wearing or holding the display generation component 120. Many of the user interfaces described with reference to one type of hardware for displaying XR content (e.g., a device on a handheld or tripod) may be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface that indicates an interaction with XR content triggered based on an interaction occurring within the space in front of a handheld or tripod-mounted device may be implemented in the same manner as an HMD where the interaction occurs within the space in front of the HMD and the response of the XR content is displayed via the HMD. Similarly, a user interface that indicates an interaction with XR content triggered based on the movement of a handheld or tripod-mounted device relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)) may be implemented in the same manner as an HMD where the movement is caused by the movement of the HMD relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eye(s), head, or hand)).
[0082] Although the relevant features of the operating environment 100 are shown in FIG. 1, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity and to not obscure more relevant aspects of the exemplary embodiments disclosed herein.
[0083] FIG. 2 is a block diagram of an example of the controller 110 according to some embodiments. Although certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity and to not obscure more appropriate aspects of the embodiments disclosed herein. Thus, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), infrared (IR), BLUETOOTH, ZIGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.
[0084] In some embodiments, one or more communication buses 204 include circuitry for interconnecting system components and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0085] Memory 220 includes high-speed random access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more magnetic storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220, or the non-transitory computer-readable storage medium of memory 220, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and an XR experience module 240.
[0086] Operating system 230 includes instructions for handling various basic system services and performing hardware-dependent tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate one or more XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for each group of one or more users). To that end, in various embodiments, XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, an adjustment unit 246, and a data transmission unit 248.
[0087] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of FIG. 1 and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data acquisition unit 241 includes instructions and / or logic therefor, and heuristics and metadata therefor.
[0088] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track at least the position / location of the display generation component 120 with respect to the scene 105 of FIG. 1 and optionally with respect to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the tracking unit 242 includes instructions and / or logic therefor, and heuristics and metadata therefor. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand with respect to the scene 105 of FIG. 1, with respect to the display generation component 120, and / or with respect to a coordinate system defined for the user's hand. The hand tracking unit 244 will be described in more detail below with respect to FIG. 4. In some embodiments, the eye tracking unit 243 is configured to track the position and movement of the user's line of sight (or more generally the user's eyes, face, or head) with respect to the scene 105 (e.g., the physical environment and / or the user (e.g., the user's hand)) or with respect to the XR content displayed via the display generation component 120. The eye tracking unit 243 will be described in more detail below with respect to FIG. 5.
[0089] In some embodiments, the adjustment unit 246 is configured to manage and adjust the XR experience presented to the user by the display generation component 120 and, optionally, by one or more of the output device 155 and / or the peripheral device 195. For that purpose, in various embodiments, the adjustment unit 246 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0090] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120 and, optionally, to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0091] Although the data acquisition unit 241, the tracking unit 242 (including, for example, the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 are shown as being present on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (including, for example, the eye tracking unit 243 and the hand tracking unit 244), the adjustment unit 246, and the data transmission unit 248 can be located within separate computing devices.
[0092] Furthermore, FIG. 2 is more intended to illustrate the functions of various features that may exist in a particular implementation, as contrasted with the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown matters can be combined, and some matters can be separated. For example, some of the functional modules separately shown in FIG. 2 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of particular functions and how the functions are assigned among them, vary depending on the implementation, and in some embodiments, they partially depend on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0093] FIG. 3 is a block diagram of an example of a display generation component 120 according to some embodiments. Although certain features are shown, those skilled in the art will understand from this disclosure that various other features are not shown for the sake of brevity so as not to obscure more suitable aspects of the embodiments disclosed herein. For that purpose, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, ASIC, FPGA, GPU, CPU, processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional inward and / or outward image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0094] In some embodiments, one or more communication buses 304 include circuitry that interconnects and controls communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of an inertial measurement unit (IMU), accelerometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, a tactile engine, one or more depth sensors (e.g., structured light, time of flight, etc.), and the like.
[0095] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emission device display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to waveguide displays, such as diffractive, reflective, polarizing, holographic, etc. For example, the display generation component 120 (e.g., HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 can present MR or VR content. In some embodiments, one or more XR displays 312 can present MR or VR content.
[0096] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of a user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand(s) and optionally the user's arm(s) (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user views when the display generation component 120 (e.g., an HMD) is not present (and may be referred to as a scene camera). One or more optional image sensors 314 can include one or more RGB cameras, one or more infrared (IR) cameras, one or more event-based cameras, and / or the like (e.g., including a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor).
[0097] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320, or the non-transitory computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and an XR presentation module 340.
[0098] The operating system 330 includes instructions for processing various basic system services and instructions for executing hardware-dependent tasks. In some embodiments, the XR presentation module 340 is configured to present XR content to a user via one or more XR displays 312. For that purpose, in various embodiments, the XR presentation module 340 includes a data acquisition unit 342, an XR presentation unit 344, an XR map generation unit 346, and a data transmission unit 348.
[0099] In some embodiments, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 of FIG. 1. For that purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0100] In some embodiments, the XR presentation unit 344 is configured to present XR content via one or more XR displays 312. For that purpose, in various embodiments, the XR presentation unit 344 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0101] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a composite reality scene or a map of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. For that purpose, in various embodiments, the XR map generation unit 346 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0102] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For that purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic therefor, as well as heuristics and metadata therefor.
[0103] The data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as being present on a single device (e.g., the display generation component 120 of FIG. 1), but in other embodiments, it should be understood that any combination of the data acquisition unit 342, the XR presentation unit 344, the XR map generation unit 346, and the data transmission unit 348 can be located within separate computing devices.
[0104] Furthermore, FIG. 3 is more intended to illustrate the functions of various features that may exist in a particular implementation, as opposed to the structural overview of the embodiments described herein. As will be recognized by those skilled in the art, the separately shown matters can be combined, and some matters can be separated. For example, several functional modules separately shown in FIG. 3 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various embodiments. The actual number of modules, as well as the specific division of a particular function and how the functions are assigned among them, vary depending on the implementation, and in some embodiments, they depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.
[0105] FIG. 4 is a schematic diagram of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 (FIG. 1) is for the scene 105 of FIG. 1 (e.g., for a portion of the physical environment surrounding the user, for the display generation component 120, or for a part of the user (e.g., the user's face, eyes, or head), and / or for a coordinate system defined with respect to the user's hand) to track the location / position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand, and is controlled by the hand tracking unit 244 (FIG. 2). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0106] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images at a resolution sufficient to distinguish the fingers and their respective positions. The image sensor 404 typically captures images of other parts of the user's body or all of the body and can have either a zoom function or a dedicated sensor with high magnification to capture hand images at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with or functions as an image sensor that captures the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment such that the field of view of the image sensor or a portion thereof is used to define an interaction space in which hand movements captured by the image sensor are processed as inputs to the controller 110.
[0107] In some embodiments, the image sensor 404 outputs a sequence of frames including 3D map data (and optionally color image data as well) to the controller 110, thereby extracting high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API) and drives the display generation component 120 accordingly. For example, the user can interact with software running on the controller 110 by moving their hand 406 and changing the pose of their hand.
[0108] In some embodiments, the image sensor 404 projects a spot pattern onto a scene that includes the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral shift of the spots of the pattern. This approach is advantageous in that the user does not need to hold or wear any kind of beacon, sensor, or other marker. This gives the depth coordinates of points in the scene relative to a predetermined reference plane at a particular distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines a series of orthogonal x, y, and z axes such that the depth coordinates of points in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) can use other 3D mapping methods such as stereoscopy or time-of-flight measurement based on single or multiple cameras or other types of sensors.
[0109] In some embodiments, the hand tracking device 140 captures and processes a temporal sequence of depth maps of the user's hand while the user moves their hand (e.g., the entire hand or one or more fingers). Software operating on a processor within the image sensor 404 and / or the controller 110 processes the 3D map data to extract hand patch descriptors within these depth maps. The software compares these descriptors to patch descriptors stored in the database 408 based on a previous learning process to estimate the hand pose in each frame. The pose typically includes the 3D locations of the user's hand joints and fingertips.
[0110] The software can also analyze the trajectories of the hand and / or fingers over multiple frames within a sequence to identify gestures. The pose estimation function described herein may be interleaved with the motion tracking function, such that patch-based pose estimation is only performed once per two (or more) frames, while tracking is used to detect changes in pose occurring over the remaining frames. Pose, motion, and gesture information is provided to an application program running on the controller 110 via the API described above. This program can, for example, move and modify the image presented on the display generation component 120 or perform other functions in response to pose and / or gesture information.
[0111] In some embodiments, gestures include air gestures. An air gesture is a gesture that is detected without (or independent of) the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140), and is based on detected movement of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the other of the user's hands, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., tap gestures including movement of the hand in a predetermined pose by a predetermined amount and / or speed, or shake gestures including a predetermined speed or amount of rotation of a part of the user's body).
[0112] In some embodiments, the input gestures used in the various examples and embodiments described herein are air gestures performed by the movement of a user's finger(s) relative to other finger(s) or portion(s) of the user's hand for interacting with an XR environment (e.g., a virtual or mixed reality environment) according to some embodiments. In some embodiments, an air gesture is a gesture that is detected without the user touching (or independently of) an input element that is part of the device, and is based on detected movement of a part of the user's body, such as movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), movement of the user's body relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of the user's other hand relative to one of the user's hands, and / or movement of the user's finger relative to another finger or portion of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture including movement of the hand in a predetermined pose with a predetermined amount and / or speed, or a shake gesture including rotation of a part of the user's body at a predetermined speed or amount).
[0113] In some embodiments where the input gesture is an air gesture (i.e., without physical contact with an input device that provides information to the computer system regarding which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., line of sight) to determine the target of the user input (e.g., in the case of direct input, as described below). Thus, in implementations that include air gestures, the input gesture is, for example, detected attention (e.g., line of sight) to a user interface element in combination with (e.g., simultaneously with) movement of the user's finger(s) and / or hand for performing pinch and / or tap inputs, as described in more detail below.
[0114] In some embodiments, input gestures directed to a user interface object are executed directly or indirectly with reference to the user interface object. For example, a user input is executed directly on a user interface object in response to the user performing an input gesture with their hand at a position corresponding to the position of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current perspective). In some embodiments, an input gesture is executed indirectly on a user interface object in accordance with a user who performs the input gesture while the position of the user's hand is not at a position corresponding to the position of the user interface object in the three-dimensional environment while detecting the user's attention (e.g., line of sight) to the user interface object. For example, in the case of a direct input gesture, the user can direct the user's input to the user interface object by starting a gesture at or near a position corresponding to the display position of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or 0 - 5 cm measured from the outer edge of the option or the central portion of the option). In the case of an indirect input gesture, the user can direct the user's input to the user interface object by paying attention to the user interface object (e.g., by gazing at the user interface object), and while paying attention to the option, the user starts the input gesture (e.g., at any position detectable by the computer system) (e.g., at a position not corresponding to the display position of the user interface object).
[0115] In some embodiments, the input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or augmented reality environment according to some embodiments. For example, the pinch inputs and tap inputs described later are performed as air gestures.
[0116] In some embodiments, the pinch input is part of an air gesture that includes one or more of a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, the pinch gesture, which is an air gesture, involves moving two or more fingers of a hand so as to contact each other, i.e., optionally including an interruption immediately after contact with each other (e.g., within 0 to 1 second). The long pinch gesture, which is an air gesture, involves moving two or more fingers of a hand so as to contact each other for at least a threshold amount of time (e.g., at least 1 second) before detecting an interruption in the contact with each other. For example, the long pinch gesture includes the user holding a pinch gesture (e.g., when two or more fingers are in contact), and the long pinch gesture continues until an interruption in the contact between two or more fingers is detected. In some embodiments, the double pinch gesture, which is an air gesture, includes two (e.g., or more) pinch inputs (e.g., performed with the same hand) that are detected continuously and directly with each other (e.g., within a predetermined period). For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., breaks the contact between two or more fingers), and then performs a second pinch input within a predetermined period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.
[0117] In some embodiments, a pinch-and-drag gesture, which is an air gesture, includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) that is performed in relation to (e.g., subsequent to) a drag input that changes the position of the user's hand from a first position (e.g., the starting position of the drag) to a second position (e.g., the ending position of the resistance). In some embodiments, the user maintains the pinch gesture while performing the drag input and releases the pinch gesture (e.g., opens two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers together and touches each other, and moves the same hand to a second position in the air with a drag gesture). In some embodiments, the pinch input is performed by the user's first hand, and the drag input is performed by the user's second hand (e.g., the user's second hand moves from the first position to the second position in the air while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture, which is an air gesture, includes an input (e.g., a pinch input and / or a tap input) that is performed using both of the user's hands. For example, the input gesture includes two (e.g., or more) pinch inputs that are performed in relation to each other (e.g., simultaneously with a predetermined period or within a predetermined period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) performed using the user's first hand, and in relation to performing the pinch input using the first hand, a second pinch input is performed using the other hand (e.g., the second hand of the user's two hands). In some embodiments, the movement between the user's two hands (e.g., to increase and / or decrease the distance or relative orientation between the user's two hands).
[0118] In some embodiments, a tap input that is performed as an air gesture (e.g., directed at a user interface element) includes a movement (s) of the user's finger (s) towards the user interface element, optionally a movement of the user's hand towards the user interface element with the user's finger (s) extended towards the user interface element, a downward movement of the user's finger (e.g., mimicking a mouse click operation or a tap on a touch screen), or other predefined movement of the user's hand. In some embodiments, a tap input that is performed as an air gesture is detected based on movement characteristics of the finger or hand that performs a tap gesture movement away from the user's perspective and / or towards an object that is the target of the tap input where the end of the movement follows. In some embodiments, an end of the movement is detected based on a change in movement characteristics of the finger or hand that performs a tap gesture (e.g., movement away from the user's perspective and / or towards the object that is the target of the tap input, reversal of the direction of movement of the finger or hand, and / or reversal of the direction of acceleration of the movement of the finger or hand).
[0119] In some embodiments, the user's attention is determined to be directed towards a portion of the three-dimensional environment based on detection of a line of sight directed towards the portion of the three-dimensional environment (optionally, without requiring other conditions). In some embodiments, in order for the device to determine that the user's attention is directed towards a portion of the three-dimensional environment, while the user's perspective is within a distance threshold from the portion of the three-dimensional environment, at least a threshold duration (e.g., dwell time) of the line of sight being directed towards the portion of the three-dimensional environment is required, and / or one or more additional conditions such as the line of sight being directed towards the portion of the three-dimensional environment are required. Based on detection of a line of sight directed towards a portion of the three-dimensional environment with one or more of the additional conditions, it is determined that the user's attention is directed towards the portion of the three-dimensional environment. If one of the additional conditions is not satisfied, the device determines that the attention is not directed towards the portion of the three-dimensional environment towards which the line of sight is directed (e.g., until one or more of the additional conditions are satisfied).
[0120] In some embodiments, the detection of the prepared state configuration of the user or a part of the user is detected by the computer system. The detection of the prepared state configuration of the hand is used by the computer system as an indication that the user is likely preparing to interact with the computer system using one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand. For example, the prepared state of the hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced to perform a pinch or grab gesture, or a pre-tap where one or more fingers are extended and the palm is facing away from the user), whether the hand is in a predetermined position relative to the user's perspective (e.g., under the user's head, above the user's waist, extended at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a particular way (e.g., moved towards the front area of the user above the user's waist and under the user's head, or away from the user's body or legs). In some embodiments, the prepared state is used to determine whether the interaction elements of the user interface respond to attention (e.g., gaze) input.
[0121] In some embodiments, the software may be downloaded in electronic form to the controller 110, for example, over a network, or alternatively may be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is likewise stored in a memory associated with the controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). The controller 110 is shown in FIG. 4 as a separate unit from the image sensor 404 by way of example, but some or all of the processing functions of the controller may be associated with the image sensor 404 by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device), or by other means. In some embodiments, at least some of these processing functions are performed by a suitable processor integrated with the display generation component 120 (e.g., in a television set, a handheld device, or a head-mounted device), or using any other suitable computerized device such as a game console or a media player. The sensing function of the image sensor 404 can likewise be integrated with a computer or other computerized device controlled by the sensor output.
[0122] FIG. 4 further includes a schematic diagram of a depth map 410 captured by an image sensor 404 according to some embodiments. The depth map includes a matrix of pixels having respective depth values, as described above. The pixels 412 corresponding to the hand 406 are segmented in this map from the background and the wrist. The luminance of each pixel in the depth map 410 is inversely proportional to the depth value, i.e., the measured z - distance from the image sensor 404, and the tone becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment components (i.e., groups of adjacent pixels) of an image having the characteristics of a human hand. These characteristics can include, for example, the overall size, shape, and movement from frame to frame of a sequence of depth maps.
[0123] FIG. 4 also schematically shows a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. In FIG. 4, the hand skeleton 414 is superimposed on the hand background 416 segmented from the original depth map. In some embodiments, the hand (e.g., knuckles, fingertips, center of the palm, the end of the hand connected to the wrist, etc.), and optionally major feature points on the wrist or arm connected to the hand are identified and placed on the hand skeleton 414. In some embodiments, the locations and movements of these major feature points over a plurality of image frames are used by the controller 110 to determine, according to some embodiments, a hand gesture or the current state of the hand performed by the hand.
[0124] FIG. 5 shows an exemplary embodiment of an eye tracking device 130 (FIG. 1). In some embodiments, the eye tracking device 130 is controlled by an eye tracking unit 243 (FIG. 2) to track the position and movement of the user's line of sight with respect to scene 105 or XR content displayed via display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, if the display generation component 120 is a head-mounted device such as a headset, helmet, goggles, or glasses, or a handheld device disposed on a wearable frame, the head-mounted device includes both a component for generating XR content for viewing by the user and a component for tracking the user's line of sight with respect to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, if the display generation component is a handheld device or an XR chamber, the eye tracking device 130 is optionally a device separate from the handheld device or XR chamber. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used with a display generation component worn on the head or a display generation component not worn on the head. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.
[0125] In some embodiments, the display generation component 120 uses a display mechanism (e.g., left and right near-eye display panels) that presents a frame including left and right images in front of the user's eyes to provide the user with a 3D virtual view. For example, the head-mounted display generation component may include left and right optical lenses (referred to herein as eyepieces) positioned between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external video cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display on which the user can directly view the physical environment and display virtual objects. In some embodiments, the display generation component projects virtual objects onto the physical environment. The virtual objects are projected, for example, onto a physical surface or as a hologram, whereby an individual can use the system to observe virtual objects superimposed on the physical environment. In such cases, separate display panels and image frames for the left and right eyes may not be required.
[0126] As shown in FIG. 5, in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) camera or a near IR (NIR) camera), and an illumination source (e.g., an IR light source or an NIR light source such as an array or a ring of LEDs) that emits light (e.g., IR light or NIR light) towards the user's eyes. The eye tracking camera may be directed towards the user's eyes to directly receive reflected IR or NIR light from the light source, or alternatively, may be directed towards a "hot" mirror disposed between the user's eyes and a display panel that reflects IR or NIR light from the eyes while allowing visual light to pass through to the eye tracking camera. The eye tracking device 130 optionally captures an image of the user's eyes (e.g., as a video stream captured at 60 to 120 frames per second (fps)), analyzes the image to generate gaze tracking information, and communicates the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by respective eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by an individual eye tracking camera and illumination source.
[0127] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LED, camera, hot mirror (if present), eyepiece, and display screen. The device-specific calibration process may be performed at a factory or another facility prior to delivery of the AR / VR device to the end user. The device-specific calibration process may be an automatic calibration process or a manual calibration process. The user-specific calibration process may include an estimation of specific user eye parameters, such as pupil location, foveal location, optical axis, visual axis, interocular distance, etc. According to some embodiments, once the device-specific and user-specific parameters for the eye tracking device 130 are determined, the images captured by the eye tracking camera are processed using a glint assist method to determine the user's current visual axis and viewpoint with respect to the display.
[0128] As shown in FIG. 5, an eye tracking device 130 (e.g., 130A or 130B) includes a line-of-sight tracking system including one or more eyepieces 520, at least one eye tracking camera 540 (e.g., an infrared (IR) or near-IR (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye(s) 592. The eye tracking camera 540 is positioned between the user's eye(s) 592 and a display 510 (e.g., a display panel on the left or right side of a head-mounted display, or a display of a handheld device, a projector, etc.), and may be directed toward a mirror 550 that reflects IR or NIR light from the eye(s) 592 while transmitting visible light (as shown, for example, at the top of FIG. 5), or may be directed toward the user's eye(s) 592 to receive the reflected IR or NIR light from the eye(s) 592 (as shown, for example, at the bottom of FIG. 5).
[0129] In some embodiments, the controller 110 renders an AR or VR frame 562 (e.g., the left and right frames of the left and right display panels) and provides the frame 562 to the display 510. The controller 110 uses the line-of-sight tracking input 542 from the eye tracking camera 540, for example, when processing the frame 562 for display, for various purposes. The controller 110 optionally estimates the user's viewpoint on the display 510 based on the line-of-sight tracking input 542 obtained from the eye tracking camera 540 using a glint assist method or other suitable method. The viewpoint estimated from the line-of-sight tracking input 542 is optionally used to determine the direction the user is currently looking.
[0130] Examples of several possible use cases of the user's current line of sight direction are described below, but this is not intended to be limiting. As an exemplary use case, the controller 110 can render virtual content differently based on the determined line of sight direction of the user. For example, the controller 110 may generate virtual content at a higher resolution in the central visual region determined from the user's current line of sight direction than in the peripheral region. As another example, the controller may position or move virtual content within the view based at least in part on the user's current line of sight direction. As another example, the controller may display specific virtual content within the view based at least in part on the user's current line of sight direction. As another exemplary use case in an AR application, the controller 110 can capture the physical environment of the XR experience and direct an external camera to focus in the determined direction. The autofocus mechanism of the external camera can then focus on an object or surface within the environment that the user is currently viewing on the display 510. As another exemplary use case, the eyepiece 520 may be a focusable lens, and the line of sight tracking information is used by the controller to adjust the focus of the eyepiece 520 so that the virtual object the user is currently viewing has appropriate binocular convergence to match the convergence of the user's eyes 592. The controller 110 can utilize the line of sight tracking information to direct and adjust the focus of the eyepiece 520 so that the nearby object the user is viewing appears at the correct distance.
[0131] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eyepieces (e.g., eyepiece(s) 520), an eye-tracking camera (e.g., eye-tracking camera(s) 540), and a light source (e.g., light source 530 (e.g., IR LED or NIR LED)) attached to a wearable housing. The light source emits light (e.g., IR light or NIR light) towards the user's eye(s) 592. In some embodiments, the light source may be arranged in a ring or circularly around each lens, as shown in FIG. 5. In some embodiments, eight light sources 530 (e.g., LEDs) are arranged around each lens 520 as an example. However, more or fewer light sources 530 may be used, and other arrangements and locations of the light sources 530 may be employed.
[0132] In some embodiments, the display 510 emits light within the visible light range and does not emit light within the IR or NIR range, so as not to introduce noise into the eye-tracking system. Note that the location and angle of the eye-tracking camera(s) 540 are given as an example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 can be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0133] An embodiment of a gaze tracking system as shown in FIG. 5 can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, extended reality, and / or augmented virtual experience.
[0134] FIG. 6 shows a glint-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a glint-assisted gaze tracking system (e.g., an eye tracking device 130 as shown in FIGS. 1 and 5). The glint-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, the glint-assisted gaze tracking system uses prior information from the previous frame when analyzing the current frame to track the pupil contour and glint within the current frame. If not in the tracking state, the glint-assisted gaze tracking system attempts to detect the pupil and glint within the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.
[0135] As shown in FIG. 6, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments, or under some conditions, not all captured frames are processed by the pipeline.
[0136] At 610, for the currently captured image, if the tracking state is yes, this method proceeds to element 640. At 610, if the tracking state is no, as shown at 620, the image is analyzed to detect the user's pupil and glint within the image. At 630, if the pupil and glint are detected normally, the method proceeds to element 640. If not detected normally, the method returns to element 610 to process the next image of the user's eyes.
[0137] At 640, when proceeding from element 610, the current frame is analyzed to track the pupil and glint based in part on previous information from the previous frame. At 640, when proceeding from element 630, the tracking state is initialized based on the detected pupil and glint within the current frame. The result of the processing at element 640 is checked to confirm that the result of the tracking or detection is reliable. For example, the result can be checked to determine whether a sufficient number of glints for performing pupil and gaze estimation are being normally tracked or detected in the current frame. At 650, if the result is not reliable, the tracking state is set to no at element 660 and the method returns to element 610 to process the next image of the user's eyes. At 650, if the result is reliable, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's viewpoint.
[0138] FIG. 6 is intended to function as an example of an eye tracking technique that can be used in a particular implementation. As will be recognized by those skilled in the art, other eye tracking techniques that currently exist or may be developed in the future can be used in computer system 101 to provide an XR experience to a user according to various embodiments, instead of, or in combination with, the glint-assisted eye tracking technique described herein.
[0139] In the present disclosure, various input methods are described with respect to interaction with a computer system. If one example is provided using one input device or input method and another example is provided using another input device or input method, it should be understood that each example may be compatible with and optionally utilize the input device or input method described in another example. Similarly, various output methods are described with respect to interaction with a computer system. If one example is provided using one output device or output method and another example is provided using another output device or output method, it should be understood that each example may be compatible with and optionally utilize the output device or output method described in another example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. If one example is provided using interaction with a virtual environment and another example is provided using a mixed reality environment, it should be understood that each example may be compatible with and optionally utilize the method described in another example. Thus, the present disclosure discloses embodiments that are combinations of features of multiple examples without comprehensively listing all features of the embodiments in the description of each exemplary embodiment. User Interface and Related Processes
[0140] Next, attention is paid to embodiments of a user interface ("UI") and related processes that may be implemented on a computer system, such as a portable multifunctional device or a head-mounted device, that communicate with a display generation component and (optionally) one or more cameras and one or more input devices.
[0141] Figures 7A - 7Q illustrate exemplary techniques for capturing and / or displaying media in various environments, according to some embodiments. FIG. 8 is a flowchart of a method for capturing media, according to various embodiments. FIG. 9 is a flowchart of a method for displaying a preview of media, according to various embodiments. FIG. 10 is a flowchart of a method for displaying previously captured media. The user interfaces of FIGS. 7A - 7Q are used to illustrate the processes described below, including the processes of FIGS. 8, 9, and 10.
[0142] Figures 7A - 7Q illustrate exemplary techniques for capturing and browsing media, according to some embodiments. The schematic diagrams and user interfaces of FIGS. 7A - 7Q are used to illustrate the processes described below, including the processes of FIGS. 8, 9, and 10.
[0143] FIG. 7A shows a user 712 holding a computer system 700 that includes a display 702 in a physical environment (e.g., a room in a house). The physical environment includes a couch 709a, a photograph 709b, a first individual 709c1, a second individual 709c2, a television 709d, and a table 709e. The display 702 presents a representation of the physical environment 704 (e.g., using "pass-through video" as described above). The user 712 holds the computer system 700 such that the couch 709a, the photograph 709b, the first individual 709c1, and the second individual 709c2 are visible from the user's perspective, which is determined based on the location of a portion of the computer system 700 that includes one or more cameras that acquire visual information about the physical environment and generate a virtual environment based on the visual information about the visual environment for use in a virtual pass-through. In the embodiments of FIGS. 7A-7Q, the user's perspective corresponds to the field of view of one or more cameras (e.g., a camera on the back of the computer system 700) that communicate with the computer system 700. Thus, in the case of a virtual pass-through, as the computer system 700 is moved throughout the physical environment, the field of view of the one or more cameras changes, thereby changing the user's perspective. Since the couch 709a, the photograph 709b, the individual 709c1, and the second individual 709c2 are visible from the user's perspective in FIG. 7A, the display 702 includes depictions of the couch 709a, the photograph 709b, the first individual 709c1, and the second individual 709c2. When the user 712 looks at the display 702, the user 712 can see a representation of the physical environment 704 along with one or more virtual objects that the computer system 700 can display (e.g., as shown in FIGS. 7B-7Q). Thus, the computer system 700 presents an extended reality environment via the display 702.
[0144] Computer system 700 is a tablet in FIG. 7A, but in some embodiments, computer system 700 can be one or more other devices such as a handheld device (e.g., a smartphone) and / or a head-mounted device. In some embodiments, when computer system 700 is a head-mounted device, the representation of physical environment 704 is an extended reality environment. In some embodiments, the representation of physical environment 704 is an extended reality environment, but the representation of physical environment 704 includes immersive visual characteristics including the display of depth data (e.g., the foreground and background of the representation of physical environment 704 are displayed differently to present a depth visual effect when viewed by a user of computer system 700). In some embodiments, computer system 700 includes one or more components of computer system 101, and / or display 702 includes components of display generation component 120. In some embodiments, display 702 presents a representation of a virtual environment (e.g., instead of the physical environment in FIG. 7A).
[0145] Figures 7B - 7E illustrate a method for capturing spatial (e.g., immersive) media. In Figures 7B - 7E, computer system 700 remains in the physical environment shown in Figure 7A, as depicted in schematic 701, which will be described in more detail below. In Figures 7B - 7E, computer system 700 is shown here in an enlarged view to better illustrate the content visible on display 702. As shown in Figure 7B, computer system 700 displays a control center virtual object 707 (e.g., in response to a swipe gesture on display 702 executed by user 712). The control center virtual object 707 includes a plurality of virtual objects. Each virtual object included in the control center virtual object 707 is selectable. When each virtual object included in the control center virtual object 707 is selected, it causes computer system 700 to perform respective operations (e.g., change the playback state of computer system 700, change the volume at which computer system 700 outputs audio, display the applications currently installed on computer system 700, and / or any other appropriate operations).
[0146] As shown in Figure 7B, computer system 700 presents a representation of physical environment 704. The representation of physical environment 704 corresponds to the user's perspective (e.g., the representation of physical environment 704 includes the content visible from the user's perspective). That is, when the user's perspective changes, the representation of physical environment 704 changes based on the change in the user's perspective of computer system 700. In some embodiments, the representation of physical environment 704 is a pass - through representation of at least a portion of the physical environment surrounding computer system 700.
[0147] As shown in FIG. 7B, the representation of the physical environment 704 is visually contrasted with the display of the control center virtual object 707. The representation of the physical environment 704 includes a first amount of shading / blurring, and the display of the control center virtual object 707 is displayed with a second amount of shading / blurring (e.g., no shading / blurring) that is different from the first amount of shading / blurring. In some embodiments, the representation of the physical environment 704 is not contrasted with the display of the control center virtual object 707. In some embodiments, the representation of the physical environment 704 has no amount of blurring / shading.
[0148] FIGS. 7B-7Q include a schematic diagram 701 of a physical environment. The computer system 700 is represented by an indication 703 within the schematic diagram 701. That is, the location and orientation of the indication 703 in the schematic diagram 701 represent the location and orientation of the computer system 700 within the physical environment. The schematic diagram 701 shows the physical environment shown in FIG. 7A, but this is merely an example, and it should be recognized that the techniques described herein can be associated with other types of physical environments. The schematic diagram 701 is merely a visual aid. The computer system 700 does not display the schematic diagram 701.
[0149] FIG. 7B shows a computer system 700 having a hardware button 711a (e.g., a hardware input device / machine) (e.g., a physical input device) and a hardware button 711b. Further, FIG. 7B shows a body part 712a of user 712. The body part 712a represents one of the fingers of user 704 (e.g., the user's index finger, ring finger, little finger, middle finger, or thumb). In some embodiments, the representation of body part 712a can be any other part of the user 704's body that can activate the hardware button 711a or the hardware button 711b (e.g., the wrist, arm, hand, and / or any other suitable body part). In FIG. 7B, the computer system 700 detects the activation of the hardware button 711a by the body part 712a, or the computer system 700 detects an input 750b directed at the camera virtual object 707a. In some embodiments, the input 750b is a tap input on the camera virtual object 707a (e.g., an air tap within the space corresponding to the location of the display of the camera virtual object 707a). In some embodiments, the input 750b is a gaze input (e.g., a sustained gaze) directed at the direction of the display of the camera virtual object 707a. In some embodiments, the input 750b is an air tap input combined with the detection of a gaze in the direction of the display of the camera virtual object 707a. In some embodiments, the input 750b is a gaze and a blink directed at the direction of the display of the camera virtual object 707a.
[0150] As shown in FIG. 7C, computer system 700 displays media capture preview 708, timer virtual object 713, camera shutter virtual object 714, repositioning virtual object 716, deletion virtual object 719, and photo well virtual object 715 in response to activation of hardware button 711a by body part 712a or detection of input 750b directed at camera virtual object 707a. Computer system 700 displays media capture preview 708 so as to be overlaid on top of a representation of physical environment 704. As shown in FIG. 7C, the display of media capture preview 708 is smaller than the representation of physical environment 704 (e.g., occupies less space on display 702). In some embodiments, computer system 700 is a head-mounted device that presents a representation of physical environment 704 along with one or more virtual objects that computer system 700 displays via a display generation component that surrounds (or substantially surrounds) the user's field of view. In embodiments where computer system 700 is an HMD, the user's viewpoint is locked in the forward direction of the user's head, such that the representation of physical environment 704 and one or more virtual objects such as media capture preview 708 shift as the user's head moves (e.g., because computer system 700 also moves as the user's head moves).
[0151] The timer virtual object 713, the camera shutter virtual object 714, the repositioning virtual object 716, the deletion virtual object 719, and the photo well virtual object 715 are all fixed to the media capture preview 708. That is, the display locations of the timer virtual object 713, the camera shutter virtual object 714, the repositioning virtual object 716, and the photo well virtual object 715 are associated with the display location of the media capture preview 708. In some embodiments, when the display location of the media capture preview 708 changes, the display locations of the timer virtual object 713, the camera shutter virtual object 714, the repositioning virtual object 716, the deletion virtual object 719, and the photo well virtual object 715 change (see, for example, FIGS. 7F-7G). As shown in FIG. 7C, the computer system 700 displays the media capture preview 708 above / on top of the camera shutter virtual object 714 and at the center of the display 702. As shown in FIG. 7C, the media capture preview 708 includes a portion of the representation of the physical environment 704 that was visible before the computer system 700 displayed the media capture preview 708. For example, in FIG. 7B (e.g., before the computer system 700 displays the media capture preview 708), the representation of the physical environment includes the couch 709a, the photo 709b, the first individual 709c1, and the second individual 709c2. Thus, as shown in FIG. 7C, the media capture preview 708 includes the depiction of the couch 709a1, the picture 709b1, the first individual 709c3, and the second individual 709c4. The media capture preview 708 provides a preview of the content to be captured in response to the computer system 700 detecting a request to capture media. The content displayed within the media capture preview 708 is based on the field of view of one or more cameras communicating with the computer system 700 (e.g., the content displayed within the media capture preview 708 is within the field of view of one or more cameras communicating with the computer system 700).The content displayed within the media capture preview 708 changes based on the change in the field of view of one or more cameras. In some embodiments, the computer system 700 includes two cameras, and the content displayed within the media capture preview 708 is the content that falls within the field of view of both cameras, which enables the capture of immersive content.
[0152] In FIG. 7C, the user's perspective corresponding to the representation of the physical environment 704 has a wider viewing angle range than the angular range of the field of view corresponding to the media capture preview 708. As a result, the representation of the physical environment 704 shows a greater amount of the physical environment than the amount of the physical environment shown within the media capture preview 708 (e.g., in the media capture preview 708, only a portion of the sofa is visible, whereas in the representation of the physical environment 704, the entire sofa is visible). Thus, the media captured while the media capture preview 708 appears as shown in FIG. 7C will include only a portion of the sofa, not the entire sofa. In some embodiments, the representation of the physical environment included in the media capture preview 708 is displayed at a first scale, and the representation of the physical environment 704 is presented at a second scale that is larger than the first scale. In some embodiments, the computer system 700 includes two cameras having different but overlapping fields of view, the media capture preview 708 represents the portion of the physical environment common to the fields of view of both cameras (e.g., the portion where the FOVs overlap), and the representation of the physical environment 704 includes the content within the FOV of the first camera and / or the second camera of the two cameras (e.g., both overlapping and non - overlapping). In some embodiments, the display of the media capture preview 708 of the computer system 700 includes the content included in the representation of the physical environment 704. In some embodiments, the representation of the physical environment 704 is displayed from an immersive perspective view, and the content included in the display of the media capture preview 708 is displayed from a non - immersive perspective view.
[0153] As shown in FIG. 7C, the media capture preview 708 is displayed with a visual appearance that does not include dimming and / or blurring, and the representation of the physical environment 704 is displayed with dimming and / or blurring. This provides a contrast between the display of the media capture preview 708 and the representation of the physical environment 704. In some embodiments, the representation of the physical environment 704 is not dimmed and / or blurred before the computer system 700 detects the input 750b or before the computer system 700 detects the activation by the body part 712a of the hardware button 711a in FIG. 7B. In some embodiments, the representation of the physical environment 704 is dimmed and / or blurred (e.g., faded out) in response to the computer system 700 detecting the input 750b or in response to the computer system 700 detecting the activation by the body part 712a of the hardware button 711a.
[0154] Regarding the virtual objects fixed to the display of the media capture preview 708, the timer virtual object 713 provides an indication of the amount of time (e.g., minutes, seconds, hours) elapsed since the computer system 700 started the media capture process. The photo well virtual object 715 includes a representation of the most recently captured media item (e.g., a still photo or video). In some embodiments, the photo well virtual object 715 includes a representation of the most recently captured media item captured by the computer system 700. In some embodiments, the photo well virtual object 715 includes a representation of the most recently captured media item captured by an external device communicating with the computer system 700. As shown in FIG. 7C, the photo well virtual object 715 includes a representation of a water jet. Thus, the most recently captured media item includes a depiction of a water jet.
[0155] The selection of the camera shutter virtual object 714 initiates a process on the computer system 700 to capture media including the content shown within the media capture preview 708. Repositioning the virtual object 716 enables the user 712 to reposition the location of the display of the media capture preview 708. For example, moving the display location of the repositioning virtual object 716 to the left results in the display of the media capture preview 708 moving to the left. Selecting the erase virtual object 719 causes the computer system 700 to cease displaying the media capture preview 708. In some embodiments, when the media capture preview 708 ceases to be displayed, the representation of the physical environment 704 ceases to blur and / or ceases to have shadows.
[0156] In FIG. 7C, the computer system 700 detects the activation of the hardware button 711a by the body part 712a, or the computer system 700 detects the input 750c directed to the camera shutter virtual object 714. In some embodiments, the input 750c is a tap on the camera shutter virtual object 714 (e.g., an air tap within the space corresponding to the location of the display of the camera shutter virtual object 714). In some embodiments, the input 750c is a line-of-sight (e.g., persistent line-of-sight) input directed toward the direction of the display of the camera shutter virtual object 714. In some embodiments, the input 750c is an air tap input combined with the detection of a line of sight in the direction of the display of the camera shutter virtual object 714. In some embodiments, the input 750c is a line of sight and a blink directed toward the direction of the display of the camera shutter virtual object 714. In some embodiments, the activation of the hardware button 711a or the input 750c by the body part 712a is a long press (e.g., press and hold) (e.g., the duration of the activation of the hardware button or the input 750c by the body part 712a lasts for several seconds). In some embodiments, the activation of the hardware button 711 or the input 750c by the body part 712a is a short press (e.g., press and release) (e.g., the duration of the activation of the hardware button 711a or the input 750c by the body part 712a is less than 1 second). In some embodiments, a specific air gesture recognized as a request to capture media (e.g., as described above with respect to the selection of virtual objects in the XR environment) is detected.
[0157] In FIG. 7D, in response to detecting the activation of the body portion 712a of the hardware button 711a or in response to detecting the input 750c, the computer system 700 starts a media capture process. In FIG. 7D, a determination is made that the activation of the hardware button 711 or the input 750c by the body portion 712a is a long press. Since a determination is made that the activation of the body portion 712a of the hardware button 711 or the input 750c is a long press, video media is captured (e.g., not still media). The media capture process records the content displayed in the media capture preview 708 while the media capture process is in progress. In some embodiments, the field of view of one or more cameras communicating with the computer system 700 is changed during the media capture process, changing what is displayed in the media capture preview 708 and changing the content being captured by the media capture process. In some embodiments, in accordance with a determination that the activation of the hardware button 711 or the input 750c by the body portion 712a is a short press, still media (e.g., a photo) is captured via the media capture process.
[0158] As shown in FIG. 7D, the display of the timer virtual object 713 shows "00:05" (e.g., 5 seconds). The timer virtual object 713 in FIG. 7D indicates that 5 seconds have elapsed since the computer system 700 started the media capture process. Further, as shown in FIG. 7D, the display of the camera shutter virtual object 714 includes a square. The display of the camera shutter virtual object 714 having a square indicates that the computer system 700 is currently recording video media. In some embodiments, the shape, size, and / or color of the camera shutter virtual object 714 are updated to indicate that the computer system 700 is currently recording video media. In FIG. 7D, the computer system 700 detects the activation of the hardware button 711a by the body part 712a, or the computer system 700 detects the input 750d directed at the camera shutter virtual object 714. In some embodiments, the input 750d is a tap input on the camera shutter virtual object 714 (e.g., an air tap within the space corresponding to the location of the display of the camera shutter virtual object 714). In some embodiments, the input 750b is a gaze input directed in the direction of the display of the camera shutter virtual object 714. In some embodiments, the input 750d is an air tap combined with the detection of a gaze in the direction of the display of the camera shutter virtual object 714. In some embodiments, the input 750d is a gaze and a blink directed in the direction of the display of the camera shutter virtual object 714.
[0159] In FIG. 7E, computer system 700 stops the media capture process in response to activation by body portion 712a of hardware button 711a, or detection of input 750d directed to camera shutter virtual object 714. Since computer system 700 is no longer executing the media capture process, the display of camera shutter virtual object 714 in FIG. 7E does not include a square. As shown in FIG. 7E, computer system 700 displays photo well virtual object 715 along with a representation of the video captured in FIGS. 7C-7D (e.g., video of the physical environment). Further, as shown in FIG. 7E, since computer system 700 is no longer executing the media capture process, the display of timer virtual object 713 is shown as 0:00.
[0160] As shown in FIG. 7E, schematic diagram 701 includes movement indicator 721. Movement indicator 721 indicates that computer system 700 has begun to move within the physical environment. In FIG. 7E, computer system 700 begins to move horizontally to the right within the physical environment.
[0161] FIGS. 7F-7H show a method of computer system 700 that displays media capture preview 708 following (e.g., trailing) the movement of computer system 700. FIGS. 7E-7G show one continuous horizontal movement of computer system 700 to the right within the physical environment. In some embodiments, computer system 700 is moved horizontally left, up, and / or down within the physical environment. In some embodiments, computer system 700 is moved in combinations of different directions (e.g., up and left and / or down and right).
[0162] In FIG. 7F, computer system 700 is positioned horizontally to the right of a previous position of computer system 700 (e.g., the position of computer system 700 in FIG. 7E). The movement of computer system 700 changes the user's perspective. As described above, the representation of physical environment 704 corresponds to the portion of the physical environment visible from the user's perspective. Thus, when the user's perspective changes, the representation of physical environment 704 also changes accordingly. Further, as described above, media capture preview 708 corresponds to the field of view of one or more cameras that communicate with computer system 700. The movement of computer system 700 changes the field of view of one or more cameras. Thus, when the field of view of one or more cameras changes, the content displayed within media capture preview 708 also changes accordingly.
[0163] As shown in FIG. 7F, the computer system 700 displays the media capture preview 708 off-center (e.g., to the right, left, above, or below the center of the display 702). In FIG. 7F, the computer system 700 is moving within the physical environment at a first speed (e.g., 1 foot per second, 3 feet per second, 5 feet per second), and the computer system 700 displays the media capture preview 708 as if it were moving at a second speed slower than the first speed (e.g., a second speed relative to the physical environment) (e.g., if the computer system is moving at 5 feet per second, the media capture preview moves at 3 feet per second). The difference between the speed at which the computer system 700 moves within the physical environment and the speed at which the computer system 700 displays the media capture preview 708 as moving relative to the physical environment creates a lag visual effect that shows the media capture preview 708 lagging (e.g., trailing) the movement of the computer system 700. From a strict perspective of the display 702, the media capture preview 708 is moved on the screen in a direction opposite to the direction of movement of the computer system 700 so as to create the visual effect of the media capture preview 708 moving slower than the computer system 700 relative to the physical environment. In some embodiments, while the computer system 700 is moving, the representation of the physical environment 704 has a first set of parallax characteristics, and the computer system 700 displays the media capture preview 708 having a second set of parallax characteristics different from the first set of parallax characteristics. For example, in FIG. 7F, as the computer system 700 moves within the physical environment, there is a first amount of shift between the depictions of the first individual 709c3 and the second individual 709c4 and the depiction of the picture 709b1 within the media capture preview 708, and there is a second amount of shift between the depictions of the first individual 709c1 and the second individual 709c2 and the depiction of the picture 709b within the representation of the physical environment 704.The representation of the physical environment 704 has a different set of parallax characteristics than the media capture preview 708, so the first shift amount between the depiction of the first individual 709c3 and the second individual 709c4 within the media capture preview 708 and the depiction of the picture 709b1 is different from the shift amount between the depiction of the first individual 709c1 and the second individual 709c2 and the depiction of the picture 709b in the representation of the physical environment 704. In some embodiments, the media capture preview 708 is displayed without any parallax effect (e.g., the first shift amount is zero), while the representation of the physical environment 704 includes a non-zero degree of shift, such that the representation of the physical environment 704 is perceived as having depth, while the media capture preview 708 does not (e.g., appears flat). In some embodiments, while the computer system 700 is moving, the computer system 700 applies a first image stabilization technique to the representation of the physical environment 704, and the computer system 700 applies a second different stabilization technique to the display of the media capture preview 708. In some embodiments, while the computer system 700 is moving, the computer system 700 applies an amount of digital image stabilization to the representation of the physical environment 704 that is less than the amount of digital stabilization that the computer system 700 applies to the media capture preview 708.
[0164] In FIG. 7F, computer system 700 is moving horizontally to the right. Because computer system 700 is moving horizontally to the right, computer system 700 displays media capture preview 708 to the left of the center of display 702. In some embodiments, computer system 700 is moved horizontally to the left within the physical environment, whereby computer system 700 displays media capture preview 708 to the right of the center of display 702. In some embodiments, computer system 700 is moved vertically upward within the physical environment, whereby computer system 700 displays media capture preview 708 below the center of display 702. In some embodiments, computer system 700 is moved vertically downward within the physical environment, whereby computer system 700 displays media capture preview 708 above the center of display 702. In some embodiments, computer system 700 is moved forward (e.g., into the page) within the physical environment (e.g., in the z direction), whereby computer system 700 displays media capture preview 708 as larger than the representation of physical environment 704 for a period of time before transitioning the size of media capture preview 708 to have the same relative size to the representation of physical environment 704 as before the start of the movement. In such embodiments, the rate at which media capture preview 708 transitions between the two sizes lags behind the speed of the forward movement of computer system 700. In some embodiments, computer system 700 is moved backward (e.g., out of the page) within the physical environment, whereby the computer system displays media capture preview 708 as smaller than the representation of physical environment 704 for a period of time before transitioning the size of media capture preview 708 to have the same relative size to the representation of physical environment 704 as before the start of the movement. In such embodiments, the rate at which media capture preview 708 transitions between the two sizes lags behind the speed of the forward movement of computer system 700.In some embodiments, as the computer system 700 moves within a physical environment, the computer system 700 detects its movement with a first amount of tracking lag (e.g., measured as a function of distance over time). For example, while the computer system 700 is located at the position indicated by the indication 703 of the schematic diagram 701 of FIG. 7F, the computer system 700 detects that it has been positioned at the position indicated by the previous position indicator 717 of FIG. 7F for an amount of time corresponding to the first tracking lag. In some embodiments, a second speed at which the computer system 700 is displayed as moving the media capture preview 708 is configured to be less than the first amount of tracking lag in order to exhibit a lag visual effect (e.g., to reduce visual artifacts).
[0165] As shown in FIG. 7F, the schematic diagram 701 includes a previous position indicator 717. The previous position indicator 717 indicates a previous position of the computer system 700 (e.g., the position of the computer system 700 in FIG. 7E). In FIG. 7F, as shown by the schematic diagram 701 including a movement indication 721, the computer system 700 continues to move horizontally to the right within the physical environment.
[0166] In FIG. 7G, the computer system 700 is positioned further to the right within the physical environment than the position of the computer system 700 in FIG. 7F and has stopped moving. As shown in FIG. 7G, the schematic diagram 701 includes a previous position indicator 717 indicating a previous position of the computer system 700 (e.g., the position of the computer system 700 in FIG. 7F). The schematic diagram 701 also includes an indication 703 indicating the current location of the computer system 700. In FIG. 7G, since the computer system 700 is no longer moving in FIG. 7F, the schematic diagram 701 does not include a movement indicator 721.
[0167] As shown in FIG. 7G, computer system 700 displays media capture preview 708 to the left of the center of display 702. That is, in FIG. 7G, although computer system 700 is no longer moving, computer system 700 continues to display media capture preview 708 as if it were moving at a second speed (e.g., 3 feet per second when the computer system is moving at 5 feet per second) so that media capture preview 708 can “catch up” to computer system 700. In some embodiments, when computer system 700 stops moving, computer system 700 displays media capture preview 708 as if it were moving at a third speed (e.g., 5 feet per second) faster than the second speed so that media capture preview 708 is recentered more quickly than if it had been moving at the second speed.
[0168] In FIG. 7G, the representation of physical environment 704 is updated relative to the representation of physical environment 704 in FIG. 7H (e.g., the representation of physical environment 704 in FIG. 7G includes table 709e in the background and a portion of couch 709a in the foreground). As described above, the user's perspective changes based on changes in the field of view of one or more cameras communicating with computer system 700. Thus, as computer system 700 is moved throughout the physical environment (e.g., changing the field of view of one or more cameras communicating with computer system 700), the user's perspective changes and, accordingly, the representation of physical environment 704 changes. Further, in FIG. 7G, the display of media capture preview 708 is updated. As computer system 700 moves within the physical environment, the field of view of one or more cameras communicating with the computer system changes, thereby changing what is displayed within media capture preview 708.
[0169] In FIG. 7H, media capture preview 708 has caught up with a previous movement of computer system 700 (e.g., a movement of computer system 700 as described above in connection with FIGS. 7E - 7G). As shown in FIG. 7H, because media capture preview 708 has caught up with the previous movement of computer system 700, computer system 700 displays media capture preview 708 at the center of display 702. The display of media capture preview 708 on computer system 700 at the center of display 702 makes the first part of the representation of physical environment 704 (e.g., the torso part of the second individual 709c2) that was previously visible (e.g., visible in FIG. 7G) not visible (e.g., because media capture preview 708 is overlaid and displayed on top of the first part), and makes the second part of the representation of physical environment 704 (e.g., the torso of the first individual 709c1) that was not previously visible visible (e.g., because media capture preview 708 is no longer overlaid and displayed on top of the second part). In FIG. 7H, computer system 700 detects an input 750h directed towards photo well virtual object 715. In some embodiments, input 750h is a line - of - sight input (e.g., a sustained line of sight) directed towards the display direction of photo well virtual object 715. In some embodiments, input 750h is a tap on photo well virtual object 715 (e.g., an air tap within the space corresponding to the location of the display of the photo well virtual object). In some embodiments, input 750h is an air tap combined with the detection of a line of sight towards the display direction of photo well virtual object 715. In some embodiments, input 750h is a line of sight and a blink directed towards the display direction of photo well virtual object 715.
[0170] As shown in FIG. 7I, in response to detecting input 750h, computer system 700 displays previously captured media item 730. The previously captured media item 730 is the most recently captured media item captured by computer system 700 (e.g., the media item captured in FIGS. 7C and 7D). In some embodiments, the previously captured media item 730 is the most recently captured media item captured by an external device communicating with computer system 700 (e.g., a device separate from computer system 700). The previously captured media item 730 is displayed as a square. In some embodiments, the previously captured media item 730 is displayed as a rectangle, triangle, or any other suitable shape.
[0171] Computer system 700 displays previously captured media item 730 together with library virtual object 731, repositioning virtual object 732, deletion virtual object 733, identifier virtual object 734, sharing virtual object 735, and projection shape virtual object 736. In some embodiments, the virtual objects listed above are fixed to the display of previously captured media item 730 (e.g., as described above in the description of FIG. 7C). In some embodiments, library virtual object 731, repositioning virtual object 732, deletion virtual object 733, identifier virtual object 734, sharing virtual object 735, and projection shape virtual object 736 are displayed in a spatial configuration different from the spatial configuration shown in FIG. 7I. In some embodiments, in response to detecting input 750h, computer system 700 displays a user interface including a subset of the virtual objects shown in FIG. 7I.
[0172] By selecting the library virtual object 731, multiple representations of previously captured media items are displayed. In some embodiments, the library virtual object 731 is displayed simultaneously with the media capture preview 708. In some embodiments, the display of multiple representations of previously captured media items replaces the display of the previously captured media item 730. Repositioning the virtual object 732 enables the user to reposition the location of the display of the representation of the previously captured media item 730 in the same way that the media capture preview 708 can be repositioned using the repositioning virtual object 716. Selecting the erase virtual object 733 causes the computer system 700 to stop displaying the previously captured media item 730. In some embodiments, selecting the erase virtual object 733 causes the computer system 700 to stop displaying the library virtual object 731, the repositioning virtual object 732, the erase virtual object 733, the identifier virtual object 734, the share virtual object 735, and the projection shape virtual object 736. In some embodiments, selecting the erase virtual object 733 causes the computer system 700 to display the media capture preview 708. The identifier virtual object 734 provides an indication of where and when the previously captured media item 730 was captured. In some embodiments, the identifier virtual object 734 provides different information for the previously captured media item 730 (e.g., the resolution of the previously captured media item, the date on which the previously captured media item was captured). Selecting the share virtual object 735 initiates a process on the computer system 700 to share the previously captured media item 730 with an external device (e.g., a device separate from the computer system 700). In FIG. 7I, the computer system 700 detects an input 750i directed to the projection shape virtual object 736.In some embodiments, input 750i is a line-of-sight input directed in the direction of display of the projected shape virtual object 736. In some embodiments, input 750i is a tap input on the projected shape virtual object 736 (e.g., an air tap in space corresponding to the location of the display of the projected shape virtual object 736). In some embodiments, input 750i is an air tap input combined with detection of a line of sight in the direction of display of the projected shape virtual object 736. In some embodiments, input 750i is a line of sight and a blink directed in the direction of the display of the projected shape virtual object 736.
[0173] As shown in FIG. 7J, in response to detecting input 750i, computer system 700 displays a representation of the previously captured media item 730 as a circle. As shown in FIG. 7J, computer system 700 displays the previously captured media item 730 as a circle, but computer system 700 also displays library virtual object 731, repositioning virtual object 732, deletion virtual object 733, identifier virtual object 734, sharing virtual object 735, and projected shape virtual object 736 as described above in the description of FIG. 7I. In some embodiments, in response to detecting input 750i, computer system 700 displays the previously captured media item 730 as a three-dimensional sphere. As shown in FIG. 7I, the schematic diagram includes a movement indicator 721. Movement indicator 721 indicates that computer system 700 has begun to move (e.g., horizontally to the left) within the physical environment. In FIG. 7J, as shown by schematic diagram 701 including movement indication 721, computer system 700 begins to move horizontally to the left within the physical environment so as to return to the initial position of computer system 700 (e.g., the position of computer system 700 in FIG. 7A).
[0174] Figures 7K through 7Q illustrate ways to interact with a three-dimensional representation of a previously captured media item. Figures 7K through 7Q include an elapsed time indication 744. The elapsed time indication 744 is a visual aid that indicates the amount of time elapsed between each figure. The elapsed time indication 744 is not displayed by computer system 700. In Figure 7K, computer system 700 is located at an initial location within a physical environment (e.g., the location of computer system 700 in Figures 7A through 7E). Between Figures 7J and 7K, computer system 700 receives a request to display a spatial capture virtual object 740. In response to receiving a request to display spatial capture virtual object 740, as shown in Figure 7K, computer system 700 displays spatial capture virtual object 740 overlaid on a representation of physical environment 704. The display of spatial capture virtual object 740 obscures portions of the representation of physical environment 704. Spatial capture virtual object 740 is a representation of a previously captured video media item. In some embodiments, computer system 700 is a head-mounted device that presents spatial capture virtual object 740 as part of an extended reality environment, and while computer system 700 is in motion, spatial capture virtual object 740 is presented with a parallax effect (e.g., as discussed above with respect to Figure 7F) that results in a media item represented by spatial capture virtual object 740 having a certain amount of depth between the foreground and background of the content within the media item. In some embodiments, spatial capture virtual object 740 is a representation of a previously captured still media. In some embodiments, the previously captured video media item was captured using one or more cameras that communicate with computer system 700. In some embodiments, the request to display spatial capture virtual object 740 includes one or more inputs directed to one or more hardware buttons that communicate with computer system 700.In some embodiments, the request to display the spatial capture virtual object 740 includes one or more inputs directed to one or more virtual objects displayed by the computer system 700. In some embodiments, the request to display the spatial capture virtual object 740 includes detection of a line of sight (e.g., by the user) in a direction corresponding to one of the sub-representations of the previously captured media item. In embodiments where the computer system 700 is an HMD, the user's viewpoint is locked in the forward direction of the user's head, and thus the representation of the physical environment 704 and one or more virtual objects such as the spatial capture virtual object 740 shift as the user's head moves (e.g., because the computer system 700 also moves as the user's head moves).
[0175] The computer system 700 displays a spatial capture virtual object 740 at a first location within a representation of the physical environment 704 corresponding to a first location (e.g., the location indicated by the spatial capture indicator 740a in the schematic diagram 701) within a physical environment. The computer system 700 displays the spatial capture virtual object 740 when the first location of the physical environment is within the field of view of one or more cameras that communicate with the computer system 700. The display of the spatial capture virtual object 740 is locked to the environment at the first location within the representation of the physical environment 704. That is, the location of the display of the spatial capture virtual object 740 within the representation of the physical environment 704 does not change even if the real-world position of the computer system 700 changes. However, the display of the spatial capture virtual object 740 is updated as the user's viewpoint changes within the physical environment. For example, as the user's viewpoint approaches the first location within the physical environment, the computer system 700 displays the spatial capture virtual object 740 larger (e.g., to provide the visual effect that the user is approaching the spatial capture virtual object 740). In contrast, when the user's viewpoint moves away from the first location within the physical environment, the computer system 700 displays the spatial capture virtual object 740 smaller (e.g., to provide the visual effect that the user is moving further away from the spatial capture virtual object 740).
[0176] As shown in FIG. 7K, in response to receiving a request to display a spatial capture virtual object 740, computer system 700 also displays a plurality of sub - representations of previously captured media items 743 simultaneously with the display of the spatial capture virtual object 740. As shown in FIG. 7K, computer system 700 displays a plurality of sub - representations of previously captured media items 743 below the display of the spatial capture virtual object 740. Each sub - representation of the previously captured media represents a respective media item (e.g., video media or still media) previously captured by computer system 700 or an external device communicating with computer system 700. The sub - representation of the previously captured media item 743a, which is displayed at the center of the plurality of sub - representations of the previously captured media items 743, corresponds to the media item that is focused and displayed (e.g., represented by the spatial capture virtual object 740). In some embodiments, the user can navigate through the plurality of sub - representations of the previously captured media items 743 and select the sub - representation that is focused and displayed at the location of the spatial capture virtual object 740. In some embodiments, the user can switch which sub - representation is focused and displayed by performing a motion input (e.g., a pinch - and - drag gesture). In some embodiments, while the spatial capture virtual object 740 is being displayed, computer system 700 enlarges the size of the displayed spatial capture virtual object 740 in response to computer system 700 detecting (e.g., via one or more cameras communicating with computer system 700) that the user has performed a pinch gesture. In some embodiments, computer system 700 reduces the size of the display of the spatial capture virtual object 740 in response to computer system 700 detecting (e.g., via one or more cameras communicating with computer system 700) that the user has performed a de - pinch gesture.
[0177] As shown in FIG. 7K, schematic diagram 701 includes a spatial capture indicator 740a. The spatial capture virtual object 740 does not physically exist within the physical environment. Rather, the spatial capture virtual object 740 is a virtual object that is only displayed within the representation of the physical environment 704. The spatial capture indicator 740a indicates the spatial orientation / position of the display of the spatial capture virtual object 740 within the representation of the physical environment 704, and also indicates the location where the spatial capture virtual object 740 is environment-locked.
[0178] In FIG. 7K, while the spatial capture virtual object 740 is being displayed, the computer system 700 plays the video media item represented by the spatial capture virtual object 740. When the computer system 700 displays the spatial capture virtual object 740, the computer system 700 automatically (e.g., without intervening user input) starts playing the video media item represented by the spatial capture virtual object 740. In some embodiments, the playback of the video media item represented by the spatial capture virtual object 740 is not automatically played. In some embodiments, the computer system 700 outputs spatial audio as part of playing the video media item represented by the spatial capture virtual object 740. In some embodiments, the video media item represented by the spatial capture virtual object 740 is a stereoscopic video media item. A stereoscopic video media item presents two different images of the same scene to the user. In some embodiments, the first image is from a first camera and the second image is from a second camera, and the first camera and the second camera have slightly different perspectives of the scene. As a result, the first image varies slightly from the second image. When the two images are viewed by the user, they are superimposed on each other, generating an illusion of depth in the resulting image. In some embodiments, using the techniques described above in connection with FIG. 7J, the user can change the shape in which the computer system 700 displays the spatial capture virtual object 740. In some embodiments, the user can change the shape of the spatial capture virtual object 740 from a flattened stereoscopic projection to a spherical stereoscopic projection and vice versa. In some embodiments, the video media item represented by the spatial capture virtual object 740 is played on an external electronic device that cannot play stereoscopic video.In some embodiments, spatial audio is used when a video media item represented by the spatial capture virtual object 740 is played on an external electronic device that cannot play stereoscopic video. In some embodiments, in accordance with a determination that a user of the computer system 700 has an interpupillary distance different from the default interpupillary distance setting of the computer system 700, the computer system 700 plays the video media item with a shift (e.g., the computer system 700 plays the video media item represented by the spatial capture virtual object 740 at a scale different from the scale at which the video media item was captured). The interpupillary distance is the distance between the centers of the pupils of an individual's eyes (e.g., in millimeters). In some embodiments, in accordance with a determination that a user of the computer system 700 has an interpupillary distance different from the default interpupillary distance setting of the computer system 700, the computer system 700 changes (e.g., shifts) the first image and / or the second image of the stereoscopic video image based on the difference between the user's interpupillary distance and the interpupillary distance setting of the computer system 700 without changing the scale of the playback of the stereoscopic video media item.
[0179] The video media item represented by the spatial capture virtual object 740 is a video of an individual moving along a trajectory. In FIG. 7K, the computer system 700 plays the video media item represented by the spatial capture virtual object 740 from a non-immersive perspective view. The playback of content from a non-immersive perspective view cannot be presented from multiple perspective views in response to detected changes in the orientation / location of the computer system 700. When the video media is presented from a non-immersive perspective view, it is presented from only one perspective view regardless of whether the orientation / location of the computer system 700 changes. In FIG. 7K, the computer system 700 is located at a first distance (e.g., virtual distance) away from a first location within the physical environment where the computer system 700 displays the spatial capture virtual object 740 in the representation of the physical environment 704. In FIG. 7K, as shown by the schematic diagram 701 including the movement indication 721, the computer system 700 begins to move towards the first location of the physical environment.
[0180] As shown in FIG. 7L, the elapsed time indication 744 is shown as "00:01". Therefore, one second has elapsed between FIGS. 7K and 7L. In FIG. 7L, the computer system 700 continues to display the spatial capture virtual object 740 within the representation of the physical environment 704. Further, in FIG. 7L, the computer system 700 is positioned closer to a first location within the physical environment than the position of the computer system 700 in FIG. 7K. Therefore, in FIG. 7L, since the computer system 700 is located closer to the first location of the physical environment, the computer system 700 increases the size of the display of the spatial capture virtual object 740 (e.g., compared to the size of the display of the spatial capture virtual object 740 in 7K). As described above, as the user's perspective changes, the display of the spatial capture virtual object 740 changes based on the changing perspective of the user. The computer system 700 displays the spatial capture virtual object 740 larger, so less of the representation of the physical environment 704 is visible. In some embodiments, the computer system 700 moves backward within the physical environment, which decreases the size of the display of the spatial capture virtual object 740. In some embodiments, while the computer system 700 moves laterally in the physical environment, the representation of the physical environment 704 is presented with respective visual parallax effects, and the computer system 700 displays the spatial capture virtual object 740 with respective visual parallax effects.
[0181] In FIG. 7L, the computer system 700 continues to play the video media represented by the spatial capture virtual object 740. Thus, in FIG. 7L, the computer system 700 updates the representation of the video media represented by the spatial capture virtual object 740 (e.g., the representation of the video media shows the user as having progressed along a trajectory) to indicate that the playback of the video media has advanced by 1 second (e.g., the amount of time elapsed between FIGS. 7K and 7L). In FIG. 7L, as indicated by the movement indication 721 in the schematic diagram 701, the computer system 700 continues to move towards the first location in the physical environment.
[0182] As shown in FIG. 7M, the elapsed time indication 744 is shown as "00:05". Thus, 4 seconds have elapsed between FIGS. 7L and 7M. In FIG. 7M, the computer system 700 is positioned closer to the first location in the physical environment than the position of the computer system 700 in FIG. 7L. In FIG. 7M, since the computer system 700 is positioned closer to the first location in the physical environment, the computer system 700 increases the size of the display of the spatial capture virtual object 740 (e.g., compared to the size of the display of the spatial capture virtual object 740 in FIG. 7L). Since the computer system 700 displays the spatial capture virtual object 740 larger, less of the physical environment 704 is visible.
[0183] In FIG. 7M, computer system 700 continues to play the video media item represented by spatial capture virtual object 740. Thus, in FIG. 7M, computer system 700 displays the frame of the representation of the video media item 4 seconds after the frame of the video media shown in FIG. 7L, indicating that the playback of the video media has advanced 4 seconds. In some embodiments, computer system 700 stops displaying spatial capture virtual object 740 when computer system 700 moves past the first location in the physical environment. In some embodiments, computer system 700 is a head-mounted device that presents spatial capture virtual object 740 as part of an extended reality environment, and the video media item represented by spatial capture virtual object 740 becomes more immersive as computer system 700 approaches the first location within the physical environment (e.g., has a greater depth between foreground and background and / or greater responsiveness to shifts in the orientation of computer system 700). In some embodiments, in FIG. 7M, spatial capture virtual object 740 is displayed in a full-screen configuration. While spatial capture virtual object 740 is displayed in a full-screen configuration, the representation of physical environment 704 is visually obscured (e.g., the representation of physical environment 704 is blurred, dimmed, and / or painted black). In some embodiments, spatial capture virtual object 740 is displayed in a full-screen configuration, but the computer displays a different virtual environment (e.g., a virtual theater environment and / or a virtual drive-in environment) instead of the representation of physical environment 704. In FIG. 7M, as indicated by movement indication 721 in schematic 701, computer system 700 begins to rotate in a clockwise direction (e.g., 90 degrees clockwise) within the physical environment.
[0184] In FIG. 7N, relative to the user's perspective in FIG. 7M, the user's perspective has rotated 90 degrees clockwise. As described above, the movement of the computer system 700 changes the user's perspective, which changes the representation of the physical environment 704. Thus, in FIG. 7N, the representation of the physical environment 704 corresponds to the change in the user's perspective. In FIG. 7N, a first individual 709c1, a second individual 709c2, a photograph 709b, and a couch 709a, which are part of the physical environment, are not visible from the user's perspective in FIG. 7N. Thus, the representation of the physical environment 704 in FIG. 7N does not include depictions of the couch 709a, the picture 709b, the first individual 709c1, and the second individual 790c2. Rather, the representation of the physical environment 704 in 7N includes a depiction of the right side of the physical environment.
[0185] In FIG. 7N, a first location of the physical environment (e.g., the location where the computer system 700 displays the spatial capture virtual object 740 within the physical representation of the physical environment 704) is not within the field of view of one or more cameras that communicate with the computer system 700. Thus, in FIG. 7N, the computer system 700 does not display the spatial capture virtual object 740. In some embodiments, when the spatial capture virtual object 740 is displayed as a virtual environment configuration within a full-screen display that does not represent the physical environment 704 (e.g., as described above in connection with FIG. 7M), the computer system displays a perspective view of the virtual environment corresponding to a user's perspective that rotates 90° clockwise from the initial perspective view of the virtual environment. In FIG. 7N, as indicated by the movement indication 721 in the schematic 701, the computer system 700 begins to rotate counterclockwise (e.g., 90 degrees to the left).
[0186] In FIG. 7O, with respect to the user's perspective in FIG. 7N, the user's perspective has rotated 90 degrees counterclockwise. As shown in FIG. 7O, the elapsed time indication 744 is shown as "00:06". Therefore, one second has elapsed between FIGS. 7M and 7O. In FIG. 7O, the first location of the physical environment (e.g., the location where the computer system 700 displays the spatial capture virtual object 740 within the representation of the physical environment 704) is within the field of view of one or more cameras that communicate with the computer system 700. Therefore, as shown in FIG. 7O, the computer system 700 displays the spatial capture virtual object 740 within the representation of the physical environment 704. In FIG. 7O, the computer system 700 continues to play the video media represented by the spatial capture virtual object 740. Therefore, in FIG. 7O, the computer system 700 displays the frame of the video media item one second after the frame of the video media item shown in FIG. 7M, indicating that the playback of the video media item has advanced by one second.
[0187] In FIG. 7O, as shown by the schematic diagram 701 including the movement indication 721, the computer system 700 begins to move towards the first location of the physical environment.
[0188] In FIG. 7P, as indicated by the positioning of indication 703 within schematic diagram 701, computer system 700 is at a first location within the physical environment (e.g., the location where computer system 700 displays spatial capture virtual object 740 in the representation of physical environment 704). Since computer system 700 is positioned at the first location within the physical environment, computer system 700 plays video media represented by spatial capture virtual object 740 from an immersive perspective view. Immersive visual content is visual content that includes content of a plurality of perspective views captured from the same first point (e.g., location) within the physical environment at a given point in time. Reproduction (e.g., playback) of content from an immersive (e.g., first-person) perspective view includes playing content from a perspective view that matches the first viewpoint and can provide a plurality of different perspective views (e.g., fields of view) in response to user input. In some embodiments, while computer system 700 is at the first location within the physical environment, computer system 700 is moved to an updated position forward within the physical environment (e.g., beyond the first location within the physical environment), whereby the computer system displays spatial capture virtual object 740 from a non-immersive perspective view at a location within the representation of physical environment 704 corresponding to a location in front of the updated position of computer system 700 within the physical environment. In some embodiments, while computer system 700 is at the first location within the physical environment, computer system 700 is moved backward within the physical environment (e.g., moved such that computer system 700 is positioned in front of the first location within the physical environment), and computer system 700 is caused to display spatial capture virtual object 740 from a non-immersive perspective view at a location within the representation of physical environment 704 corresponding to the first location. Thus, by changing the position of computer system 700 within the physical environment, the user has the ability to control when computer system 700 displays media items from an immersive or non-immersive perspective view.
[0189] In FIG. 7P, the playback of the video media represented by the spatial capture virtual object 740 from the immersive perspective view in computer system 700 of FIG. 7O occupies the entire display 702. While computer system 700 is playing back the video media represented by the spatial capture virtual object 740 from the immersive perspective view, the representation of the physical environment 704 is not visible. In some embodiments, the representation of the physical environment 704 is visible while computer system 700 is playing back the video media represented by the spatial capture virtual object 740 from the immersive perspective view.
[0190] As shown in FIG. 7P, the time indication is shown as 00:07. Thus, one second has elapsed between FIGS. 7O and 7P. In FIG. 7P, computer system 700 continues to play back the video media represented by the spatial capture virtual object 740. Thus, in FIG. 7P, computer system 700 displays a frame of the video media that is one second before the frame of the video media shown in FIG. 7O in order to indicate that the playback of the video media has advanced by one second (e.g., the spatial capture virtual object 740 shows the user as having advanced along a trajectory). In some embodiments, in response to detecting that the computer system is at a first location within the physical environment, the computer system restarts the playback of the video media represented by the spatial capture virtual object 740 from the beginning. In FIG. 7P, as indicated by the movement indication 721 within the schematic view 701, computer system 700 begins to rotate clockwise (e.g., 90 degrees to the right).
[0191] In FIG. 7Q, the user's perspective is rotated 90 degrees clockwise in the physical environment with respect to the user's perspective in FIG. 7P. In FIG. 7Q, as indicated by indication 703 within schematic diagram 701, computer system 700 is located at a first location within the physical environment. As shown in FIG. 7Q, in response to the user's perspective being rotated 90 degrees clockwise in the physical environment, computer system 700 displays the playback of video media represented by spatial capture virtual object 740 from a perspective different from the perspective from which computer system 700 displays video media in FIG. 7P. More specifically, computer system 700 displays the video media from a perspective facing individually to the right side of the trajectory shown in the video media. As described above, while computer system 700 is positioned at the first location within the physical environment, computer system 700 displays video media represented by spatial capture virtual object 740 from an immersive perspective. That is, the perspective of the playback of the media item represented by spatial capture virtual object 740 changes based on the change in the user's perspective in the physical environment.
[0192] Additional explanations regarding FIGS. 7A - 7Q are provided below with reference to methods 800, 900, and 1000 described with respect to FIGS. 7A - 7Q.
[0193] FIG. 8 is a flow diagram of an exemplary method 800 for capturing media according to some embodiments. In some embodiments, method 800 is performed by a computer system (e.g., computer system 101 and / or computer system 700 of FIG. 1) (e.g., smartphone, tablet, and / or head-mounted device) that includes a display generation component (e.g., display generation component 120 of FIGS. 1, 3, and 4) (e.g., head-up display, display, display controller, touch-sensitive display system, display (e.g., integrated and / or connected), 3D display, transparent display, projector, touch screen, and / or projector), and one or more cameras (e.g., a camera pointing below the user's hand (e.g., color sensor, infrared sensor, and other depth sensing cameras), or a camera pointing forward from the user's head), and optionally, a physical input mechanism. In some embodiments, the computer system communicates with one or more gaze tracking sensors (e.g., optical and / or IR cameras configured to track the direction of the user's gaze and / or the user's attention of the computer system). In some embodiments, the first camera has a field of view (FOV) that is outside of at least a portion of the FOV of the second camera. In some embodiments, the second camera has a FOV that is outside of at least a portion of the FOV of the first camera. In some embodiments, the first camera is positioned on the side of the computer system opposite the side of the computer system where the second camera is positioned. In some embodiments, method 800 is stored on a non-transitory (or transitory) computer-readable storage medium and is managed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 of FIG. 1). Some operations of method 800 are optionally combined and / or the order of some operations is optionally changed.
[0194] The computer system displays a first user interface overlaid on a representation of a physical environment (e.g., 704) via a display generation component (e.g., 702), while the representation of the physical environment changes as a portion of the physical environment corresponding to the representation of the physical environment changes and / or the user's (e.g., 712) viewpoint changes (e.g., captured by one or more cameras communicating with the computer system) (e.g., a mixed reality and / or extended reality environment that is a representation of the physical environment), and detects (802) a request to display a media capture user interface (e.g., selection of 712a of 711a in FIG. 7B and / or 750b). In some embodiments, as part of receiving a request to display a media capture user interface, the computer system detects an input (e.g., press, swipe, and / or tap) on a hardware button. In some embodiments, as part of receiving a request to display a media capture user interface, the computer system detects the user's gaze and / or one or more spoken command inputs at one or more locations of the first user interface (e.g., via one or more microphones communicating with the computer system).
[0195] In response to detecting a request to display a media capture user interface (e.g., capturing immersive or semi-immersive visual media, capturing three-dimensional media, capturing three-dimensional stereo media, and / or capturing spatial media), the computer system, via a display generation component (e.g., after starting media capture), displays (804), along with content (e.g., a live preview that changes the field of view of one or more cameras and / or a change in the appearance of the physical environment) (e.g., immersive media content as described above in connection with FIG. 7K) that is updated as a portion of the physical environment in a portion of the field of view of one or more cameras changes, a media capture preview (e.g., 708) that includes a representation (e.g., a virtual representation and / or a virtual object) of a portion of the field of view of one or more cameras.
[0196] The media capture preview indicates (806) the boundary of the media to be captured in response to detecting a media capture input (e.g., activation of 711a of 712a in FIG. 7D and / or 750d) while the media capture user interface is being displayed.
[0197] The media capture preview is displayed while a first portion (e.g., the portion of 709b that includes 704) of a representation of the physical environment is visible (e.g., unobstructed) (808), where the first portion of the representation of the physical environment was visible (e.g., unobstructed) prior to detecting a request to display the media capture user interface (e.g., as described above in connection with FIG. 7C).
[0198] The media capture preview is displayed (e.g., overlaid thereon) instead of the second portion of the representation of the physical environment (e.g., the portion 709c1 including the portion 704) (as described above in connection with FIG. 7C), and the first portion of the representation of the physical environment is updated (810) when the portion of the physical environment corresponding to the first portion of the representation of the physical environment changes and / or the user's viewpoint changes (e.g., one or more other portions of the representation of the physical environment remain visible while the media capture preview is being displayed and / or is visible). In some embodiments, the computer system 700 displays a media capture preview having a black border. In some embodiments, in response to detecting a request to capture media, the computer system begins capturing media (e.g., three-dimensional media, three-dimensional, stereo media, and / or spatial media). In some embodiments, the first content captured by the first camera is different from the second content captured by the second camera. In some embodiments, the media capture preview is displayed at the center (e.g., the center) of the computer system (e.g., at the center of the display generation component), near the nose of the user wearing the computer system, at the center of the three-dimensional representation of the physical environment. In some embodiments, the media capture preview is displayed between one or more virtual objects (e.g., a time-lapse virtual object, a capture control virtual object, a camera roll virtual object, a close button virtual object). In some embodiments, the computer system displays the media capture preview at the center of the display above the shutter button virtual object (as described above in connection with FIG. 7C). In some embodiments, the captured media is played back using one or more techniques as described above in connection with FIGS. 7K-7Q. In some embodiments, the media capture preview includes content of the physical environment, and the content is visible in the representation of the physical environment before the media capture preview is displayed.In some embodiments, while playing back in a non-immersive perspective view, multiple perspective views from points within the physical environment other than the first point can be displayed in response to user input. Displaying a media capture preview that indicates the boundaries of the media to be captured in response to detecting a media capture input provides the user with improved feedback regarding which content at the user's viewpoint will be captured. Displaying the media capture preview while a portion of the representation of the physical environment is visible provides the user with the ability to better compose and capture the desired media while also maintaining recognition of the physical environment, which improves the media capture operation and reduces the risk of missing a transient event that might be captured if the capture operation were inefficient or difficult to use. Improving the media capture operation improves the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).
[0199] In some embodiments, the representation of the physical environment (e.g., 704 of FIG. 7B) is a pass-through representation of the real-world environment of the computer system (e.g., the portion of the real-world environment surrounding the computer system) (e.g., as seen in FIG. 7A) (e.g., the physical environment (e.g., outside the computer system)), where the pass-through representation is virtual (e.g., a representation of camera image data captured by one or more cameras integrated into the computer system) and / or a light pass-through (e.g., light that passes directly through a portion of the system (e.g., a transparent portion) to the user). In some embodiments, in response to a change in the position and / or orientation of the computer system, the pass-through representation is updated to reflect the change in the position and / or orientation of the computer system. Providing a pass-through representation of the real-world environment of the computer system provides visual feedback to the user regarding the location (e.g., position and / or orientation) of the computer system within the real-world environment, which provides improved visual feedback, particularly when the pass-through representation is visible while a media capture preview is being displayed.
[0200] In some embodiments, the representation of the physical environment (e.g., 704) (e.g., the content within the representation) is visible (e.g., displayed or viewable via an optical passthrough) at a first scale (e.g., 1:1, 1:2, 1:4, and / or any other suitable scale), and the representation of a portion of the field of view of one or more cameras included in the media capture preview (e.g., 708) (e.g., a second portion of the representation of the physical environment) (e.g., the content within the representation) is displayed at a second scale, and the first scale is larger than the second scale (e.g., as described above in connection with FIG. 7C) (e.g., the same object that appears in both representations appears larger within the representation of the physical environment). Displaying the media capture preview at a scale smaller than the scale at which the representation of the physical environment is visible provides the user with improved visual feedback regarding which representation is which, reducing user confusion. Doing so also saves display space and enables both to be viewable together, which provides the user with the ability to better compose and capture the desired media while maintaining recognition of the physical environment, which reduces the risk of missing a transient event that could be captured when the media capture operation is inefficient or difficult to use, improving the media capture operation, improving the system's operability, and making the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).
[0201] In some embodiments, the representation of the physical environment (e.g., 704 in FIG. 7G) (e.g., individual parts) includes first content (e.g., the right half of 709E in FIG. 7G) that is not included in the representation of the field of view included in the media capture preview (e.g., 708 in FIG. 7G) (e.g., outside of it) (e.g., not displayed within the media capture preview). Having a representation of the physical environment that includes content not included in the media capture preview provides the user with the ability to view additional content that may be included in (but is not currently included in) the media capture when the user shifts their perspective, and also provides the user with a greater awareness of their current physical environment while composing the media capture. Doing so improves the media capture operation and reduces the risk of missing transient events and / or content that may be missed when the capture operation is inefficient or difficult to use. Improving the media capture operation improves the system's operability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).
[0202] In some embodiments, the representation of the physical environment (e.g., 704 in FIG. 7C) includes a portion of the physical environment (e.g., a first portion, a particular portion) (e.g., different from the first portion of the representation of the physical environment (e.g., the angular range of the second portion of the physical environment is narrower than the angular range of the first portion of the physical environment)) (e.g., 709c1 in FIG. 7C), and the representation of the field of view of one or more cameras included in the media capture preview (e.g., 708 in FIG. 7C) includes a portion of the physical environment (e.g., 709c3 in FIG. 7C) (e.g., the media capture preview includes an object that is also included in the representation of the physical environment). In some embodiments, the second portion of the physical environment included in the representation of the physical environment has a different visual appearance (e.g., blur and / or dimming) from the second portion of the physical environment included in the media capture preview.
[0203] In some embodiments, the computer system communicates (e.g., directly (e.g., wired communication) and / or wirelessly) with a physical input mechanism (e.g., 711a or 711b) (e.g., a hardware button) (e.g., a hardware input device / machine) (e.g., a physical input device), and a request to display a media capture user interface includes activation (e.g., actuation and / or selection (e.g., pressing a button)) of the physical input mechanism (e.g., activation of 712a of 711a in FIG. 7B). In response to detecting the activation of the physical input mechanism, displaying a media capture preview that includes a representation (e.g., a virtual representation and / or a virtual object) of a portion of the field of view of one or more cameras enables the computer system to perform a display operation that provides the user with greater control over the computer system without displaying additional controls, which provides additional control options without cluttering the user interface.
[0204] In some embodiments, while displaying a media capture preview, the computer system detects an input (e.g., activation of 750c in FIG. 7C and / or 711a of 712a) corresponding to a request to capture media (e.g., via one or more input devices communicating with the computer system). In response to detecting an input corresponding to a request to capture media, according to a determination that the input corresponding to the request to capture media is of a first type (e.g., a single quick press of a hardware button (e.g., a virtual shutter button), a tap gesture on a touch-sensitive surface, or a quick air gesture (e.g., an air gesture having a shorter duration than a second type of input described below) (e.g., pinch and release or any other suitable air gesture as described above with respect to selection of a virtual object in an XR environment)), the computer system starts a process to capture a first type of media content (e.g., still media (e.g., a photo)) (e.g., using one or more cameras of the computer system) as described above in connection with FIG. 7C. According to a determination that the input corresponding to the request to capture media is of a second type (e.g., different from the first type) (e.g., a long press of a button, a sustained line of sight directed at the display of a virtual shutter object, a touch-and-hold gesture at a location in space corresponding to the display of a virtual shutter object, a touch-and-hold gesture on a virtual shutter button, a sustained air gesture (e.g., a pinched hold), or any other suitable air gesture as described above with respect to selection of a virtual object in an XR environment), the computer system starts a process to capture a second type of media content (e.g., video) (e.g., different from the first type of media content) (e.g., using one or more cameras of the computer system) as described above in connection with FIG. 7C.Based on whether an input type of a short press or a long press (e.g., short press or long press) is received, starting either a first type of media capture process or a second type of media capture process enables a computer system to provide a user with additional control options regarding the type of media capture process that the computer system executes without displaying additional controls, which provides additional control options without confusing the user interface.
[0205] In some embodiments, the representation of the portion of the field of view of one or more cameras included in the media capture preview (e.g., 708) has a first set of visual parallax characteristics, and the representation of the physical environment has a second set of visual parallax characteristics that are different from the first set of visual parallax characteristics (e.g., as described above in connection with FIG. 7F). In some embodiments, the media capture preview includes a first foreground and a first background. In some embodiments, as the portion of the physical environment within the portion of the field of view of one or more cameras changes, there is a first shift amount between the first foreground and the first background. In some embodiments, the representation of the physical environment includes a second foreground and a second background. In some embodiments, as the portion of the physical environment within the portion of the field of view of one or more cameras changes, there is a second shift amount (e.g., different from the first shift amount) between the second foreground and the second background (e.g., as the portion of the physical environment within the portion of the field of view of one or more cameras changes, the first foreground moves at a first speed relative to the first background, and the second foreground moves at a second speed (different from the first speed) relative to the second background). Providing a first set of visual parallax characteristics to the representation of the portion of the field of view of one or more cameras included in the media capture preview and providing a second set of visual parallax characteristics to the representation of the physical environment provides visual feedback to the user regarding whether the computer system is in motion, provides feedback regarding which representation is the preview of the media capture and which representation is the preview of the physical environment, and provides improved visual feedback.
[0206] In some embodiments, the representation of the physical environment (e.g., 704) is an immersive perspective view (e.g., first-person perspective) (e.g., the representation of the environment is presented from a plurality of perspective views in response to detecting a change in the orientation of the user and / or computer system), and the representation of the portion of the field of view of one or more cameras included in the media capture preview is a non-immersive perspective view (e.g., third-person perspective) (e.g., as described above with respect to FIG. 7C). In some embodiments, the immersive perspective view includes the content of the representation of the environment that includes the content of a plurality of perspective views captured from the same point (e.g., location) within the environment. Providing a representation of the physical environment from the immersive perspective view and providing a portion of the field of view of one or more cameras included in the media capture preview from the non-immersive perspective view provides improved visual feedback regarding which representation is which, and also provides feedback regarding how the captured media can be previewed when it is displayed non-immersively (e.g., in a non-immersive photo album).
[0207] In some embodiments, before detecting a request to display a media capture user interface, a third portion of a representation of the physical environment (e.g., the whole of the representation of the environment), which may be the same as or different from the first portion (e.g., a first portion of the representation of the environment that is less than the whole of the representation of the environment), has a first visual appearance that includes visual characteristics (e.g., dimming, blurring, and / or overlay filters) having a first magnitude (e.g., no amount, a low non-zero amount such as 0%, 10%, or 20% of a maximum amount). In some embodiments, in response to detecting a request to display a media capture user interface (e.g., selection of 750b in FIG. 7B and / or 711a of 712a), the computer system changes a third portion of the representation of the physical environment to have a second visual appearance that includes visual features having a second magnitude (e.g., different from the first visual appearance), as described above in connection with FIG. 7C. In some embodiments, the second magnitude of the visual characteristics is different from the first magnitude of the visual characteristics (e.g., greater than or less than the first magnitude), as described above in connection with FIG. 7C. In some embodiments, the magnitude of the visual characteristics is the amount of dimming applied to the third portion of the representation, and before receiving the request, the first magnitude is such that no dimming is applied and the second amount is a non-zero level of dimming applied. Changing the visual appearance of the third portion of the representation of the physical environment in response to detecting a request to display a media capture user interface provides visual feedback to the user regarding the state of the computer system (e.g., that the computer system has detected a request to display a media capture user interface), and also emphasizes the media capture preview and provides improved visual feedback.
[0208] In some embodiments, displaying a media capture preview includes displaying the media capture preview (e.g., 708) and a spatial relationship (e.g., the spatial relationship includes a predetermined distance and relative orientation at which one or more virtual objects are displayed relative to the display of the media capture preview, and / or a position (e.g., top, bottom, and / or side) at which one or more virtual objects are displayed relative to the display of the media capture preview) (e.g., each virtual object of the one or more virtual objects has an individual spatial relationship with the media capture preview) of one or more virtual objects (e.g., 713, 714, 715, and 716) (e.g., virtual objects that, when selected, cause the computer system to perform operations (e.g., media-related operations such as capturing a photo, capturing a video, displaying a previously captured photo, and / or any other suitable operation)), including (e.g., being displayed as part of it) (e.g., being incorporated therein), where the display of the one or more virtual objects surrounds the display of the media capture preview, the media capture preview and the one or more virtual objects are displayed at a first display location (e.g., the locations of 713, 714, 715, and 716 in FIGS. 7C - 7E), and while the media capture preview and the one or more virtual objects are being displayed at the first display location, the computer system detects a change in the pose of the user's viewpoint (e.g., via one or more sensors in communication with the computer system) (as described above with respect to FIG. 7E). In some embodiments, in response to detecting a change in the pose of the user's viewpoint, the computer system displays the media capture preview and the one or more virtual objects at a second display location that is different from the first location (e.g., at a different location on the display generation component) (e.g., the locations of 708, 713, 714, 715, and 716 in FIG. 7F), and the computer system maintains the spatial relationship between the display of the media capture preview and the one or more virtual objects.In response to detecting a change in the pose of the viewpoint, maintaining the spatial relationship between the display of the media capture preview and one or more virtual objects provides the user with visual feedback that enables the user to easily identify the display position of the one or more virtual objects, which provides improved visual feedback. Displaying one or more virtual objects based on the display of the computer system's media capture preview causes the computer system to perform a display operation that provides the user with additional control options related to the media capture without requiring further user input, which reduces the number of inputs required to perform the operation.
[0209] In some embodiments, the one or more virtual objects include a time-elapsed virtual object (e.g., 713) that provides an indication of the time (e.g., seconds, minutes, and / or hours) elapsed since the start of a process (e.g., video recording) for capturing media. Displaying a time-elapsed virtual object that indicates the amount of time elapsed since the computer system started a process for capturing media provides visual feedback regarding the video media recording state of the computer system and provides improved visual feedback.
[0210] In some embodiments, one or more virtual objects include a shutter button virtual object (e.g., 714) (e.g., a software shutter button), and when the shutter button virtual object is selected (e.g., selected via detection of the user's line of sight (e.g., line of sight and dwell) directed at the shutter button virtual object, and in some embodiments, in combination with detection of the user performing one or more gestures (e.g., an air pinch gesture, a de-pinch air gesture, an air tap, and / or an air swipe) (as described above with reference to selection of virtual objects in an XR environment), and / or via detection of a tap on the shutter button virtual object), it causes the initiation of a process for capturing media (e.g., causes a computer system to initiate a process for capturing media (e.g., still media or video media)) (e.g., using one or more cameras communicating with the computer system to capture media). In some embodiments, the display of the shutter button virtual object is updated to indicate that the computer system is recording video media. Displaying a shutter button virtual object fixed to the display of the media capture preview provides visual feedback regarding the display state of the computer system (e.g., the computer system is currently displaying a media capture preview), and also provides an action that can be performed by interacting with the object, which provides improved visual feedback.
[0211] In some embodiments, when one or more virtual objects are selected (e.g., when selected via detection of a user's line of sight (e.g., line of sight and dwell) directed at a camera roll virtual object), they include a camera roll virtual (e.g., 715) object. In some embodiments, in combination with detection of a user performing one or more gestures (e.g., an air pinch gesture, a de-pinch air gesture, an air tap, and / or an air swipe) (e.g., as described above with reference to selection of virtual objects in an XR environment), and / or when selected via detection of a tap on the camera roll virtual object, a display of a previously captured media item is caused (e.g., via a display generation component) (e.g., to cause a computer system to display a previously captured media item) (e.g., still media or video media) (e.g., a previously captured media item using one or more cameras communicating with the computer system) (e.g., a previously captured media item using an external device (e.g., a smartphone) (e.g., a device separate from the computer system)). In some embodiments, in response to detecting a selection of the camera roll virtual object, the computer system displays a user interface that includes a subset of one or more virtual objects. Displaying a camera roll virtual object that is fixed to the display of the media capture preview provides visual feedback regarding the display state of the computer system (e.g., the computer system is currently displaying a media capture preview), and also provides an action that can be performed by interacting with the object, which provides improved visual feedback.
[0212] In some embodiments, while displaying a media capture preview, the computer system detects a selection of a camera roll virtual object (e.g., 750h). In some embodiments, in response to detecting a camera roll virtual object (e.g., 715 in FIG. 7H), the computer system displays, via a display generation component, a representation of a previously captured media item (e.g., 730) (e.g., by stopping the display of the media capture preview). In some embodiments, displaying a representation of previously captured media is a first deletion virtual object (e.g., 733), which, when selected (e.g., selected via detection of a user's line of sight (e.g., line of sight and dwell) directed at the location of the display of the first deletion virtual object, and in some embodiments, in combination with detection of a user performing one or more gestures (e.g., an air pinch gesture, a de-pinch air gesture, an air tap, and / or an air swipe), and / or selected via a user tapping on the first deletion virtual object (e.g., as described above with reference to selection of virtual objects in an XR environment)), causes the display of the representation of the previously captured media item to be stopped (e.g., causes the computer system to stop displaying the previously captured media item), a share virtual object (e.g., 735), which, when selected (e.g., selected via detection of a user's line of sight (e.g., line of sight and dwell) directed at the location of the display of the share virtual object, and in some embodiments, when selected in combination with detection of a user performing one or more gestures (e.g., an air pinch gesture, a de-pinch air gesture, an air tap, and / or an air swipe) (e.g., as described above with reference to selection of virtual objects in an XR environment), and / or selected via a user tapping on the share virtual object), causes the start of a process for sharing the representation of the previously captured media item (e.g., sharing the representation with contacts (e.g., external users) stored in the computer system) (e.g.,Cause a computer system to initiate a process for sharing a representation of a previously captured media item), a shared virtual object, a media library virtual object (e.g., 731), which when selected (e.g., selected via detection of a user's line of sight (e.g., line of sight and dwell) directed at the location of the display of the media library virtual object, and in some embodiments, in combination with detection of a user performing one or more gestures (e.g., air pinch gesture, depinch air gesture, air tap, and / or air swipe) (e.g., as described above with reference to selection of virtual objects in an XR environment), and / or selected via a user tap on the media library virtual object)), causes the display of a plurality of previously captured media items (e.g., cause the computer system to display a plurality of previously captured media items (e.g., non-immersive and / or immersive media items)), a media library virtual object, and / or (e.g., as described above with reference to selection of virtual objects in an XR environment) display a resize virtual object (e.g., 736) indicating that the size of the representation of the previously captured media item changes (e.g., increases in size or decreases in size) based on detection of one or more gestures (e.g., gestures on a touch-sensitive display and / or air gestures (e.g., air pinch and drag gesture, air swipe, and / or air tap)) (e.g., detected by one or more cameras communicating with the computer system) (e.g., displayed simultaneously). In some embodiments, the magnitude of the size change is based on the characteristics of the gesture (e.g., the amount of size change is based on the distance of the drag operation of the pinch and drag gesture). In some embodiments, selection of a first erasure virtual object causes the media capture preview to be redisplayed. Displaying a plurality of virtual objects based on the display of a previously captured media item provides visual feedback regarding the display state of the computer system (e.g., the computer system is currently displaying a representation of a previously captured media item),It also provides operations that can be performed by interacting with the object, which provides improved visual feedback.
[0213] In some embodiments, a computer system receives a set of one or more inputs that includes an input corresponding to a camera roll virtual object (e.g., 715). In some embodiments, the set of one or more inputs is a selection of a camera roll virtual object and a subsequent selection of a virtual object corresponding to a subset of previously captured media items (e.g., all media items captured from an immersive perspective view). In some embodiments, in response to receiving the set of one or more inputs, the computer system displays a representation (e.g., a still image and / or video) of a first previously captured media item (e.g., 730 of FIGS. 7I and 7J) of the plurality of previously captured media items at a first location (e.g., a central location on a display generation component) (e.g., the first representation of the first previously captured media item is selectable) (e.g., the first previously captured media item is captured by the computer system), and while the representation of the first previously captured media item is being displayed at the first location, the computer system receives a request to navigate to a different previously captured media item of the plurality of previously captured media items (e.g., an air pinch gesture, a de-pinch air gesture, an air tap, and / or an air swipe) (e.g., as described above with respect to selection of virtual objects in an XR environment). In some embodiments, in response to receiving the request to navigate to a different previously captured media item of the plurality of previously captured media items, the computer system replaces the display of the representation of the first previously captured media item at the first location with a display of a representation of a second previously captured media item (e.g., different from the first previously captured media item) of the plurality of previously captured media items (e.g., as described above in connection with FIG. 7K).In some embodiments, replacing the display of a representation of a first previously captured media item includes ceasing to display the first previously captured media item. In some embodiments, while a second previously captured media item is being displayed at a first location, the first previously captured media item is displayed (e.g., at a second location different from the first location). In some embodiments, the representation is displayed instead of a portion of the representation of the physical environment. In some embodiments, while a second previously captured media item is being displayed, the computer system receives a second request to navigate to a different (e.g., different from the first previously captured media item and the second previously captured media item) previously captured media item, the second request being a repetition of the initial request (e.g., the initial request and the second request include the same type of gesture), and / or the second request includes a gesture that is the opposite of the gesture included in the initial request (e.g., performed in the opposite direction). In some embodiments, a request to navigate to a previously captured media item among a plurality of previously captured media items is an air gesture (as described above with reference to the selection of virtual objects in an XR environment), and the second previously captured media item is selected by the computer system based on the magnitude of the air gesture (the direction of the air gesture, the speed of the air gesture, and / or the intensity of the air gesture). Replacing the display of the representation of the first previously captured media item at the first location with the display of the representation of the second previously captured media item provides visual feedback regarding the state of the computer system (e.g., that the computer system received a request to navigate to a different previously captured media item while the first previously captured media item was being displayed), which provides improved visual feedback.
[0214] In some embodiments, one or more virtual objects include a second deletion virtual object (e.g., 719), and when the second deletion virtual object is selected (e.g., selected via detection of a user's line of sight (e.g., line of sight and dwell) directed at the display location of the second deletion virtual object, and in some embodiments, in combination with detection of a user performing one or more gestures (e.g., an air pinch gesture, a de-pinch air gesture, an air tap, and / or an air swipe) as described above with reference to selection of virtual objects in an XR environment, and / or selected via a user tap on the second deletion virtual object), it causes the display of the media capture preview (e.g., 708) to be aborted (e.g., causes the computer system to abort the display of the media capture preview). In some embodiments, aborting the display of the media capture preview results in the display of a portion (e.g., a second portion) of the representation of the environment that was not displayed while the media capture preview was being displayed. Displaying the second deletion virtual object fixed to the display of the media capture preview provides visual feedback regarding the display state of the computer system (e.g., the computer system is currently displaying the media capture preview), and also provides an action that can be performed by interacting with the object, which provides improved visual feedback.
[0215] In some embodiments, one or more virtual objects include a repositioning virtual object (e.g., 716), and while a media capture preview (e.g., 708) is displayed at a first location within a media capture user interface, the computer system detects a set of one or more inputs (e.g., gesture(s) and / or air gesture(s) on a touch-sensitive surface) that include an input corresponding to the repositioning virtual object. In some embodiments, in response to detecting a set of one or more inputs that include an input corresponding to the repositioning virtual object, the computer system moves the media capture preview from the first location to a second location (e.g., the media capture preview is moved along the x-axis, y-axis, and / or z-axis of the media capture user interface, based on the direction, magnitude, and / or speed of the request to move the display of the repositioning virtual object). In some embodiments, the set of one or more inputs includes an input that selects the repositioning virtual object and one or more inputs that specify a subsequent target location and / or direction of movement. Displaying a repositioning object fixed to the display of the media capture preview provides visual feedback regarding the display state of the computer system (e.g., the computer system is currently displaying the media capture preview), and also provides an action that can be performed by interacting with the object, which provides improved visual feedback.
[0216] In some embodiments, while a media capture preview (e.g., 708) is being displayed, the computer system simultaneously displays, via a display generation component, a media library virtual object (e.g., 731), and the media library virtual object, when selected (e.g., selected via detection of a user's line of sight (e.g., line of sight and dwell) directed at the location of the display of the media library virtual object, and in some embodiments, in combination with detection of a user performing one or more gestures (e.g., air pinch gesture, de-pinch air gesture, air tap, and / or air swipe) (e.g., as described above with reference to selection of virtual objects within an XR environment), and / or via a user tap on the media library virtual object), causes the display of a plurality of previously captured media items (e.g., as described above in connection with FIG. 7C), causing the computer system to display a plurality of previously captured media items. In some embodiments, selection of the media library virtual object causes the display of the media library virtual object to be aborted. In some embodiments, selection of the media library virtual object causes the display of the media capture preview to be aborted. In some embodiments, selection of the media library virtual object causes a plurality of previously captured media items to be displayed below, above, and / or beside the display of the media capture preview. Simultaneously displaying a media library virtual object (e.g., displayed while a media capture preview is being displayed) provides visual feedback regarding the display state of the computer system (e.g., the computer system is currently displaying a media capture preview), and also provides an action that can be performed by interacting with the object, which provides improved visual feedback.
[0217] In some embodiments, a portion of the field of view of one or more cameras included in a media capture preview (e.g., 708) (e.g., the horizontal and / or vertical fields of view of one or more cameras) has a first viewing angle range (e.g., 0-45°, 0-90°, 40-180°, or any other suitable angular range) (e.g., the portion of the field of view of one or more cameras included in the media capture preview is based on the overlap of the fields of view of one or more cameras) (e.g., the media capture preview includes only data from the overlapping portions of the fields of view of one or more cameras) (e.g., the media capture preview does not include data from the non-overlapping portions of the fields of view of one or more cameras), the representation of the physical environment (e.g., 704) represents a second field of view of one or more cameras having a second viewing angle range (e.g., 0-45°, 0-90°, 40-180°, or any other suitable angular range) (e.g., the second field of view includes non-overlapping data from one or more cameras), and the first angular range is narrower (e.g., smaller) than the second angular range (e.g., the second angle is wider than the first angular range) (e.g., as described above in connection with FIG. 7C). In some embodiments, the first angular range is a subset of the second angular range. Displaying a representation of the physical environment that includes a wider field of view than the media capture preview provides the user with the ability to view additional content that may be included in (but is not currently included in) the media capture preview when the user shifts their viewpoint, and also provides the user with a greater awareness of their current physical environment while configuring the media capture. Doing so improves the media capture operation and reduces the risk of missing transient events and / or content that may be missed if the capture operation is inefficient or difficult to use. Improving the media capture operation improves the system's operability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).
[0218] In some embodiments, the representation of a portion of the field of view of one or more cameras included in a media capture preview (e.g., 708) is of a first content (e.g., 709c3 and 709c4 in FIG. 7F) (e.g., three-dimensional content) within the field of view of a first camera of the one or more cameras (e.g., content captured by the first camera in response to detecting a request to capture media, and / or content stored and / or remembered by the computer system), and (e.g., as described above in FIG. 7C) the field of view of a second camera of the one or more cameras (e.g., content captured by the second camera in response to detecting a request to capture media and / or content stored and / or remembered by the computer system). In some embodiments, a portion of the physical environment that is within the field of view of the first camera but not within the field of view of the second camera is not included in the media capture preview. In some embodiments, a portion of the physical environment that is within the field of view of the first camera but not within the field of view of the second camera is not included in the media capture preview but is part of the representation of the physical environment.
[0219] In some embodiments, aspects / operations of methods 800, 900, 1000, 1200, 1400, and 1500 may be exchanged, substituted, and / or added among these methods. For example, a media item captured in method 800 may be displayed as part of method 1000. For the sake of brevity, their details are not repeated here.
[0220] Figure 9 is a flow diagram of an exemplary method 900 for displaying a preview of media, according to some embodiments. In some embodiments, method 900 communicates with a display generation component (e.g., display generation components 120 and / or 702 of FIGS. 1, 3, and 4) (e.g., a display controller, a touch-sensitive display system, a display (e.g., integrated and / or connected), a 3D display, a transparent display, a projector, a heads-up display, and / or a head-mounted display) and one or more cameras (e.g., a camera pointing below the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head) in a computer system (e.g., computer system 101 and / or computer system 700 of FIG. 1) (e.g., a smartphone, a tablet, and / or a head-mounted device). In some embodiments, the computer system communicates with one or more gaze-tracking sensors (e.g., optical and / or IR cameras configured to track the direction of the user's gaze and / or the user's attention on the computer system). In some embodiments, the computer system includes a first camera. In some embodiments, method 900 is stored on a non-transitory (or transitory) computer-readable storage medium and is managed by instructions executed by one or more processors of the computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 of FIG. 1). Some operations of method 900 are optionally combined and / or the order of some operations is optionally changed.
[0221] The computer system displays an extended reality user interface including a preview (e.g., a real-time preview) of the field of view of one or more cameras (e.g., 708 of 7E) (e.g., virtual cameras or physical cameras configured to communicate with the computer system) overlaid on a first portion of a three-dimensional environment (e.g., 704) (e.g., an environment visible on a display generation component) that is visible from the perspective of a user (e.g., 712) (e.g., and / or the computer system) while the perspective of the user (e.g., 712) (e.g., and / or the computer system) is in a first pose (e.g., orientation and / or position within a physical environment and / or virtual environment) (e.g., and / or while the presence of the user's head is detected). The preview includes a representation (e.g., a three-dimensional representation, a spatial representation, and / or a two-dimensional representation) of the first portion of the three-dimensional environment and is displayed in an individual spatial configuration relative to the user's perspective (e.g., 708 of FIGS. 7C-7E) (902). In some embodiments, the field of view of the first camera is a portion of the field of view of the first camera and not the entire field of view of the first camera. In some embodiments, the preview includes a representation of at least a portion of the field of view of a second camera (e.g., different from the first camera) that includes a representation of a portion of the physical environment. In some embodiments, the field of view of the second camera is a portion of the field of view of the second camera. In some embodiments, the preview is displayed at the center (e.g., the center) of the computer system (e.g., at the center of the display generation component), near the nose of the user wearing the computer system, and at the center of the three-dimensional representation of the physical environment.
[0222] The computer system detects (904) a change in the pose of the user's perspective (and / or the user's head) from a first pose to a second pose different from the first pose (e.g., as described above in connection with FIGS. 7E-7F). In some embodiments, while the computer system and / or the user's perspective is in the second pose, the field of view of the first camera is directed towards a second portion of the physical environment and / or the second portion of the physical environment is visible in the user's perspective while the first camera is in the second pose (e.g., while displaying an extended reality environment user interface including a preview of the field of view of a first camera overlaid on a first location on the extended reality user interface).
[0223] In response to detecting a change in the pose of the user's perspective from a first pose to a second pose, the computer system shifts (e.g., pans, changes, and / or updates for) the preview of the field of view of one or more cameras away from an individual spatial configuration (e.g., 708 in FIGS. 7F and 7G) with respect to the user's perspective in a direction determined based on the change in the pose of the user's perspective from the first pose to the second pose (e.g., based on the direction and / or speed of the change in the pose of the perspective). The shift of the preview of the field of view of one or more cameras is performed at a first speed (e.g., 708 in FIGS. 7F and 7G). While the preview of the field of view of one or more cameras is shifting based on the change in the pose of the user's perspective, the representation of the three-dimensional environment changes (906) based on the change in the pose of the user's perspective at a second speed (e.g., continuing to update the preview at the same rate as the movement of the perspective is updated) that is different from the first speed (e.g., faster or slower than the first speed), as described above with respect to FIG. 7F. In some embodiments, the preview is updated to include a representation of a second portion of the environment at the first speed while being moved at the second speed. In some embodiments, displaying one or more objects within the representation of the field of view is updated at the same rate as one or more objects within the perspective of the physical environment. In some embodiments, in response to detecting a change in the orientation of the user's perspective from a first orientation to a second orientation, a first portion of the environment becomes non-visible. In some embodiments, the first portion of the environment is not encompassed by, surrounds, and is separated from a second portion of the environment, is different from the second portion of the environment, includes the second portion of the environment, and / or vice versa. In some embodiments, the locations included in the first portion of the environment are different from the locations included in the first portion of the environment.Shifting the preview of the field of view of one or more cameras is based on a change in the pose of the user's perspective at a first speed, while the representation of the three-dimensional environment changes based on a change in the pose of the user's perspective at a second speed different from the first speed (e.g., in response to detecting a change in the pose of the user's perspective from a first pose to a second pose), enabling the computer system to automatically perform operations that reduce the amount of movement between elements presented to the user and reduce the likelihood of dizziness, which performs the operation when a set of conditions is met without requiring further user input. Doing so can also prompt the user to reduce changes in perspective while in the media capture mode, since the user may tend to maintain focus on the preview. Reducing changes in perspective while capturing media can improve the media capture operation and reduce the risk of missing transient events and / or not being able to capture content that may be missed if the capture operation is inefficient or difficult to use. Improving the media capture operation improves the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).
[0224] In some embodiments, a computer system (e.g., 700) tracks changes in the pose of a user's (e.g., 712) viewpoint with a first amount of tracking lag (e.g., the delay between an actual change in viewpoint and the computer system being configured to track the amount / degree of its movement (e.g., a delay of 0.1, 0.2, 0.3, 0.4, or 0.5 meters per second)), and the first speed introduces an amount of visual delay in updating the location of a preview of the field of view of one or more cameras (e.g., 708) that is greater than the amount of visual delay in updating the location of a preview of the field of view of one or more cameras based on (e.g., solely based on) a first amount of detection tracking lag (e.g., as described above in connection with FIG. 7F) (e.g., the detection tracking lag is 0.1 meters per second and the first speed is 0.09 meters per second). In some embodiments, the maximum speed of a shift of a preview of the field of view of one or more cameras is limited (e.g., capped) to a value that is less than the amount of tracking lag (e.g., the current amount). In some embodiments, the amount of tracking lag is proportional to the rate of the pose of the viewpoint). Maintaining the first speed below the amount of tracking lag reduces the occurrence of confusion in the display of the preview caused by the tracking lag and automatically adjusts the speed of movement of the preview based on system parameters without requiring further user input.
[0225] In some embodiments, while the preview of the field of view of one or more cameras is shifting at a first rate (e.g., 708 in FIGS. 7F and 7G), the computer system detects that the user's viewpoint is changing poses by less than a first threshold amount (e.g., changing at a rate less than the threshold or not changing (e.g., a stationary pose) (e.g., the user is currently positioned in a second pose)). In some embodiments, in response to detecting that the user's viewpoint is changing poses by less than a threshold amount (e.g., as described above in connection with FIG. 7G), the computer system shifts the preview of the field of view of one or more cameras towards an individual spatial configuration at a third rate that is faster than the first rate (e.g., in response to detecting that the user has stopped moving, the preview of one or more cameras "snaps" back to its original position having an individual spatial configuration relative to the user's viewpoint). In some embodiments, the computer system shifts the preview of the field of view of one or more cameras in the opposite direction (e.g., opposite the direction of the shift away from the individual spatial configuration) when the computer system shifts the preview of the field of view of one or more cameras towards an individual spatial configuration. In response to detecting that the user's viewpoint is changing poses by less than a threshold amount, shifting the preview of the field of view of one or more cameras towards an individual spatial configuration causes the computer system to display the preview of the field of view at a location on the display that is the center of the user's field of view, which performs the operation without requiring further user input.
[0226] In some embodiments, while the preview of the field of view of one or more cameras is shifting at a first rate (e.g., 708 in FIGS. 7F and 7G), the computer system detects that the user's perspective is changing poses (e.g., changing at a rate less than a second threshold amount or not changing (e.g., a static pose)) (e.g., a non-zero amount) by less than a second threshold amount (e.g., as described above in connection with FIG. 7G). In some embodiments, in response to detecting that the user's perspective is changing poses by less than a second threshold amount, the computer system stops shifting the preview of the field of view of one or more cameras away from an individual spatial configuration (e.g., as described above in FIG. 7H) (e.g., automatically (e.g., without user input)), and the computer system displays a preview of the field of view of one or more cameras having an individual spatial configuration with respect to the user's perspective (e.g., 708 in FIG. 7H) (e.g., the preview of the field of view of one or more cameras is displayed at the position where the preview of the field of view was displayed before the computer system detected the movement of the change in the pose of the user's perspective) (e.g., the preview of the field of view of one or more cameras is displayed overlaid on a first portion of the three-dimensional environment). In some embodiments, after stopping the shift of the preview of the field of view of one or more cameras, the preview of the field of view of one or more cameras shifts away from the individual spatial configuration a second time at a rate different from the rate at which the representation of the three-dimensional environment changes in response to the computer system detecting that the user's perspective is changing poses a second time at a rate greater than the second threshold amount. Stopping the shift of the preview of the field of view of one or more cameras in response to detecting that the user's perspective is changing poses by less than a second threshold amount causes the computer system to display the preview of the field of view at the location on the display that is the center of the user's field of view, which performs the operation without requiring further user input.
[0227] In some embodiments, the change in pose (e.g., as described above in connection with FIG. 7F) includes a lateral movement along the plane of the user's perspective (e.g., corresponding thereto) (e.g., the change in pose of the user's perspective corresponds to the user's perspective moving left and / or right from the user's perspective). Shifting the preview of the field of view of one or more cameras based on the change in pose of the user's perspective in the lateral direction provides visual feedback regarding the direction in which the computer system is moving, which provides improved visual feedback.
[0228] In some embodiments, the change in pose (e.g., as described above in connection with FIG. 7F) includes a longitudinal movement along the plane of the user's perspective (e.g., corresponding thereto) (e.g., the change in pose of the user's perspective corresponds to the user's perspective moving up and / or down). In some embodiments, while the user is in a first pose, the user's head is located at a first distance from the ground, and while the user is in a second pose, the user's head is located at a second distance greater / smaller than the first distance from the ground). Shifting the preview of the field of view of one or more cameras based on the change in pose of the user's perspective in the longitudinal direction provides visual feedback regarding the direction in which the computer system is moving, which provides improved visual feedback.
[0229] In some embodiments, the pose change (e.g., as described above in connection with FIG. 7F) includes a forward and backward movement perpendicular to the plane of the user's view (e.g., corresponding thereto) (e.g., detecting a change in the pose of the user's view corresponds to the user's view moving forward and / or backward). In some embodiments, while the user's view is in the first pose, the computer system moves backward within the physical environment, which causes the computer system to stop displaying a preview of the field of view of one or more cameras. Shifting the preview of the field of view of one or more cameras in the forward and backward directions based on a change in the pose of the user's view provides visual feedback regarding the direction in which the computer system is moving, which provides improved visual feedback.
[0230] In some embodiments, the preview of the field of view of one or more cameras (e.g., 708) does not overlay a second portion of the three-dimensional environment (e.g., the portion of 704 that includes the second individual 709c2 in FIG. 7F) that is visible from the user's view. In some embodiments, the three-dimensional physical environment is the real-world environment in which the user is currently located. In some embodiments, the camera preview covers at least a portion of the view of the physical world, while at least a portion of the view of the physical world is displayed near or adjacent to the edge of the camera preview. Displaying a preview of the field of view of one or more cameras while at least a portion of the three-dimensional environment is visible provides the user with the ability to better configure and capture the desired media while also maintaining awareness of the physical environment, which improves the media capture operation and reduces the risk of missing a transient event that may be missed if the capture operation is inefficient or difficult to use. Improving the media capture operation improves the operability of the system and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).
[0231] In some embodiments, the representation of the three-dimensional environment included in the media capture preview (e.g., 708) changes based on changes in the pose of the user's viewpoint (as described above in connection with, e.g., FIG. 7G), and while detecting a change in the pose of the user's viewpoint from a first pose to a second pose (as described above in connection with, e.g., FIG. 7F), the computer system, as described above in connection with, e.g., FIG. 7F, performs a first visual stabilization on the representation of the three-dimensional environment (e.g., automatically (e.g., without user input) (e.g., the first visual stabilization corresponds to digital image stabilization)) (e.g., does not perform a first visual stabilization on the representation of the three-dimensional environment) to average (e.g., attenuate) unconscious head movements (e.g., head moments made by the user that do not correspond to a change in the pose of the user's viewpoint from a first pose to a second pose) included in the preview of the field of view of one or more cameras. In some embodiments, the first visual stabilization is an optical image stabilization technique that adjusts one or more glass elements within one or more cameras communicating with the computer system based on changes in the pose of the user's viewpoint. In some embodiments, the first visual stabilization includes a digital stabilization technique that includes zooming and / or cropping the representation of the three-dimensional environment included in the media capture preview based on changes in the pose of the user's viewpoint. Applying the first visual stabilization to the representation of the three-dimensional environment included in the preview of the field of view of one or more cameras causes the computer system to automatically perform a display operation that increases the sharpness of the representation of the three-dimensional environment included in the field of view of one or more cameras that the user sees without requiring user input, which reduces the number of inputs required to perform the operation. Applying the first visual stabilization to the representation of the three-dimensional environment improves the media capture operation and reduces the risk of not being able to capture transient events that may be missed when the capture operation is inefficient or difficult to use.Improving media capture operations improves system operability and makes the user-system interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device).
[0232] In some embodiments, performing the first stabilization includes applying a first amount of visual stabilization to a representation of the three-dimensional environment included in a preview of the field of view of one or more cameras (e.g., 708), while detecting a change in the pose of the user's viewpoint from a first pose to a second pose (e.g., as described above in connection with FIGS. 7E-7G), the computer system performs a second visual stabilization (e.g., the second visual stabilization corresponds to digital visual stabilization) on a second portion of the representation of the three-dimensional environment (e.g., a representation of the three-dimensional environment not included in the preview of the field of view of one or more cameras) (e.g., via the computer system) (e.g., automatically (e.g., without user input)) (e.g., does not perform the second visual stabilization on the preview of the field of view of one or more cameras). In some embodiments, the second visual stabilization applies a second amount of visual stabilization that is less than the first amount of visual stabilization to the second portion of the representation of the three-dimensional environment (e.g., as described above in connection with FIG. 7F) (e.g., the representation of the three-dimensional environment included in the preview of the field of view of one or more cameras is sharper (e.g., sharper) than the representation of the three-dimensional environment). In some embodiments, the second visual stabilization is the same process (e.g., method) as the first visual stabilization. In some embodiments, the second visual stabilization is performed while the first visual stabilization is being performed. Applying the second visual stabilization to the second portion of the representation of the three-dimensional environment while detecting a change in the pose of the user's viewpoint causes the computer system to automatically perform a display operation that increases the sharpness with which the user views the second portion of the representation of the three-dimensional environment without requiring user input, which reduces the number of inputs required to perform the operation. Applying less stabilization to the representation of the three-dimensional environment provides more accurate visual feedback with respect to the current physical environment, which improves the visual feedback.
[0233] In some embodiments, aspects / operations of methods 800, 900, 1000, 1200, 1400, and 1500 may be exchanged, substituted, and / or added among these methods. For example, the media capture preview shown in method 800 may be optionally shifted using the method described in method 900. For the sake of brevity, their details are not repeated here.
[0234] FIG. 10 is a flowchart of an exemplary method 1000 for displaying previously captured media according to some embodiments. In some embodiments, method 1000 is executed on a computer system (e.g., computer system 101 and / or computer system 700 of FIG. 1) (e.g., smartphone, tablet, and / or head-mounted device) that communicates with a display generation component (e.g., the display generation component 120 of FIGS. 1, 3, and 4 and / or FIG. 702) (e.g., display controller, touch-sensitive display system, display (e.g., integrated and / or connected), 3D display, transparent display, projector, head-up display, and / or head-mounted display). In some embodiments, the computer system communicates with one or more eye-tracking sensors (e.g., an optical camera and / or an IR camera configured to track the direction of the gaze of a user of the computer system). In some embodiments, method 1000 is stored in a non-transitory (or transitory) computer-readable storage medium and is controlled by instructions executed by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., control 110 of FIG. 1). Some operations of method 1000 may optionally be combined and / or the order of some operations may optionally be changed.
[0235] While displaying an extended reality environment user interface, the computer system provides a first set of visual cues (e.g., visual indications of depth, visual indications of perspective views, visual responses to shifts in the orientation of the user's perspective) such that when viewed from individual ranges of one or more perspectives, the user is at least partially surrounded by content (e.g., immersive or semi-immersive visual media, three-dimensional media, three-dimensional stereo media, and / or spatial media) (e.g., video content that can be presented from multiple perspectives in response to detected changes in the orientation of the user and / or the computer system). Detect (1002) a request to display captured media (e.g., as described above in connection with FIG. 7K) that includes immersive content (e.g., 730). In some embodiments, the media content is immersive or semi-immersive visual media. In some embodiments, the immersive or semi-immersive visual media is visual media that includes content for a plurality of perspective views captured from the same first point (e.g., location) within the physical environment at a given point in time. In some embodiments, playing (e.g., playing back) visual media from an immersive (e.g., first-person) perspective view includes playing media content from a perspective view that coincides with the first point within the physical environment at a given point in time, and can provide a plurality of different perspective views (e.g., fields of view) from all points within the physical environment in response to user input. In some embodiments, playing visual media from a non-immersive (e.g., third-person) perspective view includes playing media content from a perspective view other than the first point within the physical environment (e.g., a shift in the perspective view (e.g., occurring in response to user input)). In some embodiments, while playing in a non-immersive perspective view, a plurality of perspective views corresponding to the first point within the physical environment are not displayed in response to user input.In some embodiments, when a portion of a request to play captured media is detected, the computer system detects the user's attention (e.g., the user's line of sight and / or the user's perspective) to one or more locations on the computer system (e.g., one or more locations corresponding to one or more virtual objects), detects one or more inputs to one or more hardware input mechanisms positioned on and / or coupled to the computer system, and / or detects one or more voice commands directed to playing the captured media. In some embodiments, the media is captured using one or more techniques as described above in connection with FIGS. 7C-7D. In some embodiments, the captured media is not immersive (e.g., not immersive visual media and / or not semi-immersive visual media). In some embodiments, non-immersive video media is video content that cannot be presented from multiple perspectives in response to detected changes in the orientation of the user and / or the computer system. In some embodiments, non-immersive video media can be presented in only one perspective (e.g., a first-person perspective or a third-person perspective), regardless of whether the computer system detects a change in the orientation of the computer system and / or the user of the computer system.
[0236] In response to detecting a request to display captured media, the computer system causes the captured media to be displayed (1004) (displaying the captured media being played back) as a three-dimensional representation (e.g., 740) (e.g., a non-static representation) of the captured media at a location (e.g., the location of 740a within 701) within the three-dimensional environment selected by the computer system such that the user's first viewpoint (e.g., 712) is outside the individual range of one or more viewpoints (e.g., a non-immersive perspective view) (e.g., a third-person perspective view). In some embodiments, the three-dimensional representation of the captured media replaces one or more portions of a representation of a previously displayed (e.g., displayed prior to receiving a request to play back the captured media) environment (e.g., a virtual environment and / or a physical environment) (e.g., using one or more techniques as described above in connection with FIG. 7C). In some embodiments, the computer system causes the three-dimensional representation of the captured media to be displayed on top of a shutter button virtual object and / or at the center of a display generation component (e.g., as described above in connection with FIG. 7C). In some embodiments, in response to detecting a request to display captured media, the computer system causes a plurality of representations of a previously captured media item to be displayed simultaneously with the three-dimensional representation of the captured media. In some embodiments, in response to receiving a request to play back the captured media, the computer system causes the display of a representation of the captured media (e.g., a static representation) to be aborted. Displaying the captured media at a location that satisfies a set of defined conditions as a three-dimensional representation (e.g., the location is outside the individual range of one or more viewpoints) automatically enables the computer system to present the three-dimensional representation from a non-immersive perspective view without user input, which reduces the number of inputs required to perform an operation. Doing so also provides improved visual feedback that the initially displayed viewpoint is not within the individual range of one or more viewpoints.
[0237] In some embodiments, the location selected by the computer system is a location within a physical environment (e.g., the location of 740a at 701), the three-dimensional representation of the captured media is an environment-locked virtual object, and while displaying the three-dimensional representation of the captured media, the computer system detects that the user's viewpoint has changed (e.g., as described above in connection with FIGS. 7K - 7P) (e.g., the user moves left / right, up / down, and / or forward / backward within the real-world environment) (e.g., the user looks up, down, right, or left) (e.g., the field of view of one or more cameras integrated with the computer system changes) (e.g., detected via one or more cameras integrated with the computer system) (e.g., detected via an external device (e.g., a smartwatch, one or more cameras, and / or an external computer system) communicating (e.g., wirelessly) with the computer system). In some embodiments, in response to detecting that the user's viewpoint has changed, the computer system maintains the display of the three-dimensional representation of the captured media at the location within the three-dimensional environment selected by the computer system (e.g., as described above in connection with FIG. 7K). In some embodiments, the three-dimensional representation is environment-locked. In some embodiments, in response to detecting that the user's viewpoint has changed, the size of the three-dimensional representation of the captured media changes (e.g., the three-dimensional representation of the captured media is displayed larger and / or smaller). In some embodiments, in response to detecting that the user's viewpoint has changed, the three-dimensional representation of the captured media is no longer displayed. Maintaining the display of the three-dimensional representation of the captured media at the location within the three-dimensional environment selected by the computer system in response to detecting that the user's viewpoint has changed allows the computer system to automatically perform a display operation that moves and maintains the view of the three-dimensional representation of the captured media without the user needing to provide user input, which reduces the number of inputs required to perform the operation.
[0238] In some embodiments, the user's first perspective corresponds to a first perspective location that is at a first distance (e.g., 10 meters, 5 meters, or 3 meters) from a location selected by the computer system, and while displaying a three-dimensional representation of media captured at the location selected by the computer system (e.g., the location 740 in FIG. 7K) and while the user is at the first perspective location, the computer system detects a change in the pose of the user's perspective with respect to a second perspective of the user that corresponds to a second perspective location that is at a second distance (e.g., 5 meters, 3 meters, or 1 meter) from the location selected by the computer system (e.g., a repositioning of the user's entire body, a repositioning of a part of the user's body (e.g., head, hand, arm, and / or leg), a repositioning of the user's head) (e.g., the change in the user's positioning is detected by one or more cameras and / or external devices of the user as described above), and the second distance is less than the first distance (e.g., as described above in FIG. 7L). In some embodiments, in response to detecting a change in the pose of the user's perspective with respect to the user's second perspective, the computer system displays a three-dimensional representation having a second set of visual cues (e.g., 740 in FIG. 7L) that the user is at least partially surrounded by the content. In some embodiments, the second set of visual cues includes at least a second visual cue that the user is at least partially surrounded by content that is not provided while displaying the three-dimensional representation when viewed from the user's first perspective (e.g., 740 in FIG. 7L). In some embodiments, as the user's perspective approaches the location selected by the computer system, the three-dimensional representation of the captured media becomes more immersive (although less immersive than when the user's perspective is within the range of one or more perspectives that provide the first set of visual cues).In some embodiments, the second set of visual cues does not include at least a first visual cue (e.g., a visual indication of depth, a visual indication of a perspective view, a visual response to a shift in the orientation of the user's viewpoint) that the user is at least partially surrounded by the content included in the first set of visual cues. In some embodiments, the second set of visual cues does not include at least a first visual cue (e.g., a visual indication of depth, a visual indication of a perspective view, a visual response to a shift in the orientation of the user's viewpoint) that the user is at least partially surrounded by the content included in the first set of visual cues. As the user approaches a location selected by the computer system, by displaying a three-dimensional representation using different sets of visual cues, the computer system can automatically perform a display operation that enables the user to perceive the three-dimensional representation of the captured media differently based on the location of the user's viewpoint relative to the location selected by the computer, without requiring user input, and this operates without requiring further user input. As the user approaches a location selected by the computer system, displaying a three-dimensional representation using different sets of visual cues provides the user with visual feedback regarding what actions the user needs to take such that the three-dimensional representation is displayed from an immersive perspective view, providing improved visual feedback.
[0239] In some embodiments, while the user is at a second viewpoint location and while a three-dimensional representation of the captured media is being displayed, the computer system detects a change in the pose of the user's viewpoint with respect to a third viewpoint corresponding to a third viewpoint location (e.g., as described above in connection with FIG. 7M). In some embodiments, in response to detecting a change in the pose of the user's viewpoint with respect to the user's third viewpoint, a first set of display criteria that includes a first criterion that is satisfied when the distance between the third viewpoint location and a location selected by the computer system is greater than a first threshold distance (e.g., 0.1 meter, 0.5 meter, 1 meter, 2 meters, 3 meters) (e.g., the user passes a location selected by the computer system), the computer system stops displaying the three-dimensional representation of the captured media (e.g., as described above in connection with FIG. 7M). In some embodiments, the change in pose from the user's second viewpoint to the user's third viewpoint is in the same direction (e.g., along the same plane) (along the same path) as the change in pose from the user's first viewpoint to the user's second viewpoint (e.g., the directional component of the vector for the change in pose from the user's first viewpoint to the user's second viewpoint is the same as the directional component of the vector for the change in pose from the user's second viewpoint to the user's third viewpoint). In some embodiments, in response to detecting a change in the pose of the user's viewpoint with respect to the user's third viewpoint, the computer system maintains the display of the three-dimensional representation of the captured media according to a determination that the first set of display criteria is not satisfied. In some embodiments, the computer system stops displaying the three-dimensional representation of the captured media when the distance from the third viewpoint location to a location selected by the computer system is greater than the first threshold distance.In some embodiments, the first set of display criteria includes criteria that are met when a location selected by the computer system is not visible from a third perspective of the user (e.g., the location selected by the computer system is no longer in front of the user). Ceasing to display the three-dimensional representation according to a determination that the distance between the third perspective location and the location selected by the computer system is greater than a first threshold distance provides visual feedback regarding the location (e.g., the position of the computer system relative to the position selected by the computer system), which provides improved visual feedback.
[0240] In some embodiments, while the user is at the second viewpoint location and while the three-dimensional representation of the captured media is being displayed, the computer system detects a change in the pose of the user's viewpoint with respect to the user's fourth viewpoint corresponding to the fourth viewpoint location (e.g., as described above in connection with FIG. 7P). In some embodiments, in response to detecting a change in the pose of the user's viewpoint with respect to the user's fourth viewpoint, a second set of display criteria, including a second criterion that is satisfied when the distance between the fourth viewpoint location and the location selected by the computer system is less than a second threshold distance (e.g., 0.1 meter, 0.5 meter, 1 meter, 2 meters, or 3 meters) (e.g., the user is moving towards the location selected by the computer system), is satisfied. According to the determination that the second set of display criteria is satisfied, the computer system displays the three-dimensional representation of the captured media at an individual location selected by the computer system that is farther from the fourth viewpoint location than the location selected by the computer system (e.g., as described above in connection with FIG. 7P) (e.g., the three-dimensional representation of the captured media is moved away from the user's current location). In some embodiments, the individual location is a location within the physical environment that is visible from the user's fourth viewpoint. In some embodiments, the second set of display criteria includes a criterion that is satisfied when the location selected by the computer system is not visible from the user's fourth viewpoint (e.g., the location selected by the computer system is no longer in front of the user). In some embodiments, the individual location is a location that is visible from the user's fourth viewpoint. In some embodiments, the computer system displays the three-dimensional representation of the captured media at the individual location selected by the computer system when the distance from the fourth viewpoint location to the location selected by the computer system is greater than the second threshold distance.In some embodiments, in accordance with a determination that the viewpoint location has moved past the location of the three-dimensional representation (e.g., when the distance between the fourth viewpoint location and the location selected by the computer system is greater than (or less than) a third threshold distance), the computer system displays the three-dimensional representation of the captured media at an individual location selected by the computer system that is farther from the fourth viewpoint location than from the location selected by the computer system. In some embodiments, the change in pose from the user's second viewpoint to the user's fourth viewpoint is in the same direction as the change in pose from the user's first viewpoint to the user's second viewpoint (e.g., the directional components of the vector of the change in pose from the user's first viewpoint to the user's second viewpoint are the same as the directional components of the vector of the change in pose from the user's second viewpoint to the user's fourth viewpoint). Displaying the three-dimensional representation of the captured media at an individual location selected by the computer system that is farther from the fourth viewpoint location than from the location selected by the computer system when certain specified conditions are met automatically changes the display of the three-dimensional representation so that the three-dimensional representation of the captured media is readily viewable by the user, which performs an operation when a set of conditions is met without requiring further user input. Displaying the three-dimensional representation of the captured media at an individual location captured from a location farther from the fourth viewpoint location than from the location selected by the computer system provides visual feedback to the user regarding the location of the computer system (e.g., the position of the computer system relative to the position selected by the computing system), which provides improved visual feedback.
[0241] In some embodiments, the user's first perspective corresponds to a fifth perspective location that is a fifth distance from the location selected by the computer system, and displays a three-dimensional representation of the media captured at the location selected by the computer. While the user is at the fifth perspective location, the computer system detects a change in the pose of the user's perspective with respect to the user's sixth perspective corresponding to a sixth perspective location that is a sixth distance from the location selected by the computer system (e.g., repositioning of the user's entire body, repositioning of a part of the user's body (e.g., head, hand, arm, leg), repositioning of the user's head) (e.g., the change in the user's positioning is detected by one or more cameras and / or external devices of the user as described above) (e.g., as described above in connection with FIG. 7P). In some embodiments, in response to detecting a change in the pose of the user's perspective to the user's sixth perspective, the computer system displays a three-dimensional representation having a third set of visual cues that the user is at least partially surrounded by the content. In some embodiments, as the user's perspective moves further from the location selected by the computer system, the three-dimensional representation of the captured media becomes less immersive (although less immersive than when the user's perspective is within the range of one or more perspectives that provide the first set of visual cues). In some embodiments, the third set of visual cues does not include at least a third visual cue that the user is at least partially surrounded by the content included in the first set of visual cues. In some embodiments, the third set of visual cues does not include at least a fourth visual cue that the user is at least partially surrounded by the content provided while displaying the three-dimensional representation when viewed from the...
Claims
1. A method, comprising: in a computer system communicating with a display generation component, one or more cameras, and one or more input devices, displaying, via the display generation component, a user interface, wherein the user interface includes: a representation of a physical environment, wherein a first portion of the representation of the physical environment is inside a capture region of the one or more cameras and a second portion of the representation of the physical environment is outside the capture region of the one or more cameras; and a viewfinder including a boundary; while displaying the user interface, detecting, via the one or more input devices, a first request to capture media; in response to detecting the first request to capture media, capturing, using the one or more cameras, a first media item including at least the first portion of the representation of the physical environment; changing an appearance of the viewfinder, the changing of the appearance of the viewfinder including: changing an appearance of a first portion of content within a threshold distance of a first side of the boundary of the viewfinder while maintaining a display of at least a portion of the representation of the physical environment within the viewfinder; and changing an appearance of a second portion of content within a threshold distance of a second side of the boundary of the viewfinder, different from the first side of the boundary of the viewfinder, while maintaining a display of at least a portion of the representation of the physical environment within the viewfinder. A method according to claim 1, wherein the first media item is a stereoscopic media item. A method according to claim 1, wherein prior to detecting the first request to capture media, the viewfinder is displayed in a first appearance, and changing the appearance of the viewfinder includes displaying the viewfinder in a second appearance different from the first appearance, and the method further includes changing the appearance of the viewfinder from the second appearance to the first appearance after the viewfinder has been displayed in the second appearance for a period of time. The method according to claim 1, further comprising
4. The method according to claim 1, wherein the appearance of the first portion of the content and the appearance of the second portion of the content are changed in the same manner.
5. Changing the appearance of the viewfinder includes changing the appearance of a third portion of the content that is within a threshold distance on a third side of the boundary of the viewfinder, and the appearance of the first portion of the content, the second portion of the content, and the third portion of the content are changed in the same manner. The method according to claim 4.
6. The boundary of the viewfinder is a reticle virtual object, and before the first requirement for capturing media is detected, the reticle virtual object is displayed with a first appearance, and the method The method according to claim 1, further comprising changing the appearance of the reticle virtual object from the first appearance to a second appearance different from the first appearance in response to detecting the first requirement for capturing media.
7. Displaying the viewfinder includes displaying one or more elements within the viewfinder, and changing the appearance of the reticle virtual object includes changing the appearance of the reticle virtual object with respect to the one or more elements displayed within the viewfinder. The method according to claim 6.
8. Displaying the user interface includes displaying one or more corners, and changing the appearance of the viewfinder Changing the appearance of the one or more corners in a first manner; and Changing the appearance of at least the first side of the boundary of the viewfinder in a second manner different from the first manner. The method according to claim 6.
9. Changing the appearance of the viewfinder includes changing a first set of one or more optical characteristics of a first portion of the content within the viewfinder. The method according to claim 1.
10. The first set of one or more optical characteristics includes the contrast of the content within the viewfinder. The method according to claim 9.
11. The method according to claim 9, wherein the first set of one or more optical properties includes the luminance of the content in the viewfinder.
12. The method according to claim 9, wherein the first set of one or more optical properties includes the translucency of the content in the viewfinder.
13. The method according to claim 9, wherein the first set of the one or more optical properties includes the size of the content in the viewfinder.
14. The method according to claim 1, wherein displaying the user interface includes displaying a first set of virtual control objects within the boundary of the viewfinder.
15. The first set of virtual objects includes a first proximity virtual object, and the method further includes: detecting an input corresponding to the selection of the first proximity virtual object; aborting the display of the user interface in response to detecting the input corresponding to the selection of the first proximity virtual object. The method according to claim 14.
16. The first set of virtual objects includes a first media review virtual object, and the method further includes: detecting an input corresponding to the selection of the first media review virtual object; displaying one or more representations of previously captured media items in response to detecting the input corresponding to the selection of the first media review virtual object. The method according to claim 14.
17. The first set of virtual objects includes a recording time virtual object indicating the amount of time elapsed since the computer system started a first video capture operation. The method according to claim 14.
18. The method according to claim 1, wherein displaying the user interface includes displaying a second set of virtual control objects outside the boundary of the viewfinder.
19. The second set of virtual control objects includes a set of camera mode virtual objects respectively corresponding to the respective operation modes of the one or more cameras. The set of camera mode virtual objects includes a first camera mode virtual object and a second camera mode virtual object, and the method includes: Detecting an input corresponding to the selection of an individual camera mode virtual object within the set of the camera mode virtual objects; In response to detecting the input corresponding to the selection of the individual camera mode virtual object within the set of the camera mode virtual objects, Configuring the one or more cameras to operate in a first mode according to a determination that the input corresponds to the selection of the first camera mode virtual object; Configuring the one or more cameras to operate in a second mode according to a determination that the input corresponds to the selection of the second camera mode virtual object; The method according to claim 18, further comprising.
20. The second set of virtual control objects includes a second proximity virtual object, and the method includes Detecting an input corresponding to the selection of the second proximity virtual object; Aborting the display of the user interface in response to detecting the input corresponding to the selection of the second proximity virtual object; The method according to claim 18, further comprising.
21. Displaying a representation of the first media item that fades into the display of the user interface after the first media item is captured; The method according to claim 1, further comprising.
22. The displaying of the representation of the first media item that fades into the display of the user interface is Displaying the representation of the first media item transitioning from a first size to a second size, wherein the first size is larger than the second size; Displaying the representation of the first media item moving from a first location within the user interface to a second location within the user interface, wherein the second location corresponds to a corner of the viewfinder; The method according to claim 21, comprising.
23. Before displaying the representation of the first media item, the viewfinder includes a second reticle virtual object, and the displaying of the representation of the first media item replaces the displaying of the second reticle virtual object. The method according to claim 21.
24. Displaying the user interface includes displaying a second media review virtual object at a third location within the user interface, the second media review virtual object being displayed in a third size, and displaying the representation of the first media item that fades into the user interface is displaying the representation of the first media item transitioning from a fourth size to a fifth size, the fourth size being larger than the fifth size and the fourth size being larger than the third size, and displaying the representation of the first media item as moving from a fourth location within the user interface to the third location within the user interface. The method according to claim 21 includes these steps. **Claim 25** The method according to claim 1, wherein the first media item is a still image or a video. **Claim 26** detecting a second request to capture media after changing the appearance of the viewfinder; in response to detecting the second request to capture media; displaying a first type of feedback according to a determination that the second request to capture media corresponds to a request to capture a still image; displaying a second type of feedback different from the first type of feedback according to a determination that the second request to capture media corresponds to a request to capture a video; and The method according to claim 25 further includes these steps. **Claim 27** The first request to capture media corresponds to a first type of input, the first media item is a still image, and the first type of input corresponds to a short press of a first hardware input mechanism. The method according to claim 1. **Claim 28** detecting a third request to capture media after changing the appearance of the viewfinder; In response to detecting the third request for capturing media, and in accordance with the determination that the third request for capturing media is a second type of input, where the second type of input corresponds to a long press of a second hardware input mechanism, capturing a video media item, The method according to claim 1, further comprising. **Claim 29** The third request for capturing media corresponds to the second type of input, and the method In response to detecting the third request for capturing media, displaying an indication that the capture of the video media item corresponds to a second video capture operation, and displaying the indication Displaying the indication in a first appearance according to a determination that a set of criteria is not met, Displaying the indication in a second appearance according to a determination that the set of criteria is met, Including, During the indication being displayed, aborting the detection of the second type of input, In response to aborting the detection of the second type of input, Aborting the execution of the second video capture operation according to a determination that the set of criteria was not met before aborting the detection of the second type of input, Continuing the execution of the second video capture operation according to a determination that the set of criteria was met before aborting the detection of the second type of input, The method according to claim 28, further comprising. **Claim 30** The indication indicates the amount of time elapsed since the computer system started the second video capture operation, and displaying the indication Displaying the indication in the first appearance, During the indication being displayed in the first appearance, detecting that the set of criteria is met, Changing the appearance of the indication from the first appearance to the second appearance in response to detecting that the set of criteria is met, the method according to claim 29. **Claim 31** Detecting an input corresponding to activation of a third hardware input mechanism while the computer system is capturing the video media item; Aborting the capture of the video media item in response to detecting the input corresponding to activation of the third hardware input mechanism; The method according to claim 28, further comprising.
32. The method according to claim 1, wherein the first side of the boundary and the second side of the boundary are on opposite sides of the boundary.
33. The method according to claim 1, wherein the first portion of the content is at the threshold distance from the first side of the boundary of the viewfinder, and the second portion of the content is at the threshold distance from the second side of the boundary of the viewfinder.
34. Displaying the user interface includes displaying a third reticle virtual object, and the third reticle virtual object indicates the capture area of the one or more cameras. The method according to claim 1.
35. A computer program for causing a computer to execute the method according to any one of claims 1 to 34.
36. A memory storing the computer program according to claim 35; One or more processors capable of executing the computer program stored in the memory; A computer system comprising: The computer system is configured to communicate with a display generation component, one or more input devices, and one or more cameras. Computer system.
37. Means for executing the method according to any one of claims 1 to 34 A computer system comprising.
Citation Information
Patent Citations
Imaging device, control method, and program
JP2014107836A
Imaging control device, imaging device, control method, program, and storage medium
JP2018117186A
Imaging device, image processing device, image processing program recording medium
WO2012001947A1
Method for providing binocular stereoscopic image, observation device, and camera unit
WO2016208539A1