Devices, methods, and graphical user interfaces for capturing media using camera application
By detecting user gaze and gestures and dynamically adjusting control salience, the inefficiency and energy waste of existing camera applications when capturing media are solved, achieving more efficient user interaction and energy management.
Patent Information
- Application Number
- CN202480018857.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2024-03-21
- Publication Date
- 2025-10-31
AI Technical Summary
When using camera applications to capture media in existing technologies, user interaction methods are cumbersome, inefficient, and lack feedback, resulting in wasted computer system energy and cognitive burden on users, especially in battery-powered devices.
By detecting user gaze and gestures, the system dynamically displays control salience, providing an improved user interface and feedback, reducing user input, and optimizing energy usage.
It improves the efficiency and intuitiveness of user interaction and saves energy on computing devices, especially the battery life of wearable devices.
Smart Images

Figure CN120883172A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to the following applications: U.S. Provisional Patent Application Serial No. 63 / 453,708, filed March 21, 2023, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING MEDIA WITH A CAMERA APPLICATION"; U.S. Provisional Patent Application Serial No. 63 / 470,878, filed June 3, 2023, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING MEDIA WITH A CAMERA APPLICATION"; and U.S. Provisional Patent Application Serial No. 63 / 470,878, filed July 23, 2023, entitled "DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING MEDIA WITH A CAMERA". The contents of U.S. Provisional Patent Application Serial No. 63 / 528,409 entitled “DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING MEDIA WITH A CAMERA APPLICATION” are incorporated herein by reference in their entirety; and U.S. Provisional Patent Application Serial No. 63 / 537,801 entitled “DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING MEDIA WITH A CAMERA APPLICATION”, filed September 11, 2023; U.S. Provisional Patent Application Serial No. 63 / 548,166 entitled “DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING MEDIA WITH A CAMERA APPLICATION”, filed November 10, 2023; and U.S. Provisional Patent Application Serial No. 63 / 548,166 entitled “DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR CAPTURING MEDIA WITH A CAMERA APPLICATION”, filed January 31, 2024. The contents of each of these patent applications, in U.S. Patent Application Serial No. 18 / 429,138, entitled "CAMERA APPLICATION", are incorporated herein by reference in their entirety. Technical Field
[0003] This disclosure relates in its entirety to computer systems that provide computer-generated experiences in communication with display generation components, a first camera, and an optional second camera, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. Background Technology
[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices for computer systems and other electronic computing devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) are used to interact with the virtual / augmented reality environment. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention
[0005] Some methods and interfaces for capturing media using camera applications (e.g., when interacting with an environment that includes at least some virtual elements such as the application, augmented reality, mixed reality, and / or virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback about the computer system's state (e.g., the computer system's readiness for capturing media, the orientation of the camera used for media capture, and / or the current capture quality), systems that excessively blur the environment when a user attempts to capture media, and systems with complex, verbose, and / or error-prone inputs for controlling media capture impose a significant cognitive burden on the user and degrade the media capture experience. Furthermore, these methods take longer than necessary, thus wasting the computer system's energy. This latter consideration is particularly important in battery-powered devices.
[0006] Therefore, there is a need for computer systems with improved methods and interfaces to provide users with a computer-generated experience, making media capture using computer systems more efficient and intuitive for users. Such methods and interfaces optionally complement or replace conventional methods of media capture using camera applications. These methods and interfaces create a more efficient human-computer interface by helping users understand the relationship between the input provided and the device's response to that input, thereby reducing the quantity, extent, and / or nature of user input.
[0007] The disclosed system reduces or eliminates the aforementioned defects and other problems associated with the user interface of a computer system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop, tablet, or handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touchscreen" or "touchscreen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to display generation components, the computer system also has one or more output devices, including one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, memory, and one or more modules, and a program or instruction set stored in memory for performing multiple functions. In some implementations, the user interacts with the GUI through touch and gestures of a stylus and / or fingers on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (as captured by a camera and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some implementations, functions performed through interaction optionally include image editing, drawing, presentations, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a transient and / or non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0008] There is a need for electronic devices with improved methods and interfaces for capturing media using camera applications. Such methods and interfaces can complement or replace conventional methods of capturing media using camera applications. These methods and interfaces provide users with improved feedback on the status of the computer system, reduce the amount, extent, and / or nature of user input, and result in a more efficient human-computer interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charging. These methods and interfaces reduce energy usage, thereby reducing heat generated by the computing device, which is particularly important for wearable computing devices such as head-mounted displays (HMDs), where excessive heat generation can lead to user discomfort even if the device components operate well within their operating parameter limits.
[0009] In some implementations, the computer system displays a set of controls (e.g., transmission controls and / or other types of controls) associated with controlling the playback of media content in response to the detection of a user's gaze and / or gesture. In some implementations, the computer system initially displays a first set of controls in a desalience state (e.g., with reduced visual salience) in response to the detection of a first input, and then displays a second set of controls (optionally including additional controls) in an increased salience state in response to the detection of a second input. In this way, the computer system optionally provides feedback to the user that the display of controls has begun to invoke the controls without unduly distracting the user from the content (e.g., by initially displaying the controls in a less visually salience manner), and then displays the controls in a more visually salience manner based on the detection of user input indicating that the user wishes to interact further with the controls, to allow for easier and more accurate interaction with the computer system.
[0010] According to some embodiments, a method is described that is executed at a computer system communicating with a display generation component and a first camera. The method includes: when a first user interface including a camera viewfinder is displayed via the display generation component: detecting a first input; and in response to detecting the first input: initiating the capture of first media content using the first camera based on determining that the gaze of a user of the computer system is directed towards a corresponding area of the camera viewfinder when the first input is detected; and abandoning the initiation of the capture of the first media content based on determining that the gaze of the user of the computer system is not directed towards a corresponding area of the camera viewfinder when the first input is detected.
[0011] According to some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a first camera. The one or more programs include instructions for: when a first user interface including a camera viewfinder is displayed via the display generating component: detecting a first input; and in response to detecting the first input: initiating the capture of first media content using the first camera based on determining that the user's gaze is directed towards a corresponding area of the camera viewfinder when the first input is detected; and abandoning the initiation of the capture of the first media content based on determining that the user's gaze is not directed towards a corresponding area of the camera viewfinder when the first input is detected.
[0012] According to some embodiments, a transient computer-readable storage medium is described. This transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a first camera. The one or more programs include instructions for: when a first user interface including a camera viewfinder is displayed via the display generating component: detecting a first input; and in response to detecting the first input: initiating the capture of first media content using the first camera based on determining that the user's gaze is directed towards a corresponding area of the camera viewfinder when the first input is detected; and abandoning the initiation of the capture of the first media content based on determining that the user's gaze is not directed towards a corresponding area of the camera viewfinder when the first input is detected.
[0013] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and a first camera. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting a first input when a first user interface including a camera viewfinder is displayed via the display generation component; and in response to detecting the first input: initiating the capture of first media content using the first camera based on determining that the user's gaze is directed towards a corresponding area of the camera viewfinder when the first input is detected; and abandoning the initiation of the capture of the first media content based on determining that the user's gaze is not directed towards a corresponding area of the camera viewfinder when the first input is detected.
[0014] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and a first camera, and the computer system includes means for performing the following operations when a first user interface including a camera viewfinder is displayed via the display generation component: detecting a first input; and in response to detecting the first input: initiating the capture of first media content using the first camera based on determining that the user's gaze is directed towards a corresponding area of the camera viewfinder when the first input is detected; and abandoning the initiation of the capture of the first media content based on determining that the user's gaze is not directed towards a corresponding area of the camera viewfinder when the first input is detected.
[0015] According to some embodiments, a computer program product is described. The computer program product is configured to be executed by one or more processors of a computer system communicating with a display generating component and a first camera. The one or more programs include instructions for: when a first user interface including a camera viewfinder is displayed via the display generating component: detecting a first input; and in response to detecting the first input: initiating the capture of first media content using the first camera based on determining that the user's gaze is directed towards a corresponding area of the camera viewfinder when the first input is detected; and abandoning the initiation of the capture of the first media content based on determining that the user's gaze is not directed towards a corresponding area of the camera viewfinder when the first input is detected.
[0016] According to some embodiments, a method is described that is executed at a computer system communicating with a display generation component and a first camera. The method includes: displaying a first user interface via the display generation component a camera preview including at least a portion of the field of view of the first camera; detecting a change in the orientation of the field of view of the first camera relative to a corresponding orientation; and in response to detecting the change in orientation: displaying a first indicator representing the orientation of the field of view of the first camera according to determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion that is met when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
[0017] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and a first camera. The one or more programs include instructions for: displaying a first user interface via the display generation component a camera preview including at least a portion of the field of view of the first camera; detecting a change in the orientation of the field of view of the first camera relative to a corresponding orientation; and in response to detecting the change in orientation: displaying a first indicator representing the orientation of the field of view of the first camera according to determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion satisfied when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
[0018] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and a first camera. The one or more programs include instructions for: displaying a first user interface via the display generation component a camera preview including at least a portion of the field of view of the first camera; detecting a change in the orientation of the field of view of the first camera relative to a corresponding orientation; and in response to detecting the change in orientation: displaying a first indicator representing the orientation of the field of view of the first camera according to determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion satisfied when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
[0019] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and a first camera. The computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying a first user interface via the display generation component a camera preview including at least a portion of the field of view of the first camera; detecting a change in the orientation of the field of view of the first camera relative to a corresponding orientation; and in response to detecting the change in orientation: displaying a first indicator representing the orientation of the field of view of the first camera according to determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion satisfied when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
[0020] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and a first camera, and the computer system includes: means for performing the following operations: displaying a first user interface via the display generation component a camera preview including at least a portion of the field of view of the first camera; means for performing the following operations: detecting a change in the orientation of the field of view of the first camera relative to a corresponding orientation; and means for performing the following operations: in response to detecting a change in orientation: displaying a first indicator representing the orientation of the field of view of the first camera according to determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion satisfied when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
[0021] According to some embodiments, a computer program product is described. The computer program product is configured to be executed by one or more processors of a computer system communicating with a display generation component and a first camera. The one or more programs include instructions for: displaying a first user interface via the display generation component a camera preview including at least a portion of the field of view of the first camera; detecting a change in the orientation of the field of view of the first camera relative to a corresponding orientation; and in response to detecting the change in orientation: displaying a first indicator representing the orientation of the field of view of the first camera according to determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion satisfied when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
[0022] According to some embodiments, a method is described that is executed at a computer system communicating with a display generation component and a plurality of cameras including a first camera and a second camera. The method includes: displaying a capture preview for spatial media capture via the display generation component, wherein capture input detected while displaying the capture preview causes the computer system to capture media from the first and second cameras to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, which, when viewed simultaneously, form the illusion of a spatial representation of the fields of view of the plurality of cameras; detecting the position of an individual within the fields of view of the plurality of cameras while displaying the capture preview for spatial media capture; and, in response to detecting the position of the individual within the fields of view of the plurality of cameras, displaying a prompt via the display generation component to change the distance between the individual and the plurality of cameras based on determining that the individual's position relative to the fields of view of the plurality of cameras does not meet the criteria for capturing spatial media at a threshold quality level.
[0023] According to some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and a plurality of cameras including a first camera and a second camera. The one or more programs include instructions for: displaying a capture preview for spatial media capture via the display generation component, wherein capture input detected while displaying the capture preview causes the computer system to capture media from the first and second cameras to generate a spatial media item comprising one or more images for the right eye and one or more images for the left eye, which, when viewed simultaneously, form the illusion of a spatial representation of the fields of view of the plurality of cameras; detecting the position of an individual within the fields of view of the plurality of cameras while displaying the capture preview for spatial media capture; and, in response to detecting the position of the individual within the fields of view of the plurality of cameras, displaying a prompt via the display generation component to change the distance between the individual and the plurality of cameras based on determining that the individual's position relative to the fields of view of the plurality of cameras does not meet the criteria for capturing spatial media at a threshold quality level.
[0024] According to some embodiments, a transient computer-readable storage medium is described. This transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generation component and a plurality of cameras including a first camera and a second camera. The one or more programs include instructions for: displaying a capture preview for spatial media capture via the display generation component, wherein capture input detected while displaying the capture preview causes the computer system to capture media from the first and second cameras to generate a spatial media item comprising one or more images for the right eye and one or more images for the left eye, which, when viewed simultaneously, form the illusion of a spatial representation of the fields of view of the plurality of cameras; detecting the position of an individual within the fields of view of the plurality of cameras while displaying the capture preview for spatial media capture; and, in response to detecting the position of the individual within the fields of view of the plurality of cameras, displaying a prompt via the display generation component to change the distance between the individual and the plurality of cameras based on determining that the individual's position relative to the fields of view of the plurality of cameras does not meet the criteria for capturing spatial media at a threshold quality level.
[0025] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and a plurality of cameras, including a first camera and a second camera, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying a capture preview for spatial media capture via the display generation component, wherein capture input detected while displaying the capture preview will cause the computer system to capture media from the first camera and the second camera to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, which, when viewed simultaneously, form the illusion of a spatial representation of the field of view of the plurality of cameras; detecting the position of an individual within the field of view of the plurality of cameras while displaying the capture preview for spatial media capture; and, in response to detecting the position of the individual within the field of view of the plurality of cameras, displaying a prompt via the display generation component to change the distance between the individual and the plurality of cameras based on determining that the individual's position relative to the field of view of the plurality of cameras does not meet the criteria for capturing spatial media at a threshold quality level.
[0026] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and a plurality of cameras, including a first camera and a second camera. The computer system includes: means for performing the following operations: displaying a capture preview for spatial media capture via the display generation component, wherein capture input detected while displaying the capture preview causes the computer system to capture media from the first and second cameras to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, which, when viewed simultaneously, form the illusion of a spatial representation of the fields of view of the plurality of cameras; means for performing the following operations: detecting the position of an individual within the fields of view of the plurality of cameras while displaying the capture preview for spatial media capture; and means for performing the following operations: in response to detecting the position of the individual within the fields of view of the plurality of cameras, displaying a prompt via the display generation component to change the distance between the individual and the plurality of cameras based on determining that the individual's position relative to the fields of view of the plurality of cameras does not meet the criteria for capturing spatial media at a threshold quality level.
[0027] According to some embodiments, a computer program product is described. The computer program product is configured to be executed by one or more processors of a computer system communicating with a display generation component and a plurality of cameras including a first camera and a second camera. The one or more programs include instructions for: displaying a capture preview for spatial media capture via the display generation component, wherein capture input detected while displaying the capture preview will cause the computer system to capture media from the first and second cameras to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, which, when viewed simultaneously, form the illusion of a spatial representation of the fields of view of the plurality of cameras; detecting the position of an individual within the fields of view of the plurality of cameras while displaying the capture preview for spatial media capture; and, in response to detecting the position of the individual within the fields of view of the plurality of cameras, displaying a prompt via the display generation component to change the distance between the individual and the plurality of cameras based on determining that the individual's position relative to the fields of view of the plurality of cameras does not meet the criteria for capturing spatial media at a threshold quality level.
[0028] According to some embodiments, a method is described that is executed at a computer system having a display generation component and one or more sensors, the one or more sensors including one or more cameras. The method includes: capturing video media using the one or more cameras; and while capturing the video media: detecting movement of the one or more cameras via the one or more sensors; and in response to detecting movement of the one or more cameras: displaying, via the display generation component, the movement of a visual indicator relative to a display reference object based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, wherein displaying the movement of the visual indicator includes: based on determining that the movement of the one or more cameras is movement in a first direction of camera movement, the display visual indicator moves relative to the display reference object in the first direction of indicator movement; and based on determining that the movement of the one or more cameras is movement in a second direction of camera movement different from the first direction of camera movement, the display visual indicator moves relative to the display reference object in the second direction of indicator movement different from the first direction of indicator movement.
[0029] According to some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with display generation components and one or more sensors, the one or more sensors including one or more cameras. The one or more programs include instructions for: detecting movement of the one or more cameras via the one or more sensors; and in response to detecting movement of the one or more cameras: displaying, via the display generation components, movement of a visual indicator relative to a display reference object based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, wherein displaying the movement of the visual indicator includes: based on determining that the movement of the one or more cameras is movement in a first direction of camera movement, displaying the visual indicator relative to the display reference object in a first direction of indicator movement; and based on determining that the movement of the one or more cameras is movement in a second direction of camera movement different from the first direction of camera movement, displaying the visual indicator relative to the display reference object in a second direction of indicator movement different from the first direction of indicator movement.
[0030] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with display generation components and one or more sensors, the one or more sensors including one or more cameras. The one or more programs include instructions for: capturing video media using the one or more cameras; and when capturing video media: detecting movement of the one or more cameras via the one or more sensors; and in response to detecting movement of the one or more cameras: displaying a visual indicator relative to a display reference object via the display generation components based on determining that the movement of the one or more cameras satisfies a set of one or more movement criteria, wherein the movement of the visual indicator includes: moving the visual indicator relative to the display reference object in a first direction of indicator movement based on determining that the movement of the one or more cameras is movement in a first direction of camera movement; and moving the visual indicator relative to the display reference object in a second direction of indicator movement, different from the first direction of indicator movement, based on determining that the movement of the one or more cameras is movement in a second direction of camera movement different from the first direction of camera movement.
[0031] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and one or more sensors, including one or more cameras, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: capturing video media using the one or more cameras; and when capturing the video media: detecting movement of the one or more cameras via the one or more sensors; and in response to detecting movement of the one or more cameras: displaying a visual indicator relative to a display reference object via the display generation component based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, wherein the movement of the visual indicator includes: moving the visual indicator relative to the display reference object in a first direction of indicator movement based on determining that the movement of the one or more cameras is movement in a first direction of camera movement; and moving the visual indicator relative to the display reference object in a second direction of indicator movement, different from the first direction of indicator movement, based on determining that the movement of the one or more cameras is movement in a second direction of camera movement different from the first direction of camera movement.
[0032] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and one or more sensors, including one or more cameras, and the computer system includes: means for performing the following operations: capturing video media using the one or more cameras; and means for performing the following operations: when capturing the video media: detecting movement of the one or more cameras via the one or more sensors; and in response to detecting movement of the one or more cameras: displaying a visual indicator relative to a display reference object via the display generation component based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, wherein the movement of the visual indicator includes: moving the visual indicator relative to the display reference object in the first direction of indicator movement based on determining that the movement of the one or more cameras is movement in a first direction of camera movement; and moving the visual indicator relative to the display reference object in the second direction of indicator movement, different from the first direction of indicator movement, based on determining that the movement of the one or more cameras is movement in a second direction of camera movement different from the first direction of camera movement.
[0033] According to some embodiments, a computer program product is described. The computer program product is configured to be executed by one or more processors of a computer system communicating with a display generation component and one or more sensors, the one or more sensors including one or more cameras. The computer program product includes instructions for: capturing video media using the one or more cameras; and when capturing the video media: detecting movement of the one or more cameras via the one or more sensors; and in response to detecting movement of the one or more cameras: displaying a visual indicator relative to a display reference object via the display generation component based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, wherein the movement of the visual indicator includes: moving the visual indicator relative to the display reference object in a first direction of indicator movement based on determining that the movement of the one or more cameras is movement in a first direction of camera movement; and moving the visual indicator relative to the display reference object in a second direction of indicator movement, different from the first direction of indicator movement, based on determining that the movement of the one or more cameras is movement in a second direction of camera movement different from the first direction of indicator movement.
[0034] According to some embodiments, a method is described that is executed at a computer system communicating with a display generation component. The method includes: while playback of a video media item is in progress, wherein playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item; changing the visual salience of the video media item relative to the boundary region based on a representation of movement corresponding to a viewpoint of the video media item occurring when the video media item is captured; wherein changing the visual salience of the video media item relative to the boundary region based on the representation of movement corresponding to the viewpoint of the video media item includes: changing the visual salience of the video media item relative to the boundary region to a first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount; and changing the visual salience of the video media item relative to the boundary region to a second relative visual salience level different from the first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount.
[0035] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component. The one or more programs include instructions for: when playback of a video media item is in progress, wherein playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item; changing the visual salience of the video media item relative to the boundary region based on a representation of movement corresponding to the viewpoint of the video media item occurring when the video media item is captured; wherein changing the visual salience of the video media item relative to the boundary region based on the representation of movement corresponding to the viewpoint of the video media item includes: changing the visual salience of the video media item relative to the boundary region to a first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount; and changing the visual salience of the video media item relative to the boundary region to a second relative visual salience level different from the first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount.
[0036] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component. The one or more programs include instructions for: when playback of a video media item is in progress, wherein playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item; changing the visual salience of the video media item relative to the boundary region based on a representation of movement corresponding to the viewpoint of the video media item occurring when the video media item is captured; wherein changing the visual salience of the video media item relative to the boundary region based on the representation of movement corresponding to the viewpoint of the video media item includes: changing the visual salience of the video media item relative to the boundary region to a first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount; and changing the visual salience of the video media item relative to the boundary region to a second relative visual salience level different from the first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount.
[0037] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when playback of a video media item is in progress, wherein playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item, changing the visual salience of the video media item relative to the boundary region based on a representation of movement corresponding to the viewpoint of the video media item occurring when the video media item is captured, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of movement corresponding to the viewpoint of the video media item includes: changing the visual salience of the video media item relative to the boundary region to a first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount; and changing the visual salience of the video media item relative to the boundary region to a second relative visual salience level different from the first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount.
[0038] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and a first camera, and the computer system includes means for performing the following operations: when playback of a video media item is in progress, wherein playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item, changing the visual salience of the video media item relative to the boundary region based on a representation of movement corresponding to the viewpoint of the video media item occurring when the video media item is captured, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of movement corresponding to the viewpoint of the video media item includes: changing the visual salience of the video media item relative to the boundary region to a first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount; and changing the visual salience of the video media item relative to the boundary region to a second relative visual salience level different from the first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount.
[0039] According to some embodiments, a computer program product is described. The computer program product is configured to be executed by one or more processors of a computer system communicating with a display generation component. The one or more programs include instructions for: while playback of a video media item is in progress, wherein playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item; changing the visual salience of the video media item relative to the boundary region based on a representation of movement corresponding to the viewpoint of the video media item occurring when the video media item is captured; wherein changing the visual salience of the video media item relative to the boundary region based on the representation of movement corresponding to the viewpoint of the video media item includes: changing the visual salience of the video media item relative to the boundary region to a first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount; and changing the visual salience of the video media item relative to the boundary region to a second relative visual salience level different from the first relative visual salience level based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount.
[0040] According to some embodiments, a method is described that is executed at a computer system communicating with a display generation component and one or more cameras. The method includes: when capturing spatial video media of an environment using one or more cameras, wherein the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint, the first and second video components forming an illusion of a spatial representation of the environment when viewed simultaneously; displaying, via the display generation component, a virtual indicator element in the environment representing an anchor position corresponding to a respective viewpoint of the spatial video media, wherein the virtual indicator element is displayed while the environment is visible via the display generation component; detecting a first change in the viewpoint from which the spatial video media is captured when the virtual indicator element is displayed while the environment is visible via the display generation component; and, in response to detecting the first change in the viewpoint from which the spatial video media is captured, changing the appearance of the virtual indicator element to indicate the corresponding viewpoint of the spatial video media.
[0041] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with a display generating component and one or more cameras. The one or more programs include instructions for: when capturing spatial video media of an environment using one or more cameras, wherein the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint, the first and second video components forming an illusion of a spatial representation of the environment when viewed simultaneously; displaying, via the display generating component, a virtual indicator element in the environment representing an anchor position corresponding to a corresponding viewpoint of the spatial video media, wherein the virtual indicator element is displayed while the environment is visible via the display generating component; detecting a first change in the viewpoint from which the spatial video media is captured when the virtual indicator element is displayed while the environment is visible via the display generating component; and, in response to detecting the first change in the viewpoint from which the spatial video media is captured, changing the appearance of the virtual indicator element to indicate the corresponding viewpoint of the spatial video media.
[0042] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras. The one or more programs include instructions for: when capturing spatial video media of an environment using one or more cameras, wherein the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint, the first and second video components forming an illusion of a spatial representation of the environment when viewed simultaneously; displaying, via the display generating component, a virtual indicator element in the environment representing an anchor position corresponding to a corresponding viewpoint of the spatial video media, wherein the virtual indicator element is displayed while the environment is visible via the display generating component; detecting a first change in the viewpoint from which the spatial video media is captured when the virtual indicator element is displayed while the environment is visible via the display generating component; and, in response to detecting the first change in the viewpoint from which the spatial video media is captured, changing the appearance of the virtual indicator element to indicate the corresponding viewpoint of the spatial video media.
[0043] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and one or more cameras, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when capturing spatial video media of an environment using the one or more cameras, wherein the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint, the first and second video components forming an illusion of a spatial representation of the environment when viewed simultaneously; displaying, via the display generation component, virtual indicator elements in the environment representing anchor positions corresponding to the respective viewpoints of the spatial video media, wherein the virtual indicator elements are displayed while the environment is visible via the display generation component; when the virtual indicator elements are displayed while the environment is visible via the display generation component, detecting a first change in the viewpoint from which the spatial video media is captured; and in response to detecting the first change in the viewpoint from which the spatial video media is captured, changing the appearance of the virtual indicator elements to indicate the respective viewpoints corresponding to the spatial video media.
[0044] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component and one or more cameras, and the computer system includes: means for performing the following operations: when capturing spatial video media of an environment using one or more cameras, wherein the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint, the first and second video components forming an illusion of a spatial representation of the environment when viewed simultaneously; displaying, via the display generation component, a virtual indicator element in the environment representing an anchor position corresponding to a respective viewpoint of the spatial video media, wherein the virtual indicator element is displayed while the environment is visible via the display generation component; means for performing the following operations: when the virtual indicator element is displayed while the environment is visible via the display generation component, detecting a first change in the viewpoint from which the spatial video media is captured; and means for performing the following operations: in response to detecting the first change in the viewpoint from which the spatial video media is captured, changing the appearance of the virtual indicator element to indicate a corresponding viewpoint of the spatial video media.
[0045] According to some embodiments, a computer program product is described. The computer program product is configured to be executed by one or more processors of a computer system communicating with a display generating component and one or more cameras. The one or more programs include instructions for: when capturing spatial video media of an environment using one or more cameras, wherein the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint, the first and second video components forming an illusion of a spatial representation of the environment when viewed simultaneously; displaying, via the display generating component, a virtual indicator element in the environment representing an anchor position corresponding to a corresponding viewpoint of the spatial video media, wherein the virtual indicator element is displayed while the environment is visible via the display generating component; detecting a first change in the viewpoint from which the spatial video media is captured when the virtual indicator element is displayed while the environment is visible via the display generating component; and, in response to detecting the first change in the viewpoint from which the spatial video media is captured, changing the appearance of the virtual indicator element to indicate the corresponding viewpoint of the spatial video media.
[0046] According to some embodiments, a method is described that is executed at a computer system communicating with a display generation component. The method includes: when displaying a representation of a spatial media item via the display generation component, wherein the spatial media item includes a first component corresponding to a right-eye viewpoint and a second component different from the first component corresponding to a left-eye viewpoint, the first and second components forming an illusion of a spatial representation when viewed simultaneously; simultaneously displaying a spatial viewing indicator having a first appearance along with a representation of the spatial media item based on determining that the spatial media item satisfies one or more sets of stability criteria; and abandoning the simultaneous display of the spatial viewing indicator having the first appearance along with a representation of the spatial media item based on determining that the spatial media item does not satisfy the one or more sets of stability criteria.
[0047] According to some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component. The one or more programs include instructions for: when displaying a representation of a spatial media item via the display generating component, wherein the spatial media item includes a first component corresponding to a right-eye viewpoint and a second component different from the first component corresponding to a left-eye viewpoint, the first and second components forming an illusion of spatial representation when viewed simultaneously; simultaneously displaying a spatial viewing indicator with a first appearance along with a representation of the spatial media item based on determining that the spatial media item satisfies one or more sets of stability criteria; and abandoning the simultaneous display of the spatial viewing indicator with the first appearance along with a representation of the spatial media item based on determining that the spatial media item does not satisfy the one or more sets of stability criteria.
[0048] According to some embodiments, a transient computer-readable storage medium is described. The transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component. The one or more programs include instructions for: when displaying a representation of a spatial media item via the display generating component, wherein the spatial media item includes a first component corresponding to a right-eye viewpoint and a second component different from the first component corresponding to a left-eye viewpoint, the first and second components forming an illusion of spatial representation when viewed simultaneously; simultaneously displaying a spatial viewing indicator having a first appearance along with a representation of the spatial media item based on determining that the spatial media item satisfies one or more sets of stability criteria; and abandoning the simultaneous display of the spatial viewing indicator having the first appearance along with a representation of the spatial media item based on determining that the spatial media item does not satisfy the one or more sets of stability criteria.
[0049] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when displaying a representation of a spatial media item via the display generation component, wherein the spatial media item includes a first component corresponding to a right-eye viewpoint and a second component different from the first component corresponding to a left-eye viewpoint, the first component and the second component forming an illusion of spatial representation when viewed simultaneously; simultaneously displaying a spatial viewing indicator having a first appearance along with a representation of the spatial media item based on determining that the spatial media item satisfies one or more sets of stability criteria; and abandoning the simultaneous display of the spatial viewing indicator having the first appearance along with a representation of the spatial media item based on determining that the spatial media item does not satisfy the one or more sets of stability criteria.
[0050] According to some embodiments, a computer system is described. The computer system is configured to communicate with a display generation component, and the computer system includes means for performing the following operations: when displaying a representation of a spatial media item via the display generation component, wherein the spatial media item includes a first component corresponding to a right-eye viewpoint and a second component different from the first component corresponding to a left-eye viewpoint, the first and second components forming an illusion of a spatial representation when viewed simultaneously: simultaneously displaying a spatial viewing indicator having a first appearance with the representation of the spatial media item based on determining that the spatial media item satisfies one or more sets of stability criteria; and abandoning the simultaneous display of the spatial viewing indicator having the first appearance with the representation of the spatial media item based on determining that the spatial media item does not satisfy the one or more sets of stability criteria.
[0051] According to some embodiments, a computer program product is described. The computer program product is configured to be executed by one or more processors of a computer system communicating with a display generation component. The one or more programs include instructions for: when displaying a representation of a spatial media item via the display generation component, wherein the spatial media item includes a first component corresponding to a right-eye viewpoint and a second component different from the first component corresponding to a left-eye viewpoint, the first and second components forming an illusion of spatial representation when viewed simultaneously; simultaneously displaying a spatial viewing indicator with a first appearance along with a representation of the spatial media item based on determining that the spatial media item satisfies one or more sets of stability criteria; and abandoning the simultaneous display of the spatial viewing indicator with the first appearance along with a representation of the spatial media item based on determining that the spatial media item does not satisfy the one or more sets of stability criteria.
[0052] In some embodiments, a method is described. In some embodiments, the method is performed at a computer system communicating with one or more display generating components and one or more cameras. The method includes: when the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and when the computer system is not capturing spatial media: based on determining that the orientation of the computer system is outside a threshold orientation range, outputting a first prompt prompting the user to rotate the computer system to within the threshold orientation range.
[0053] In some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more display generating components and one or more cameras. The one or more programs include instructions for: when the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and when the computer system is not capturing spatial media: based on determining that the orientation of the computer system is outside a threshold orientation range, outputting a first prompt prompting the user to rotate the computer system to within the threshold orientation range.
[0054] In some embodiments, a transient computer-readable storage medium is described. This transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more display generating components and one or more cameras. The one or more programs include instructions for: when the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and when the computer system is not capturing spatial media: based on determining that the orientation of the computer system is outside a threshold orientation range, outputting a first prompt prompting the user to rotate the computer system within the threshold orientation range.
[0055] In some embodiments, a computer system is described. The computer system is configured to communicate with one or more display generating components and one or more cameras, and includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and when the computer system is not capturing spatial media: based on determining that the orientation of the computer system is outside a threshold orientation range, outputting a first prompt prompting the user to rotate the computer system into the threshold orientation range.
[0056] In some embodiments, a computer system is described. The computer system is configured to communicate with one or more display generating components and one or more cameras, and includes means for performing the following operations: when the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and when the computer system is not capturing spatial media: based on determining that the orientation of the computer system is outside a threshold orientation range, outputting a first prompt prompting the user to rotate the computer system into the threshold orientation range.
[0057] In some embodiments, a computer program product is described. This computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more display generating components and one or more cameras. The one or more programs include instructions for: when the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and when the computer system is not capturing spatial media: based on determining that the orientation of the computer system is outside a threshold orientation range, outputting a first prompt prompting the user to rotate the computer system into the threshold orientation range.
[0058] In some embodiments, a method is described. In some embodiments, the method is performed at a computer system communicating with one or more display generation components and one or more cameras. The method includes: displaying a first user interface corresponding to a camera application of the computer system via one or more display generation components; and when displaying the first user interface corresponding to the camera application of the computer system: based on determining that the computer system is associated with a head-mounted device separate from the computer system, providing a spatial media capture mode option corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first visual component and the second visual component forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and based on determining that the computer system is not associated with a head-mounted device separate from the computer system, abandoning the provision of the spatial media capture mode option.
[0059] In some embodiments, a non-transitory computer-readable storage medium is described. This non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more display generating components and one or more cameras. The one or more programs include instructions for: displaying a first user interface corresponding to a camera application of the computer system via the one or more display generating components; and when displaying the first user interface corresponding to the camera application of the computer system: based on determining that the computer system is associated with a head-mounted device separate from the computer system, providing a spatial media capture mode option corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and based on determining that the computer system is not associated with a head-mounted device separate from the computer system, abandoning the provision of the spatial media capture mode option.
[0060] In some embodiments, a transient computer-readable storage medium is described. This transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system communicating with one or more display generating components and one or more cameras. The one or more programs include instructions for: displaying a first user interface corresponding to a camera application of the computer system via the one or more display generating components; and when displaying the first user interface corresponding to the camera application of the computer system: based on determining that the computer system is associated with a head-mounted device separate from the computer system, providing a spatial media capture mode option corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and based on determining that the computer system is not associated with a head-mounted device separate from the computer system, abandoning the provision of the spatial media capture mode option.
[0061] In some embodiments, a computer system is described. The computer system is configured to communicate with one or more display generation components and one or more cameras, and includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying a first user interface corresponding to a camera application of the computer system via the one or more display generation components; and when displaying the first user interface corresponding to the camera application of the computer system: based on determining that the computer system is associated with a head-mounted device separate from the computer system, providing a spatial media capture mode option corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and based on determining that the computer system is not associated with a head-mounted device separate from the computer system, abandoning the provision of the spatial media capture mode option.
[0062] In some embodiments, a computer system is described. The computer system is configured to communicate with one or more display generation components and one or more cameras, and includes: means for performing the following operations: displaying a first user interface corresponding to a camera application of the computer system via the one or more display generation components; and means for performing the following operations: when displaying the first user interface corresponding to the camera application of the computer system: based on determining that the computer system is associated with a head-mounted device separate from the computer system, providing a spatial media capture mode option corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and based on determining that the computer system is not associated with a head-mounted device separate from the computer system, abandoning the provision of the spatial media capture mode option.
[0063] In some embodiments, a computer program product is described. This computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with one or more display generating components and one or more cameras. The one or more programs include instructions for: displaying a first user interface corresponding to a camera application of the computer system via the one or more display generating components; and when displaying the first user interface corresponding to the camera application of the computer system: based on determining that the computer system is associated with a head-mounted device separate from the computer system, providing a spatial media capture mode option corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first and second visual components forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and based on determining that the computer system is not associated with a head-mounted device separate from the computer system, abandoning the provision of the spatial media capture mode option.
[0064] It should be noted that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not exhaustive; in particular, many additional features and advantages will be apparent to those skilled in the art from the accompanying drawings, description, and claims. Furthermore, it should be pointed out that the language used in this specification has been chosen in principle for readability and instruction purposes, and such choice may not be necessary to depict or define the subject matter of the invention. Attached Figure Description
[0065] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein similar reference numerals in all the drawings indicate corresponding parts.
[0066] Figure 1A This is a block diagram illustrating the operating environment of a computer system used to provide XR experiences in some implementation schemes.
[0067] Figures 1B to 1P It is used in Figure 1A Examples of computer systems that provide XR experiences in the operating environment.
[0068] Figure 2 This is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience in some implementations.
[0069] Figure 3 This is a block diagram illustrating the display generation component of a computer system configured in some implementations to provide a visual component of an XR experience to a user.
[0070] Figure 4 This is a block diagram illustrating a hand tracking unit of a computer system configured to capture user gesture input in some implementations.
[0071] Figure 5 This is a block diagram illustrating an eye-tracking unit of a computer system configured to capture a user's gaze input in some implementations.
[0072] Figure 6 This is a flowchart illustrating a flare-assisted gaze tracking pipeline in some implementations.
[0073] Figures 7A to 7AB Example techniques for displaying a user interface for gaze-activated media captures are illustrated in some implementations.
[0074] Figure 8 This is a flowchart illustrating a method for displaying a user interface for gaze-activated media capture in some implementations.
[0075] Figures 9A to 9I Example techniques for displaying camera previews with flush indicators are illustrated in some implementations.
[0076] Figure 10 This is a flowchart of a method for displaying a camera preview with a flush indicator, according to some implementation schemes.
[0077] Figures 11A1 to 11H Example techniques for displaying camera previews for spatial media capture with prompts for improved capture quality are illustrated in some implementations.
[0078] Figure 12 This is a flowchart illustrating a method for camera previewing for spatial media capture, showing hints for improved capture quality in some implementations.
[0079] Figures 13A to 13P Example techniques for displaying a camera preview for media capture with a camera movement indicator are illustrated in some implementations.
[0080] Figure 14 This is a flowchart illustrating a method for camera previewing spatial media capture with a camera movement indicator in some implementations.
[0081] Figures 15A to 15N Examples of techniques used in some implementations to modify video playback to improve viewing comfort are illustrated.
[0082] Figure 16 This is a flowchart of a method for modifying video playback to improve viewing comfort in some implementation schemes.
[0083] Figures 17A to 17R Example techniques for displaying camera previews for media capture with viewpoint stability guidance are illustrated in some implementations.
[0084] Figure 18 This is a flowchart illustrating a method for camera previewing media capture with viewpoint stability guidance in some implementations.
[0085] Figures 19A to 19M Example techniques are illustrated in some implementations for presenting viewing settings for media playback based on media stability characteristics.
[0086] Figure 20 This is a flowchart of a method for presenting media playback viewing settings based on media stability characteristics in some implementation schemes.
[0087] Figures 21A to 21O Example techniques for displaying one or more camera user interfaces for capturing media are illustrated in some implementations.
[0088] Figure 22 This is a flowchart of a method for providing a spatial media capture mode in some implementations.
[0089] Figure 23 This is a flowchart of a method for providing spatial media capture mode options in some implementations. Detailed Implementation
[0090] In some implementations, this disclosure relates to a user interface for providing extended reality (XR) experiences to users.
[0091] The systems, methods, and GUIs described in this article improve the use of camera applications to capture media in a variety of ways.
[0092] In some implementations, when a user interface for media capture is displayed, the computer system tracks the user's gaze as it detects potential user input (e.g., hardware button presses, input on touch-sensitive surfaces, and / or air gestures). If the user is looking at a specific area of the user interface (e.g., the central area), such as the location of or near a media capture enablement indicator displayed at the center of the user interface, the computer system initiates media capture; otherwise, it does not initiate media capture if the user is not looking at that specific area. By initiating media capture only when the user is looking at a specific area when input is detected, unintended and / or unwanted media capture is reduced, allowing the user to freely interact with the computer system and / or the environment without accidentally capturing media. Furthermore, initiating media capture only when the user is looking at a specific area when input is detected allows the user to capture media effectively and intuitively when needed, for example, by looking at a specific area and providing input to quickly enable media capture during fleeting media capture opportunities.
[0093] In some implementations, the computer system detects how the orientation of the camera used for media capture (e.g., the orientation of the camera's field of view) changes relative to the target orientation (e.g., an orientation aligned with the horizon of the environment). In response to detecting a change in camera orientation, if the camera orientation differs from the target orientation by more than a threshold amount, the computer system displays an alignment indicator representing the camera orientation. Displaying the alignment indicator when the camera orientation, and therefore the orientation of any media captured by the camera, differs from the target orientation by more than the threshold amount reduces unintended and / or unwanted media capture by warning the user that the captured media will not behave in an aligned manner. Additionally, displaying the alignment indicator allows the user to efficiently and intuitively synthesize media captures with the desired orientation.
[0094] In some implementations, the computer system displays a user interface for spatial media capture, where multiple cameras generate different images for the user's right and left eyes, thereby creating an appearance / illusion of depth in the captured media (e.g., three-dimensional, where the relative distance of an object from the capture plane is perceptible) (e.g., the different images generated for the user's right and left eyes simulate different images of the physical environment received at the right and left eyes due to differences in eye position). The computer system detects the position of an individual in the media capture relative to the multiple cameras and determines whether the individual's relative position will adversely affect the quality of the spatial media capture (e.g., the appearance / illusion of depth). For example, if the individual is too close to the multiple cameras, the images generated for the left and right eyes will differ too much to create a high-quality appearance / illusion of depth, while if the individual is too far from the multiple cameras, the images generated for the left and right eyes will not differ enough to create a high-quality appearance / illusion of depth. If the computer system determines that the individual's relative position will adversely affect the quality of the spatial media capture, the computer system displays a prompt to the user to change the distance from the individual. The display of distance adjustment prompts reduces unintended and / or unwanted media capture by warning users when capturing spatial media will not result in a high-quality depth appearance / illusion. Additionally, the display of flush indicators allows users to efficiently and intuitively synthesize spatial media captures of the desired quality.
[0095] In some implementations, the computer system displays content in a first area of the user interface. In some implementations, while the computer system is displaying content and when a first set of controls is not displayed in a first state, the computer system detects a first input from a first portion of the user interface. In some implementations, in response to detecting the first input, and based on determining that the user's gaze was directed to a second area of the user interface when the first input was detected, the computer system displays one or more controls in the first state in the user interface, and based on determining that the user's gaze was not directed to the second area of the user interface when the first input was detected, the computer system abandons displaying the one or more controls in the first state.
[0096] In some implementations, the computer system displays content in a user interface. In some implementations, when displaying content, the computer system detects a first input based on movement of a first portion of the user's body. In some implementations, in response to detecting the first input, the computer system displays a first set of one or more controls in the user interface, wherein the first set of one or more controls is displayed in a first state and shown within a first area of the user interface. In some implementations, when the first set of one or more controls is displayed in the first state: based on determining that one or more first criteria are met, including criteria met when the user's attention is directed to the first area of the user interface based on movement of a second portion of the user body that is different from the first portion of the user body, the computer system switches from displaying the first set of one or more controls in the first state to displaying a second set of one or more controls in a second state, wherein the second state is different from the first state.
[0097] Figures 1A to 6 A description of a sample computer system for providing XR experiences to users is provided. Figures 7A to 7AB Example techniques for displaying a user interface for gaze-activated media captures are illustrated in some implementations. Figure 8 This is a flowchart illustrating a method for displaying a user interface for gaze-activated media capture in some implementations. Figures 7A to 7AB The user interface in the example is used to demonstrate Figure 8 The process in. Figures 9A to 9I Example techniques for displaying camera previews with flush indicators are illustrated in some implementations. Figure 10 This is a flowchart of a method for displaying a camera preview with a flush indicator, according to some implementation schemes. Figures 9A to 9I The user interface in the example is used to demonstrate Figure 10 The process in. Figures 11A1 to 11H Example techniques for displaying camera previews for spatial media capture with prompts for improved capture quality are illustrated in some implementations. Figure 12 This is a flowchart illustrating a method for camera previewing for spatial media capture, showing hints for improved capture quality in some implementations. Figures 13A to 13P Example techniques for displaying a camera preview for media capture with a camera movement indicator are illustrated in some implementations. Figure 14 This is a flowchart illustrating a method for media capture with a camera movement indicator in some implementations. Figures 13A to 13P The user interface in the example is used to demonstrate Figure 14 The process in. Figures 15A to 15N Examples of techniques used in some implementations to modify video playback to improve viewing comfort are illustrated. Figure 16 This is a flowchart of a method for modifying video playback to improve viewing comfort in some implementation schemes. Figures 15A to 15N The user interface in the example is used to demonstrate Figure 16 The process in. Figures 17A to 17R Example techniques for displaying camera previews for media capture with viewpoint stability guidance are illustrated in some implementations. Figure 18 This is a flowchart illustrating a method for camera previewing media capture with viewpoint stability guidance in some implementations. Figures 17A to 17R The user interface in the example is used to demonstrate Figure 18 The process in. Figures 19A to 19M Example techniques are illustrated in some implementations for presenting viewing settings for media playback based on media stability characteristics. Figure 20 This is a flowchart of a method for presenting media playback viewing settings based on media stability characteristics in some implementation schemes. Figures 19A to 19M The user interface in the example is used to demonstrate Figure 20 The process in. Figures 21A to 21O Example techniques for displaying one or more camera user interfaces for capturing media are illustrated in some implementations. Figure 22 This is a flowchart of a method for providing a spatial media capture mode in some implementations. Figure 23 This is a flowchart of a method for providing spatial media capture mode options in some implementations. Figures 21A to 21O The user interface in the example is used to demonstrate Figure 22 and Figure 23 The process in.
[0098] The processes described below enhance device operability and (e.g., by helping users provide appropriate input and reducing user errors when operating / interacting with the device) make the user-device interface more efficient through various technologies. These include providing users with improved visual feedback, reducing the amount of input required to perform operations, providing additional control options without cluttering the user interface with additional display controls, performing operations when a set of conditions are met without further user input, improving privacy and / or security, providing a richer, more detailed, and / or more realistic user experience while saving storage space, and / or additional technologies. These technologies also reduce power consumption and extend device battery life by enabling users to use the device faster and more efficiently. This saves battery power and, therefore, weight, and improves the device's ergonomics. These technologies also enable real-time communication, allow the use of fewer and / or less precise sensors, resulting in a more compact, lighter, and cheaper device, and enabling the device to be used in a variety of lighting conditions. These technologies reduce energy consumption, thereby reducing the heat generated by the device. This is especially important for wearable devices, where if a device generates too much heat, even when operating entirely within the parameters of its components, it can become uncomfortable for the user to wear.
[0099] Furthermore, in methods described herein where one or more steps depend on the satisfaction of one or more conditions, it should be understood that the method may be repeated in multiple repetitions such that, during the repetitions, all conditions determining the steps in the method are satisfied in different repetitions of the method. For example, if the method requires performing a first step (if the condition is satisfied) and a second step (if the condition is not satisfied), those skilled in the art will know that the stated steps are repeated until both the conditions are satisfied and not satisfied (in no particular order). Thus, a method described as having one or more steps depending on the satisfaction of one or more conditions can be rewritten as a method that repeats until each condition described in the method is satisfied. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing discretionary operations based on the satisfaction of the corresponding one or more conditions, and thus to determine whether possible conditions have been satisfied without explicitly repeating the steps of the method until all conditions determining the steps in the method are satisfied. Those skilled in the art will also understand that, similar to methods having discretionary steps, a system or computer-readable storage medium may repeat the steps of the method multiple times as needed to ensure that all discretionary steps have been performed.
[0100] In some implementation schemes, such as Figure 1A As shown, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted display (HMD), a monitor, a projector, a touchscreen, etc.), one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., image sensors, light sensors, depth sensors, haptic sensors, orientation sensors, proximity sensors, temperature sensors, position sensors, motion sensors, speed sensors, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input devices 125, output devices 155, sensors 190, and peripheral devices 195 are integrated with the display generation component 120 (e.g., in a head-mounted or handheld device).
[0101] In describing XR experiences, various terms are used to distinguish several related but different environments that a user can sense and / or interact with (e.g., interacting with input detected by the computer system 101 that generates the XR experience, causing the computer system to generate audio, visual, and / or haptic feedback corresponding to various inputs provided to the computer system 101). The following is a subset of these terms:
[0102] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. Physical environments, such as physical parks, include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment through senses such as sight, touch, hearing, taste, and smell.
[0103] Extended Reality: Conversely, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In XR, a subset of a person's physical motion, or a representation thereof, is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one physical law. For example, an XR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the properties of virtual objects in the XR environment can be done in response to a representation of physical motion (e.g., a voice command). People can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with audio objects that create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. For example, audio objects can enable audio transparency, which selectively introduces ambient sounds from the physical environment, with or without computer-generated audio. In some XR environments, people can sense and / or interact only with audio objects.
[0104] Examples of XR include virtual reality and mixed reality.
[0105] Virtual Reality (VR): A virtual reality (VR) environment is a simulated environment designed to be entirely based on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. Examples of virtual objects include trees, buildings, and computer-generated images representing human avatars. A person can sense and / or interact with virtual objects in a VR environment through the simulation of their presence within the computer-generated environment and / or through the simulation of a subset of their physical movements within the computer-generated environment.
[0106] Mixed Reality: Compared to VR environments, which are designed to be entirely based on computer-generated sensory input, mixed reality (MR) environments refer to simulated environments designed to incorporate sensory input from the physical environment, or its representations, in addition to computer-generated sensory input (e.g., virtual objects). On the virtual continuum, a mixed reality environment is any state between, but not limited to, a purely physical environment as one end and a virtual reality environment as the other. In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. Additionally, some electronic systems used to present an MR environment can track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical objects or their representations from the physical environment). For example, a system can cause motion so that virtual trees appear stationary relative to the physical ground.
[0107] Examples of mixed reality include augmented reality and augmented virtual reality.
[0108] Augmented Reality (AR): An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are overlaid on a physical environment or a representation of the physical environment. For example, an electronic system for presenting an AR environment may have a transparent or semi-transparent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or semi-transparent display, allowing a person to perceive the virtual objects overlaid on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via the images or videos of the physical environment and perceives the virtual objects overlaid on the physical environment. As used herein, video of the physical environment displayed on an opaque display is referred to as “pass-through video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Alternatively, the system may have a projection system that projects virtual objects onto the physical environment, such as as a hologram or onto a physical surface, allowing a person to perceive the virtual objects superimposed on the physical environment. Augmented reality environments also refer to simulated environments in which the representation of the physical environment is transformed by computer-generated sensory information. For example, in providing pass-through video, the system may transform one or more sensor images to apply a selected viewpoint (e.g., viewpoint) different from the viewpoint captured by the imaging sensor. As another example, the representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions of it, such that the modified portions can be representative but not realistic versions of the original captured image. Furthermore, the representation of the physical environment can be transformed by graphically removing or blurring portions of it.
[0109] Augmented Virtual: An augmented virtual (AV) environment is a simulated environment in which a virtual or computer-generated environment combines one or more sensory inputs from a physical environment. Sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park could have virtual trees and virtual buildings, but a person's face could be realistically reproduced from an image taken of a physical person. Similarly, virtual objects could adopt the shape or color of a physical object imaged by one or more imaging sensors. Furthermore, virtual objects could use shadows that correspond to the sun's position within the physical environment.
[0110] In augmented reality, mixed reality, or virtual reality environments, a view of the three-dimensional environment is visible to the user. This view is typically visible to the user via a virtual viewport through one or more display generating components (e.g., a display providing stereoscopic content to different eyes of the same user), which has a viewport boundary that defines the extent of the three-dimensional environment visible to the user via the one or more display generating components. In some embodiments, the area defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size of one or more display generating components, optical properties or other physical characteristics, and / or the position and / or orientation of one or more display generating components relative to the user's eyes). The viewport and viewport boundary typically move with the movement of one or more display generating components (e.g., with the user's head for head-mounted devices, or with the user's hand for handheld devices such as tablets or smartphones). The user's viewpoint determines what is visible within the viewport. The viewpoint typically specifies its position and orientation relative to the 3D environment, and as the viewpoint moves, the view of the 3D environment also moves within the viewport. For head-mounted devices, the viewpoint is usually based on the position and orientation of the user's head, face, and / or eyes to provide a perceptibly accurate view of the 3D environment that provides an immersive experience when the user is using the head-mounted device. For handheld or fixed devices, the viewpoint moves with the movement of the handheld or fixed device and / or with changes in the user's positioning relative to the handheld or fixed device (e.g., the user moves towards, away from, up, down, right, and / or left). For devices that include display generation components with virtual pass-through, a portion of the physical environment visible (e.g., displayed and / or projected) via one or more display generation components is based on the field of view of one or more cameras communicating with the display generation components, which typically move with the movement of the display generation components (e.g., for head-mounted devices, moving with the movement of the user's head, or for handheld devices such as tablets or smartphones, moving with the movement of the user's hand), because the user's viewpoint moves with the movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual objects are updated based on the movement of the user's viewpoint)).For a display generating component with optical transparency, portions of the physical environment visible through one or more display generating components (e.g., optically visible through one or more portions or fully transparent portions of the display generating component) are based on the user's field of view through the portion or fully transparent portion of the display generating component (e.g., for a head-mounted device, it moves with the movement of the user's head, or for a handheld device such as a tablet or smartphone, it moves with the movement of the user's hand), because the user's viewpoint moves with the movement of the user's field of view through the portion or fully transparent portion of the display generating component (and the appearance of one or more virtual objects is updated based on the user's viewpoint).
[0111] In some embodiments, the representation of the physical environment (e.g., displayed via virtual passthrough or optical passthrough) may be partially or completely occluded by the virtual environment. In some embodiments, the amount of virtual environment displayed (e.g., the amount of physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally results in more virtual environment being displayed, replacing and / or occluding more physical environment, and decreasing the immersion level optionally results in less virtual environment being displayed, thereby revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, displayed with increased transparency), and one or more third background objects are de-emphasized. In some embodiments, the level of immersion includes the associated degree to which virtual content (e.g., a virtual environment and / or virtual content) displayed by the computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual environment, optionally including the number of items of the displayed background content and / or the displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed via the display generating component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed via the display generating component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content in a representation of the physical environment). In some implementations, the background content includes a user interface (e.g., a user interface corresponding to an application generated by a computer system), virtual objects not associated with or included in the virtual environment and / or virtual content (e.g., files generated by the computer system or representations of other users), and / or real objects (e.g., transparent objects representing real objects in the user's surrounding physical environment, visible such that they are displayed via display generation components and / or via transparent or semi-transparent components of the display generation components, because the computer system does not obscure / impede their visibility through the display generation components). In some implementations, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real objects are displayed in an unobstructed manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with the background content, which is optionally displayed at full brightness, color, and / or semi-transparency.In some implementations, at higher immersion levels (e.g., a second immersion level above the first immersion level), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full-screen or fully immersive mode). Alternatively, a virtual environment displayed at a medium immersion level is displayed simultaneously with darkened, blurred, or otherwise de-emphasized background content. In some implementations, the visual characteristics of background objects differ between background objects. For example, at a particular immersion level, one or more first background objects are visually de-emphasized more than one or more second background objects (e.g., dimmed, blurred, and / or displayed with increased transparency), and one or more third background objects are de-emphasized. In some implementations, zero immersion or a zero immersion level corresponds to a virtual environment that is de-emphasized, and instead, a representation of the physical environment (optionally having one or more virtual objects, such as an application, window, or virtual 3D object) is displayed, without being occluded by the virtual environment. Using physical input elements to adjust immersion levels provides a quick and efficient way to adjust immersion, which enhances the operability of computer systems and makes user-device interfaces more efficient.
[0112] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same location and / or position within the user's viewpoint, the virtual object remains viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the direction forward of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); therefore, the user's viewpoint remains fixed even when the user's gaze shifts without moving the user's head. In embodiments where the computer system has a display generating component (e.g., a display screen) that can be repositioned relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the display generating component of the computer system. For example, a viewpoint-locked virtual object displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint, even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation of a viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In an implementation where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, so the virtual object is also referred to as a "head-locked virtual object".
[0113] Environment-locked visual objects: When a computer system displays a virtual object at a location and / or position within the user's viewpoint, the virtual object is environment-locked (or, "world-locked"), the location and / or position being based on a location and / or object in a three-dimensional environment (e.g., a physical or virtual environment) (e.g., selected and / or anchored to that location and / or object with reference to it). As the user's viewpoint moves, the location and / or object in the environment relative to the user's viewpoint changes, causing the environment-locked virtual object to appear at different locations and / or positions within the user's viewpoint. For example, an environment-locked virtual object locked to a tree immediately in front of the user appears at the center of the user's viewpoint. When the user's viewpoint shifts to the right (e.g., the user's head turns to the right) so that the tree is now centered to the left in the user's viewpoint (e.g., the tree's position shifts in the user's viewpoint), the environment-locked virtual object locked to the tree appears centered to the left in the user's viewpoint. In other words, the position and / or orientation of an environment-locked virtual object displayed in the user's viewpoint depends on the position to which the virtual object is locked and / or the orientation and / or orientation of the object within the environment. In some implementations, the computer system uses a stationary frame of reference (e.g., a coordinate system anchored to a fixed position and / or object in the physical environment) to determine the position of the environment-locked virtual object displayed in the user's viewpoint. An environment-locked virtual object may be locked to a stationary part of the environment (e.g., a floor, wall, table, or other stationary object), or it may be locked to a movable part of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body that moves independently of the user's viewpoint, such as a hand, wrist, arm, or foot), causing the virtual object to move with the viewpoint or that part of the environment to maintain a fixed relationship between the virtual object and that part of the environment.
[0114] In some implementations, environment-locked or viewpoint-locked virtual objects exhibit lazy following behavior, reducing or delaying their movement relative to the movement of a reference point they are following. In some implementations, when exhibiting lazy following behavior, the computer system intentionally delays the movement of the virtual object when movement of the reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) is detected. For example, when the reference point (e.g., a portion of the environment or the viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up). In some implementations, when the virtual object exhibits lazy following behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold amount, such as 0 to 5 degrees or 0 to 50 cm). For example, when the reference point (e.g., the portion of the environment to which the virtual object is locked or the viewpoint) moves by a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment to which the virtual object is locked), and when the reference point (e.g., the portion of the environment to which the virtual object is locked or the viewpoint) moves by a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases to above a threshold (e.g., a "lazy following" threshold), because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some implementations, maintaining a substantially fixed position of the virtual object relative to a reference point includes displaying the virtual object within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the reference point).
[0115] Hardware: Many different types of electronic systems enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays shaped as lenses designed to be placed over a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet devices, and desktop / laptop computers. Head-mounted systems may include speakers and / or other audio output devices integrated into the system for providing audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or semi-transparent display instead of an opaque display. A transparent or semi-transparent display may have a medium through which light representing an image is directed to the human eye. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, a transparent or semi-transparent display may be configured to selectively become opaque. Projection-based systems may employ retinal projection techniques that project graphic images onto the human retina. Projection systems may also be configured to project virtual objects into a physical environment, such as as holograms or onto a physical surface. In some embodiments, controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 2The controller 110 is described in more detail. In some embodiments, the controller 110 is a computing device located locally or remotely relative to scene 105 (e.g., physical environment). For example, the controller 110 is a local server located within scene 105. Alternatively, the controller 110 is a remote server (e.g., a cloud server, central server, etc.) located outside scene 105. In some embodiments, the controller 110 is communicatively coupled to display generation components 120 (e.g., HMD, monitor, projector, touchscreen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., physical enclosure) of display generation component 120 (e.g., HMD or portable electronic device including display and one or more processors), one or more input devices in input device 125, one or more output devices in output device 155, one or more sensors in sensor 190, and / or one or more peripheral devices in peripheral device 195, or shares the same physical housing or support structure with one or more of the aforementioned devices.
[0116] In some embodiments, the display generation component 120 is configured to provide an XR experience to a user (e.g., at least the visual component of the XR experience). In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. The following is relative to... Figure 3 The display generation component 120 is described in more detail. In some embodiments, the functionality of the controller 110 is provided by and / or combined with the display generation component 120.
[0117] According to some implementation schemes, the display generation component 120 provides an XR experience to the user when the user is virtually and / or physically present in scene 105.
[0118] In some embodiments, the display generating component is worn on a part of the user's body (e.g., on his / her head, his / her hand, etc.). Thus, the display generating component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generating component 120 surrounds the user's field of view. In some embodiments, the display generating component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device having a display facing the user's field of view and a camera facing scene 105. In some embodiments, the handheld device is optionally placed within a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generating component 120 is an XR chamber, housing, or room configured to present XR content, wherein the user does not wear or hold the display generating component 120. Many user interfaces described with reference to one type of hardware used for displaying XR content (e.g., a handheld device or a tripod-mounted device) can be implemented on another type of hardware used for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface illustrating interaction with XR content triggered by an interaction occurring in the space in front of a handheld device or tripod-mounted device can be similarly implemented using an HMD, where the interaction occurs in the space in front of the HMD and the response to the XR content is displayed via the HMD. Similarly, a user interface illustrating interaction with XR content triggered by movement of a handheld device or tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented using an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).
[0119] Despite Figure 1A The relevant features of the operating environment 100 are shown in this disclosure, but those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and so as not to obscure further relevant aspects of the exemplary embodiments disclosed herein.
[0120] Figures 1A to 1PVarious examples of computer systems for performing methods and providing audio, visual, and / or haptic feedback as part of the user interface described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., first and second display components 1-120a, 1-120b and / or first and second optical modules 11.1.1-104a and 11.1.1-104b) for displaying virtual elements and / or representations of the physical environment to a user of the computer system, the virtual elements and / or the representations of the physical environment optionally generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216 to make it easier for a user who would otherwise use glasses or contact lenses to correct their vision to view the user interface, the one or more corrective lenses optionally being removably attached to one or more optical modules in the optical modules. While many user interfaces illustrated herein represent a single view of the user interface, user interfaces in HMDs optionally employ two optical modules (e.g., first display component 1-120a and second display component 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b) for display, one optical module for the user's right eye and a different optical module for the user's left eye, presenting slightly different images to the two different eyes to generate the illusion of stereoscopic depth. A single view of the user interface is typically a right-eye view or a left-eye view, and the depth effect is explained in text or using other diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display component 1-108) for displaying status information of the computer system to the user of the computer system (when the computer system is not worn) and / or to others near the computer system, the status information optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., one or more sensors in sensor assemblies 1-356, and / or...) for detecting information about the physical environment of the device. Figure 1I This information can be used (optionally in conjunction with one or more illuminators, such as...) Figure 1IThe illuminator described herein generates a digital pass-through image, captures visual media corresponding to the physical environment (e.g., photographs and / or videos), or determines the pose (e.g., position and / or orientation) of physical objects and / or surfaces in the physical environment, enabling virtual objects to be placed based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors (e.g., sensor assemblies 1-356 and / or...) for detecting hand position and / or movement. Figure 1I One or more sensors), which can be used (optionally in conjunction with one or more illuminators, such as Figure 1I The illuminator 6-124 described herein determines when one or more air gestures are performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., ...). Figure 1I Eye-tracking and gaze-tracking sensors in the system), these sensors can be used (optionally combined with one or more lights, such as...) Figure 10The light (11.3.2-110) in the image determines attention or gaze position and / or gaze movement, which may optionally be used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine user facial expressions and / or hand movements for generating an avatar or representation of the user, such as an anthropomorphic avatar or representation for real-time communication sessions, wherein the avatar has facial expressions, hand movements, and / or body movements detected by the user based on or similar to the device. Gaze and / or attention information may optionally be combined with hand tracking information to determine user interaction with one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132 and / or dials or buttons 1-328), knobs (e.g., first buttons 1-128, buttons 11.1.1-114 and / or dials or buttons 1-328), digital crowns (e.g., pressable and twistable or rotatable first buttons 1-128, buttons 11.1.1-114 and / or dials or buttons 1-328), touchpads, touchscreens, keyboards, mice and / or other input devices. One or more buttons (e.g., first buttons 1-128, buttons 11.1.1-114, second buttons 1-132, and / or dials or buttons 1-328) are optionally used to perform system operations, such as recentering content in the user-visible 3D environment of the device, displaying the main user interface for launching an application, initiating a real-time communication session, or initiating the display of a virtual 3D background. Knobs or digital crowns (e.g., pressable and twistable or rotatable first buttons 1-128, buttons 11.1.1-114, and / or dials or buttons 1-328) are optionally rotatable to adjust parameters of the visual content, such as the level of immersion of the virtual 3D environment (e.g., the extent to which the virtual content occupies the user's viewport in the 3D environment) or other parameters associated with the 3D environment and the virtual content displayed via optical modules (e.g., first display components 1-120a and second display components 1-120b and / or first optical modules 11.1.1-104a and second optical modules 11.1.1-104b).
[0121] Figure 1BExamples of head-mounted display (HMD) devices 1-100 configured to be worn by a user and provide virtual and altered / mixed reality (VR / AR) experiences are illustrated in front, top, and perspective views. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strip assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 secured at either end to the electronic strip assembly 1-104. The electronic strip assembly 1-104 and the strap 1-106 may be part of a retention assembly configured to wrap around the user's head to hold the display unit 1-102 against the user's face.
[0122] In at least one example, the band assembly 1-106 may include a first band 1-116 configured to wrap around the back of the user's head and a second band 1-117 configured to extend above the top of the user's head. As shown, the second band may extend between the first electronic band 1-105a and the second electronic band 1-105b of the electronic band assembly 1-104. The band assembly 1-104 and the band assembly 1-106 may be part of a fixing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.
[0123] In at least one example, the fixing mechanism includes a first electronic strip 1-105a, which includes a first proximal end 1-134 coupled to a display unit 1-102 (e.g., a housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite to the first proximal end 1-134. The fixing mechanism may also include a second electronic strip 1-105b, which includes a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite to the second proximal end 1-138. The fixing mechanism may also include a first strip 1-116 and a second strip 1-117, the first strip including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second strip extending between the first electronic strip 1-105a and the second electronic strip 1-105b. Strips 1-105a to b and strip 1-116 may be coupled via a connecting mechanism or component 1-114. In at least one example, the second strip 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between a first proximal end 1-134 and a first distal end 1-136, and a second end 1-148 coupled to the second electronic strip 1-105b between a second proximal end 1-138 and a second distal end 1-140.
[0124] In at least one example, the first electronic strip and the second electronic strips 1-105a to b comprise plastic, metal, or other structural materials forming the shape of the substantially rigid strips 1-105a to b. In at least one example, the first strip and the second strips 1-116, 1-117 are formed of an elastic flexible material (including woven textiles, rubber, etc.). The first strip 1-116 and the second strip 1-117 may be flexible enough to conform to the shape of the user's head when wearing the HMD 1-100.
[0125] In at least one example, one or more of the first electronic stripe and the second electronic stripe 1-105a to b may define an inner stripe volume and include one or more electronic components disposed within the inner stripe volume. In one example, such as Figure 1B As shown, the first electronic strip 1-105a may include electronic components 1-112. In one example, electronic components 1-112 may include a speaker. In another example, electronic components 1-112 may include computing components, such as a processor.
[0126] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is located in... Figure 1B The section marked 1-152 with dashed lines is because the display assembly 1-108 is configured to obscure the first opening 1-152 from a view when the HMD 1-100 is assembled. The housing 1-150 may also define a rearward second opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover disposed in or across the front opening 1-152 to obscure the front opening 1-152 and a display screen (shown in other figures). In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 in general, has a curvature configured to follow the curvature of the user's face. The display screen of display component 1-108 can be bent as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, wherein display unit 1-102 is pressed.
[0127] In at least one example, the housing 1-150 may define a first hole 1-126 between a first opening 1-152 and a second opening 1-154, and a second hole 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first hole 1-128, and a second button 1-132 disposed in the second hole 1-130. The first button 1-128 and the second button 1-132 are pressable through their respective holes 1-126 and 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 may be a rotary dial and a pressable button. In at least one example, the first button 1-128 is a pressable and rotary dial button, and the second button 1-132 is a pressable button.
[0128] Figure 1C A rear perspective view of HMD 1-100 is illustrated. HMD 1-100 may include a light seal 1-110 extending rearwardly around the periphery of housing 1-150 of display assembly 1-108, as shown. The light seal 1-110 may be configured to extend from housing 1-150 to the user's face, surrounding the user's eyes, to block external light from being visible. In one example, HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b disposed at or within a rearwardly facing second opening 1-154 defined by housing 1-150 and / or disposed within the internal volume of housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a to b may include a corresponding display screen 1-122a, 1-122b configured to project light toward the user's eyes in a rearward direction through the second opening 1-154.
[0129] In at least one example, reference Figure 1B and Figure 1C Both, the display assembly 1-108 can be a front-facing display assembly including a display screen configured to project light in a first forward direction, and the rear display screens 1-122a to b can be configured to project light in a second rearward direction opposite to the first direction. As described above, the light seal 1-110 can be configured to block light from outside the HMD 1-100 from reaching the user's eyes, including a component made of... Figure 1B The front perspective view shows the light projected by the front display screen of the display assembly 1-108. In at least one example, the HMD 1-100 may also include a curtain 1-124 that blocks the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a to b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.
[0130] Figure 1B and Figure 1C Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1D to 1F Any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1D to 1F Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1B and Figure 1C Examples of devices, features, components, and parts are shown.
[0131] Figure 1D An exploded view of an example HMD 1-200 including its various parts or components, separated according to the modularity and selective coupling of these components. For example, HMD 1-200 may include a strip 1-216 selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first strip and the second strips 1-205a to 1-205b are removably coupled to a display unit 1-202.
[0132] Furthermore, HMD 1-200 may include a light-sealing member 1-210 configured to be removably coupled to display unit 1-202. HMD 1-200 may also include a lens 1-218, which may be removably coupled to display unit 1-202, for example, on a first display assembly and a second display assembly including a display screen. Lens 1-218 may include a custom prescription lens configured for vision correction. As noted, in Figure 1D The exploded view shows that each component described above can be removably coupled, attached, reattached, and replaced to update the component, or replaced for different users. For example, belts such as belt 1-216, light seals such as light seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a to b can be replaced according to the user, so that these parts are customized to fit and correspond to a single user of HMD 1-200.
[0133] Figure 1D Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1B , Figure 1C and Figures 1E to 1FAny other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1B , Figure 1C and Figures 1E to 1F Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1D Examples of devices, features, components, and parts are shown.
[0134] Figure 1E An exploded view illustrating an example of a display unit 1-306 of an HMD is shown. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320, which includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.
[0135] In at least one example, the display unit 1-306 may further include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the positioning of the display screens 1-322a to b of the display unit 1-320 relative to the frame 1-350. In at least one example, the display unit 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a to b has at least one motor, such that the motor is capable of translating the display screens 1-322a to b to match the interpupillary distance of the user's eyes.
[0136] In at least one example, display unit 1-306 may include a dial or button 1-328 that is pressable relative to frame 1-350 and accessible to a user outside frame 1-350. Button 1-328 may be electrically connected to motor assembly 1-362 via a controller, such that button 1-328 can be operated by a user to cause the motor of motor assembly 1-362 to adjust the positioning of display screens 1-322a to b.
[0137] Figure 1E Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1B to 1D and Figure 1F Any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1B to 1D and Figure 1FAny of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1E Examples of devices, features, components, and parts are shown.
[0138] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positioning of the first display sub-assemblies 1-420a and 1-420b of the rear display assembly 1-421, including a first and second corresponding display screen for interpupillary adjustment, as described above.
[0139] Figure 1F The exploded views shown in this article refer to the various parts, systems, and components. Figures 1B to 1E The following figures, which are referenced in this disclosure, provide a more detailed description. Figure 1F The display unit 1-406 shown can be connected with Figures 1B to 1E The fastening mechanism assembly and integration shown includes electronic strips, belts, and other components including light seals, connecting assemblies, etc.
[0140] Figure 1F Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1B to 1E Any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1B to 1E Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1F Examples of devices, features, components, and parts are shown.
[0141] Figure 1G An exploded perspective view of the front cover assembly 3-100 of the HMD device described herein is illustrated, for example... Figure 1G The front cover assembly 3-1 of the HMD 3-100 shown herein or any other HMD device shown and described herein. Figure 1GThe front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or “cover”), an adhesive layer 3-106, a display assembly 3-108 including a biconvex lens panel or array 3-110, and a structural decorative element 3-112. The adhesive layer 3-106 secures the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the decorative element 3-112. The decorative element 3-112 secures various components of the front cover assembly 3-100 to the frame or base of the HMD device.
[0142] In at least one example, such as Figure 1G As shown, the transparent cover 3-102, the protective cover 3-104, and the display assembly 3-108 including a biconvex lens array 3-110 can be bent to adapt to the curvature of a user's face. The transparent cover 3-102 and the protective cover 3-104 can be bent in two or three dimensions, for example, vertically in and out of the Z-plane along the Z direction, and horizontally in and out of the Z-plane along the X direction. In at least one example, the display assembly 3-108 may include the biconvex lens array 3-110 and a display panel with pixels configured to project light through the protective cover 3-104 and the transparent cover 3-102. The display assembly 3-108 can be bent in at least one direction (e.g., the horizontal direction) to adapt to the curvature of a user's face from one side (e.g., the left) to the other (e.g., the right). In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in the following figures, but may include the biconvex lens array 3-110 and the display layer) may be similarly or concentrically curved in the horizontal direction to accommodate the curvature of the user's face.
[0143] In at least one example, the cover 3-104 may include a transparent or translucent material through which the display component 3-108 projects light. In one example, the cover 3-104 may include one or more opaque portions, such as opaque ink-printed portions or other opaque film portions on the back of the cover 3-104. When the HMD device is worn, the rear surface may be the surface of the cover 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the cover 3-104 opposite the rear surface. In at least one example, one or more opaque portions of the cover 3-104 may include peripheral portions that visually conceal any components surrounding the outer periphery of the display screen of the display component 3-108. In this way, the opaque portions of the cover conceal any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the cover 3-104, including electronic components, structural components, etc.
[0144] In at least one example, the housing 3-104 may define one or more transparent aperture portions 3-120 through which sensors can transmit and receive signals. In one example, portion 3-120 is an aperture through which sensors can extend or transmit and receive signals. In one example, portion 3-120 is a transparent portion, or a portion more transparent than the surrounding translucent or opaque portion of the housing, through which sensors can transmit and receive signals through the housing and via transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.
[0145] Figure 1G Any of the features, components, and / or parts shown herein (including their arrangement and configuration) may be included, individually or in any combination, in any other example of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included, individually or in any combination. Figure 1G Examples of devices, features, components, and parts are shown.
[0146] Figure 1H An exploded view of an example HMD device 6-100 is shown. The HMD device 6-100 may include a sensor array or system 6-102, which includes one or more sensors, cameras, projectors, etc., mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 on which one or more sensors of the sensor system 6-102 may be fixed / secured.
[0147] Figure 1I A portion of an HMD device 6-100, including a front transparent cover 6-104 and a sensor system 6-102, is illustrated. The sensor system 6-102 may include multiple different sensors, transmitters, and receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is illustrated on the front of the sensor system 6-102 to illustrate the relative positioning of the various sensors and transmitters and the orientation of each sensor / transmitter in system 6-102. As referenced herein, "side," "side," "lateral," "horizontal," and other similar terms refer to... Figure 1J The orientation or direction indicated by the X-axis. Terms such as "vertical," "upward," "downward," and similar terms refer to the orientation or direction indicated by... Figure 1JThe orientation or direction indicated by the Z-axis. Terms such as "frontward," "rearward," "forward," "backward," and similar terms refer to the orientation or direction indicated by the Z-axis. Figure 1J The orientation or direction indicated by the Y-axis shown.
[0148] In at least one example, a transparent cover 6-104 may define the front outer surface of an HMD device 6-100, and a sensor system 6-102, including various sensors and their components, may be positioned behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through it, including both light detected by the sensor system 6-102 and light emitted therefrom.
[0149] As described elsewhere herein, the HMD device 6-100 may include one or more controllers, which include processors for electrically coupling various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as displays. Furthermore, as will be shown in more detail below with reference to other accompanying drawings, various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to the HMD device 6-100. Figure 1I Various structural frame components, brackets, etc., not shown. For clarity, Figure 1I The components of the sensor system 6-102 are shown, which are not attached to or electrically coupled to other components.
[0150] In at least one example, the device may include one or more controllers having a processor configured to execute instructions stored on a memory component electrically coupled to the processor. These instructions may include, or cause the processor to execute, one or more algorithms for self-correcting the angle and position of the various cameras described herein as the camera's initial position, angle, or orientation is affected by collisions or deformations due to accidental drop events or other events over time.
[0151] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. System 6-102 may include two scene cameras 6-102, respectively positioned on either side of the nose bridge or arched structure of the HMD device 6-100, such that each of the two cameras 6-106 approximately corresponds to the positioning of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and, when the HMD device 6-100 is used, provide images and content for MR video pass-through to a display screen facing the user's eyes. The scene cameras 6-106 may also be used for environment and object reconstruction.
[0152] In at least one example, sensor system 6-102 may include a first depth sensor 6-108 that is generally forward-pointing in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction as well as user hand and body tracking. In at least one example, sensor system 6-102 may include a second depth sensor 6-110 centrally located along the width of HMD device 6-100 (e.g., along the X-axis). For example, the second depth sensor 6-110 may be located above the central bridge of the nose or on an adapter structure above the nose when the user wears HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction as well as hand and body tracking. In at least one example, the second depth sensor may include a LiDAR sensor.
[0153] In at least one example, the sensor system 6-102 may include a depth projector 6-112, which is typically forward-facing to project electromagnetic waves (e.g., in the form of a predetermined spot pattern) into or within the field of view of the user and / or scene camera 6-106, or into or beyond the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a spot pattern, which are reflected from an object and back into the aforementioned depth sensors, including depth sensors 6-108 and 6-110. In at least one example, the depth projector 6-112 may be used for environment and object reconstruction, as well as hand and body tracking.
[0154] In at least one example, the sensor system 6-102 may include a downward-facing camera 6-114, whose field of view is generally directed downwards relative to the HMD device 6-100 on the Z-axis. In at least one example, the downward-facing camera 6-114 may be positioned as shown on the left and right sides of the HMD device 6-100 and used for hand and body tracking, head-mounted device tracking, and facial avatar detection and creation for displaying a user avatar on the front display screen of the HMD device 6-100 as described elsewhere herein. For example, the downward-facing camera 6-114 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.
[0155] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be positioned as shown on the left and right sides of the HMD device 6-100 and used for hand and body tracking, head-mounted device tracking, and facial avatar detection and creation for displaying a user avatar on the front display screen of the HMD device 6-100 as described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. Used for hand and body tracking, head-mounted device tracking, and facial avatar creation.
[0156] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right views along the X-axis or in a direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, head-mounted device tracking, and facial avatar detection and reconstruction.
[0157] In at least one example, the sensor system 6-102 may include multiple eye-tracking and gaze-tracking sensors for determining identity, status, and the user's gaze direction during and / or prior to use. In at least one example, the eye / gaze-tracking sensor may include a nose-eye camera 6-120 positioned on either side of the user's nose and adjacent to the user's nose when wearing the HMD device 6-100. The eye / gaze sensor may also include a bottom eye camera 6-122 positioned below the respective user's eye for capturing images of the eye for use in facial avatar detection and creation, gaze tracking, and iris identification functions.
[0158] In at least one example, sensor system 6-102 may include an infrared illuminator 6-124 that is pointed outward from HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of sensor system 6-102. In at least one example, sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, flicker sensor 6-126 may detect the refresh rate of the overhead light to avoid display flicker. In one example, infrared illuminator 6-124 may include a light-emitting diode and may be specifically designed for low-light environments to illuminate a user's hands and other objects in low light for detection by the infrared sensors of sensor system 6-102.
[0159] In at least one example, multiple sensors (including scene camera 6-106, downward camera 6-114, chin camera 6-116, side camera 6-118, depth projector 6-112, and depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for sizing, thereby improving the hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, as described above and Figure 1I The downward-facing camera 6-114, the jaw camera 6-116, and the side camera 6-118 shown can be wide-angle cameras capable of operating in both the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, and 6-118 can operate solely in black-and-white light detection to simplify image processing and achieve sensitivity.
[0160] Figure 1I Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1J to 1L Any other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figures 1J to 1L Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1I Examples of devices, features, components, and parts are shown.
[0161] Figure 1JA lower perspective view of an example HMD 6-200 including a cover or shield 6-204 fixed to a frame 6-230 is shown. In at least one example, a sensor 6-203 of a sensor system 6-202 may be disposed around the periphery of the HMD 6-200 such that the sensor 6-203 is disposed outwardly around the periphery of the display area or region 6-232 so as not to obstruct the view of the displayed light. In at least one example, the sensor may be disposed behind the shield 6-204 and aligned with a transparent portion of the shield, thereby allowing light to pass back and forth through the shield 6-204 by the sensor and the projector. In at least one example, an opaque ink or other opaque material or film / layer may be disposed on the shield 6-204 around the display area 6-232 to conceal components of the HMD 6-200 outside the display area 6-232 rather than through a transparent portion defined by the opaque portion through which the sensor and the projector transmit and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass through the display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area surrounding the periphery of the display and the shield 6-204.
[0162] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent areas 6-209 through which the sensor 6-203 of the sensor system 6-202 transmits and receives signals. In the illustrated examples, the sensor 6-203 of the sensor system 6-202, which transmits and receives signals through the shield 6-204, or more specifically through the transparent area 6-209 defined by the opaque portion 6-207 of the shield 6-204, may include... Figure 1I The examples illustrate those same or similar sensors, such as depth sensors 6-108 and 6-110, depth projector 6-112, first scene camera and second scene camera 6-106, first downward camera and second downward camera 6-114, first side camera and second side camera 6-118, and first infrared illuminator and second infrared illuminator 6-124. These sensors also... Figure 1K and Figure 1L The example is shown. Other sensors, sensor types, number of sensors, and their relative positioning can be included in one or more other examples of the HMD.
[0163] Figure 1J Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1I and Figures 1K to 1LAny other example of the devices, features, components, and parts shown and described herein. Similarly, refer to... Figure 1I and Figures 1K to 1L Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1J Examples of devices, features, components, and parts are shown.
[0164] Figure 1K A front view of a portion of an example of an HMD device 6-300, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330, is shown. Figure 1K The examples shown do not include a front cover or shield to illustrate brackets 6-336 and 6-338. For example, Figure 1J The shield 6-204 shown includes an opaque portion 6-207 that visually covers / blocks the view of anything outside the display / display area 6-334 (e.g., radially / peripherally outside the display / display area), including the sensor 6-303 and the bracket 6-338.
[0165] In at least one example, various sensors of sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, scene camera 6-306 includes strict tolerances for angles relative to each other. For example, the tolerance for the mounting angle between two scene cameras 6-306 may be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such strict tolerances, in one example, scene camera 6-306 may be mounted to bracket 6-338 instead of a housing. The bracket may include a cantilever on which scene camera 6-306 and other sensors of sensor system 6-302 may be mounted to maintain their positioning and orientation in the event of a drop event caused by a user that results in any deformation of other brackets 6-226, housing 6-330, and / or housing.
[0166] Figure 1K Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1I to 1J and Figure 1L This is illustrated in any of the other examples of devices, features, components, and parts described herein. Similarly, refer to... Figures 1I to 1J and Figure 1L Any of the features, components, and / or parts shown or described (including their arrangement and configuration) may be included individually or in any combination. Figure 1K Examples of devices, features, components, and parts are shown.
[0167] Figure 1LA bottom view illustrating an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is shown. The sensor system 6-402 may be similar to other sensor systems described above and elsewhere herein, including references to… Figures 1I to 1K As described above. In at least one example, the jaw camera 6-416 may be oriented downwards to capture images of the user's lower facial features. In one example, the jaw camera 6-416 may be directly coupled to a frame or housing 6-430 or one or more internal brackets that are directly coupled to the frame or housing 6-430. The frame or housing 6-430 may include one or more holes / openings 6-415 through which the jaw camera 6-416 transmits and receives signals.
[0168] Figure 1L Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figures 1I to 1K This is illustrated in any of the other examples of devices, features, components, and parts described herein. Similarly, refer to... Figures 1I to 1K Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination. Figure 1L Examples of devices, features, components, and parts are shown.
[0169] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is illustrated. This IPD adjustment system includes a first optical module and a second optical module 11.1.1-104a-104a-105a-106 ... In at least one example, buttons 11.1.1-114 can be electrically communicated with the first motor and the second motors 11.1.1-110a to b via a processor or other circuit components to activate the first motor and the second motors 11.1.1-110a to b and respectively cause the first optical module and the second optical modules 11.1.1-104a to b to change their positions relative to each other.
[0170] In at least one example, the first and second optical modules 11.1.1-104a to b may include corresponding display screens configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, the user can manipulate (e.g., press and / or rotate) buttons 11.1.1-114 to activate positional adjustment of the optical modules 11.1.1-104a to b to match the interpupillary distance of the user's eyes. The optical modules 11.1.1-104a to b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD, such that the optical modules 11.1.1-104a to b can be adjusted to match the IPD.
[0171] In one example, a user can manipulate buttons 11.1.1-114 to cause automatic positional adjustment of the first and second optical modules 11.1.1-104a to b. In another example, a user can manipulate buttons 11.1.1-114 to cause manual adjustment, moving the optical modules 11.1.1-104a to b further or closer (e.g., when the user rotates buttons 11.1.1-114 in one way or another) until the user visually matches their own IPD. In one example, manual adjustment is communicated electronically via one or more circuits, and the power for moving the optical modules 11.1.1-104a to b via motors 11.1.1-110a to b is supplied by a power source. In another example, the adjustment and movement of the optical modules 11.1.1-104a to b via actuating buttons 11.1.1-114 are mechanically actuated via moving buttons 11.1.1-114.
[0172] Figure 1M Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, individually or in any combination, in any other example of the devices, features, components, and parts shown and described herein. Similarly, any of the features, components, and / or parts shown and described, and described herein, with reference to any other shown and described figures (including their arrangement and configuration), may be included, individually or in any combination, in any other example of the devices, features, components, and / or parts shown and / or described herein. Figure 1M Examples of devices, features, components, and parts are shown.
[0173] Figure 1N A front perspective view of a portion of HMD 11.1.2-100 is shown, including an outer structural frame 11.1.2-102 defining first and second holes 11.1.2-106a, 11.1.2-106b, and an inner or intermediate structural frame 11.1.2-104. Holes 11.1.2-106a to b are located in... Figure 1NThe holes 11.1.2-106a to b are shown in dashed lines because viewing the HMD 11.1.2-100 may be obstructed by one or more other components coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as illustrated. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second holes 11.1.2-106a to b.
[0174] Mounting brackets 11.1.2-108 may include intermediate or central portions 11.1.2-109 coupled to the inner frame 11.1.2-104. In some examples, the intermediate or central portions 11.1.2-109 may not be the geometric center or middle of the brackets 11.1.2-108. Instead, the intermediate / central portions 11.1.2-109 may be positioned between a first cantilever extension arm and a second cantilever extension arm extending away from the intermediate portions 11.1.2-109. In at least one example, mounting bracket 108 includes first cantilever arms 11.1.2-112 and second cantilever arms 11.1.2-114 extending away from the intermediate portions 11.1.2-109 of the mounting brackets 11.1.2-108 coupled to the inner frame 11.1.2-104.
[0175] like Figure 1N As shown, the outer frame 11.1.2-102 may define a curved geometry on its lower side to adapt to the user's nose when the user wears the HMD 11.1.2-100. This curved geometry may be referred to as the bridge of the nose 11.1.2-111 and is centrally located on the lower side of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the inner frame 11.1.2-104 between holes 11.1.2-106a and b, such that the cantilever 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the central portion 11.1.2-109 to complement the nose bridge geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to adapt to the user's nose, as described above. The geometry of the bridge of the nose 11.1.2-111 adapts to the nose, as it provides a curvature that conforms to the shape of the user's nose, offering a comfortable fit from above, above, and around.
[0176] The first cantilever 11.1.2-112 may extend in a first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108, and the second cantilever 11.1.2-114 may extend in a second direction opposite to the first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as “cantilever” or “cantilever” arms because each arm 11.1.2-112, 11.1.2-114 includes free distal ends 11.1.2-116, 11.1.2-118, respectively, which are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, arms 11.1.2-112 and 11.1.2-114 extend from the middle section 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102 and 11.1.2-104 are not attached.
[0177] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include multiple sensors 11.1.2-110a-f. Each of the multiple sensors 11.1.2-110a-f may include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a-f may be used for object recognition in three-dimensional space, making it important to maintain the precise relative positioning of two or more of the multiple sensors 11.1.2-110a-f. The cantilever nature of the mounting bracket 11.1.2-108 protects the sensors 11.1.2-110a-f from damage and displacement in the event of an accidental drop by the user. Because the sensors 11.1.2-110a-f cantilevered on the arms 11.1.2-112 and 11.1.2-114 of the mounting bracket 11.1.2-108, the stress and deformation of the internal frame and / or the external frames 11.1.2-104 and 11.1.2-102 are not transmitted to the cantilever arms 11.1.2-112 and 11.1.2-114, and therefore do not affect the relative position of the sensors 11.1.2-110a-f coupled to / mounted to the mounting bracket 11.1.2-108.
[0178] Figure 1NAny of the features, components, and / or parts shown herein (including their arrangement and configuration) may be included individually or in any combination of any other examples of the devices, features, components, and other examples described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination of any other examples of the devices, features, components, and other examples described herein. Figure 1N Examples of devices, features, components, and parts are shown.
[0179] Figure 10 An example of optical modules 11.3.2-100 for use in electronic devices, such as HMDs, including the HMD devices described herein, is illustrated. As shown in one or more other examples described herein, optical module 11.3.2-100 may be one of two optical modules within an HMD, wherein each optical module is aligned to project light toward a user's eye. In this way, a first optical module may project light toward a user's first eye via a display screen, and a second optical module of the same device may project light toward a user's second eye via another display screen.
[0180] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a tube or optical module tube. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward the user's eyes when the HMD to which the display module 11.3.2-100 belongs is worn during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.
[0181] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eye during use. In at least one example, the optical module 11.3.2-100 may also include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the camera 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light-emitting diodes (LEDs) or other lights configured to project light toward the user's eyes when the HMD is worn. The individual lights 11.3.2-110 in the light strips 11.3.2-108 may be spaced apart around the light strips 11.3.2-108, and are therefore uniformly or non-uniformly spaced around the displays 11.3.2-104 at various locations on the light strips 11.3.2-108 and around the displays 11.3.2-104.
[0182] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light onto the user's eyes through the viewing opening 11.3.2-101. In one example, a camera 11.3.2-106 is configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.
[0183] As mentioned above, Figure 10 Each of the components and features of the optical modules 11.3.2-100 shown can be replicated in another (e.g., a second) optical module set up with the HMD to interact with the user's other eye (e.g., projecting light and capturing images).
[0184] Figure 10 Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included individually or in any combination. Figure 1P Any of the other examples of devices, features, components, and parts shown or otherwise described herein. Similarly, refer to... Figure 1P Or any of the features, components and / or parts (including their arrangement and configuration) shown or described herein may be included individually or in any combination thereof. Figure 10Examples of devices, features, components, and parts are shown.
[0185] Figure 1P A cross-sectional view of an example optical module 11.3.2-200 is shown, which includes a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or channel 11.3.2-212 and a second aperture or channel 11.3.2-214. Channels 11.3.2-212 and 11.3.2-214 can be configured to slidably engage corresponding tracks or guides of an HMD device to allow the optical module 11.3.2-200 to be adjusted and positioned relative to the user's eye to match the user's interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage the guide rod to secure the optical module 11.3.2-200 in the appropriate position within the HMD.
[0186] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and positioned between the display assembly 11.3.2-204 and the user's eye when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eye. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, lenses 11.3.2-216 are positioned above light strips 11.3.2-208 and one or more eye-tracking cameras 11.3.2-206, such that cameras 11.3.2-206 are configured to capture an image of a user's eye through lenses 11.3.2-216, and light strips 11.3.2-208 include lamps configured to project light onto the user's eye through lenses 11.3.2-216 during use.
[0187] Figure 1P Any of the features, components, and / or parts shown herein (including their arrangement and configuration) may be included individually or in any combination of the devices, features, components, and parts described herein and in any other example of the other examples. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included individually or in any combination of the other examples of the devices, features, components, and parts described herein. Figure 1P Examples of devices, features, components, and parts are shown.
[0188] Figure 2This is a block diagram of an example controller 110 in some implementations. Although some specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the implementations disclosed herein. Therefore, as a non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.
[0189] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.
[0190] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores programs, modules, and data structures, or subsets thereof, including optional operating system 230 and XR experience module 240.
[0191] Operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., single XR experiences for one or more users, or multiple XR experiences for corresponding groups of one or more users). To this end, in various embodiments, XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0192] In some implementations, the data acquisition unit 241 is configured to acquire data from at least... Figure 1A The display generation unit 120, and optionally acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0193] In some implementations, the tracking unit 242 is configured to map scene 105, and the tracking at least shows the generated component 120 relative to... Figure 1A The tracking unit 242 tracks the location / position of scenario 105, and optionally the location / position relative to one or more of the tracking input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the location / position of one or more portions of the user's hand, and / or the location / position of one or more portions of the user's hand relative to... Figure 1A The movement of scene 105 relative to the display generating component 120 and / or relative to a coordinate system (defined relative to the user's hand). The following refers to the movement relative to... Figure 4 The hand tracking unit 244 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the user's gaze (or more broadly, the user's eyes, face, or head) relative to scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to XR content displayed via display generation component 120. The following description is relative to... Figure 5 The eye-tracking unit 243 is described in more detail.
[0194] In some implementations, coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by display generation component 120, and optionally by one or more of output device 155 and / or peripheral device 195. To this end, in various implementations, coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0195] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0196] Although the data acquisition unit 241, tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), coordination unit 246, and data transmission unit 248 are shown residing on a single device (e.g., controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, tracking unit 242 (e.g., including eye tracking unit 243 and hand tracking unit 244), coordination unit 246, and data transmission unit 248 may reside in a separate computing device.
[0197] also, Figure 2 This is used more as a functional description of various features that may exist in a particular specific implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, individually shown items can be combined, and some items can be separated. For example, Figure 2 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0198] Figure 3This is a block diagram illustrating examples of generating component 120 in some implementation schemes. Although some specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features have not been illustrated for the sake of brevity and to avoid obscuring further relevant aspects of the embodiments disclosed herein. Therefore, as a non-limiting example, in some embodiments, the display generating component 120 (e.g., HMD) includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZigBee, and / or similar interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these components and various other components.
[0199] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include inertial measurement units (IMUs), accelerometers, gyroscopes, thermometers, one or more physiological sensors (e.g., blood pressure monitors, heart rate monitors, blood oxygen sensors, blood glucose sensors, etc.), one or more microphones, one or more speakers, haptic engines, and / or one or more depth sensors (e.g., structured light, time-of-flight, etc.).
[0200] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, the one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS), and / or similar display types. In some embodiments, the one or more XR displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, display generation component 120 (e.g., HMD) includes a single XR display. In another example, display generation component 120 includes XR displays for each of the user's eyes. In some embodiments, the one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, the one or more XR displays 312 are capable of presenting either MR or VR content.
[0201] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward in order to acquire image data corresponding to the scene that the user would see in the absence of a display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.
[0202] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optional operating system 330 and XR rendering module 340.
[0203] Operating system 330 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To this end, in various embodiments, XR rendering module 340 includes a data acquisition unit 342, an XR rendering unit 344, an XR mapping generation unit 346, and a data transmission unit 348.
[0204] In some implementations, the data acquisition unit 342 is configured to acquire data from at least... Figure 1A The controller 110 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). To this end, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0205] In some implementations, the XR rendering unit 344 is configured to render XR content via one or more XR displays 312. To this end, in various implementations, the XR rendering unit 344 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0206] In some implementations, the XR mapping generation unit 346 is configured to generate XR maps based on media content data (e.g., 3D maps of mixed reality scenes or maps in which computer-generated objects can be placed to generate extended reality physical environments). To this end, in various implementations, the XR mapping generation unit 346 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0207] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, the data transmission unit 348 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.
[0208] Although the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data transmission unit 348 are shown residing in a single device (e.g., Figure 1A The display generation unit 120 is located on the display generation unit, but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR rendering unit 344, the XR mapping generation unit 346, and the data transmission unit 348 may be located in a separate computing device.
[0209] also, Figure 3 This is more of a functional description of various features that may exist in a particular specific implementation, and differs from the structural diagrams of the implementations described herein. As those skilled in the art will recognize, individually shown items can be combined, and some items can be separated. For example, Figure 3 Some functional modules shown individually may be implemented in a single module, and the various functions of a single functional block may be implemented in various implementations through one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some implementations, it depends in part on the specific combination of hardware, software and / or firmware chosen for that particular implementation.
[0210] Figure 4 This is a schematic illustration of an example embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A Controlled by hand tracking unit 244 Figure 2 The hand tracking device 140 tracks the location / position of one or more parts of a user's hand, and / or the movement of one or more parts of the user's hand relative to scene 105 of FIG. 1 (e.g., relative to a portion of the user's surrounding physical environment, relative to display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system (defined relative to the user's hand)). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0211] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures at least three-dimensional scene information including the human user's hand 406. The image sensor 404 captures images of the hand at sufficient resolution to distinguish the fingers and their corresponding positions. The image sensor 404 typically captures images of other parts of the user's body, or possibly all parts of the body, and may have scaling capabilities or be a dedicated sensor with increased magnification to capture images of the hand at the desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of scene 105, or serves as the image sensor for capturing the physical environment of scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in a way that uses the field of view of the image sensor 404 or a portion thereof to define an interaction space in which hand movements captured by the image sensor are considered input to the controller 110.
[0212] In some implementations, image sensor 404 outputs a sequence of frames containing 3D image data (and, in addition, possibly color image data) to controller 110, which extracts high-level information from the image data. This high-level information is typically provided via an application programming interface (API) to an application running on the controller, which in turn drives display generation component 120. For example, a user can interact with software running on controller 110 by moving his hand 406 and changing his hand pose.
[0213] In some embodiments, image sensor 404 projects a speckle pattern onto a scene containing hand 406 and captures an image of the projected pattern. In some embodiments, controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) via triangulation based on the lateral offset of the specks in the pattern. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene relative to a predetermined reference plane at a specific distance from image sensor 404. In this disclosure, it is assumed that image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, image sensor 404 (e.g., a hand-tracking device) may use other 3D mapping methods, such as stereo imaging or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.
[0214] In some implementations, hand tracking device 140 captures and processes time-series depth maps containing the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in image sensor 404 and / or controller 110 processes the 3D map data to extract image block descriptors of the hand from these depth maps. The software may match these descriptors with image block descriptors stored in database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D position of the user's hand joints and fingertips.
[0215] The software can also analyze the trajectories of the hand and / or fingers across multiple frames in a sequence to identify gestures. The pose estimation function described herein can be alternated with motion tracking, such that patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to find pose changes occurring in the remaining frames. Pose, motion, and gesture information is provided to an application running on controller 110 via the aforementioned API. This application can, for example, move and modify the image presented on display generation unit 120 in response to the pose and / or gesture information, or perform other functions.
[0216] In some implementations, gestures include air gestures. An air gesture is a gesture detected without the user touching an input element that is part of the device (e.g., computer system 101, one or more input devices 125 and / or hand tracking device 140) (or independent of an input element that is part of the device) and based on the detected movement of a part of the user's body (e.g., head, one or two arms, one or two hands, one or more fingers and / or one or two legs) through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., including a tapping gesture in which the hand moves a predetermined amount and / or speed in a predetermined pose, or a shaking gesture including a predetermined speed or amount of rotation of a part of the user's body)).
[0217] In some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures for interacting with an XR environment (e.g., a virtual or mixed reality environment) performed by the movement of a user's finger relative to other fingers (or portions of the user's hand). In some embodiments, air gestures are detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and are based on the detected movement of a portion of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or portion of the user's hand), and / or absolute movement of a portion of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a portion of the user's body)).
[0218] In some implementations where the input gesture is an air gesture (e.g., where the input device provides information to the computer system about which user interface element is the target of the user input in the absence of physical contact, such as contact with a user interface element displayed on a touchscreen, or contact with a mouse or touchpad to move the cursor to the user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Therefore, in implementations involving air gestures, for example, the input gesture is combined with (e.g., simultaneously) movement of the user's fingers and / or hand to detect attention (e.g., gaze) toward a user interface element to perform pinch and / or tap input, as described below.
[0219] In some implementations, input gestures directed to a user interface object are performed, either directly or indirectly, by referencing the user interface object. For example, user input is performed directly on the user interface object when the user's hand performs an input gesture at a location corresponding to the user interface object's position in the three-dimensional environment (e.g., determined based on the user's current viewpoint). In some implementations, when user attention to the user interface object (e.g., gazing) is detected, input gestures are performed indirectly on the user interface object, based on the fact that the user's hand is not positioned at a location corresponding to the user interface object's position in the three-dimensional environment at the time the user performs the input gesture. For example, for direct input gestures, the user can guide their input to the user interface object by initiating a gesture at or near a location corresponding to the user interface object's display position (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm measured from the outer edge or center of the option). For indirect input gestures, the user can guide their input to the user interface object by focusing on it (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location not corresponding to the user interface object's display position).
[0220] In some implementations, the input gestures (e.g., air gestures) used in the various examples and implementations described herein include pinch and tap inputs used in some implementations for interacting with virtual or mixed reality environments. For example, the pinch and tap inputs described below are performed as air gestures.
[0221] In some implementations, pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other, i.e., optionally followed by an immediate (e.g., within 0 to 1 second) interruption of contact. A long pinch gesture as an air gesture includes the movement of two or more fingers of the hand to contact each other for at least a threshold amount of time (e.g., at least 1 second) before an interruption of contact is detected. For example, a long pinch gesture includes the user holding a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some implementations, a double pinch gesture as an air gesture includes two (e.g., more) pinch inputs (e.g., performed by the same hand) that are detected consecutively with each other immediately (e.g., within a predefined time period). For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts the contact between two or more fingers), and performs a second pinch input within a predefined time period after releasing the first pinch input (e.g., within 1 second or 2 seconds).
[0222] In some embodiments, pinch and drag gestures as air gestures include pinch gestures (e.g., pinching gestures or long pinch gestures) performed in conjunction with (e.g., following) drag input that changes the user's hand position from a first position (e.g., the start position of the drag) to a second position (e.g., the end position of the drag). In some embodiments, the user holds the pinch gesture while performing drag input and releases the pinch gesture (e.g., opening two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and drag input are performed by the same hand (e.g., the user pinches two or more fingers together to touch each other and uses the drag gesture to move the same hand to the second position in the air). In some embodiments, the pinch input is performed by the user's first hand and the drag input is performed by the user's second hand (e.g., the user's second hand moves in the air from the first position to the second position while the user continues pinch input with the user's first hand). In some embodiments, input gestures as air gestures include inputs performed using both of the user's hands (e.g., pinch and / or tap input). For example, input gestures include two (e.g., more) pinch inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For instance, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch-and-drag input) is performed using the user's first hand, and a second pinch input is performed using the other hand (e.g., a second hand among the user's two hands). In some implementations, there is movement between the user's two hands (e.g., increasing and / or decreasing the distance or relative orientation between the user's two hands).
[0223] In some embodiments, a tap input performed as an air gesture (e.g., pointing at a user interface element) includes movement of a user's finger toward the user interface element, movement of the user's hand toward the user interface element (optionally, the user's finger extends toward the user interface element), downward movement of the user's finger (e.g., mimicking a mouse click or a tap on a touchscreen), or other predefined movements of the user's hand. In some embodiments, the tap input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tap gesture movement, which is the finger or hand moving away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on changes in the movement characteristics of the finger or hand performing the tap gesture (e.g., the end of movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of finger or hand movement, and / or a reversal of the acceleration direction of finger or hand movement).
[0224] In some implementations, the user's attention is determined to be directed to a portion of the 3D environment based on the detection of a gaze directed to that portion of the 3D environment (optionally, no other conditions are required). In some implementations, the user's attention is determined to be directed to that portion of the 3D environment based on the detection of a gaze directed to that portion of the 3D environment using one or more additional conditions, such as requiring the gaze to be directed to that portion of the 3D environment for at least a threshold duration (e.g., dwell time) and / or requiring the gaze to be directed to that portion of the 3D environment when the user's viewpoint is within a distance threshold from that portion of the 3D environment, so that the device determines that the user's attention is directed to that portion of the 3D environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the 3D environment to which the gaze is directed (e.g., until the one or more additional conditions are met).
[0225] In some implementations, the detection of the readiness configuration of a user or a portion of a user is performed by a computer system. The detection of the hand's readiness configuration is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein). For example, the readiness of the hand is determined based on whether it has a predetermined hand shape (e.g., a pre-pinch shape with the thumb and one or more fingers extended and spaced apart in preparation for a pinch or grasping gesture, or a pre-tap with one or more fingers extended and the back of the hand facing the user), whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or whether the hand has moved in a specific manner (e.g., moving towards an area in front of the user above the user's waist and below the user's head, or moving away from the user's body or legs). In some implementations, the readiness state is used to determine whether an interactive element of the user interface responds to attentional (e.g., gaze) input.
[0226] In scenarios where input is described by reference to air gestures, it should be understood that hardware input devices attached to or held by one or both of the user's hands can be used to detect such gestures. Optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers and / or one or more inertial measurement units can be used to track the spatial positioning of the hardware input device, and the positioning and / or movement of the hardware input device can be used in place of the positioning and / or movement of one or both hands in relation to the corresponding air gesture. In scenarios describing input using air gestures, it should be understood that similar gestures can be detected using hardware input devices attached to or held by one or both of the user's hands. User input can be detected using controls contained within the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger covers that can detect the positioning or changes in positioning of parts of the hand and / or fingers relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls. User input using controls contained within the hardware input device replaces hand and / or finger gestures such as air taps or air pinches in the corresponding air gesture. For example, a selection input described as being performed using an air tap or air pinch input can alternatively be detected using button presses, taps on touch-sensitive surfaces, presses on pressure-sensitive surfaces, or other hardware inputs. As another example, motion input described as being performed using air pinch and drag can be optionally detected based on interaction with hardware input controls (such as pressing and holding a button, touching a touch on a touch-sensitive surface, pressing a pressure-sensitive surface, or other hardware input following movement of a hardware input device (e.g., a hand associated with the hardware input device) through space). Similarly, two-handed input involving movement of hands relative to each other can be performed using an air gesture and a hardware input device not in the hand performing the air gesture, two hardware input devices held in different hands, or two air gestures performed by different hands using air gestures and / or inputs detected by one or more of the aforementioned hardware input devices.
[0227] In some embodiments, the software may be downloaded to controller 110 electronically, for example, via a network, or alternatively, may be provided on a tangible, non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, database 408 is also stored in memory associated with controller 110. Alternatively or additionally, some or all of the described functions of the computer may be implemented in dedicated hardware, such as custom or semi-custom integrated circuits or programmable digital signal processors (DSPs). Although in Figure 4The controller 110 is shown, but for example, as a separate unit from the image sensor 404, some or all of the controller's processing functions may be performed by a suitable microprocessor and software, or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand-tracking device), or by other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or with any other suitable computerized device (such as a game console or media player). The sensing function of the image sensor 404 may also be integrated into a computer or other computerized device controlled by the sensor output.
[0228] Figure 4 Also included is a schematic representation of a depth map 410 captured by image sensor 404 in some embodiments. As described above, the depth map comprises a matrix of pixels with corresponding depth values. Pixel 412 corresponding to hand 406 has been segmented from the background and wrist in this map. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from image sensor 404), where gray shadows become darker as depth increases. Controller 110 processes these depth values to identify and segment components of the image that exhibit human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and frame-to-frame motion from the depth map sequence.
[0229] Figure 4 The hand skeleton 414, which the controller 110 ultimately extracts from the depth map 410 of the hand 406, is also schematically illustrated in some embodiments. Figure 4 In this configuration, the hand skeleton 414 is superimposed on the hand background 416, which has already been segmented from the original depth map. In some embodiments, key feature points of the hand, and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, fingertips, the center of the palm, the end of the hand connecting to the wrist, etc.), are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points across multiple image frames to determine, in some embodiments, the gesture performed by the hand or the current state of the hand.
[0230] Figure 5 An eye-tracking device 130 is illustrated. Figure 1A Example implementation of ). In some implementations, the eye-tracking device 130 comprises an eye-tracking unit 243 ( Figure 2The eye-tracking device 130 controls the positioning and movement of a user's gaze relative to scene 105 or relative to XR content displayed via display generation component 120. In some embodiments, the eye-tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as a head-mounted device, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye-tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye-tracking device 130 is optionally a separate device from the handheld device or XR room. In some embodiments, the eye-tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye-tracking device 130 is optionally used in conjunction with a display generation component that is also head-mounted or not head-mounted. In some embodiments, the eye-tracking device 130 is not a head-mounted device and is optionally used in conjunction with head-mounted display generation components. In some embodiments, the eye-tracking device 130 is not a head-mounted device and is optionally part of non-head-mounted display generation components.
[0231] In some embodiments, the display generation unit 120 uses display mechanisms (e.g., a left near-eye display panel and a right near-eye display panel) to display frames including left and right images in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, the head-mounted display generation unit may include left and right optical lenses (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation unit may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation unit may have a transparent or semi-transparent display on which virtual objects are displayed, allowing the user to view the physical environment directly through the transparent or semi-transparent display. In some embodiments, the display generation unit projects virtual objects onto the physical environment. The virtual objects may be projected, for example, onto a physical surface or as holograms, allowing an individual to observe virtual objects superimposed on the physical environment using the system. In this case, separate display panels and image frames for the left and right eyes may not be necessary.
[0232] like Figure 5As shown, in some embodiments, eye-tracking device 130 (e.g., gaze tracking device) includes at least one eye-tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera) and an illumination source (e.g., an array or ring of IR or NIR light sources, such as LEDs) that emits light (e.g., IR or NIR light) toward the user's eye. The eye-tracking camera may be pointed at the user's eye to receive IR or NIR light reflected directly from the eye, or alternatively, it may be pointed at "hot" mirrors located between the user's eye and the display panel, which reflect the IR or NIR light from the eye back to the eye-tracking camera while allowing visible light to pass through. Eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps), analyzes these images to generate gaze tracking information, and transmits the gaze tracking information to controller 110. In some embodiments, the user's two eyes are tracked separately using corresponding eye-tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked using corresponding eye-tracking cameras and illumination sources.
[0233] In some implementations, a device-specific calibration procedure is used to calibrate the eye-tracking device 130 to determine parameters for the eye-tracking device in a specific operating environment 100, such as the 3D geometry and parameters of the LEDs, camera, thermal mirror (if present), eye lenses, and display. The device-specific calibration procedure can be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration procedure can be automated or manual. User-specific calibration procedures may include estimating eye parameters for a particular user, such as pupil position, foveal position, optical axis, visual axis, interocular distance, etc. In some implementations, once the device-specific and user-specific parameters for the eye-tracking device 130 are determined, a flash-assisted method can be used to process images captured by the eye-tracking camera to determine the user's current visual axis and gaze point relative to the display.
[0234] like Figure 5As shown, the eye-tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system. The gaze tracking system includes at least one eye-tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is performed, and an illumination source 530 (e.g., an array or ring of IR or NIR light sources, such as NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye-tracking camera 540 may be pointed toward a mirror 550 located between the user's eye 592 and a display 510 (e.g., the left or right display panel of a head-mounted display, or the display of a handheld device, projector, etc.). These mirrors reflect the IR or NIR light from the eye 592 while allowing visible light to pass through. Figure 5 (as shown in the top portion), or alternatively, it can be pointed towards the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion), Figure 5 (As shown in the bottom part).
[0235] In some implementations, controller 110 renders AR or VR frames 562 (e.g., left and right frames for the left and right display panels) and provides frames 562 to display 510. Controller 110 uses gaze tracking input 542 from eye-tracking camera 540 for various purposes, such as processing frame 562 for display. Controller 110 optionally estimates the user's gaze point on display 510 based on the gaze tracking input 542 obtained from eye-tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated based on gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0236] The following describes several possible use cases for the user's current gaze direction and is not intended to be limiting. As an example use case, controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, controller 110 may generate virtual content at a higher resolution in the concave region determined according to the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content in the view based at least partially on the user's current gaze direction. As yet another example, the controller may display specific virtual content in the view based at least partially on the user's current gaze direction. As another example use case in an AR application, controller 110 may guide an external camera used to capture the physical environment of an XR experience to focus in the determined direction. The external camera's autofocus mechanism may then focus on an object or surface in the environment that the user is currently looking at on display 510. As another example use case, eye lens 520 may be a focusable lens, and the controller uses gaze tracking information to adjust the focus of eye lens 520 so that the virtual object the user is currently looking at has appropriate convergence / divergence to match the convergence of the user's eyes 592. The controller 110 can use gaze tracking information to guide the eye lens 520 to adjust its focus so that the nearby object that the user is looking at appears at the correct distance.
[0237] In some embodiments, the eye-tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lens 520), an eye-tracking camera (e.g., eye-tracking camera 540), and a light source (e.g., illumination source 530 (e.g., IR or NIR LED)) mounted in a wearable housing. The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each lens in the head-mounted device, such as... Figure 5 As shown in the diagram. In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.
[0238] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. It should be noted that the position and angle of the eye-tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye-tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at different wavelengths (e.g., 940 nm) may be used on each side of the user's face.
[0239] like Figure 5 The illustrated gaze tracking system implementation can be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide users with computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experiences.
[0240] Figure 6 Examples of flash-assisted gaze tracking pipelines are provided. In some embodiments, the gaze tracking pipeline uses a flash-assisted gaze tracking system (e.g., such as...). Figure 1A and Figure 5 The illustrated eye-tracking device 130) is used to implement this. The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and flash in the current frame. When not in tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in tracking state.
[0241] like Figure 6 As shown, the gaze-tracking camera captures left and right images of the user's left and right eyes. The captured images are then fed into a gaze-tracking pipeline for processing to begin at 610. As indicated by the arrow returning to element 600, the gaze-tracking system can continue capturing images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be fed into the pipeline for processing. However, in some embodiments or under certain conditions, not all captured frames are processed by the pipeline.
[0242] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, the image is analyzed to detect the user's pupil and flash, as indicated at 620. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.
[0243] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flashes in part based on previous information from the previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupils and flashes detected in the current frame. The processing result at element 640 is checked to verify that the tracking or detection result is credible. For example, the result may be checked to determine whether a sufficient number of pupils and flashes used for gaze estimation were successfully tracked or detected in the current frame. At 650, if the result is not credible, the tracking state is set to no at element 660, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is credible, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.
[0244] Figure 6 This is intended as an example of an eye-tracking technology that can be used in a particular specific implementation. As those skilled in the art will recognize, in some implementations, other existing or future eye-tracking technologies may be used in computer system 101 in place of or in combination with the flash-assisted eye-tracking technology described herein to provide an XR experience to a user.
[0245] In some implementations, a portion of the captured real-world environment 602 is used to provide an XR experience to the user, such as a mixed reality environment in which one or more virtual objects are overlaid on a representation of the real-world environment 602.
[0246] Therefore, this description describes some embodiments of a three-dimensional environment (e.g., an XR environment) that includes representations of real-world objects and virtual objects. For example, the three-dimensional environment optionally includes a representation of a table existing in a physical environment, which is captured and displayed in the three-dimensional environment (e.g., actively displayed via a camera and display of a computer system or passively displayed via a transparent or semi-transparent display of a computer system). As previously described, the three-dimensional environment is optionally a mixed reality system, wherein the three-dimensional environment is based on a physical environment captured by one or more sensors of a computer system and displayed via a display generation component. As a mixed reality system, the computer system is optionally capable of selectively displaying portions and / or objects of the physical environment such that the corresponding portions and / or objects of the physical environment appear as if they exist in the three-dimensional environment displayed by the computer system. Similarly, the computer system is optionally capable of displaying virtual objects in the three-dimensional environment to appear as if the virtual objects exist in the real world (e.g., the physical environment) by placing virtual objects in the three-dimensional environment at corresponding locations in the real world that have corresponding positions in the three-dimensional environment. For example, the computer system optionally displays a vase such that the vase appears as if a real vase were placed on top of a table in the physical environment. In some implementations, a corresponding location in the three-dimensional environment has a corresponding location in the physical environment. Therefore, when a computer system is described as displaying a virtual object at a corresponding location relative to a physical object (e.g., such as at or near a user's hand or at or near a physical table), the computer system displays the virtual object at a specific location in the three-dimensional environment such that it appears as if the virtual object were at or near a physical object in the physical environment (e.g., the virtual object is displayed in the three-dimensional environment at a location in the physical environment that would be displayed if the virtual object were a real object at that specific location).
[0247] In some implementations, real-world objects that exist in a physical environment and are displayed in a 3D environment (e.g., and / or visible via display-generated components) can interact with virtual objects that exist only in the 3D environment. For example, the 3D environment may include a table and a vase placed on top of the table, where the table is a view (or representation) of a physical table in the physical environment, and the vase is a virtual object.
[0248] In a three-dimensional environment (e.g., a real environment, a virtual environment, or a hybrid environment including both real and virtual objects), an object is sometimes referred to as having depth or simulated depth, or as being visible, displayed, or placed at different depths. In this context, depth refers to a dimension other than height or width. In some embodiments, depth is defined relative to a fixed set of coordinates (e.g., where a room or object has a height, depth, and width defined relative to a fixed set of coordinates). In some embodiments, depth is defined relative to a user's position or viewpoint, in which case the depth dimension varies based on the user's position and / or the position and angle of the user's viewpoint. In some embodiments where depth is defined relative to the user's location relative to a surface of the environment (e.g., the surface of the environment's floor or ground), objects further away from the user along lines extending parallel to the surface are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from the user's position and parallel to the surface of the environment (e.g., depth is defined in a cylindrical or substantially cylindrical coordinate system, where the user's position is at the center of a cylinder extending from the user's head toward the user's feet). In some embodiments where depth is defined relative to the user's viewpoint (e.g., a direction relative to a point in space that determines which part of the environment is visible via a head-mounted device or other display), objects further away from the user's viewpoint along a line extending parallel to the user's viewpoint are considered to have greater depth in the environment, and / or the depth of an object is measured along an axis extending outward from a line extending from and parallel to the user's viewpoint (e.g., defining depth in a spherical or substantially spherical coordinate system, where the origin of the viewpoint is at the center of a sphere extending outward from the user's head). In some embodiments, depth is defined relative to a user interface container (e.g., a window or application displaying application and / or system content), where the user interface container has a height and / or width, and depth is a dimension orthogonal to the height and / or width of the user interface container. In some implementations, when a depth is defined relative to a user interface container, when the container is placed in a three-dimensional environment or initially displayed (e.g., such that the container's depth dimension extends outward away from the user or the user's viewpoint), the container's height and / or width are typically orthogonal or substantially orthogonal to a straight line extending from the user's location (e.g., the user's viewpoint or the user's position) to the user interface container (e.g., the center of the user interface container or another feature point of the user interface container). In some implementations, when a depth is defined relative to a user interface container, the object's depth relative to the user interface container refers to the object's positioning along the depth dimension of the user interface container. In some implementations, multiple different containers may have different depth dimensions (e.g., different depth dimensions extending away from the user or the user's viewpoint in different directions and / or from different starting points).In some implementations, when depth is defined relative to a user interface container, the orientation of the depth dimension remains constant relative to the user interface container as the position of the user interface container changes, as the user and / or the user's viewpoint changes (e.g., when multiple different viewers are viewing the same container in a 3D environment, such as during a collaborative session and / or when multiple participants are in a real-time communication session with shared virtual content including the container). In some implementations, for curved containers (e.g., containers including those with curved surfaces or curved content regions), the depth dimension optionally extends into the surface of the curved container. In some cases, z-interval (e.g., the distance between two objects in the depth dimension), z-height (e.g., the distance of one object from another in the depth dimension), z-position (e.g., the position of an object in the depth dimension), z-depth (e.g., the position of an object in the depth dimension), or simulated z-dimensionality (e.g., depth used as a dimension of an object, a dimension of the environment, an orientation in space, and / or an orientation in simulated space) are used to refer to the concept of depth as described above.
[0249] In some implementations, a user may optionally be able to interact with virtual objects in a three-dimensional environment using one or both hands, as if the virtual objects were real objects in the physical environment. For example, as described above, one or more sensors of the computer system may optionally capture one or both of the user's hands and display a representation of the user's hands in the three-dimensional environment (e.g., in a manner similar to displaying real-world objects in the three-dimensional environment described above). Alternatively, in some implementations, the user's hands may be seen via the display generating component, through the ability to see the physical environment through the user interface, due to the transparency / semi-transparency of a portion of the user interface being displayed by the display generating component, or due to the projection of the user interface onto a transparent / semi-transparent surface or onto the user's eyes or into the user's field of view. Thus, in some implementations, the user's hands are displayed at corresponding locations in the three-dimensional environment and are treated as if they were objects in the three-dimensional environment that can interact with virtual objects in the three-dimensional environment, as if these virtual objects were physical objects in the physical environment. In some implementations, the computer system may update the display of the user's hand representation in the three-dimensional environment in conjunction with the movement of the user's hands in the physical environment.
[0250] In some embodiments described below, the computer system optionally determines the “effective” distance between a physical object in the physical world and a virtual object in a three-dimensional environment, for example, to determine whether a physical object is directly interacting with a virtual object (e.g., whether a hand is touching, grasping, holding, or within a threshold distance of a virtual object). For example, a hand directly interacting with a virtual object optionally includes one or more of the following: a finger pressing a virtual button, a user’s hand grasping a virtual vase, a user’s hand clasped together to pinch / hold the application’s user interface, and two fingers performing any other type of interaction described herein. For example, the computer system optionally determines the distance between a user’s hand and a virtual object when determining whether and / or how a user is interacting with a virtual object. In some embodiments, the computer system determines the distance between a user’s hand and a virtual object by determining the distance between the position of a hand in the three-dimensional environment and the position of the virtual object of interest in the three-dimensional environment. For example, a user's one or both hands are located at a specific location in the physical world. The computer system optionally captures the one or both hands and displays them at a specific corresponding location in a three-dimensional environment (e.g., the location where the hand would be displayed in the three-dimensional environment if it were a virtual hand rather than a physical hand). Optionally, the location of the hand in the three-dimensional environment is compared with the location of a virtual object of interest in the three-dimensional environment to determine the distance between the user's one or both hands and the virtual object. In some embodiments, the computer system optionally determines the distance between a physical object and a virtual object by comparing locations in the physical world (e.g., rather than comparing locations in the three-dimensional environment). For example, when determining the distance between a user's one or both hands and a virtual object, the computer system optionally determines the corresponding location of the virtual object in the physical world (e.g., the location where the virtual object would be located in the physical world if it were a physical object rather than a virtual object), and then determines the distance between the corresponding physical location and the user's one or both hands. In some embodiments, the same technique is optionally used to determine the distance between any physical object and any virtual object. Therefore, as described herein, when determining whether a physical object is in contact with a virtual object or whether a physical object is within a threshold distance of a virtual object, the computer system may optionally perform any of the techniques described above to map the position of the physical object to the three-dimensional environment and / or map the position of the virtual object to the physical environment.
[0251] In some implementations, the same or similar techniques are used to determine where and what the user's gaze is directed at, and / or where and what the physical stylus held by the user is pointing at. For example, if the user's gaze is directed at a specific location in the physical environment, the computer system optionally determines a corresponding location in the three-dimensional environment (e.g., a virtual location of the gaze), and if a virtual object is located at that corresponding virtual location, the computer system optionally determines that the user's gaze is directed at that virtual object. Similarly, the computer system may optionally be able to determine the direction in which the stylus is pointing in the physical environment based on the orientation of the physical stylus. In some implementations, based on this determination, the computer system determines a corresponding virtual location in the three-dimensional environment corresponding to the location pointed at by the stylus in the physical environment, and optionally determines that the stylus is pointing at the corresponding virtual location in the three-dimensional environment.
[0252] Similarly, the embodiments described herein may refer to the location of a user (e.g., a user of a computer system) in a three-dimensional environment and / or the location of the computer system in a three-dimensional environment. In some embodiments, the user of the computer system is holding, wearing, or otherwise located at or near the computer system. Thus, in some embodiments, the location of the computer system serves as a proxy for the location of the user. In some embodiments, the location of the computer system and / or the user in the physical environment corresponds to a corresponding location in the three-dimensional environment. For example, the location of the computer system would be a location in the physical environment (and its corresponding location in the three-dimensional environment) that, if the user stands at, facing the corresponding portion of the physical environment visible via the display generation component, the user will see from that location objects in the physical environment that are positioned, oriented, and / or sized (e.g., in an absolute sense and / or relative to each other) in the same way as objects displayed or visible in the three-dimensional environment by or via the display generation component of the computer system. Similarly, if the virtual objects displayed in a 3D environment are physical objects in the physical environment (e.g., physical objects placed in the physical environment at the same location as these virtual objects in the 3D environment, and physical objects in the physical environment having the same size and orientation as in the 3D environment), then the position of the computer system and / or the user is the position from which the user will see these virtual objects in the physical environment at the same location, orientation, and / or size (e.g., in an absolute sense and / or relative to each other and real-world objects) as the virtual objects displayed in the 3D environment by the display generation components of the computer system.
[0253] In this disclosure, various input methods are described in relation to interaction with a computer system. When an example is provided using one input device or method, and another example is provided using another input device or method, it should be understood that each example is compatible with and optionally utilizes the input device or method described with respect to the other example. Similarly, various output methods are described in relation to interaction with a computer system. When an example is provided using one output device or method, and another example is provided using another output device or method, it should be understood that each example is compatible with and optionally utilizes the output device or method described with respect to the other example. Similarly, various methods are described in relation to interaction with a virtual or mixed reality environment via a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with and optionally utilizes the methods described with respect to the other example. Therefore, this disclosure discloses embodiments that are combinations of features of a plurality of examples without exhaustively listing all features of the embodiments in the description of each example embodiment.
[0254] User interface and related processes
[0255] Now turn attention to implementations of user interfaces (“UIs”) and associated processes that can be implemented on computer systems (such as portable multifunction devices or head-mounted devices) that communicate with display generation components and at least one camera.
[0256] Figures 7A to 7AB Exemplary methods for capturing gaze-related media according to some implementation schemes are illustrated. Figure 8 This is a flowchart of an exemplary method 800 for capturing gaze-related media according to some implementation schemes. Figures 7A to 7C The user interface in the document is used to illustrate the process described below, including Figure 8 The process in.
[0257] Figures 7A to 7B Examples are shown from the back (e.g., Figure 7A ) and from the front (e.g., Figure 7B View computer system 700, which includes device 702 (e.g., tablet computer). Device 702 includes, for example, Figure 7B The display 708, visible on its front side, and as shown Figure 7A The first camera 704A and the second camera 704B, visible from the rear side of the device, are shown. Figure 7AIn this configuration, the first camera 704A and the second camera 704B are spaced apart from each other but face substantially the same direction (e.g., away from the user when the user is viewing the display 708). The computer system 700 includes multiple input devices, including hardware buttons 706 and a touch-sensitive surface of the display 708 of the device 702. In some embodiments, the computer system 700 includes one or more user devices, such as mobile phones, tablets, and / or laptops. In some embodiments, the computer system 700 includes one or more features of the computer system 101 as described above, such as a controller 110 and a display generation component 120.
[0258] exist Figure 7B At this location, computer system 700 displays media capture user interface 710 (e.g., a user interface for a camera application) via display 708. In some embodiments, display 708 is a transparent or semi-transparent display, such that the user-perceived media capture user interface 710 of computer system 700 is overlaid on a transparent view of the physical environment. In some embodiments, display 708 is an opaque display, and media capture user interface 710 is an overlaid transparent video (e.g., image data captured using at least one of multiple cameras (e.g., a first camera 704A and / or a second camera 704B)). Although Figures 7A to 7AB The techniques used in device 702 (tablet computer) are illustrated, but these techniques are also applicable to head-mounted devices, such as those described with respect to computer system 101. In some embodiments where computer system 700 includes a head-mounted device, computer system 700 optionally includes two displays (e.g., one display for each eye of the user of computer system 700), for example, each display showing the field of view of a respective camera among multiple cameras (e.g., first camera 704A and / or second camera 704B) and / or different portions or versions of the media capture user interface 710 to create a depth (e.g., three-dimensional) effect. In some embodiments where computer system 700 is a head-mounted device, first camera 704A and second camera 704B are oriented in substantially the same direction as the user's eyes when looking straight ahead. Additionally, first camera 704A and second camera 704B have different but overlapping fields of view, thereby allowing the capture of spatial media (e.g., virtual reality photographs and / or videos that can present the appearance / illusion of depth).
[0259] like Figure 7B As illustrated, the media capture user interface 710 includes a camera viewfinder 712 that transparently covers a first portion of a representation of the field of view of the first camera 704A. In some embodiments, the first portion of the field of view of the first camera 704A framed by the camera viewfinder 712 represents or approximates the current portion of the field of view of the first camera 704A that will be included (e.g., captured) in the media capture.
[0260] The media capture user interface 710 also includes a boundary 714 and a darkened area 716, the boundary visually indicating the edge of the camera viewfinder 712, and the darkened area covering a second portion of the representation of the field of view of the first camera 704A that falls outside the camera viewfinder 712. In some embodiments, the boundary 714 visually indicates the edge of the camera viewfinder 712 by modifying (e.g., blurring, darkening, occluding, or otherwise visually indicating the edge of the camera viewfinder) the appearance of a third portion of the representation of the field of view of the first camera 704A. For example, the appearance of the third portion of the representation of the field of view of the first camera 704A may be modified to create a soft gradient between the camera viewfinder 712 and the darkened area 716, thereby causing a vignetting effect on the first portion of the representation of the field of view of the first camera 704A.
[0261] The media capture user interface 710 also includes a shutter power indicator 718 displayed in the central area of the camera viewfinder 712. Two concentric rings included in the shutter power indicator 718 are at least partially transparent or translucent, such that the portion of the field of view of the first camera 704A represented below the concentric rings remains at least partially visible to the user. The media capture user interface 710 also includes an option power indicator 720, a first status indicator 722A, and a second status indicator 722B (e.g., regarding...). Figures 9A to 9I (Detailed description). In some embodiments, the computer system 700 alters the appearance of one or more elements of the media capture user interface 710 based on the portion of the field of view of the first camera 704A covered by one or more elements (e.g., shutter power indicator 718, option power indicator 720, first status indicator 722A, second status indicator 722B). For example, the computer system 700 may detect or sample the color of a portion of the field of view of the first camera 704A covered by the option power indicator 720, and thus change the color of the option power indicator 720 to match or contrast with the sampled color.
[0262] like Figure 7C As illustrated, when the media capture user interface 710 is displayed, the computer system 700 detects initial input (e.g., input 728A and / or input 728B) by detecting the position of the user's hand (e.g., finger 724 and / or hand 726). For example, input 728A is detected when finger 724 is detected in a position near or above (but not pressed) hardware button 706 (e.g., using a capacitive sensor in hardware button 706 or another sensor of computer system 700), and / or input 728B is detected when hand 726 is detected in an elevated position (e.g., using multiple cameras and / or one or more sensors of computer system 700).
[0263] like Figure 7DAs illustrated, in response to the detection of initial input (e.g., input 728A and / or input 728B), computer system 700 alters the appearance of one or more elements of media capture user interface 710 to indicate that computer system 700 has entered a media capture readiness state. Changing the appearance of media capture user interface 710 includes further darkening the darkened area 716 (e.g., with...). Figure 7C Compared to the amount of darkening in the darkened area 716, this increases the contrast between the darkened area 716 and the camera viewfinder 712. Changing the appearance of the media capture user interface 710 also includes reducing the semi-transparency of the concentric rings of the shutter power indicator 718 (e.g., increasing the opacity, for example, compared to the amount of darkening in the darkened area 716 in the viewfinder 712), thereby increasing the contrast between the darkened area 716 and the camera viewfinder 712. Figure 7C Compared to the semi-transparency of concentric rings in the image, this increases the visual salience of the shutter power indicator 718. In some embodiments, the appearance of other elements of the media capture user interface 710 (e.g., option power indicator 720, first status indicator 722A, or second status indicator 722B) also changes in response to the detection of initial input, for example, by varying the semi-transparency of the concentric rings in the image. Figure 7C The way elements are displayed in the middle can be used to increase or decrease their salience.
[0264] Figure 7E and Figure 7F This illustrates how the computer system 700 changes various elements of the media capture user interface 710 in response to the movement of the device 702. For example... Figure 7E As illustrated, when the media capture user interface 710 is displayed, the computer system 700 detects a rightward movement of the device 702. In response to detecting the movement of the device 702, the computer system 700 shifts the camera viewfinder 712 to the left of the display 708, thereby creating an inertial appearance for the camera viewfinder 712 (e.g., displaying the camera viewfinder 712 as if it were "connected" to the display 708 by a spring rather than by a fixed component). The computer system 700 shifts the boundary 714 and shutter indication 718 to the left along with the camera viewfinder 712, such that the boundary 714 and shutter indication 718 remain fixed within the reference frame of the camera viewfinder 712. The computer system 700 does not shift the option indication 720, the first status indicator 722A, and the second status indicator 722B on the display 708, such that the option indication 720, the first status indicator 722A, and the second status indicator 722B remain fixed within the reference frame of the display 708. Figure 7FAs illustrated, after the movement of device 702 stops, computer system 700 recenters camera viewfinder 712, boundary 714, and shutter power indicator 718 within the reference frame of display 708. Computer system 700 continues to display the darkened area 716 of the representation covering the field of view of first camera 704A that falls outside camera viewfinder 712, but the shape of darkened area 716 changes accordingly as camera viewfinder 712 shifts and is recentered.
[0265] like Figure 7G As illustrated, when the media capture user interface 710 is displayed, the computer system 700 detects potential media capture inputs, such as button press input 730A, air gesture input 730B (e.g., air gestures performed with the thumb and forefinger of hand 726, such as air pinch or air tap gestures), and / or tap input 730C (e.g., tap inputs on the touch-sensitive surface of display 708). Upon detecting a potential media capture input, the computer system 700 further detects the user's gaze 732 at a location on the right side of the media capture user interface that is not at or near the shutter power indicator 718. Since the gaze 732 is not pointed at or near the shutter power indicator 718 when the potential media capture input is detected (e.g., because the gaze 732 is not in the central area of the camera viewfinder 712), the computer system 700 does not initiate media capture in response to the potential media capture input. In some embodiments, in response to the potential media capture input and based on the location of the user's gaze pointing to the gaze 732, the computer system 700 performs non-media capture operations (e.g., adjusting white balance or focus).
[0266] like Figure 7H As illustrated, computer system 700 detects gaze 732 pointing towards shutter power indicator 718 (e.g., gaze 732 is now located in the central area of camera viewfinder 712). In response to detecting gaze 732 pointing towards shutter power indicator 718, computer system 700 changes the appearance of shutter power indicator 718, thereby increasing the visual salience of shutter power indicator 718 (e.g., as shown by...). Figure 7H The image depicts a shutter speed indicator displaying 718. Figure 7G The difference in representation (as shown in the comparison) indicates that the user can initiate media capture (e.g., by performing a potential media capture input). Changing the appearance of the shutter power indicator 718 includes further reducing the semi-transparency of the concentric rings of the shutter power indicator 718 (e.g., increasing the opacity) (e.g., compared to...). Figure 7G The semi-transparency of the concentric rings is compared to that of the background, and the size of the concentric rings is slightly reduced (e.g., the rings are squeezed together). In some embodiments, changing the appearance of the shutter energy representation 718 includes increasing the brightness and / or contrast of the concentric rings relative to the background.
[0267] like Figure 7I1 As illustrated, when the shutter indication 718 with increased visual salience is displayed, the computer system 700 detects a second potential media capture input, such as a button press input 736A of hardware button 706, an air gesture input 736B (e.g., an air gesture performed with the thumb and forefinger of hand 726, such as an air pinch gesture), and / or a tap input 736C (e.g., a tap input on the touch-sensitive surface of display 708). Upon detection of the second potential media capture input, gaze 732 is directed towards the shutter indication 718 (e.g., gaze 732 is in the central area of the camera viewfinder 712). Therefore, in response to the second potential media capture input, the computer system 700 initiates media capture. In some embodiments, the computer system 700 initiates capture of both photographic and video media, and later determines (e.g., as described below) which capture mode the user requested. In some implementations, media capture is spatial media capture, performed using both a first camera 704A and a second camera 704B to capture virtual reality media (e.g., photographs and / or videos) that can present an appearance / illusion of depth. The computer system 700 further reduces the translucency and size of the concentric rings of the shutter power representation 718 in response to a second potential media capture input, thereby reflecting the state of the detected second potential media capture input (e.g., further increasing the salience of the concentric rings and squeezing them closer together as the hardware button 706 is further pressed and / or the air gesture input 736B is further pinched together).
[0268] In some implementation schemes, Figure 7I1 The technologies and user interfaces described in the text are by Figures 1A to 1P One or more of the devices described herein shall be used to provide this. Figure 7I2 Examples are given (for example, such as...) Figures 7A to 7I1 The media capture user interface X710 described herein is displayed on the display module X702 of a head-mounted device (HMD) X700. In some embodiments, the device X700 includes a pair of display modules that provide stereoscopic content to the different eyes of the same user. For example, the HMD X700 includes display module X702 (which provides content to the user's left eye) and a second display module (which provides content to the user's right eye). In some embodiments, the second display module displays an image that is slightly different from that of display module X702 to generate the illusion of stereoscopic depth.
[0269] like Figure 7I2As illustrated, when the shutter power indicator X718 with increased visual salience is displayed, the HMD X700 detects a second potential media capture input, such as a button press input X736A of the hardware button X706, an air gesture input X736B (e.g., an air gesture performed with the thumb and forefinger of X750A, such as an air pinch gesture, as described in further detail below), and / or a tap input X736C (e.g., a tap input on the touch-sensitive surface of the display 708). Upon detection of the second potential media capture input, gaze X732 is directed towards the shutter power indicator X718 (e.g., gaze X732 is in the central area of the camera viewfinder X712). Therefore, in response to the second potential media capture input, the HMD X700 initiates media capture. In some embodiments, the HMD X700 initiates capture of both photographic and video media, and later determines (e.g., as described below) which capture mode the user requested. In some implementations, media capture is spatial media capture (e.g., performed using cameras such as a first camera 704A and a second camera 704B) to capture virtual reality media (e.g., photos and / or videos) that can present the appearance / illusion of depth. The HMD X700 further reduces the translucency and size of the concentric rings of the shutter energy representation X718 in response to a second potential media capture input, thereby reflecting the state of the detected second potential media capture input (e.g., further increasing the salience of the concentric rings and squeezing them closer together as the hardware button X706 is further pressed and / or the air gesture input X736B is further pinched together).
[0270] In some embodiments, the HMD X700 detects a second potential media capture input based on an air gesture input X736B performed by a user of the HMD X700. In some embodiments, the HMD X700 detects the user's hand X750A (e.g., the user's left or right hand or both hands) and determines whether the movement of the hand X750A performs a predetermined air gesture corresponding to the second potential media capture input. In some embodiments, the predetermined air gesture includes a pinch gesture. In some embodiments, a pinch gesture includes detecting the movement of fingers X750C and thumb X750D toward each other.
[0271] Figures 1B to 1PAny of the features, components, and / or parts shown (including their arrangement and configuration) may be included in the HMD X700 individually or in any combination. For example, in some embodiments, the HMD X700 includes any of the features, components, and / or parts of HMD 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100 individually or in any combination. In some embodiments, the display module X702 includes, individually or in any combination, display units 1-102, 1-202, 1-306, 1-406, display generating component 120, display screens 1-122a-b, first rear display screen 1-322a and second rear display screen 1-322b, display 11.3.2-104, first display component 1-120a and second display component 1-120b, display component 1-320, and display component 1-421. The first display sub-assembly 1-420a and the second display sub-assembly 1-420b, display assembly 3-108, display assembly 11.3.2-204, first optical module 11.1.1-104a and the second optical module 11.1.1-104b, optical module 11.3.2-100, optical module 11.3.2-200, biconvex lens array 3-110, display area or display zone 6-232 and / or display / display area 6-334 features, components and / or parts. In some embodiments, the HMD X700 includes a sensor X704, which, individually or in any combination, includes any feature, component, and / or part of any of the following: sensor 190, sensor 306, image sensor 314, image sensor 404, sensor assemblies 1-356, sensor assemblies 1-456, sensor systems 6-102, sensor systems 6-202, sensor 6-203, sensor systems 6-302, sensor 6-303, sensor systems 6-402, and / or sensors 11.1.2-110a-f. In some embodiments, the input device X703 and / or the hardware button X706, individually or in any combination, includes any feature, component, and / or part of any of the following: first buttons 1-128, buttons 11.1.1-114, second buttons 1-132, and / or a dial or button 1-328. In some implementations, the HMD X700 includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback (e.g., audio output), which is optionally generated based on detected events and / or user input detected by the HMD X700.
[0272] like Figure 7JAs illustrated, a second potential media capture input is released before (e.g., immediately) a first threshold time period has elapsed (e.g., immediately). This could be a finger 724 lifting off a hardware button 706, a hand 726 opening an air pinch gesture, or a tap input 736C lifting off a touch-sensitive surface of the display 708. Therefore, the computer system 700 captures photographic media (e.g., still photographs and / or media captures of a limited duration from content before and / or after the capture input is detected (e.g., before and / or after the air pinch gesture is released, the air tap gesture is detected, or a button press is detected or released), such as capturing a short animated photograph of several frames while taking a photograph, thus creating a "live" effect. When capturing photographic media with a "live" effect, it optionally includes stereoscopic depth content such that when the photographic media with a "live" effect is displayed or played back, the photographic media with a "live" effect plays a sequence of frames (e.g., one or more frames before the frame captured when the capture input is detected and / or one or more frames after the frame captured when the capture input is detected). In some implementations, playing photo media with a "live" effect includes playing multiple frames for a first eye and multiple frames for a second eye, wherein the frames for the first eye and the frames for the second eye are time-synchronized, such that the "live" effect also includes stereoscopic depth information (e.g., minute differences in the frames for the first eye displayed simultaneously with the frames for the second eye to create the illusion of stereoscopic depth). In some implementations, the user has the option to enable or disable frame capture before and / or after the capture input to generate photo media with a "live" effect. In some implementations, the user has the option to suppress the display of the "live" effect and / or discard additional frames used to generate the "live" effect. In some implementations, additional frames between captured frames are interpolated to give the "live" effect a smoother appearance, to give an appearance already captured at a higher frame rate, and / or to generate a slow-motion effect during movement in the "live" effect. In some implementations, the frame sequence before or after the frames captured when the capture input is detected has a lower resolution and / or smaller size than the frames captured when the capture input is detected. In an implementation where computer system 700 initiates capture of both still media and video media, computer system 700 discards any captured video media, thereby determining the requested photographic media based on a rapid / immediate release of a second potential media capture input. Computer system 700 further reduces the semi-transparency of the concentric rings of shutter energy indication 718 (e.g., the concentric rings may temporarily become opaque at this stage), thereby indicating that photographic media capture has been initiated, and increases the size of the concentric rings (e.g., returning to...). Figure 7F The dimensions of the concentric rings (indicated by the size of the rings) thus indicate that the second potential media capture input has been released. Figure 7K As illustrated, when temporarily reducing the semi-transparency (e.g., as shown in the example), Figure 7J After (as illustrated), the computer system 700 increases the semi-transparency of the concentric rings of the shutter power indicator 718, for example, before detecting media capture input and / or initiating capture (e.g., in...). Figure 7G The translucency level of the concentric rings at (location).
[0273] Computer system 700 also initiates photo media capture animation, such as Figures 7J to 7N As illustrated, the darkened area 716 expands to close around the shutter power indicator 718 and then retracts from the shutter power indicator 718 (e.g., simulating the closing and opening of a physical camera shutter). In some embodiments, as the darkened area 716 completely closes over the shutter power indicator 718 (e.g., as...), Figure 7L As illustrated, the camera viewfinder 712 and border 714 may be invisible during certain portions of the media capture animation. The photo media capture animation includes darkening of the darkened area 716 during the photo media capture animation (e.g., with...). Figures 7I1 to 7I2 The darkened area 716 in the middle is darkened compared to the darkened area in the middle.
[0274] like Figures 70 to 7Q As illustrated, after completing photo media capture (e.g., taking a photo), computer system 700 displays the captured media icon 738, which is a thumbnail of the captured photo media. Figure 7O As illustrated, the captured media icon 738 is initially displayed at a first size and a first analog exposure level (e.g., at a first brightness), such that the captured media icon 738 initially appears as an opaque white icon. Figure 7P As illustrated, after a short period of time, the captured media icon 738 is displayed at a second smaller size and a second lower simulated exposure level, making the thumbnail of the photographic media capture visible to the user (e.g., displayed at a brightness closer to the actual brightness of the photographic media capture). Figure 7Q As illustrated, if the computer system 700 does not detect a user gazing at or otherwise interacting with the captured media icon 738 within a threshold time period (e.g., a few seconds), the computer system 700 increases the semi-transparency of the captured media icon 738 (e.g., decreases the opacity), thereby reducing the visual salience of the captured media icon 738.
[0275] Figures 7R to 7U An example is shown where computer system 700 initiates media capture in a timer delay mode (e.g., using a self-timer). Figure 7RWhen media capture is initiated in timer delay mode, in response to the detection of a second potential media capture input (e.g., a button press input 736A of hardware button 706, an air gesture input 736B, and / or a tap input 736C), computer system 700 initiates a three-second media capture timer delay (e.g., instead of as...). Figures 7I1 to 7J The example illustrates initiating media capture immediately upon detecting a potential media capture input. (For instance...) Figures 7S to 7U As illustrated, during the three-second duration of the media capture timer delay, computer system 700 animates shutter power display 718 to indicate the past time of the media capture timer delay: Figure 7S At this point, computer system 700 causes the space between the concentric rings of shutter power indicator 718 to appear as a complete filled ring, instead of transparently covering the background. Then, as... Figure 7T As shown, during the first second of the three-second media capture, the computer system reduces the fill portion of the space between the concentric rings of the shutter power representation 718 by one-third (e.g., displaying two-thirds of the complete ring after the first second has elapsed). Figure 7U As shown, during the second second of the three-second media capture timer delay, computer system 700 reduces the fill portion of the space between the concentric rings of shutter power representation 718 by another third (e.g., displaying one-third of the complete ring after the second second has elapsed). After the entire three-second media capture timer delay has elapsed, computer system 700 initiates media capture (e.g., as described above regarding...). Figures 7J to 7Q And / or the following about Figures 7W to 7AA The above).
[0276] like Figure 7V As illustrated, when the media capture user interface 710 is displayed, the computer system 700 detects a third potential media capture input, such as a button press input 740A of the hardware button 706 when gaze 732 is directed at the shutter capability indicator 718, an air gesture input 740B, and / or a tap input 740C (e.g., as per the context of...). Figures 7I1 to 7I2 However, the media capture input is held for a period of time instead of being released (e.g., immediately released, as mentioned above). Figure 7J When the third potential media capture input is held for more than a first threshold time period after initial detection (e.g., when the third potential media capture input is not immediately released), the circular area of the shutter power indicator 718 gradually changes to a different color (e.g., red or another color), thereby indicating that the media capture mode is changing from photo media capture mode to video media capture mode. In some embodiments, if the computer system 700 detects that the shutter power indicator 718 is still red (e.g., when...), the circular area of the shutter power indicator 718 gradually changes to a different color (e.g., red or another color), indicating that the media capture mode is changing from photo media capture mode to video media capture mode. Figure 7VAt the illustrated point, before the second threshold time period has elapsed, the third potential media capture input is released, and the computer system 700 captures the photographic media as described above.
[0277] like Figure 7W As illustrated, a third potential media capture input (e.g., a button press input 740A, an air gesture input 740B, and / or a tap input 740C of hardware button 706) is held until after a second threshold time period, at which point the shutter indicator 718 has turned completely red and begins to contract and move upward toward the corner of the camera viewfinder 712 (e.g., indicating to the user that the media capture mode has now switched from photo media capture to video media capture). Therefore, the computer system 700 captures video media. In an embodiment where the computer system 700 initiates capture of both still media and video media, the computer system 700 discards any captured photo media, thereby determining the video media requested by the user based on the extended duration of the third potential media capture input. Upon initiating video media capture, the computer system 700 displays a video capture timer 742 at the corner of the camera viewfinder 712, indicating the current past time of the video media capture.
[0278] like Figure 7X As illustrated, once the computer system 700 has initiated video media capture (e.g., by instructing the user through the color and / or movement of the shutter indication 718 and / or the appearance of the video capture timer 742), the user releases the third potential media capture input, and video media capture continues. The shutter indication 718 continues to retract and move upward toward the corner of the camera viewfinder 712, where the video capture timer 742 displays the current past time. Additionally, the computer system 700 further increases the semi-transparency of the captured media icon 738 (e.g., with...). Figure 7P Compared to the semi-transparency of the captured media icon 738, this further reduces the visual salience of the captured media icon 738 during video media capture.
[0279] like Figure 7Y As illustrated, while video media capture continues, computer system 700 displays shutter power indicator 718 as a small red dot in the corner of camera viewfinder 712, located next to video capture timer 742. In some embodiments, shutter power indicator flashes or pulses next to video capture timer 742, indicating that video media capture is currently being recorded. Although computer system 700 no longer detects the position of the user's hand (e.g., finger 724 is no longer near hardware button 706, and / or hand 726 is no longer in the field of view of multiple cameras), the computer system does not change the appearance of darkened area 716 and / or does not terminate video media capture.
[0280] like Figure 7Z As illustrated, while video media capture continues, computer system 700 detects a fourth potential media capture input, such as a button press input 744A, an air gesture input 744B, and / or a tap input 744C. In response to detecting the fourth potential media capture input, computer system 700 terminates video media capture. Figure 7AA As illustrated, upon completion of video media capture, computer system 700 again displays shutter power indicator 718 at the center of camera viewfinder 712 and stops displaying video capture timer 742 (e.g., computer system 700 changes the appearance of one or more elements of media capture user interface 710 to restore them to their original state). Figure 7D (Appearance in the image). The computer system 700 updates the captured media icon 738 to a thumbnail of the completed video media capture. In some implementations, the captured media icon 738 appears as follows: Figures 70 to 7Q The changes were made as described (e.g., initially appearing large and overexposed before being resolved to a smaller size and lower exposure, and increasing the semi-transparency of the captured media icon 738 after a threshold amount of time without user interaction).
[0281] Figures 7AA1 to 7AA9 An example is shown of a version of a media capture user interface 710, including a mode control display 746 and a video status display 750, according to some implementation schemes. For example... Figure 7AA1 As illustrated, the media capture user interface 710 includes a mode control indication 746 displayed in the lower area of the camera viewfinder 712. The mode control indication 746 includes a photo mode indication 746A and a video mode indication 746B. The computer system 700 initially displays the mode control indication 746 opaquely. The computer system 700 highlights the photo mode indication 746A with a backplate to indicate that a photo media capture mode is currently selected (e.g., the computer system 700 will now initiate photo media capture in response to appropriate media capture input). In some embodiments, the computer system 700 displays the text of the photo mode indication 746A in a different color than the text of the video mode indication 746B to indicate the currently selected media capture mode.
[0282] like Figure 7AA2 As illustrated, after a threshold amount of time (e.g., a few seconds) during which no gaze 732 pointing to the modal control display 746 is detected, the computer system 700 increases the semi-transparency of the modal control display 746 (e.g., decreases its opacity) (e.g., as described above regarding increasing the visual salience of the captured media icon 738). Figure 7AA3As illustrated, in response to detecting a gaze 732 pointing to the mode control display 746, the computer system 700 reduces the semi-transparency of the mode control display 746 (e.g., increases its opacity).
[0283] like Figure 7AA3 As illustrated, when the mode control indication 746 with a highlighted photo mode indication 746A is displayed, the computer system 700 detects a fifth potential media capture input, such as a button press input 747A of the hardware button 706, an air gesture input 747B (e.g., an air gesture performed by the thumb and forefinger of the hand 726, such as an air pinch gesture), and / or a tap input 747C (e.g., a tap input on the touch-sensitive surface of the display 708). Upon detection of the fifth potential media capture input, gaze 732 is directed towards the mode control indication 746. Therefore, in response to the fifth potential media capture input, the computer system 700 switches the currently selected media capture mode from photo media capture mode to video media capture mode.
[0284] like Figure 7AA4 As illustrated, computer system 700 uses a back panel to highlight video mode capability indicator 746B to indicate that a video media capture mode is currently selected. In some embodiments, computer system 700 changes the color of the text on photo mode capability indicator 746A and the text on video mode capability indicator 746B to indicate the currently selected media capture mode. Additionally, the computer system changes the appearance of shutter capability indicator 718 to indicate that a video media capture mode is currently selected, for example, by displaying shutter capability indicator 718 filled with a semi-transparent or opaque color (e.g., red or another color).
[0285] like Figure 7AA5 As illustrated, when the mode control indicator 746, which displays a highlighted video mode indicator 746B and a color-filled shutter mode indicator 718, is displayed, the computer system 700 detects a sixth potential media capture input, such as a button press input 748A of the hardware button 706 when gaze 732 is directed at the shutter mode indicator 718, an air gesture input 748B, and / or a tap input 748C, and initiates media capture accordingly (e.g., as per [reference to...]). Figures 7I1 to 7I2 and / or Figure 7W (As described above). With the current video media capture mode selected, computer system 700 begins capturing video media.
[0286] like Figure 7AA6As illustrated, when capturing video media, the computer system 700 stops displaying the mode control enablement display 746 and instead displays the video status enablement display 750 in the lower area of the camera viewfinder 712. The video status enablement display 750 indicates that video media is currently being captured and includes the current past time of the video capture. At the start of video capture, the computer system 700 initially displays the video status enablement display 750 in an opaque manner. Figure 7AA7 As illustrated, after a threshold amount of time (e.g., a few seconds) during which no gaze 732 directed at the video state salience representation 750 is detected, the computer system 700 increases the semi-transparency of the video state salience representation 750 (e.g., decreases its opacity) (e.g., as described above regarding changing the visual salience of the captured media icon 738 and / or the modal control salience representation 746). Figure 7AA8 As illustrated, in response to detecting a gaze 732 directed at the video state indicator 750, the computer system 700 reduces the semi-transparency of the video state indicator 750 (e.g., increases its opacity).
[0287] like Figure 7AA8 As illustrated, when a video status indicator 750 with increased visual salience is displayed, the computer system 700 detects a seventh potential media capture input, such as a button press input 752A of hardware button 706, an air gesture input 752B, and / or a tap input 752C. Upon detection of the seventh potential media capture input, gaze 732 is directed towards the video status indicator 750. Therefore, in response to the seventh potential media capture input, the computer system 700 terminates video media capture. Figure 7AA9 As illustrated, when video media capture ends, the computer system stops displaying the video status indicator 750 and again displays the mode control indicator 746 in the lower area of the camera viewfinder 712.
[0288] like Figure 7AB As illustrated, when the computer system 700 is not capturing any media, the computer system 700 no longer detects the position of the user's hand (e.g., finger 724 is no longer near hardware button 706, and / or hand 726 is no longer in the field of view of multiple cameras). In response, the computer system 700 changes the appearance of one or more elements of the media capture user interface 710, indicating that the computer system 700 is no longer in the media capture ready-to-capture state (e.g., reverts to the appearance of one or more elements of the media capture user interface 710 in the ready-to-capture state). Figure 7B (How it appears in the text).
[0289] In some implementations of the computer system 700, which is a head-mounted device, Figures 7A to 7FThe illustrated techniques are implemented such that a user prepares to capture media by moving their hands to a location that may provide capture input (e.g., air gestures and / or interaction with hardware buttons) and moving their head, thereby moving the field of view of the first camera 704A and the second camera 704B to synthesize (e.g., framing) the shot. In these embodiments, the user can use gaze, air gestures, and / or hardware button presses to control functions other than initiating media capture, such as selecting option enablement display 720 to view and change media capture settings, interacting with virtual objects in the XR environment, or interacting with a virtual assistant. Thus, when the user gazes at the central area of the camera viewfinder 712 (e.g., at or near shutter enablement display 718) and provides potential capture input, the computer system 700 implements... Figures 7H to 7AA The illustrated technique, because the user's gaze indicates a potential capture input, is intended to initiate media capture rather than control other functions.
[0290] The following is a reference about Figure 8 Method 800 describes and provides information about Figures 7A to 7AB Additional description.
[0291] Figure 8This is a flowchart of an exemplary method 800 for displaying a media capture user interface for gaze activation, according to some embodiments. In some embodiments, method 800 is performed at a computer system (e.g., 101, 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.2-100, 700, X700, and / or 702) with display generation components (e.g., 1-102, 1-120a, 1-120b, 11.1.1-104a, 11.1.1-104b, 1-108, 1-122a, 1-122b, 1-202, 1-306, 1-308, 1-320, 1-322a, 1-322b, 1-406, 1-402). 1-421, 3-108, 6-334, 11.3.2-100, 11.3.2-104, 11.3.2-200, 11.3.2-204, 708 and / or X702) (e.g., display controller; touch-sensitive display system; display (e.g., integrated and / or connected), 3D display, transparent display, projector, head-up display and / or head-mounted display) and a first camera (e.g., 6-106, 6-114, 6-116, 6-118, 6-120, 6-122, 6-306, 6-416, 11.1.1-104a-b, 11.1.2-110a-f, 11.3. 2-100, 11.3.2-106 and / or 11.3.2-206, 704A, 704B and / or X704) communication (in some embodiments, the computer system includes one or more cameras, such as a rear (user-facing) camera and a front (environment-facing) camera; in some embodiments, the first camera is a virtual camera) (in some embodiments, the computer system includes one or more sensors (e.g., 1-356, 1-456, 6-102, 6-106, 6-108, 6-110, 6-112, 6-114, 6-116, 6-118, 6-120, 6-122, 6-124, 6-126, 6 -128, 6-202, 6-203, 6-302, 6-303, 6-306, 6-402, 6-416, 11.1.1-104a, 11.1.1-104b, 11.1.2-110a-f, 11.3.2-100, 11.3.2-106, 11.3.2-206 and / or X704), such as capacitive / touch sensors, gaze sensors, etc.; in some embodiments, the computer system includes at least one hardware button (e.g., 1-128, 1-132, 11.1.1-114, 1-328, 706, X706 and / or X703) that can be dynamically mapped to different functions.In some implementations, method 800 is performed by storing in a non-transitory (or transient) computer-readable storage medium and by one or more processors of a computer system (such as one or more processors 202 of computer system 101). Figure 1A The controller 110 in the middle executes instructions to manage. Some operations in method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0292] Computer systems (e.g., 101, 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.2-100, 700, X700 and / or 702) when generated via display components (e.g., 1-102, 1-120a, 1-120b, 11.1.1-104a, 11.1.1-104b, 1-108, 1-122a, 1-122b, 1-202, 1-306, 1-3...) 08, 1-320, 1-322a, 1-322b, 1-406, 1-402, 1-421, 3-108, 6-334, 11.3.2-100, 11.3.2-104, 11.3.2-200, 11.3.2-204, 708 and / or X702) display (802) including the camera viewfinder (e.g., 712 and / or X712) (e.g., viewfinder / camera preview object, such as an object that frames / covers the area used for media capture) When a user interface (e.g., 710 and / or X710) (e.g., a camera / capture UI representing at least a portion of the field of view of a first camera) (in some embodiments, at least a portion of the environment is covered via a transparent display, transparent camera data, and / or virtual content (e.g., physical environment, virtual environment, and / or mixed reality environment)) detects (804) a first input (e.g., 730A, 730B, 730C, 736A, 736B, 736C, X736A, X736B, X736C, 740A, 740B, and / or 740C) (for triggering an activation or selection input for an operation at the device when the input is detected with the user’s attention directed to an optional user interface object; in some embodiments, pressing a hardware button; in some embodiments, gesture input, air gesture (e.g., air pinch); in some embodiments, voice input; in some embodiments, the activation input does not include location-based input, such as touch input on a touch-sensitive display or mouse click).
[0293] In response to the detection of a first input (806) and based on determining that when the first input is detected (e.g., if the user (e.g., simultaneously) requests to capture and look at an appropriate part of the UI), the computer system is pointing (e.g., 732 and / or X732) at a corresponding area of the camera viewfinder (e.g., a capture activation area of the UI (e.g., a pre-determined and / or predefined area), such as the viewfinder or the central area of the UI; in some embodiments, gaze detection is ...
Claims
1. A method, the method comprising: At the computer system that communicates with the display generation component and the first camera: When a first user interface including a camera viewfinder is displayed via the display generation component: Detect the first input; and In response to the detection of the first input: Based on determining that the user of the computer system is gazing at the corresponding area of the camera viewfinder when the first input is detected, the capture of the first media content using the first camera is initiated. as well as If it is determined that the user's gaze, as input by the computer system when the first input is detected, is not directed towards the corresponding area of the camera viewfinder, the capture of the first media content is abandoned.
2. The method according to claim 1, wherein the corresponding region is the central region within the camera viewfinder.
3. The method according to any one of claims 1 to 2, wherein the first media content includes photographic media content.
4. The method according to any one of claims 1 to 3, wherein the first media content includes video content.
5. The method according to any one of claims 1 to 4, wherein the first user interface includes a first media capture optional interface object, wherein the first media capture optional interface object is displayed in the corresponding area.
6. The method of claim 5, wherein at least a first element of the first media capture optional interface object is at least partially translucent.
7. The method according to any one of claims 5 to 6, further comprising: When the camera viewfinder is displayed, the computer system detects that the user's gaze is directed towards the corresponding area of the camera viewfinder. as well as In response to detecting that the user's gaze is directed toward the corresponding area of the camera viewfinder, a first change is made to one or more visual features of the first media capture optional interface object.
8. The method of claim 7, wherein making the first change to one or more visual features of the first media capture optional interface object includes increasing the brightness of at least a portion of the first media capture optional interface object.
9. The method of any one of claims 7 to 8, wherein making the first change to one or more visual features of the first media capture optional interface object includes reducing the size of at least a portion of the first media capture optional interface object.
10. The method according to any one of claims 5 to 9, further comprising: In response to detecting the first input, a second change is made to one or more visual features of the first media capture optional interface object based on the first input.
11. The method of claim 10, wherein the first input includes a first stage and a second stage; and wherein the second change on the one or more visual features of the first media capture optional interface object based on the first input includes: In response to the first stage of detecting the first input, a third change is made to one or more visual features of the first media capture optional interface object; as well as In response to the second stage of detecting the first input, a fourth change is made to one or more visual features of the first media capture optional interface object.
12. The method according to claim 11, further comprising: After the third change is made to one or more visual features of the first media capture optional interface object, a fifth change is made to the one or more visual features of the first media capture optional interface object based on the determination that the duration of the first phase of the first input exceeds a first duration threshold.
13. The method according to any one of claims 5 to 12, further comprising: In response to detecting the first input, the computer system determines that when the first input is detected, the user's gaze is directed towards the corresponding area of the camera viewfinder: An animation is displayed in conjunction with the first media capture optional interface object, wherein the animation indicates the current past time of the media capture delay timer.
14. The method according to any one of claims 1 to 13, wherein the first input comprises an air gesture.
15. The method according to any one of claims 1 to 14, wherein the first input includes activation of a hardware button communicating with the computer system.
16. The method according to any one of claims 1 to 15, further comprising: When the first user interface is displayed and before the first input is detected: Detect the second input; as well as In response to detecting the second input, a first change is made to one or more visual features of the first user interface.
17. The method of claim 16, wherein detecting the second input includes detecting the position of the user's hand.
18. The method of any one of claims 16 to 17, wherein making the first change to one or more visual features of the first user interface includes changing the appearance of the first user interface in the area outside the camera viewfinder.
19. The method of any one of claims 16 to 18, wherein making the first change to one or more visual features of the first user interface includes changing the appearance of the second media capture user interface object.
20. The method of any one of claims 16 to 19, wherein making the first change to the one or more visual features of the first user interface includes changing the appearance of one or more user interface elements other than the third media capture optional interface object.
21. The method according to any one of claims 16 to 20, further comprising: When the first user interface is displayed: Detect the third input; as well as In response to the detection of the third input, one or more visual features of the first user interface are changed based on the third input.
22. The method according to any one of claims 1 to 21, wherein: The camera viewfinder covers a first portion of the representation of the field of view of the first camera; and The first user interface includes a first visual indication along a first edge of the camera viewfinder, wherein the first visual indication modifies the visual appearance of a second portion of the representation of the field of view of the first camera corresponding to the position of the first visual indication.
23. The method according to any one of claims 1 to 22, wherein initiating the capture of the first media content using the first camera includes displaying a first media capture animation via the display generation component.
24. The method of claim 23, wherein displaying the first media capture animation includes darkening a second area of the first user interface and extending the darkening of a first area toward the center of the camera viewfinder.
25. The method of any one of claims 23 to 24, wherein displaying the first media capture animation comprises darkening a third area of the first user interface and shrinking the darkened third area away from the center of the camera viewfinder.
26. The method of any one of claims 23 to 25, wherein before displaying the first media capture animation, the first user interface includes a second visual indication along a second edge of the camera viewfinder, the second visual indication causing a first portion of the user interface to darken to a first darkening level; and wherein displaying the first media capture animation includes darkening at least a second portion of the user interface to a second darkening level, wherein the second darkening level appears darker than the first darkening level.
27. The method according to any one of claims 1 to 26, further comprising: In response to initiating the capture of the first media content using the first camera: Based on the determination that the first input corresponds to a request to capture media of the first type, display the media capture animation of the first type; and Based on the determination that the first input corresponds to a request to capture the second type of media, a media capture animation of the second type is displayed.
28. The method of claim 27, further comprising: In response to detecting the initial portion of the first input, a first visual indicator is displayed for a corresponding time period; in: When the first input is stopped while the first visual indicator is being continuously displayed, the first input corresponds to a request to capture the first type of media, and When the first input is stopped after the first visual indicator is stopped being displayed, the first input corresponds to a request to capture the second type of media.
29. The method according to any one of claims 1 to 28, further comprising: After capturing the first media content, a representation of the first media content is displayed via the display generation component.
30. The method of claim 29, wherein the representation of the first media content includes animation of changes in the appearance of the representation of the first media content.
31. The method according to any one of claims 29 to 30, wherein at least a portion of the representation of the first media content is at least partially transparent.
32. The method according to claim 31, further comprising: When the representation of the first media content is displayed, user interaction with the representation of the first media content is detected; as well as In response to detecting user interaction with the representation of the first media content, the transparency of the representation of the first media content is changed.
33. The method of claim 32, wherein detecting the user interaction with the representation of the first media content includes determining that the user of the computer system is gazing at the representation of the first media content.
34. The method of claim 32, wherein detecting the user interaction with the representation of the first media content includes detecting the initiation of capturing the second media content using the first camera.
35. The method according to any one of claims 1 to 34, wherein: The first user interface includes a fourth media capture optional interface object displayed at a first location in the first user interface; The capture of the first media content using the first camera includes video capture; and Initiating the capture of the first media content using the first camera includes displaying an animation of the fourth media capture optional object moving toward the second position via the display generation component.
36. The method of any one of claims 1 to 35, wherein at least a third portion of the first user interface covers at least a portion of the field of view of the first camera, and wherein displaying the first user interface includes altering one or more visual features of the at least third portion of the user interface based on the at least portion of the field of view of the first camera.
37. The method according to any one of claims 1 to 36, further comprising: Detect the movement of the field of view of the first camera in the first direction; When the movement of the field of view of the first camera in the first direction is detected, the second part of the first user interface is moved away from the initial position of the second part of the first user interface in a second direction opposite to the first direction; as well as After moving the second portion of the first user interface in the second direction, the second portion of the first user interface is moved back to the initial position of the second portion of the first user interface.
38. The method according to any one of claims 1 to 37, wherein the first user interface comprises one or more modal control user interface objects, and the method comprises: When the one or more modal control user interface objects are displayed, detect input pointing to the one or more modal control user interface objects; as well as In response to detecting input directed to one or more modal control user interface objects, switch from the current capture mode to a different capture mode.
39. The method of claim 38, wherein initiating the capture of the first media content using the first camera comprises: It is determined that the computer system is configured to capture media using a first capture mode, capturing media types corresponding to the first capture mode; as well as The computer system is determined to be configured to capture media using a second capture mode different from the first capture mode, capturing different media types that correspond to the second capture mode and are different from the media types corresponding to the first capture mode.
40. The method of any one of claims 38 to 39, wherein the one or more mode control user interface objects include elements representing the first capture mode and elements representing the second capture mode, and wherein displaying the mode control user interface object includes: Based on the determination that the computer system is configured to capture media using the first capture mode, the elements representing the first capture mode are visually emphasized in a corresponding manner; as well as Based on the determination that the computer system is configured to capture media using the second capture mode, the elements representing the second capture mode are visually emphasized in the corresponding manner.
41. The method according to any one of claims 1 to 40, further comprising: Based on the determination that a first set of criteria is met, the visual emphasis of one or more elements of a first group of the first user interface is reduced, wherein the first set of criteria includes criteria that are met when the user's gaze on the first user interface is not directed at one or more elements of the first group of the first user interface for a period of time exceeding a threshold.
42. The method according to any one of claims 1 to 41, further comprising: Based on the determination that a second set of criteria is met, visual emphasis is increased on one or more elements of a second set of criteria in the first user interface, wherein the second set of criteria includes criteria that are met when the user's gaze is directed at one or more elements of the second set of criteria in the first user interface.
43. The method according to any one of claims 1 to 42, further comprising: In response to the detection of an event, a third group of one or more elements of the first user interface are displayed with increased visual salience.
44. The method of claim 43, wherein when the media capture application is launched, the third group of one or more elements of the first user interface is displayed with increased visual salience before the visual salience of the third group of one or more elements of the first user interface is automatically reduced to a lower level of visual salience.
45. The method of any one of claims 43 to 44, wherein the occurrence of the event includes switching between the current capture mode and a third capture mode and a fourth capture mode different from the third capture mode.
46. The method of any one of claims 43 to 45, wherein the event occurrence includes a media capture event.
47. The method according to any one of claims 1 to 46, further comprising: Displays a user interface object indicating the current media capture mode; as well as When the first media content is captured: In the first user interface, status information about the media capture operation is displayed in the location previously occupied by the user interface object that was indicated as the current media capture mode.
48. The method according to any one of claims 1 to 47, the method further comprising: When the first media content is captured, an indication of the status of capturing the first media content is displayed; When the indication of the state of capturing the first media content is displayed, input pointing to the indication of the state of capturing the first media content is detected; as well as In response to the detection of the input indicating the state of capturing the first media content, the capture of the first media content is stopped.
49. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a first camera, the one or more programs including instructions for performing the method according to any one of claims 1 to 48.
50. A computer system configured to communicate with a display generation component and a first camera, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 48.
51. A computer system configured to communicate with a display generation component and a first camera, the computer system comprising: Apparatus for performing the method according to any one of claims 1 to 48.
52. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and a first camera, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 48.
53. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a first camera, the one or more programs including instructions for: When a first user interface including a camera viewfinder is displayed via the display generation component: Detect the first input; and In response to the detection of the first input: Based on determining that the user of the computer system is gazing at a corresponding area of the camera viewfinder when the first input is detected, the system initiates the capture of first media content using the first camera; and If it is determined that the user's gaze, as input by the computer system when the first input is detected, is not directed towards the corresponding area of the camera viewfinder, the capture of the first media content is abandoned.
54. A computer system configured to communicate with a display generation component and a first camera, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When a first user interface including a camera viewfinder is displayed via the display generation component: Detect the first input; and In response to the detection of the first input: Based on determining that the user of the computer system is gazing at the corresponding area of the camera viewfinder when the first input is detected, the capture of the first media content using the first camera is initiated. as well as If it is determined that the user's gaze, as input by the computer system when the first input is detected, is not directed towards the corresponding area of the camera viewfinder, the capture of the first media content is abandoned.
55. A computer system configured to communicate with a display generating component and a first camera, the computer system comprising: When a first user interface including a camera viewfinder is displayed via the display generation component: A device for performing the following operation: detecting a first input; as well as In response to the detection of the first input: A device for performing the following operation: initiating the capture of first media content using the first camera based on determining that a user of the computer system is gazing at a corresponding area of the camera viewfinder when the first input is detected; as well as A device for performing the following operation: abandoning the initiation of capturing the first media content based on determining that the user's gaze, as input by the computer system when the first input is detected, is not directed towards the corresponding area of the camera viewfinder.
56. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a first camera, the one or more programs comprising instructions for performing the following operations: When a first user interface including a camera viewfinder is displayed via the display generation component: Detect the first input; and In response to the detection of the first input: Based on determining that the user of the computer system is gazing at a corresponding area of the camera viewfinder when the first input is detected, the system initiates the capture of first media content using the first camera; and If it is determined that the user's gaze, as input by the computer system when the first input is detected, is not directed towards the corresponding area of the camera viewfinder, the capture of the first media content is abandoned.
57. A method, the method comprising: At the computer system that communicates with the display generation component and the first camera: A first user interface displays a camera preview including at least a portion of the field of view of the first camera via the display generation component; Detect the change in orientation of the field of view of the first camera relative to the corresponding orientation; as well as In response to the detection of the change in orientation: Based on the determination that a first set of criteria is met, a first indicator representing the orientation of the field of view of the first camera is displayed, wherein the first set of criteria includes a first criterion that is met when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
58. The method according to claim 57, further comprising: When the first indicator representing the orientation of the field of view of the first camera is displayed, a second change in the orientation of the field of view of the first camera relative to the corresponding orientation is detected; as well as In response to the detection of the second change in the orientation: The first indicator is stopped from being displayed if the second set of criteria is met, wherein the second set of criteria includes a second criterion that is met when the difference between the current orientation and the corresponding orientation of the field of view of the first camera is less than a second threshold amount.
59. The method of claim 58, wherein the second threshold quantity is different from the first threshold quantity.
60. The method according to any one of claims 57 to 59, wherein: The first user interface includes a first visual indication along a first edge of the camera preview; The first visual indication modifies the visual appearance of the second portion of the field of view represented by the first camera, which is located below the first visual indication. and The first visual indication is viewpoint locked.
61. The method according to any one of claims 57 to 60, wherein a first portion of the first indicator representing the orientation of the field of view of the first camera is displayed inside a first area of the first user interface, and wherein a second portion of the first indicator representing the orientation of the field of view of the first camera is displayed outside the first area of the first user interface.
62. The method of claim 61, wherein the first portion of the first indicator is displayed in an orientation that remains fixed relative to the user's viewpoint.
63. The method of claim 62, wherein the second portion of the first indicator is displayed in an orientation that remains fixed relative to one or more portions of the three-dimensional environment.
64. The method according to any one of claims 61 to 63, the method further comprising: After displaying the first indicator representing the orientation of the field of view of the first camera, a third change in the orientation of the field of view of the first camera relative to the corresponding orientation is detected; as well as In response to the detection of the third change in the orientation: Based on the determination that the difference between the current orientation and the corresponding orientation of the field of view of the first camera has increased, the angle between the first portion and the second portion of the first indicator is increased.
65. The method according to any one of claims 61 to 64, the method further comprising: When the first indicator is displayed: Based on the determination that the first set of criteria is met: The first portion of the first indicator is displayed in the first color; as well as The second portion of the first indicator is displayed in a second color different from the first color; as well as Based on the determination that a third set of criteria is met, both the first portion and the second portion are displayed in a third color, wherein the third set of criteria includes a third criterion that is met when the difference between the current orientation and the corresponding orientation of the field of view of the first camera is less than a third threshold amount.
66. The method according to any one of claims 61 to 65, wherein the first region of the first user interface is a circular region having a first diameter: When the first indicator is displayed: Based on the determination that the first set of criteria is met: The first portion of the first indicator is displayed with a width smaller than the first diameter; and The second portion of the first indicator is displayed as two or more separate elements displayed on different sides of the first portion, wherein the inner edges of the first element and the inner edges of the second element are spaced apart from the first portion of the first indicator; and Based on the determination that a fourth set of criteria is met, the inner edges of the first element and the inner edges of the second element are shifted toward the first portion, wherein the fourth set of criteria includes a fourth criterion that is met when the difference between the current orientation and the corresponding orientation of the field of view of the first camera is less than a fourth threshold amount.
67. The method of any one of claims 57 to 66, wherein displaying the first user interface includes displaying a media capture optional interface object, and wherein when the media capture optional interface object is displayed, the first indicator representing the orientation of the field of view of the first camera is executed.
68. The method of claim 67, wherein the first indicator representing the orientation of the field of view of the first camera is visually incorporated into the media capture optional interface object.
69. The method of any one of claims 67 to 68, wherein at least a portion of the first indicator overlaps with at least a portion of the media capture optional interface object.
70. The method according to any one of claims 67 to 69, the method further comprising: Detect the corresponding input; as well as Based on the corresponding input, change one or more visual features of the media capture optional interface object.
71. The method according to any one of claims 57 to 70, the method further comprising: The display generating component displays a representation of the first media content captured by the first camera, wherein the representation is displayed in an orientation aligned with a first orthogonal axis of the corresponding orientation.
72. The method according to any one of claims 57 to 71, wherein the change in the orientation of the field of view of the first camera relative to the corresponding orientation is caused by a change in the orientation of at least a portion of the body of the user of the computer system.
73. The method according to any one of claims 57 to 72, wherein the computer system further communicates with a second camera spaced apart from the first camera, and wherein the first set of criteria is satisfied when the orientation of a line formed between the first camera and the second camera is substantially parallel to a second orthogonal axis of the respective orientation.
74. The method according to any one of claims 57 to 73, wherein the first user interface includes an option enablement representation that, when selected, causes the display of one or more media capture option selectable interface objects.
75. The method of claim 74, wherein the first set of criteria includes a second set of criteria that are satisfied when a first media capture setting, which can be controlled by a first media capture option optional interface object in one or more media capture option optional interface objects, is in a first state.
76. The method of any one of claims 74 to 75, wherein the one or more media capture option optional interface objects include a photo mode optional interface object for switching multi-frame photo capture modes.
77. The method according to any one of claims 74 to 76, wherein the first user interface includes a first state indicator that indicates the current state of a second media capture option optional interface object corresponding to one or more media capture option optional interface objects, indicating the second media capture settings.
78. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a first camera, the one or more programs including instructions for performing the method according to any one of claims 57 to 77.
79. A computer system configured to communicate with a display generation component and a first camera, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 57 to 77.
80. A computer system configured to communicate with a display generation component and a first camera, the computer system comprising: Apparatus for performing the method according to any one of claims 57 to 77.
81. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and a first camera, the one or more programs comprising instructions for performing the method according to any one of claims 57 to 77.
82. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a first camera, the one or more programs including instructions for: A first user interface displays a camera preview including at least a portion of the field of view of the first camera via the display generation component; Detecting the change in orientation of the field of view of the first camera relative to a corresponding orientation; and In response to the detection of the change in orientation: Based on the determination that a first set of criteria is met, a first indicator representing the orientation of the field of view of the first camera is displayed, wherein the first set of criteria includes a first criterion that is met when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
83. A computer system configured to communicate with a display generation component and a first camera, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: A first user interface displays a camera preview including at least a portion of the field of view of the first camera via the display generation component; Detect the change in orientation of the field of view of the first camera relative to the corresponding orientation; as well as In response to the detection of the change in orientation: Based on the determination that a first set of criteria is met, a first indicator representing the orientation of the field of view of the first camera is displayed, wherein the first set of criteria includes a first criterion that is met when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
84. A computer system configured to communicate with a display generation component and a first camera, the computer system comprising: An apparatus for performing the following operations: displaying a first user interface via the display generation component a camera preview including at least a portion of the field of view of the first camera; A device for performing the following operation: detecting a change in the orientation of the field of view of the first camera relative to a corresponding orientation; as well as In response to the detection of the change in orientation: A device for performing the following operation: displaying a first indicator representing the orientation of the field of view of the first camera, based on determining that a first set of criteria is met, wherein the first set of criteria includes a first criterion that is met when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
85. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a first camera, the one or more programs comprising instructions for performing the following operations: A first user interface displays a camera preview including at least a portion of the field of view of the first camera via the display generation component; Detecting the change in orientation of the field of view of the first camera relative to a corresponding orientation; and In response to the detection of the change in orientation: Based on the determination that a first set of criteria is met, a first indicator representing the orientation of the field of view of the first camera is displayed, wherein the first set of criteria includes a first criterion that is met when the difference between the current orientation of the field of view of the first camera and the corresponding orientation exceeds a first threshold amount.
86. A method comprising: At the computer system that communicates with the display generation component and multiple cameras, including a first camera and a second camera: The display generating component displays a capture preview for spatial media, wherein a capture input detected while displaying the capture preview causes the computer system to capture media from the first camera and the second camera to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, the one or more images forming an illusion of a spatial representation of the field of view of the plurality of cameras when viewed simultaneously; When the capture preview for spatial media capture is displayed, the position of the individual within the field of view of the plurality of cameras is detected; as well as In response to detecting the position of the individual within the field of view of the plurality of cameras, and based on a criterion that the individual's position relative to the field of view of the plurality of cameras is insufficient to capture spatial media at a threshold quality level, a prompt indicating a change in the distance between the individual and the plurality of cameras is displayed via the display generation component.
87. The method of claim 86, wherein the prompt for changing the distance between the individual and the plurality of cameras includes text.
88. The method according to any one of claims 86 to 87, the method further comprising: In response to detecting the position of the individual within the field of view of the plurality of cameras, and based on determining that the individual's position relative to the field of view of the plurality of cameras meets the criterion for capturing spatial media at the threshold quality level, the prompt to change the distance between the individual and the plurality of cameras is abandoned.
89. The method of any one of claims 86 to 88, wherein the capture preview for spatial capture includes a media capture optional interface object, and wherein the prompt to change the distance between the individual and the plurality of cameras includes a first change to one or more visual features of the media capture optional interface object.
90. The method of claim 89, wherein the media capture optional interface object is displayed in the central area of the capture preview for spatial capture.
91. The method according to any one of claims 89 to 90, wherein making the first change to one or more visual features of the media capture optional interface object comprises changing a first portion of the media capture optional interface object from a first color to a second color.
92. The method according to claim 91, further comprising: When the media capture optional interface object is displayed in the second color, the first portion of the media capture optional interface object is changed from the second color to the first color for a period of time.
93. The method according to any one of claims 89 to 92, wherein: Before displaying the prompt to change the distance between the individual and the plurality of cameras, a first portion of the media capture optional object is partially transparent, such that a first portion of the camera preview used for spatial capture is visible through the first portion of the media capture optional interface object; and Making the first change to one or more visual features of the media capture optional interface object includes at least partially obscuring the first portion of the camera preview.
94. The method of claim 93, wherein obscuring the first portion of the camera preview at least partially comprises blurring the first portion of the camera preview.
95. The method according to any one of claims 93 to 94, wherein obscuring the first portion of the camera preview at least partially comprises darkening the first portion of the camera preview.
96. The method according to any one of claims 93 to 95, the method further comprising: After the first portion of the camera preview is at least partially obscured, a change in the position of the individual within the field of view of the plurality of cameras is detected; as well as In response to detecting a change in the position of the individual within the field of view of the plurality of cameras, a second change is made to one or more visual features of the media capture optional interface object.
97. The method of claim 96, wherein the second change to one or more visual features of the media capture optional interface object is based on the direction of the change of the individual's position.
98. The method of any one of claims 96 to 97, wherein the second change to one or more visual features of the media capture optional interface object is based on the magnitude of the change in the position of the individual.
99. The method of any one of claims 86 to 98, wherein the prompting to change the distance between the individual and the plurality of cameras includes a prompting to decrease the distance between the individual and the plurality of cameras.
100. The method of any one of claims 86 to 99, wherein the prompt to change the distance between the individual and the plurality of cameras includes a prompt to increase the distance between the individual and the plurality of cameras.
101. The method according to any one of claims 86 to 100, the method further comprising: The computer system detects the movement of a user in a physical environment, wherein, in response to detecting the movement of the user, the system performs the detection of the individual's position within the field of view of the plurality of cameras.
102. The method according to any one of claims 86 to 101, the method further comprising: Detecting a change in the field of view of the plurality of cameras, wherein detecting the position of the individual within the field of view of the plurality of cameras is performed in response to detecting the change in the field of view of the plurality of cameras.
103. The method according to any one of claims 86 to 102, wherein: The first camera and the second camera generate the spatial media item by capturing at least part of the content within a first set of one or more capture planes within the field of view of the first camera and the field of view of the second camera; and When the location of the individual is detected, multiple objects are within the field of view of the first camera and / or the second camera; and The plurality of objects includes a first object, which is the object among the plurality of objects that is geographically closest to the first group of one or more capture planes; and The individual is the first object.
104. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with display generation components and a plurality of cameras including a first camera and a second camera, the one or more programs including instructions for performing the method according to any one of claims 86 to 103.
105. A computer system configured to communicate with a display generation component and a plurality of cameras including a first camera and a second camera, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 86 to 103.
106. A computer system configured to communicate with a display generation component and a plurality of cameras including a first camera and a second camera, the computer system comprising: Apparatus for performing the method according to any one of claims 86 to 103.
107. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and a plurality of cameras including a first camera and a second camera, the one or more programs comprising instructions for performing the method according to any one of claims 86 to 103.
108. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and a plurality of cameras including a first camera and a second camera, the one or more programs including instructions for: The display generating component displays a capture preview for spatial media, wherein a capture input detected while displaying the capture preview causes the computer system to capture media from the first camera and the second camera to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, the one or more images forming an illusion of a spatial representation of the field of view of the plurality of cameras when viewed simultaneously; When the capture preview for spatial media capture is displayed, the position of the individual within the field of view of the plurality of cameras is detected; as well as In response to detecting the position of the individual within the field of view of the plurality of cameras, and based on a criterion that the individual's position relative to the field of view of the plurality of cameras is insufficient to capture spatial media at a threshold quality level, a prompt indicating a change in the distance between the individual and the plurality of cameras is displayed via the display generation component.
109. A computer system configured to communicate with a display generation component and a plurality of cameras including a first camera and a second camera, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: The display generating component displays a capture preview for spatial media, wherein a capture input detected while displaying the capture preview causes the computer system to capture media from the first camera and the second camera to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, the one or more images forming an illusion of a spatial representation of the field of view of the plurality of cameras when viewed simultaneously; When the capture preview for spatial media capture is displayed, the position of the individual within the field of view of the plurality of cameras is detected; as well as In response to detecting the position of the individual within the field of view of the plurality of cameras, and based on a criterion that the individual's position relative to the field of view of the plurality of cameras is insufficient to capture spatial media at a threshold quality level, a prompt indicating a change in the distance between the individual and the plurality of cameras is displayed via the display generation component.
110. A computer system configured to communicate with a display generation component and a plurality of cameras including a first camera and a second camera, the computer system comprising: An apparatus for performing the following operation: displaying a capture preview for spatial media via the display generation component, wherein a capture input detected while displaying the capture preview causes the computer system to capture media from the first camera and the second camera to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, the one or more images forming an illusion of a spatial representation of the field of view of the plurality of cameras when viewed simultaneously; A device for performing the following operation: when the capture preview for spatial media capture is displayed, detecting the position of an individual within the field of view of the plurality of cameras; as well as An apparatus for performing the following operations: in response to detecting the position of the individual in the field of view of the plurality of cameras, displaying a prompt via the display generation component indicating a change in the distance between the individual and the plurality of cameras, based on a criterion that the individual's position relative to the field of view of the plurality of cameras is insufficient to capture spatial media at a threshold quality level.
111. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and a plurality of cameras including a first camera and a second camera, the one or more programs comprising instructions for performing the following operations: The display generating component displays a capture preview for spatial media, wherein a capture input detected while displaying the capture preview causes the computer system to capture media from the first camera and the second camera to generate a spatial media item, the spatial media item including one or more images for the right eye and one or more images for the left eye, the one or more images forming an illusion of a spatial representation of the field of view of the plurality of cameras when viewed simultaneously; When the capture preview for spatial media capture is displayed, the position of the individual within the field of view of the plurality of cameras is detected; as well as In response to detecting the position of the individual within the field of view of the plurality of cameras, and based on a criterion that the individual's position relative to the field of view of the plurality of cameras is insufficient to capture spatial media at a threshold quality level, a prompt indicating a change in the distance between the individual and the plurality of cameras is displayed via the display generation component.
112. A method, the method comprising: At a computer system communicating with display generation components and one or more sensors, said one or more sensors include one or more cameras: Use one or more of the cameras to capture video media; as well as When the video media is captured: Detecting the movement of the one or more cameras via the one or more sensors; and In response to the detection of the movement of the one or more cameras: Based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, the movement of a visual indicator relative to a display reference object is displayed via the display generation component, wherein the movement of the visual indicator includes: Based on determining that the movement of the one or more cameras is a movement in a first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the first direction of indicator movement; as well as Based on the determination that the movement of the one or more cameras is a movement in a second direction of camera movement, different from the first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the second direction of indicator movement, different from the first direction of indicator movement.
113. The method of claim 112, wherein displaying the movement of the visual indicator relative to the display reference object comprises: Based on determining that the movement of the one or more cameras is a movement of a first magnitude, the visual indicator is shown to have moved a first distance relative to the display reference object; as well as Based on the determination that the movement of the one or more cameras is a movement of a second value different from the first value, the visual indicator is displayed to move a second distance relative to the display reference object, which is different from the first distance.
114. The method according to any one of claims 112 to 113, wherein the set of one or more motion criteria includes a first criterion satisfied when the amount of motion of the one or more cameras exceeds a first threshold amount.
115. The method according to any one of claims 112 to 114, wherein the set of one or more motion criteria includes a second criterion satisfied when the amount of change in the motion of the one or more cameras exceeds a first threshold change.
116. The method according to any one of claims 112 to 115, the method further comprising: The visual indicator is displayed within the display reference object.
117. The method of any one of claims 112 to 116, wherein the visual indicator is included in an optional user interface object displayed via the display generation component, the optional user interface object causing the capture of the video media to stop when selected.
118. The method according to claim 117, further comprising: When capturing the video media, detect air gesture input; as well as In response to detecting the air gesture input and based on determining that the user of the computer system is looking at the optional user interface object when the air gesture input is detected, the capture of the video media is stopped.
119. The method according to any one of claims 117 to 118, the method further comprising: When capturing the video media, detect user input that selects the optional user interface object; as well as In response to detecting user input that selects the optional user interface object: Stop displaying the optional user interface object; as well as Display a second optional user interface object that differs from the first optional user interface object, which initiates media capture when selected.
120. The method according to any one of claims 112 to 119, wherein the movement of the visual indicator includes displaying the visual indicator moving according to simulated physical properties.
121. The method according to any one of claims 112 to 120, wherein the movement of displaying the visual indicator comprises: Based on determining that the movement of the one or more cameras is a first-value movement, the spatial characteristics of the visual indicator are distorted by a first amount; as well as Based on determining that the movement of the one or more cameras is a movement of a second value different from the first value, the spatial characteristics of the visual indicator are distorted by a second value, which is different from the first value of the spatial characteristics distortion of the visual indicator.
122. The method according to any one of claims 112 to 121, wherein: The first direction of the indicator movement indicates the first direction of the camera movement, wherein the first direction of the camera movement includes a first set of one or more directions corresponding to a plurality of components of the movement of the one or more cameras; and The second direction of the indicator movement indicates the second direction of the camera movement, wherein the second direction of the camera movement includes a second set of one or more directions corresponding to the plurality of components of the movement of the one or more cameras.
123. The method of claim 122, wherein the plurality of components of the movement of the one or more cameras includes one or more translation components.
124. The method according to any one of claims 122 to 123, wherein the plurality of components of the movement of the one or more cameras includes one or more rotational components.
125. The method according to any one of claims 122 to 124, wherein one or more components of the movement of the one or more cameras are not included in the plurality of components of the movement of the one or more cameras.
126. The method according to any one of claims 112 to 125, wherein capturing the video media using the one or more cameras comprises: Generate the first video component corresponding to the right eye viewpoint; as well as A second video component, different from the first video component, is generated corresponding to the left eye viewpoint, wherein viewing the first video component and the second video component simultaneously creates the illusion of a three-dimensional representation of the video media.
127. The method according to any one of claims 112 to 126, the method further comprising: Before capturing the video media, a flush indicator is displayed, wherein the flush indicator indicates the orientation of the one or more cameras relative to a corresponding orientation.
128. The method according to any one of claims 112 to 127, the method further comprising: When the visual indicator is displayed: In response to detecting movement of the one or more cameras and determining that the amount of movement of the one or more cameras exceeds a notification threshold, a text notification is displayed.
129. The method according to claim 128, further comprising: When the text notification is displayed, a second movement of the one or more cameras is detected via the one or more sensors; as well as In response to detecting the second movement of the one or more cameras and determining that the amount of the second movement of the one or more cameras does not exceed the notification hold threshold, the display of the text notification is stopped.
130. The method of claim 129, wherein the notification retention threshold is a magnitude threshold lower than the notification threshold.
131. A non-transitory computer-readable storage medium storing one or more programs executed by one or more processors of a computer system configured to communicate with a display generation component and one or more sensors, the one or more sensors including one or more cameras, the one or more programs including instructions for performing the method according to any one of claims 112 to 130.
132. A computer system configured to communicate with a display generation component and one or more sensors, said one or more sensors including one or more cameras, said computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 112 to 130.
133. A computer system configured to communicate with a display generation component and one or more sensors, said one or more sensors including one or more cameras, said computer system comprising: Apparatus for performing the method according to any one of claims 112 to 130.
134. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more sensors, the one or more sensors including one or more cameras, the one or more programs including instructions for performing the method according to any one of claims 112 to 130.
135. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with display generation components and one or more sensors, the one or more sensors including one or more cameras, the one or more programs including instructions for: Use one or more of the cameras to capture video media; as well as When the video media is captured: Detecting the movement of the one or more cameras via the one or more sensors; and In response to the detection of the movement of the one or more cameras: Based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, the movement of a visual indicator relative to a display reference object is displayed via the display generation component, wherein the movement of the visual indicator includes: Based on determining that the movement of the one or more cameras is a movement in a first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the first direction of indicator movement; as well as Based on the determination that the movement of the one or more cameras is a movement in a second direction of camera movement, different from the first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the second direction of indicator movement, different from the first direction of indicator movement.
136. A computer system configured to communicate with a display generation component and one or more sensors, said one or more sensors including one or more cameras, said computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: Use one or more of the cameras to capture video media; as well as When the video media is captured: Detecting the movement of the one or more cameras via the one or more sensors; and In response to the detection of the movement of the one or more cameras: Based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, the movement of a visual indicator relative to a display reference object is displayed via the display generation component, wherein the movement of the visual indicator includes: Based on determining that the movement of the one or more cameras is a movement in a first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the first direction of indicator movement; as well as Based on the determination that the movement of the one or more cameras is a movement in a second direction of camera movement, different from the first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the second direction of indicator movement, different from the first direction of indicator movement.
137. A computer system configured to communicate with a display generation component and one or more sensors, said one or more sensors including one or more cameras, said computer system comprising: A device for performing the following operations: capturing video media using the one or more cameras; as well as A device for performing the following operation when capturing the video media: Detecting the movement of the one or more cameras via the one or more sensors; and In response to the detection of the movement of the one or more cameras: Based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, the movement of a visual indicator relative to a display reference object is displayed via the display generation component, wherein the movement of the visual indicator includes: Based on determining that the movement of the one or more cameras is a movement in a first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the first direction of indicator movement; as well as Based on the determination that the movement of the one or more cameras is a movement in a second direction of camera movement, different from the first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the second direction of indicator movement, different from the first direction of indicator movement.
138. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a display generation component and one or more sensors, the one or more sensors including one or more cameras, the one or more programs including instructions for: Use one or more of the cameras to capture video media; as well as When the video media is captured: Detecting the movement of the one or more cameras via the one or more sensors; and In response to the detection of the movement of the one or more cameras: Based on determining that the movement of the one or more cameras satisfies one or more sets of movement criteria, the movement of a visual indicator relative to a display reference object is displayed via the display generation component, wherein the movement of the visual indicator includes: Based on determining that the movement of the one or more cameras is a movement in a first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the first direction of indicator movement; as well as Based on the determination that the movement of the one or more cameras is a movement in a second direction of camera movement, different from the first direction of camera movement, the visual indicator is shown to move relative to the display reference object in the second direction of indicator movement, different from the first direction of indicator movement.
139. A method comprising: At the computer system communicating with the display generation component: While playback of a video media item is in progress, wherein the playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item, and altering the visual salience of the video media item relative to the boundary region based on a representation of a movement corresponding to the viewpoint of the video media item that occurred when the video media item was captured, wherein altering the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes: Based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount, the visual salience of the video media item relative to the boundary region is changed to a first relative visual salience level; and Based on the determination that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount, the visual salience of the video media item relative to the boundary region is changed to a second relative visual salience level different from the first relative visual salience level.
140. The method of claim 139, wherein: The first movement amount is a larger movement amount than the second movement amount; and Compared to the second visual saliency level, the video media is displayed less prominently relative to the boundary region at the first visual saliency level.
141. The method according to any one of claims 139 to 140, wherein the boundary region includes a pass-through region, the pass-through region including a representation of the user's physical environment.
142. The method according to any one of claims 139 to 141, the method further comprising: The detection request plays back the user input of the video media item; as well as In response to the detection of the user input, playback of the video media item is initiated.
143. The method of any one of claims 139 to 142, wherein the video media item is stored in association with the representation of the movement corresponding to the viewpoint of the video media.
144. The method of any one of claims 139 to 143, wherein the representation of the movement corresponding to the viewpoint of the video media includes movement information captured when the video media item is captured.
145. The method of any one of claims 139 to 144, wherein the representation of the movement corresponding to the viewpoint of the video media includes movement information determined after the video media item is captured.
146. The method of any one of claims 139 to 145, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item comprises: Based on the determination that the movement corresponding to the viewpoint of the video item corresponds to the first movement amount, the visual salience of the video media item relative to the boundary region is changed by a first change amount; as well as Based on the determination that the movement corresponding to the viewpoint of the video item corresponds to the second movement amount, the visual salience of the video media item relative to the boundary region is changed by a second change amount, which is different from the first change amount.
147. The method according to any one of claims 139 to 146, wherein: The first movement amount is a larger movement amount than the second movement amount; Changing the visual salience of the video media item relative to the boundary region to the first visual salience level includes displaying the boundary region occupying the first area; and Changing the visual salience of the video media item relative to the boundary region to the second visual salience level includes displaying the boundary region, which occupies a second region smaller than the first region.
148. The method according to any one of claims 139 to 146, wherein: The first movement amount is a smaller movement amount than the second movement amount; Changing the visual salience of the video media item relative to the boundary region to the first visual salience level includes displaying the boundary region occupying the third area; and Changing the visual salience of the video media item relative to the boundary region to the second visual salience level includes displaying the boundary region that occupies a fourth region larger than the third region.
149. The method according to any one of claims 139 to 148, the method further comprising: While the video media item is being played, the visual characteristics of the boundary region change over a period of time as the video plays.
150. The method according to any one of claims 139 to 149, wherein: Changing the visual salience of the video media item relative to the boundary region to the first relative visual salience level includes displaying the video media item at a first visual ratio and occupying a first area; and Changing the visual salience of the video media item relative to the boundary region to the first relative visual salience level includes displaying the video media item at the first visual scale and occupying a second region of a different size than the first region.
151. The method of any one of claims 139 to 150, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item comprises: Based on the movement of the viewpoint corresponding to the upcoming segment of the video media item, a third movement amount is determined, which changes the visual salience of the video media item relative to the boundary region to a third relative visual salience level during playback of a portion of the video media item preceding the upcoming segment.
152. The method according to any one of claims 139 to 151, wherein: Determining that the movement of the viewpoint corresponding to the video media item corresponds to the first movement amount includes determining that the amount of movement of the viewpoint corresponding to the video media item corresponds to the first movement amount value; and Determining that the movement of the viewpoint corresponding to the video media item corresponds to the second movement amount includes determining that the amount of movement of the viewpoint corresponding to the video media item corresponds to a second movement amount value that is different from the first movement amount value.
153. The method according to any one of claims 139 to 152, wherein: Determining that the movement of the viewpoint corresponding to the video media item corresponds to the first movement amount includes determining that the rate of change of the magnitude of the movement of the viewpoint corresponding to the video media item corresponds to the first rate of change; and Determining that the movement corresponding to the viewpoint of the video media item corresponds to the second movement amount includes determining that the rate of change of the amount of movement corresponding to the viewpoint of the video media item corresponds to a second rate of change that is different from the first rate of change.
154. The method of any one of claims 139 to 153, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes changing the size of the corresponding area occupied by the video media item.
155. The method according to any one of claims 139 to 154, wherein: Displaying the video media item and its boundary area outside the video media item simultaneously includes applying a blur effect to the corresponding display area; and Changing the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes changing the blur radius of the blur effect.
156. The method according to any one of claims 139 to 155, wherein: Displaying the video media item and its boundary area outside the video media item simultaneously includes applying a feathering effect to the corresponding display area; and Changing the visual salience of a video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes changing the feather radius of the feathering effect.
157. The method of any one of claims 139 to 156, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes changing the visibility of the representation of the XR environment included in the boundary region.
158. The method of claim 157, wherein the representation of the XR environment includes a representation of the physical environment.
159. The method according to any one of claims 157 to 158, wherein the representation of the XR environment includes a representation of a virtual environment.
160. The method of any one of claims 157 to 159, wherein changing the visibility of the representation of the XR environment includes changing the darkness level of the boundary region.
161. The method of any one of claims 139 to 160, wherein the video media item comprises a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint, wherein simultaneously viewing the first video component and the second video component forms the illusion of a three-dimensional representation of the video media item.
162. The method according to claim 161, further comprising: When playback of a second video media item is in progress, wherein the second video media item does not include two or more video components, which, when viewed simultaneously, create the illusion of a three-dimensional representation of the second video media item, thereby abandoning any alteration to the visual salience of the second video media item relative to the boundary region.
163. The method according to any one of claims 139 to 162, wherein playback of the video media item includes displaying the video media item as a virtual object in an XR environment.
164. The method of any one of claims 139 to 163, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item comprises: Based on the representation of the movement corresponding to the viewpoint of the video media item, the visual salience of the video media item relative to the boundary region is changed to a corresponding relative visual salience level; as well as After changing the visual salience of the video media item relative to the boundary region to a corresponding relative visual salience level based on the representation of the movement corresponding to the viewpoint of the video media item: Based on the determination that at least a threshold time has elapsed since the visual salience of the video media item relative to the boundary region was changed to the corresponding relative visual salience level, and based on the representation of the movement of the viewpoint corresponding to the video media item, the visual salience of the video media item relative to the boundary region is changed to a relative visual salience level different from the corresponding relative visual salience level. as well as If less than a threshold time has elapsed since the visual salience of the video media item relative to the boundary region was changed to the corresponding relative visual salience level, then the change of the visual salience of the video media item relative to the boundary region is abandoned.
165. The method of any one of claims 139 to 164, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item comprises: The visual salience of the video media item relative to the boundary region is changed to a corresponding relative visual salience level; as well as After changing the visual salience of the video media item relative to the boundary region to the corresponding relative visual salience level, it is detected that the movement of the viewpoint corresponding to the video media item in the corresponding portion of the video media item will change within a corresponding amount of time. as well as In response to the detection that the movement of the viewpoint corresponding to the video media item in the corresponding portion of the video media item will change within the corresponding time amount: Based on determining that the corresponding time amount is higher than a threshold time amount, the visual salience of the video media item relative to the boundary region is changed based on the movement of the viewpoint corresponding to the video media item; as well as Based on the determination that the corresponding time amount is lower than the threshold time amount, the change of the visual salience of the video media item relative to the boundary region based on the movement of the viewpoint corresponding to the video media item is abandoned.
166. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for performing the method according to any one of claims 139 to 165.
167. A computer system configured to communicate with a display generation component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 139 to 165.
168. A computer system configured to communicate with a display generation component, the computer system comprising: Apparatus for performing the method according to any one of claims 139 to 165.
169. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising instructions for performing the method according to any one of claims 139 to 165.
170. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: While playback of a video media item is in progress, wherein the playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item, and altering the visual salience of the video media item relative to the boundary region based on a representation of a movement corresponding to the viewpoint of the video media item that occurred when the video media item was captured, wherein altering the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes: Based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount, the visual salience of the video media item relative to the boundary region is changed to a first relative visual salience level; as well as Based on the determination that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount, the visual salience of the video media item relative to the boundary region is changed to a second relative visual salience level different from the first relative visual salience level.
171. A computer system configured to communicate with a display generation component, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: While playback of a video media item is in progress, wherein the playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item, and altering the visual salience of the video media item relative to the boundary region based on a representation of a movement corresponding to the viewpoint of the video media item that occurred when the video media item was captured, wherein altering the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes: Based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount, the visual salience of the video media item relative to the boundary region is changed to a first relative visual salience level; as well as Based on the determination that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount, the visual salience of the video media item relative to the boundary region is changed to a second relative visual salience level different from the first relative visual salience level.
172. A computer system configured to communicate with a display generation component, the computer system comprising: An apparatus for performing the following operation: while playback of a video media item is in progress, wherein the playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item, and changing the visual salience of the video media item relative to the boundary region based on a representation of a movement corresponding to a viewpoint of the video media item that occurs when the video media item is captured, wherein changing the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes: Based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount, the visual salience of the video media item relative to the boundary region is changed to a first relative visual salience level; and Based on the determination that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount, the visual salience of the video media item relative to the boundary region is changed to a second relative visual salience level different from the first relative visual salience level.
173. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising instructions for: While playback of a video media item is in progress, wherein the playback of the video media item includes simultaneously displaying the video media item and a boundary region located outside the video media item, and altering the visual salience of the video media item relative to the boundary region based on a representation of a movement corresponding to the viewpoint of the video media item that occurred when the video media item was captured, wherein altering the visual salience of the video media item relative to the boundary region based on the representation of the movement corresponding to the viewpoint of the video media item includes: Based on determining that the movement corresponding to the viewpoint of the video media item corresponds to a first movement amount, the visual salience of the video media item relative to the boundary region is changed to a first relative visual salience level; as well as Based on the determination that the movement corresponding to the viewpoint of the video media item corresponds to a second movement amount different from the first movement amount, the visual salience of the video media item relative to the boundary region is changed to a second relative visual salience level different from the first relative visual salience level.
174. A method, the method comprising: At the computer system that communicates with the display generation component and one or more cameras: When using one or more of the cameras to capture spatial video media of the environment The spatial video media includes a first video component corresponding to the right eye viewpoint and a second video component, different from the first video component, corresponding to the left eye viewpoint. When the first video component and the second video component are viewed simultaneously, they create the illusion of a spatial representation of the environment. Virtual indicator elements representing anchor positions corresponding to the viewpoints of the spatial video media are displayed in the environment via the display generation component, wherein the virtual indicator elements are displayed when the environment is visible via the display generation component. When the virtual indicator element is displayed in an environment visible via the display generation component, a first change in the viewpoint from which the spatial video media is captured is detected; and In response to detecting a first change in the viewpoint from which the spatial video media is captured, the appearance of the virtual indicator element is changed to indicate the corresponding viewpoint of the spatial video media.
175. The method of claim 174, wherein changing the appearance of the virtual indicator element to indicate the corresponding viewpoint corresponding to the spatial video media comprises changing the appearance of the virtual indicator element relative to the environment visible via the display generation component.
176. The method according to any one of claims 174 to 175, wherein: Displaying the virtual indicator element includes displaying the virtual indicator element at a virtual location in the environment; and The virtual indicator element is not included in the spatial video media.
177. The method of any one of claims 174 to 176, wherein changing the appearance of the virtual indicator element comprises moving the virtual indicator element to a corresponding position within an anchoring region of the environment, wherein the anchoring region is an environment-locked region that includes the anchoring position.
178. The method of claim 177, wherein moving the virtual indicator element to the corresponding position within the anchoring area of the environment comprises moving the virtual indicator element according to one or more simulated physical properties.
179. The method according to any one of claims 177 to 178, wherein: The first change in the viewpoint from which the spatial video media is captured includes a viewpoint movement of a first distance in a first direction; and Moving the virtual indicator element to the corresponding position within the anchored area of the environment includes moving the virtual indicator element a second distance in the first direction, wherein the second distance is shorter than the first distance.
180. The method according to any one of claims 174 to 179, wherein the virtual indicator element is environment-locked.
181. The method according to any one of claims 174 to 180, the method further comprising: When the virtual indicator element is displayed in an environment visible via the display generation component, a second change in the viewpoint from which the spatial video media is captured is detected; and In response to detecting a second change in the viewpoint from which the spatial video media is captured and based on determining that the viewpoint from which the spatial video media is captured satisfies a first set of one or more alignment criteria, a virtual alignment element is displayed via the display generation component when the environment is visible via the display generation component, wherein the virtual alignment element indicates the current position in the environment representing the viewpoint from which the spatial video media is captured; When the virtual alignment element is displayed in an environment visible via the display generation component, a third change in the viewpoint from which the spatial video media is captured is detected; and In response to detecting the third change in the viewpoint from which the spatial video media is captured, the virtual alignment element is moved to indicate the viewpoint, wherein the movement of the virtual alignment element is based on the third change in the viewpoint from which the spatial video media is captured.
182. The method of claim 181, wherein the first set of one or more alignment criteria includes a first criterion satisfied when the current position of the viewpoint in the environment representing the capture of the spatial video media therefrom is at least a first threshold distance from the anchor position.
183. The method according to claim 182, further comprising: When the virtual alignment element is displayed in an environment visible via the display generation component, a fourth change in the viewpoint from which the spatial video media is captured is detected; and In response to detecting the fourth change in the viewpoint from which the spatial video media is captured: The virtual alignment element is stopped from being displayed if the viewpoint from which the spatial video media is captured satisfies a second set of alignment criteria, wherein the second set of alignment criteria includes a second criterion that is satisfied when the current position of the viewpoint in the environment from which the spatial video media is captured is less than a second threshold distance from the anchor position, wherein the second threshold distance is less than the first threshold distance.
184. The method according to any one of claims 181 to 183, wherein: The third change in the viewpoint from which the spatial video media is captured includes a viewpoint movement of a third distance; and Moving the virtual alignment element based on the third change in the viewpoint from which it captures the spatial video media includes: The virtual alignment element is moved a fourth distance, wherein the fourth distance is shorter than the third distance, based on the determination that the third change from the viewpoint from which the spatial video media is captured satisfies a third set of one or more alignment criteria.
185. The method of claim 184, wherein moving the virtual alignment element based on the third change in the viewpoint from which the spatial video media is captured comprises moving the virtual alignment element according to one or more simulated physical properties.
186. The method according to claim 185, further comprising: After moving the virtual alignment element according to one or more simulated physical properties, the virtual alignment element is displayed at a first position, wherein the first position is closer to the anchor position than the current position in the environment representing the viewpoint from which the spatial video media is captured.
187. The method according to any one of claims 181 to 186, wherein displaying the virtual indicator element comprises displaying the virtual indicator element in a first color, and displaying the virtual alignment element comprises displaying the virtual alignment element in a second color different from the first color.
188. The method of any one of claims 181 to 187, wherein moving the virtual alignment element based on the third change in the viewpoint from which the spatial video media is captured comprises: When the virtual indicator element is displayed at the indicator location in the environment: The virtual alignment element is moved to the indicator position based on the determination that the viewpoint from which the spatial video media is captured satisfies a fourth set of alignment criteria, wherein the fourth set of alignment criteria includes criteria that are satisfied when the current position of the viewpoint representing the viewpoint from which the spatial video media is captured is less than a third threshold distance from the anchor position in the environment.
189. The method according to any one of claims 181 to 188, the method further comprising: When the virtual indicator element is displayed in an environment that is visible via the display generation component, a fifth change in the viewpoint from which the spatial video media is captured is detected; as well as In response to detecting a fifth change in the viewpoint from which the spatial video media is captured and based on determining that the viewpoint from which the spatial video media is captured satisfies a fifth set of alignment criteria, a virtual boundary element visually representing a predetermined threshold distance from the anchor position is displayed when the environment is visible via the display generation component, wherein the fifth set of alignment criteria includes criteria satisfied when the current position of the viewpoint representing the viewpoint from which the spatial video media is captured is at least a fourth threshold distance from the anchor position.
190. The method of claim 189, further comprising: When the virtual boundary element, visually representing the predetermined threshold distance from the anchor position, is displayed when the environment is visible via the display generation component, a sixth change in the viewpoint from which the spatial video media is captured is detected; and In response to the detection of the fifth change: Move the virtual indicator element along the first path; as well as Move the virtual boundary element along the first path.
191. The method according to any one of claims 189 to 190, the method further comprising: When the virtual alignment element and the virtual boundary element are displayed, a seventh change in the viewpoint from which the spatial video media is captured is detected; as well as In response to the detection of the seventh change: Based on the determination that the viewpoint satisfies one or more increasing criteria, wherein the one or more decreasing criteria include criteria satisfied when the distance between the current position of the viewpoint representing the spatial video media captured from it and the anchor position in the environment increases due to the seventh change of the viewpoint: Increase the opacity of the virtual alignment element at a first rate; and The opacity of the virtual boundary element is increased at a second rate.
192. The method of claim 191, wherein the first rate is different from the second rate.
193. The method according to any one of claims 189 to 192, the method further comprising: When the virtual boundary element, visually representing the predetermined threshold distance from the anchor position, is displayed when the environment is visible via the display generation component, an eighth change in the viewpoint from which the spatial video media is captured is detected; and In response to the eighth change of the viewpoint from which the spatial video media is captured and based on the determination that the distance between the current position of the viewpoint representing the spatial video media captured from it and the anchor position in the environment increases to a boundary distance range due to the eighth change of the viewpoint, the size of the virtual boundary element is increased from a first size to a second size, wherein the boundary distance range includes the predetermined threshold distance.
194. The method according to claim 193, further comprising: When the virtual boundary element is displayed at the second size, a ninth change in the viewpoint from which the spatial video media is captured is detected; as well as In response to the ninth change of the viewpoint from which the spatial video media is captured and based on the determination that the distance between the current position of the viewpoint representing the spatial video media captured from it and the anchor position in the environment has decreased to below the boundary distance range due to the ninth change of the viewpoint, the size of the virtual boundary element is reduced from the second size to the first size.
195. The method according to any one of claims 189 to 194, the method further comprising: When the virtual boundary element, visually representing the predetermined threshold distance from the anchor position, is displayed when the environment is visible via the display generation component, a tenth change in the viewpoint from which the spatial video media is captured is detected; and In response to the tenth change of the viewpoint from which the spatial video media is captured and based on the determination that the distance between the current position of the viewpoint representing the spatial video media captured from it and the anchor position in the environment exceeds the predetermined threshold distance due to the tenth change of the viewpoint, the size of the virtual boundary element is reduced.
196. The method according to any one of claims 189 to 195, the method further comprising: When the virtual boundary element, visually representing the predetermined threshold distance from the anchor position, is displayed when the environment is visible via the display generation component, an eleventh change in the viewpoint from which the spatial video media is captured is detected; and In response to the eleventh change of the viewpoint from which the spatial video media is captured and based on the determination that the distance between the current position of the viewpoint representing the spatial video media captured from it and the anchor position in the environment exceeds the predetermined threshold distance due to the eleventh change of the viewpoint, the display of the virtual boundary element is stopped.
197. The method according to any one of claims 181 to 196, wherein: Before detecting the third change in the viewpoint from which it captures the spatial video media, the virtual alignment element is displayed at a first display position relative to the viewpoint from which it captures the spatial video media; and Moving the virtual alignment element based on the third change in the viewpoint from which it captures the spatial video media includes: Based on the determination that the third change of the viewpoint satisfies one or more sets of movement criteria, the virtual alignment element is displayed at a second display position, different from the first display position, relative to the viewpoint from which the spatial video media is captured, wherein the one or more sets of movement criteria include criteria satisfied when the third change of the viewpoint from which the spatial video media is captured includes movement in at least one of a plurality of directions; and If the third change of the viewpoint does not satisfy the set of one or more movement criteria, the virtual alignment element is maintained at the first display position relative to the viewpoint from which the spatial video media is captured.
198. The method of any one of claims 174 to 197, wherein displaying the virtual indicator element comprises positioning the virtual indicator element within a virtual plane in the environment, wherein the virtual plane in the environment is spaced apart from the user by at least a threshold depth.
199. The method according to claim 198, further comprising: Detect the user's gaze; as well as The virtual plane is located based on the user's gaze.
200. The method of claim 199, wherein locating the virtual plane based on the user's gaze includes determining the convergence position of the user's gaze, wherein the virtual plane includes the convergence position.
201. The method of any one of claims 174 to 200, wherein changing the appearance of the virtual indicator element to indicate the corresponding viewpoint corresponding to the spatial video media comprises: If the distance between the current position of the viewpoint representing the capture of the spatial video media in the environment and the anchor position exceeds a second predetermined threshold distance, the display of the virtual indicator element is stopped.
202. The method of claim 201, wherein changing the appearance of the virtual indicator element to indicate the corresponding viewpoint corresponding to the spatial video media comprises: The opacity of the virtual indicator element is reduced as the distance between the current position of the viewpoint capturing the spatial video media in the environment and the anchor position increases to at least a fifth threshold distance, wherein the fifth threshold distance is less than the second predetermined threshold distance.
203. The method according to any one of claims 201 to 202, the method further comprising: After the virtual indicator element is stopped from being displayed and the viewpoint from which the spatial video media is captured is determined to satisfy one or more stability criteria, wherein the one or more stability criteria include criteria that are satisfied when the movement of the viewpoint from which the spatial video media is captured remains below a movement threshold for at least a threshold duration: A second virtual indicator element is displayed in the environment via the display generation component, representing a second anchor position corresponding to a second viewpoint of the spatial video media, wherein the second virtual indicator element is displayed when the environment is visible via the display generation component.
204. The method according to claim 203, further comprising: After stopping the display of the virtual indicator element and determining that the viewpoint from which the spatial video media is captured does not meet the set of one or more stability criteria, the display of the second virtual indicator element is abandoned.
205. The method according to any one of claims 201 to 204, the method further comprising: When the virtual indicator element is displayed in the environment that is visible through the display generation component, a second virtual alignment element is displayed through the display generation component in the same manner. as well as After stopping the display of the virtual indicator element, the second virtual alignment element continues to be displayed.
206. The method of any one of claims 174 to 205, wherein changing the appearance of the virtual indicator element to indicate the corresponding viewpoint corresponding to the spatial video media comprises: The virtual indicator element is stopped from being displayed if the distance between the current position of the viewpoint that captures the spatial video media in the environment and the anchor position decreases below a holding threshold distance.
207. The method of claim 206, wherein changing the appearance of the virtual indicator element to indicate the corresponding viewpoint corresponding to the spatial video media comprises: The opacity of the virtual indicator element is reduced as the distance between the current position of the viewpoint that captures the spatial video media in the environment and the anchor position decreases to below a sixth threshold distance, wherein the sixth threshold distance is greater than the hold threshold distance.
208. The method according to any one of claims 174 to 207, the method further comprising: When the spatial video media of the environment is captured and the viewpoint from which the spatial video media is captured satisfies the sixth set of alignment criteria, a graphic alignment indication is displayed.
209. The method according to any one of claims 174 to 208, the method further comprising: After capturing the spatial video media of the environment, a playable representation of the spatial video media is displayed.
210. The method according to claim 209, further comprising: When the playable representation of the spatial video media is displayed: Based on the playable representation that a corresponding playback mode can be used to play the spatial video media, an indication of the corresponding playback mode is displayed in a first appearance; and Based on the playable indication that the corresponding playback mode is not available for playing the spatial video media, the instruction to display the corresponding playback mode in the first appearance is abandoned.
211. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs including instructions for performing the method according to any one of claims 174 to 210.
212. A computer system configured to communicate with a display generation component and one or more cameras, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 174 to 210.
213. A computer system configured to communicate with a display generation component and one or more cameras, the computer system comprising: Apparatus for performing the method according to any one of claims 174 to 210.
214. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component and one or more cameras, the one or more programs comprising instructions for performing the method according to any one of claims 174 to 210.
215. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs including instructions for: When the one or more cameras are used to capture spatial video media of an environment, the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint. When the first video component and the second video component are viewed simultaneously, they form an illusion of a spatial representation of the environment. Virtual indicator elements representing anchor positions corresponding to the respective viewpoints of the spatial video media are displayed in the environment via the display generation component, wherein the virtual indicator elements are displayed when the environment is visible via the display generation component. When the virtual indicator element is displayed in an environment visible via the display generation component, a first change in the viewpoint from which the spatial video media is captured is detected; and In response to detecting a first change in the viewpoint from which the spatial video media is captured, the appearance of the virtual indicator element is changed to indicate the corresponding viewpoint of the spatial video media.
216. A computer system configured to communicate with a display generation component and one or more cameras, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When using one or more of the cameras to capture spatial video media of the environment The spatial video media includes a first video component corresponding to the right eye viewpoint and a second video component, different from the first video component, corresponding to the left eye viewpoint. When the first video component and the second video component are viewed simultaneously, they create the illusion of a spatial representation of the environment. Virtual indicator elements representing anchor positions corresponding to the viewpoints of the spatial video media are displayed in the environment via the display generation component, wherein the virtual indicator elements are displayed when the environment is visible via the display generation component. When the virtual indicator element is displayed in an environment visible via the display generation component, a first change in the viewpoint from which the spatial video media is captured is detected; and In response to detecting a first change in the viewpoint from which the spatial video media is captured, the appearance of the virtual indicator element is changed to indicate the corresponding viewpoint of the spatial video media.
217. A computer system configured to communicate with a display generation component and one or more cameras, the computer system comprising: An apparatus for performing the following operations: when capturing spatial video media of an environment using one or more cameras, wherein the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint, the first video component and the second video component forming an illusion of spatial representation of the environment when viewed simultaneously, and displaying virtual indicator elements in the environment representing anchor positions corresponding to the respective viewpoints of the spatial video media via the display generation component, wherein the virtual indicator elements are displayed when the environment is visible via the display generation component; A device for performing the following operation: when the virtual indicator element is displayed in an environment visible via the display generating component, detecting a first change in the viewpoint from which the spatial video media is captured; as well as A means for performing the following operation: in response to detecting a first change in the viewpoint from which the spatial video media is captured, changing the appearance of the virtual indicator element to indicate the corresponding viewpoint corresponding to the spatial video media.
218. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generating component and one or more cameras, the one or more programs comprising instructions for performing the following operations: When the one or more cameras are used to capture spatial video media of an environment, the spatial video media includes a first video component corresponding to a right-eye viewpoint and a second video component different from the first video component corresponding to a left-eye viewpoint. When the first video component and the second video component are viewed simultaneously, they form an illusion of a spatial representation of the environment. Virtual indicator elements representing anchor positions corresponding to the respective viewpoints of the spatial video media are displayed in the environment via the display generation component, wherein the virtual indicator elements are displayed when the environment is visible via the display generation component. When the virtual indicator element is displayed in an environment visible via the display generation component, a first change in the viewpoint from which the spatial video media is captured is detected; and In response to detecting a first change in the viewpoint from which the spatial video media is captured, the appearance of the virtual indicator element is changed to indicate the corresponding viewpoint of the spatial video media.
219. A method, the method comprising: At the computer system communicating with the display generation component: When a representation of a spatial media item is displayed via the display generation component, the spatial media item includes a first component corresponding to the right eye viewpoint and a second component different from the first component corresponding to the left eye viewpoint. The first component and the second component, when viewed simultaneously, create the illusion of a spatial representation. Based on the determination that the spatial media item meets one or more stability criteria, a spatial viewing indicator with a first appearance is simultaneously displayed along with the representation of the spatial media item; and If it is determined that the spatial media item does not meet the set of one or more stability criteria, the simultaneous display of the spatial viewing indicator with the first appearance and the representation of the spatial media item is abandoned.
220. The method according to claim 219, further comprising: Before displaying the representation of the spatial media item, a request to display the representation of the spatial media item is received; as well as In response to receiving the request to display the representation of the spatial media item, the representation of the spatial media item is displayed.
221. The method of any one of claims 219 to 220, wherein when the spatial viewing indicator is selected, the computer system initiates the provision of an extended representation of the spatial media item, wherein the size of the extended representation of the spatial media item exceeds the size of the representation of the spatial media item.
222. The method according to claim 221, further comprising: When the representation of the spatial media item is displayed, a first three-dimensional effect is provided to the spatial media item; as well as When the extended representation of the spatial media item is displayed, a second three-dimensional effect is provided to the spatial media item, wherein the second three-dimensional effect is different from the first three-dimensional effect.
223. The method according to any one of claims 221 to 222, further comprising: Detect the input of the selected space viewing indicator; as well as In response to detecting the input that selects the space viewing indicator, an extended representation for providing the space media item is initiated, wherein initiating the extended representation for providing the space media item includes: Based on the determination that the spatial media item does not meet the set of one or more stability criteria, an extended notification interface is displayed via the display generation component, wherein the extended notification interface indicates that the spatial media item does not meet the set of one or more stability criteria. as well as Based on the determination that the spatial media item satisfies one or more of the following stability criteria: The extended representation of the spatial media item is displayed via the display generation component; and Disallow the display of the extended notification interface.
224. The method of claim 223, wherein the extended notification interface comprises: An optional confirmation object, when selected, causes the computer system to initiate the display of the extended representation of the spatial media item; as well as An optional cancellation object, when selected, causes the computer system to abandon the display of the extended representation of the spatial media item.
225. The method according to any one of claims 219 to 224, the method further comprising: If it is determined that the spatial media item does not meet the set of one or more stability criteria, the spatial viewing indicator with a second appearance different from the first appearance will be displayed simultaneously with the representation of the spatial media item.
226. The method of claim 225, wherein: Displaying the space viewing indicator having the first appearance includes displaying the space viewing indicator at a first contrast level; and Displaying the space viewing indicator having the second appearance includes displaying the space viewing indicator at a second contrast level that is lower than the first contrast level.
227. The method according to any one of claims 225 to 226, wherein: Displaying a space viewing indicator with the second appearance includes displaying the space viewing indicator with a warning icon; and Displaying the space viewing indicator with the first appearance includes abandoning the display of the space viewing indicator with the warning icon.
228. The method of any one of claims 219 to 224, wherein abandoning the display of the space viewing indicator having the first appearance includes abandoning the simultaneous display of the space viewing indicator and the representation of the space media item.
229. The method according to claim 228, further comprising: After abandoning the simultaneous display of the spatial view indicator and the representation of the spatial media item, and when displaying the representation of the spatial media item without the spatial view indicator, an optional menu user interface object is displayed via the display generation component; Detect input that selects the optional menu user interface object; as well as In response to detecting a selection of the optional menu user interface object, an option menu interface is displayed via the display generation component, wherein displaying the option menu interface includes displaying a second space viewing indicator.
230. The method according to any one of claims 219 to 229, the method further comprising: Detect requests to edit the spatial media item; In response to the detection of the request to edit the spatial media item, a modified version of the spatial media item is generated; as well as When the modified representation of the spatial media item is displayed via the display generation component: Based on the determination that the modified version of the spatial media item meets the set of one or more stability criteria, the spatial viewing indicator having the first appearance and the representation of the modified version of the spatial media item are displayed simultaneously; as well as If it is determined that the modified version of the spatial media item does not meet the set of one or more stability criteria, the simultaneous display of the spatial viewing indicator with the first appearance and the representation of the modified version of the spatial media item is abandoned.
231. The method according to claim 230, wherein: The request to edit the spatial media item includes a request to remove a portion of the spatial media item; and The modified version of the spatial media item does not include the portion of the spatial media item.
232. The method according to any one of claims 219 to 231, the method further comprising: When the representation of the spatial media item is displayed via the display generation component, an input requesting playback of the spatial media item is detected.
233. The method according to claim 232, further comprising: In response to the input that detects a request to play the spatial media item: If the spatial media item is determined to not meet the set of one or more stability criteria, a playback notification interface is displayed, wherein the playback notification interface indicates that the spatial media item does not meet the set of one or more stability criteria. as well as Based on the determination that the spatial media item satisfies one or more of the following stability criteria: Play the aforementioned spatial media item; and Disallow the display of the playback notification interface.
234. The method according to any one of claims 232 to 233, further comprising: When the spatial media item is played, the playback of the spatial media item is modified to reduce the occurrence of movement of the viewpoint corresponding to the spatial media item when the spatial media item is captured.
235. The method according to any one of claims 219 to 234, the method further comprising: When the representation of the spatial media item is displayed, a detection request is made to view the input of a second spatial media item that is different from the spatial media item. as well as In response to the input that detects a request to view the second spatial media item: Stop displaying the representation of the spatial media item; and The representation of the second spatial media item is displayed via the display generation component.
236. The method according to any one of claims 219 to 235, wherein the set of one or more stability criteria includes criteria that are satisfied when the movement of the viewpoint corresponding to the spatial media item during capture does not exceed a threshold displacement from the anchor position representing the target viewpoint.
237. The method according to any one of claims 219 to 236, wherein the set of one or more stability criteria includes criteria satisfied when the movement of the viewpoint corresponding to the spatial media item during capture does not exceed a threshold rate of change.
238. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for performing the method according to any one of claims 219 to 237.
239. A computer system configured to communicate with a display generation component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 219 to 237.
240. A computer system configured to communicate with a display generation component, the computer system comprising: Apparatus for performing the method according to any one of claims 219 to 237.
241. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising instructions for performing the method according to any one of claims 219 to 237.
242. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs including instructions for: When a representation of a spatial media item is displayed via the display generation component, the spatial media item includes a first component corresponding to the right eye viewpoint and a second component different from the first component corresponding to the left eye viewpoint. The first component and the second component, when viewed simultaneously, create the illusion of a spatial representation. Based on the determination that the spatial media item meets one or more stability criteria, a spatial viewing indicator with a first appearance is simultaneously displayed along with the representation of the spatial media item; and If it is determined that the spatial media item does not meet the set of one or more stability criteria, the simultaneous display of the spatial viewing indicator with the first appearance and the representation of the spatial media item is abandoned.
243. A computer system configured to communicate with a display generation component, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When a representation of a spatial media item is displayed via the display generation component, the spatial media item includes a first component corresponding to the right eye viewpoint and a second component different from the first component corresponding to the left eye viewpoint. The first component and the second component, when viewed simultaneously, create the illusion of a spatial representation. Based on the determination that the spatial media item meets one or more stability criteria, a spatial viewing indicator with a first appearance is simultaneously displayed along with the representation of the spatial media item; and If it is determined that the spatial media item does not meet the set of one or more stability criteria, the simultaneous display of the spatial viewing indicator with the first appearance and the representation of the spatial media item is abandoned.
244. A computer system configured to communicate with a display generation component, the computer system comprising: An apparatus for performing the following operation: when a representation of a spatial media item is displayed via the display generating component, wherein the spatial media item includes a first component corresponding to a right-eye viewpoint and a second component different from the first component corresponding to a left-eye viewpoint, the first component and the second component forming an illusion of spatial representation when viewed simultaneously: Based on the determination that the spatial media item meets one or more stability criteria, a spatial viewing indicator with a first appearance is simultaneously displayed along with the representation of the spatial media item; and If it is determined that the spatial media item does not meet the set of one or more stability criteria, the simultaneous display of the spatial viewing indicator with the first appearance and the representation of the spatial media item is abandoned.
245. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a display generation component, the one or more programs comprising instructions for: When a representation of a spatial media item is displayed via the display generation component, the spatial media item includes a first component corresponding to the right eye viewpoint and a second component different from the first component corresponding to the left eye viewpoint. The first component and the second component, when viewed simultaneously, create the illusion of a spatial representation. Based on the determination that the spatial media item meets one or more stability criteria, a spatial viewing indicator with a first appearance is simultaneously displayed along with the representation of the spatial media item; and If it is determined that the spatial media item does not meet the set of one or more stability criteria, the simultaneous display of the spatial viewing indicator with the first appearance and the representation of the spatial media item is abandoned.
246. A method comprising: At a computer system that communicates with one or more display generation components and one or more cameras: When the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media includes a first visual component corresponding to the right eye viewpoint and a second visual component different from the first visual component and corresponding to the left eye viewpoint. When the first visual component and the second visual component are viewed simultaneously, they form the illusion of a spatial representation of the captured visual content. And when the computer system is not capturing spatial media: If the orientation of the computer system is determined to be outside the threshold orientation range, a first prompt is output to the user, suggesting that the computer system be rotated into the threshold orientation range.
247. The method according to claim 246, further comprising: When the computer system is in the spatial media capture mode and when the computer system is not capturing spatial media: Based on the determination that the orientation of the computer system is within the threshold orientation range, the first prompt is abandoned.
248. The method of any one of claims 246 to 247, wherein outputting the first prompt comprises displaying a first visual prompt via the one or more display generating components prompting the user to rotate the computer system into the threshold orientation range.
249. The method of any one of claims 246 to 248, wherein outputting the first prompt includes outputting a first audio prompt prompting the user to rotate the computer system within the threshold orientation range.
250. The method according to any one of claims 246 to 249, the method further comprising: When the computer system is in the spatial media capture mode and when the computer system is not capturing spatial media: Based on the determination that the orientation of the computer system is outside the threshold orientation range, the capture of space media is disabled; as well as When capture of spatial media is disabled, the computer system detects that it has shifted from outside the threshold orientation range to within the threshold orientation range. as well as In response to the detection that the computer system is within the threshold orientation range, capture of space media is enabled.
251. The method of any one of claims 246 to 250, wherein determining that the orientation of the computer system is outside the threshold orientation range comprises determining that a first axis defined by a first camera of the one or more cameras and a second camera of the one or more cameras other than the first camera is rotated relative to a target axis by a number of degrees greater than a threshold.
252. The method according to claim 251, wherein: The computer system includes a first side and a second side perpendicular to the first side; The first side defines the axis of the first device; The second side defines the axis of the second device; The first side is longer than the second side; and The first axis corresponds to the first device axis.
253. The method according to any one of claims 246 to 252, the method further comprising: When the computer system is in the spatial media capture mode and when the computer system is not capturing spatial media: The space photograph capture capability representation, which can be selected to capture space photographic media, is displayed via the one or more display generation components; When the spatial photo capture capability representation is displayed, one or more user inputs corresponding to the selection of the spatial photo capture capability representation are received; as well as In response to receiving one or more user inputs corresponding to a selection of the spatial photograph capture capability representation, a first spatial photograph is captured.
254. The method according to any one of claims 246 to 252, the method further comprising: When the computer system is in the spatial media capture mode and when the computer system is not capturing spatial media: The spatial video capture capability representation, which can be selected to capture spatial video media, is displayed via the one or more display generation components. When the spatial video capture capability representation is displayed, one or more user inputs corresponding to the selection of the spatial video capture capability representation are received; as well as In response to receiving one or more user inputs corresponding to a selection of the spatial video capture capability representation, a first spatial video is captured.
255. The method according to claim 254, further comprising: When the spatial video capture capability representation is displayed, a spatial photo capture capability representation that can be selected to capture spatial photo media is displayed via the one or more display generation components; When the spatial photo capture capability representation is displayed, one or more user inputs corresponding to the selection of the spatial photo capture capability representation are received; as well as In response to receiving one or more user inputs corresponding to a selection of the spatial photograph capture capability representation, a first spatial photograph is captured.
256. The method according to claim 255, further comprising: When capturing second spatial video, the spatial photo capture capability representation is displayed via the one or more display generation components; When the second spatial video is captured and the spatial photo capture capability representation is displayed, a second set of one or more user inputs corresponding to the selection of the spatial photo capture capability representation are received; as well as In response to receiving one or more user inputs from the second group corresponding to a selection of the spatial photograph capture capability representation, a second spatial photograph is captured.
257. The method according to any one of claims 246 to 256, the method further comprising: When the computer system is in the spatial media capture mode: Based on the determination that a first error criterion corresponding to a first error is met, wherein the first error is related to the capture of spatial media, an error correction prompt is output prompting the user to perform one or more actions to correct the first error.
258. The method according to claim 257, wherein: Determining that the first error criterion is met includes determining that the computer system detects light intensity less than a threshold; and The output of the error correction prompt includes outputting a notification to the user that the computer system has detected a light amount less than the threshold.
259. The method according to any one of claims 257 to 258, wherein: Determining that the first error criterion is met includes determining that at least a portion of the computer system is less than a threshold distance from one or more detected individuals captured by at least some of the one or more cameras; and The output of the error correction prompt includes outputting a distance correction prompt to the user, notifying the computer system that the distance to one or more detected individuals captured by at least some of the one or more cameras is less than the threshold distance.
260. The method according to any one of claims 246 to 259, wherein the computer system is a head-mounted device configured to be worn on a user's head.
261. The method according to any one of claims 246 to 260, wherein the computer system is not a head-mounted device.
262. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more cameras, the one or more programs including instructions for performing the method according to any one of claims 246 to 261.
263. A computer system configured to communicate with one or more display generating components and one or more cameras, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 246 to 261.
264. A computer system configured to communicate with one or more display generating components and one or more cameras, the computer system comprising: Apparatus for performing the method according to any one of claims 246 to 261.
265. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more cameras, the one or more programs comprising instructions for performing the method according to any one of claims 246 to 261.
266. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more cameras, the one or more programs including instructions for: When the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media includes a first visual component corresponding to the right eye viewpoint and a second visual component different from the first visual component and corresponding to the left eye viewpoint. When the first visual component and the second visual component are viewed simultaneously, they form the illusion of a spatial representation of the captured visual content. And when the computer system is not capturing spatial media: If the orientation of the computer system is determined to be outside the threshold orientation range, a first prompt is output to the user, suggesting that the computer system be rotated into the threshold orientation range.
267. A computer system configured to communicate with one or more display generating components and one or more cameras, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: When the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media includes a first visual component corresponding to the right eye viewpoint and a second visual component different from the first visual component and corresponding to the left eye viewpoint. When the first visual component and the second visual component are viewed simultaneously, they form the illusion of a spatial representation of the captured visual content. And when the computer system is not capturing spatial media: If the orientation of the computer system is determined to be outside the threshold orientation range, a first prompt is output to the user, suggesting that the computer system be rotated into the threshold orientation range.
268. A computer system configured to communicate with one or more display generation components and one or more cameras, the computer system comprising: A device for performing the following operations: when the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first visual component and the second visual component forming an illusion of a spatial representation of captured visual content when viewed simultaneously, and when the computer system is not capturing spatial media: If the orientation of the computer system is determined to be outside the threshold orientation range, a first prompt is output to the user, suggesting that the computer system be rotated into the threshold orientation range.
269. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more cameras, the one or more programs comprising instructions for: When the computer system is in a spatial media capture mode corresponding to spatial media, the spatial media includes a first visual component corresponding to the right eye viewpoint and a second visual component different from the first visual component and corresponding to the left eye viewpoint. When the first visual component and the second visual component are viewed simultaneously, they form the illusion of a spatial representation of the captured visual content. And when the computer system is not capturing spatial media: If the orientation of the computer system is determined to be outside the threshold orientation range, a first prompt is output to the user, suggesting that the computer system be rotated into the threshold orientation range.
270. A method comprising: At a computer system that communicates with one or more display generation components and one or more cameras: A first user interface corresponding to the camera application of the computer system is displayed via the one or more display generation components; as well as When the first user interface corresponding to the camera application of the computer system is displayed: Based on the determination that the computer system is associated with a head-mounted device separate from the computer system, a spatial media capture mode option is provided corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to the right eye viewpoint and a second visual component different from the first visual component and corresponding to the left eye viewpoint, the first visual component and the second visual component forming an illusion of spatial representation of the captured visual content when viewed simultaneously; as well as Based on the determination that the computer system is not associated with a head-mounted device separate from the computer system, the option to provide the spatial media capture mode is abandoned.
271. The method according to claim 270, further comprising: When the first user interface corresponding to the camera application of the computer system is displayed, one or more user inputs corresponding to a user request to change the current media capture mode are received; as well as In response to receiving one or more user inputs corresponding to the user request to change the current media capture mode: Based on the determination that the spatial media capture mode option is enabled, switch from the current media capture mode to the spatial media capture mode; as well as If the spatial media capture mode option is determined to be disabled, the switch from the current media capture mode to the spatial media capture mode is abandoned.
272. The method according to claim 271, further comprising: In response to receiving one or more user inputs corresponding to the user request to change the current media capture mode: If the spatial media capture mode option is determined to be disabled, the system switches from the current media capture mode to a second media capture mode that is different from both the current media capture mode and the spatial media capture mode.
273. The method of any one of claims 271 to 272, wherein the one or more user inputs corresponding to the user request to change the current media capture mode include one or more tap inputs.
274. The method of any one of claims 271 to 273, wherein the one or more user inputs corresponding to the user request to change the current media capture mode include one or more swipe inputs.
275. The method according to claim 274, wherein: The camera application includes multiple media capture modes, including the spatial media capture mode and the non-spatial still image capture mode. The multiple media capture modes are arranged in a defined order; and The spatial media capture mode is adjacent to the non-spatial still image capture mode in the defined order.
276. The method of any one of claims 270 to 275, wherein determining that the computer system is associated with the head-mounted device includes determining that both the computer system and the head-mounted device are associated with the same user account.
277. The method according to any one of claims 270 to 276, the method further comprising: When the computer system is not associated with a head-mounted device that is separate from the computer system: Receive one or more user inputs corresponding to a user request to enable the spatial media capture mode option; In response to receiving one or more user inputs corresponding to the user request to enable the spatial media capture mode option, the spatial media capture mode option is enabled; The first user interface corresponding to the camera application of the computer system is displayed via the one or more display generation components; as well as When the first user interface corresponding to the camera application of the computer system is displayed: Based on the determination that the spatial media capture mode option is enabled, the spatial media capture mode option is provided.
278. The method according to any one of claims 270 to 277, wherein: The camera application includes multiple media capture modes, including the spatial media capture mode and a first media capture mode different from the spatial media capture mode; The first user interface corresponds to the first media capture mode; and The method further includes: When the first user interface corresponding to the first media capture mode is displayed, and when the computer system is in the first media capture mode, one or more user inputs corresponding to a user request to switch from the first media capture mode to the spatial media capture mode are received; In response to receiving one or more user inputs corresponding to the user request to switch from the first media capture mode to the spatial media capture mode: The transition from the first media capture mode to the spatial media capture mode; and A spatial media capture user interface, different from the first user interface and corresponding to the spatial media capture mode, is displayed via the one or more display generation components, wherein: The spatial media capture user interface includes camera-level controls that indicate the orientation of the computer system; and The first user interface does not include the camera-level controls.
279. The method according to any one of claims 270 to 278, wherein: The camera application includes multiple media capture modes, including the spatial media capture mode and a first media capture mode different from the spatial media capture mode; When the computer system is in the first media capture mode, one or more camera options are available; and When the computer system is in the spatial media capture mode, the one or more camera options are unavailable.
280. The method of claim 279, wherein the one or more camera options include a camera flip option for switching from a first set of cameras in the one or more cameras to a second set of cameras in the one or more cameras, wherein the second set of cameras is different from the first set of cameras.
281. The method of any one of claims 279 to 280, wherein the one or more camera options include a flash option for selectively enabling and / or disabling flash features.
282. The method of any one of claims 279 to 281, wherein the one or more camera options include a frame rate option for modifying the media capture frame rate.
283. The method of any one of claims 279 to 282, wherein the one or more camera options include scaling options for modifying the media capture scaling level.
284. The method according to any one of claims 270 to 283, the method further comprising: A media library user interface, comprising representations of multiple media items, is displayed via the one or more display generation components, including displaying a representation of a first corresponding media item within the media library user interface, wherein: Based on determining that the first corresponding media item is spatial media, the representation of the first corresponding media item is displayed in a first manner; and Based on the determination that the first corresponding media item is not spatial media, the representation of the first corresponding media item is displayed in a second manner different from the first manner.
285. The method according to claim 284, wherein: The representation of the first corresponding media item displayed in the first manner includes displaying the representation of the first corresponding media item having a first visual badge indicating that the first corresponding media item is spatial media; and The representation of displaying the first corresponding media item in the second manner includes displaying the representation of the first corresponding media item without the first visual badge.
286. The method according to any one of claims 284 to 285, wherein: The representation of the first corresponding media item displayed in the first manner includes displaying the representation of the first corresponding media item in a first media set indicating that the first corresponding media item is spatial media, wherein the first media set does not include non-spatial media items; and The representation of the first corresponding media item displayed in the second manner includes displaying the representation of the first corresponding media item in a second media set different from the first media set, the second media set including non-spatial media items not included in the first media set.
287. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more cameras, the one or more programs including instructions for performing the method according to any one of claims 270 to 286.
288. A computer system configured to communicate with one or more display generation components and one or more cameras, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 270 to 286.
289. A computer system configured to communicate with one or more display generation components and one or more cameras, the computer system comprising: Apparatus for performing the method according to any one of claims 270 to 286.
290. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more cameras, the one or more programs comprising instructions for performing the method according to any one of claims 270 to 286.
291. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more cameras, the one or more programs including instructions for: A first user interface corresponding to the camera application of the computer system is displayed via the one or more display generation components; and When the first user interface corresponding to the camera application of the computer system is displayed: Based on the determination that the computer system is associated with a head-mounted device separate from the computer system, a spatial media capture mode option is provided corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first visual component and the second visual component forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and Based on the determination that the computer system is not associated with a head-mounted device separate from the computer system, the option to provide the spatial media capture mode is abandoned.
292. A computer system configured to communicate with one or more display generating components and one or more cameras, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: A first user interface corresponding to the camera application of the computer system is displayed via the one or more display generation components; and When the first user interface corresponding to the camera application of the computer system is displayed: Based on the determination that the computer system is associated with a head-mounted device separate from the computer system, a spatial media capture mode option is provided corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to the right eye viewpoint and a second visual component different from the first visual component and corresponding to the left eye viewpoint, the first visual component and the second visual component forming an illusion of spatial representation of the captured visual content when viewed simultaneously; as well as Based on the determination that the computer system is not associated with a head-mounted device separate from the computer system, the option to provide the spatial media capture mode is abandoned.
293. A computer system configured to communicate with one or more display generating components and one or more cameras, the computer system comprising: A device for performing the following operations: displaying a first user interface corresponding to a camera application of the computer system via the one or more display generating components; as well as A device for performing the following operation: when displaying the first user interface corresponding to the camera application of the computer system: Based on the determination that the computer system is associated with a head-mounted device separate from the computer system, a spatial media capture mode option is provided corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to the right eye viewpoint and a second visual component different from the first visual component and corresponding to the left eye viewpoint, the first visual component and the second visual component forming an illusion of spatial representation of the captured visual content when viewed simultaneously; as well as Based on the determination that the computer system is not associated with a head-mounted device separate from the computer system, the option to provide the spatial media capture mode is abandoned.
294. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more cameras, the one or more programs comprising instructions for: A first user interface corresponding to the camera application of the computer system is displayed via the one or more display generation components; and When the first user interface corresponding to the camera application of the computer system is displayed: Based on the determination that the computer system is associated with a head-mounted device separate from the computer system, a spatial media capture mode option is provided corresponding to a spatial media capture mode for capturing spatial media, the spatial media including a first visual component corresponding to a right-eye viewpoint and a second visual component different from the first visual component and corresponding to a left-eye viewpoint, the first visual component and the second visual component forming an illusion of a spatial representation of the captured visual content when viewed simultaneously; and Based on the determination that the computer system is not associated with a head-mounted device separate from the computer system, the option to provide the spatial media capture mode is abandoned.