Devices, methods, and graphical user interfaces for interacting with augmented reality experience
By communicating with the display generation components and input devices in a computer system, the representation of multiple augmented reality experiences is realized simultaneously in a three-dimensional environment, and the representations are selectively displayed or hidden according to user input, the problems of interaction complexity and inefficiency in the prior art are solved, and the efficiency and intuitiveness of user interaction are improved.
Patent Information
- Application Number
- CN202510584157.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-15
- Filing Date
- 2023-09-21
- Publication Date
- 2025-06-20
AI Technical Summary
The methods and interfaces for interacting with extended real-life experience in the prior art have problems such as insufficient feedback, complex operations, cumbersome and error-prone, resulting in a large cognitive burden on users and low interaction efficiency.
By communicating with the display generation components and input devices in a computer system, a representation of multiple augmented reality experiences simultaneously is realized in a three-dimensional environment, and selectively display or hide these representations according to user input, simplifying the user interaction process.
It improves the efficiency and intuitiveness of users' interaction with computer systems, reduces the number and complexity of user input, and enhances the immersion and interactivity of the extended reality experience.
Smart Images

Figure CN120179076A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application for invention titled "Devices, Methods, and Graphical User Interfaces for Interacting with Extended Reality Experiences" with the application date of September 21, 2023, application number 202380067076.3. Technical Field
[0002] This application claims priority to the following patent applications: U.S. Patent Application No. 18 / 369,075, titled "Devices, Methods, and Graphical User Interfaces for Interacting with Extended Reality Experiences," filed on September 15, 2023; U.S. Provisional Patent Application No. 63 / 538,453, titled "Devices, Methods, and Graphical User Interfaces for Interacting with Extended Reality Experiences," filed on September 14, 2023; and U.S. Provisional Patent Application No. 63 / 409,184, titled "Devices, Methods, and Graphical User Interfaces for Interacting with Extended Reality Experiences," filed on September 22, 2022. The entire content of each of these patent applications is incorporated herein by reference in its entirety. Technical Field
[0004] The present disclosure generally relates to computer systems that communicate with one or more display generation components and one or more input devices to provide computer-generated experiences, including but not limited to electronic devices that provide virtual reality experiences and mixed reality experiences via a display. Background Art
[0005] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or augment the physical world. Input devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) for computer systems and other electronic computing devices are used to interact with the virtual / augmented reality environment. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). SUMMARY OF THE INVENTION
[0006] Some methods and interfaces for interacting with an environment that includes at least some virtual elements (e.g., an application, an augmented reality environment, a mixed reality environment, and a virtual reality environment) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems in which virtual object manipulation is complex, tedious, and error-prone impose a significant cognitive burden on the user and detract from the experience of the virtual / augmented reality environment. In addition, these methods take longer than necessary, thereby wasting the energy of the computer system. This latter consideration is particularly important in battery-powered devices.
[0007] Accordingly, there is a need for computer systems with improved methods and interfaces to provide computer-generated experiences (such as, for example, an extended reality experience) to a user such that the interaction of the user with the computer system is more effective and intuitive for the user. Such methods and interfaces optionally supplement or replace conventional methods for providing an extended reality experience to a user. Such methods and interfaces form a more effective human-machine interface by helping the user understand the connection between the inputs provided and the device's response to those inputs, thereby reducing the quantity, degree, and / or nature of the inputs from the user.
[0008] The above-mentioned deficiencies and other problems associated with the user interface of a computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device such as a watch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to a display generation component, the computer system also has one or more output devices, which include one or more haptic output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, a program or set of instructions stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movements of the user's eyes and hands in space relative to the GUI (and / or the computer system) or the user's body (as captured by cameras and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed through the interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, playing games, making and receiving phone calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured to be executed by one or more processors.
[0009] There is a need for electronic devices with improved methods and interfaces for interacting with extended reality experiences. Such methods and interfaces can supplement or replace conventional methods for interacting with extended reality experiences. Such methods and interfaces reduce the amount, degree, and / or nature of input from the user and result in a more efficient human-machine interface. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges.
[0010] According to some embodiments, a method is described. The method includes: at a computer system in communication with one or more display generation components and one or more input devices: simultaneously display, via the one or more display generation components, representations of a plurality of augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receive a first user input via the one or more input devices while the representations of the plurality of augmented reality experiences are simultaneously displayed in the three-dimensional environment; and in response to receiving the first user input: stop displaying the representations of one or more of the plurality of augmented reality experiences; and display the first augmented reality experience in the three-dimensional environment via the one or more display generation components based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0011] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: simultaneously display, via the one or more display generation components, representations of a plurality of augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receive a first user input via the one or more input devices while the representations of the plurality of augmented reality experiences are simultaneously displayed in the three-dimensional environment; and in response to receiving the first user input: stop displaying the representations of one or more of the plurality of augmented reality experiences; and display the first augmented reality experience in the three-dimensional environment via the one or more display generation components based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0012] According to some embodiments, a transient computer-readable storage medium is described. In some embodiments, the transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receiving, via the one or more input devices, a first user input while the representations of the multiple augmented reality experiences are simultaneously displayed in the three-dimensional environment; and in response to receiving the first user input: stopping the display of the representations of one or more of the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0013] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receiving, via the one or more input devices, a first user input while the representations of the multiple augmented reality experiences are simultaneously displayed in the three-dimensional environment; and in response to receiving the first user input: stopping the display of the representations of one or more of the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0014] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: means for simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; means for receiving, via the one or more input devices, a first user input when the representations of the multiple augmented reality experiences are simultaneously displayed in the three-dimensional environment; and means for performing the following operations in response to receiving the first user input: stopping the display of the representations of one or more of the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0015] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generation components and one or more input devices, the one or more programs including instructions for: simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receiving, via the one or more input devices, a first user input when the representations of the multiple augmented reality experiences are simultaneously displayed in the three-dimensional environment; and in response to receiving the first user input: stopping the display of the representations of one or more of the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0016] According to some embodiments, a method is described. The method includes: at a computer system in communication with one or more display generation components and one or more input devices: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0017] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0018] According to some embodiments, a transitory computer-readable storage medium is described. In some embodiments, the transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0019] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0020] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: means for receiving a first sequence of one or more user inputs via a first physical control; and means for performing the following operations in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0021] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generation components and one or more input devices, the one or more programs including instructions for the following operations: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0022] According to some embodiments, a method is described. The method includes: at a computer system in communication with one or more display generation components and one or more input devices: when a view of a three-dimensional environment in which the computer system is located is visible, detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment; and in response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: simultaneously displaying, via the one or more display generation components and with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences that can be displayed by the computer system.
[0023] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: when a view of a three-dimensional environment in which the computer system is located is visible, detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment; and in response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: simultaneously displaying, via the one or more display generation components and with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences that can be displayed by the computer system.
[0024] According to some embodiments, a transitory computer-readable storage medium is described. In some embodiments, the transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: when a view of a three-dimensional environment in which the computer system is located is visible, detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment; and in response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: simultaneously displaying, via the one or more display generation components and with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences that can be displayed by the computer system.
[0025] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: when a view of a three-dimensional environment in which the computer system is located is visible, detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment; and in response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: simultaneously displaying, via the one or more display generation components and with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences that can be displayed by the computer system.
[0026] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: means for detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment when a view of the three-dimensional environment in which the computer system is located is visible; and means for performing the following operations in response to detecting the first set of conditions in the three-dimensional environment: simultaneously displaying, via the one or more display generation components and with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences that can be displayed by the computer system.
[0027] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generation components and one or more input devices, the one or more programs including instructions for: when a view of a three-dimensional environment in which the computer system is located is visible, detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment; and in response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: simultaneously displaying, via the one or more display generation components and with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences that can be displayed by the computer system.
[0028] According to some embodiments, a method is described. The method includes: at a computer system in communication with one or more display generation components and one or more input devices: detecting, via the one or more input devices, a user's gaze corresponding to a first display position of the one or more display generation components; in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components, displaying, via the one or more display generation components, a first object; when the first object is being displayed, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying, via the one or more display generation components, a movement of the first object; and after displaying the movement of the first object: performing a first operation based on determining that the user's gaze satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and refraining from performing the first operation based on determining that the user's gaze does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0029] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a user's gaze corresponding to a first display position of the one or more display generation components; in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components, displaying, via the one or more display generation components, a first object; when the first object is being displayed, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying, via the one or more display generation components, a movement of the first object; and after displaying the movement of the first object: performing a first operation based on determining that the user's gaze satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and refraining from performing the first operation based on determining that the user's gaze does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0030] According to some embodiments, a transient computer-readable storage medium is described. In some embodiments, the transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generation components and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a user's gaze corresponding to a first display position of the one or more display generation components; in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components, displaying a first object via the one or more display generation components; when the first object is being displayed, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying a movement of the first object via the one or more display generation components; and after displaying the movement of the first object: performing a first operation based on determining that the user's gaze satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and aborting the performance of the first operation based on determining that the user's gaze does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0031] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: detecting, via the one or more input devices, a user's gaze corresponding to a first display position of the one or more display generation components; in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components, displaying a first object via the one or more display generation components; when the first object is being displayed, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying a movement of the first object via the one or more display generation components; and after displaying the movement of the first object: performing a first operation based on determining that the user's gaze satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and aborting the performance of the first operation based on determining that the user's gaze does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0032] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: means for detecting, via the one or more input devices, a user's gaze corresponding to a first display position of the one or more display generation components; means for displaying a first object via the one or more display generation components in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components; means for detecting that a first set of criteria is met when the first object is being displayed; means for displaying a movement of the first object via the one or more display generation components in response to detecting that the first set of criteria is met; and means for performing the following operations after displaying the movement of the first object: performing a first operation based on determining that the user's gaze meets a second set of criteria indicating gaze tracking of the movement of the first object; and abandoning performance of the first operation based on determining that the user's gaze does not meet the second set of criteria indicating gaze tracking of the movement of the first object.
[0033] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generation components and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a user's gaze corresponding to a first display position of the one or more display generation components; displaying a first object via the one or more display generation components in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components; detecting that a first set of criteria is met when the first object is being displayed; displaying a movement of the first object via the one or more display generation components in response to detecting that the first set of criteria is met; and, after displaying the movement of the first object: performing a first operation based on determining that the user's gaze meets a second set of criteria indicating gaze tracking of the movement of the first object; and abandoning performance of the first operation based on determining that the user's gaze does not meet the second set of criteria indicating gaze tracking of the movement of the first object.
[0034] According to some embodiments, a method is described. The method includes: at a computer system in communication with one or more display generation components and one or more input devices: displaying virtual content via the one or more display generation components; when the virtual content is being displayed, detecting a first gesture in front of the face of a user of the computer system via the one or more input devices; and in response to detecting the first gesture: stopping the display of at least a portion of the virtual content based on determining that the first gesture in front of the user's face meets a first set of criteria; and maintaining the display of the virtual content based on determining that the first gesture in front of the user's face does not meet the first set of criteria.
[0035] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; when the virtual content is being displayed, detecting a first gesture in front of the face of a user of the computer system via the one or more input devices; and in response to detecting the first gesture: stopping the display of at least a portion of the virtual content based on determining that the first gesture in front of the user's face meets a first set of criteria; and maintaining the display of the virtual content based on determining that the first gesture in front of the user's face does not meet the first set of criteria.
[0036] According to some embodiments, a transitory computer-readable storage medium is described. In some embodiments, the transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; when the virtual content is being displayed, detecting a first gesture in front of the face of a user of the computer system via the one or more input devices; and in response to detecting the first gesture: stopping the display of at least a portion of the virtual content based on determining that the first gesture in front of the user's face meets a first set of criteria; and maintaining the display of the virtual content based on determining that the first gesture in front of the user's face does not meet the first set of criteria.
[0037] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; when displaying the virtual content, detecting a first gesture in front of the face of a user of the computer system via the one or more input devices; and in response to detecting the first gesture: stopping display of at least a portion of the virtual content based on determining that the first gesture in front of the user's face meets a first set of criteria; and maintaining display of the virtual content based on determining that the first gesture in front of the user's face does not meet the first set of criteria.
[0038] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: means for displaying virtual content via the one or more display generation components; means for detecting, when displaying the virtual content, a first gesture in front of the face of a user of the computer system via the one or more input devices; and means for performing the following operations in response to detecting the first gesture: stopping display of at least a portion of the virtual content based on determining that the first gesture in front of the user's face meets a first set of criteria; and maintaining display of the virtual content based on determining that the first gesture in front of the user's face does not meet the first set of criteria.
[0039] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generation components and one or more input devices, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; when displaying the virtual content, detecting a first gesture in front of the face of a user of the computer system via the one or more input devices; and in response to detecting the first gesture: stopping display of at least a portion of the virtual content based on determining that the first gesture in front of the user's face meets a first set of criteria; and maintaining display of the virtual content based on determining that the first gesture in front of the user's face does not meet the first set of criteria.
[0040] Note that the various embodiments described above can be combined with any other embodiments described herein. The features and advantages described in this specification are not exhaustive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art from the drawings, the specification, and the claims. In addition, it should be noted that the language used in this specification has been selected for readability and guidance purposes, and may not have been selected to depict or define the subject matter of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the following drawings, in which like reference numerals indicate corresponding parts in all the drawings.
[0042] Figure 1A is a block diagram illustrating an operating environment of a computer system for providing an XR experience according to some embodiments.
[0043] Figures 1B to 1P is for providing an example of a computer system for an XR experience in an Figure 1A operating environment.
[0044] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience according to some embodiments.
[0045] Figure 3 is a block diagram illustrating a display generation component of a computer system configured to provide a visual component of an XR experience to a user according to some embodiments.
[0046] Figure 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture a user's gesture input according to some embodiments.
[0047] Figure 5 is a block diagram illustrating an eye tracking unit of a computer system configured to capture a user's gaze input according to some embodiments.
[0048] Figure 6 is a flowchart illustrating a flash-assisted gaze tracking pipeline according to some embodiments.
[0049] Figures 7A to 7K Illustrates example techniques for navigating an extended reality experience according to some embodiments.
[0050] Figure 8 is a flowchart of a method for navigating an extended reality experience according to various embodiments.
[0051] Figure 9Flowchart of a method for navigating an extended reality experience according to various embodiments.
[0052] Figures 10A to 10G Illustrates example techniques for providing suggestions related to an extended reality experience according to some embodiments.
[0053] Figure 11 Flowchart of a method for providing suggestions related to an extended reality experience according to various embodiments.
[0054] Figures 12A to 12K Illustrates example techniques for gaze-based interaction according to some embodiments.
[0055] Figure 13 Flowchart of a method for gaze-based interaction according to some embodiments.
[0056] Figures 14A to 14L Illustrates example techniques for interacting with virtual content according to various embodiments.
[0057] Figure 15 Flowchart of a method for interacting with virtual content according to some embodiments. Detailed Description
[0058] According to some embodiments, the present disclosure relates to a user interface for providing a user with an extended reality (XR) experience.
[0059] The systems, methods, and GUIs described herein improve user interface interactions with extended reality environments and other virtual content in a variety of ways.
[0060] In some embodiments, a computer system simultaneously displays representations of multiple augmented reality experiences in a three-dimensional environment, the representations including a first representation of a first augmented reality experience and a second representation of a second augmented reality experience. When the representations of multiple augmented reality experiences are simultaneously displayed, the computer system receives a first user input. In response to receiving the first user input, the computer system stops displaying the representation of one or more of the multiple augmented reality experiences and, based on the direction and / or magnitude of the first user input, displays the first augmented reality experience or the second augmented reality experience. The computer system thus provides the user with the ability to switch between different augmented reality experiences in an intuitive and efficient manner.
[0061] In some embodiments, a computer system receives a first sequence of one or more user inputs via a first physical control. In some embodiments, the first physical control is a rotatable and pressable physical control such that a user can provide rotational input and press input via the first physical control. In response to receiving the first sequence of one or more user inputs, the computer system displays a first extended reality experience or a second extended reality experience based on the direction and / or magnitude of the first sequence of one or more user inputs. The computer system thus provides the user with the ability to switch between different augmented reality experiences in an intuitive and efficient manner.
[0062] In some embodiments, the computer system detects a first set of conditions in a three-dimensional environment in which the computer system is located. For example, in various embodiments, the computer system detects one or more visible objects, audio content, or other conditions in the physical environment in which the computer system is located. In response to detecting the first set of conditions in the three-dimensional environment, the computer system displays a first suggestion corresponding to a first augmented reality experience. The first augmented reality experience is selected from a plurality of augmented reality experiences that can be displayed by the computer system. For example, in some embodiments, the first augmented reality experience is selected based on the first set of conditions in the three-dimensional environment. The computer system thus provides the user with suggestions for potentially relevant augmented reality experiences based on the conditions detected by the computer system.
[0063] In some embodiments, the computer system detects a user's gaze corresponding to a first display position of one or more display generation components. In response to detecting the user's gaze corresponding to the first display position, the computer system displays a first object. For example, the computer system displays a fixation target that the user aims to track with his or her eye tracking. When the first object is displayed, the computer system detects that a first set of criteria is met, and in response to detecting that the first set of criteria is met, the computer system displays the movement of the first object. If the user successfully tracks the movement of the first object with his or her gaze, the computer system performs a first operation, and if the user does not successfully track the movement of the first object with his or her gaze, the computer system does not perform the first operation. For example, in some embodiments, the user can unlock the computer system by using gaze input to track the movement of the first object. The computer system thus provides the user with an intuitive and efficient way to perform operations such as unlocking the computer system using gaze input.
[0064] In some embodiments, the computer system displays virtual content. When the virtual content is displayed, the computer system detects a first gesture in front of the face of a user of the computer system. If the first gesture meets a first set of criteria, the computer system stops displaying at least a portion of the virtual content. In this way, the computer system allows the user to quickly and easily clear some or all of the virtual content using a gesture.
[0065] Figures 1A to 6 Provides a description of an example computer system for providing an XR experience to a user. Figures 7A to 7K Illustrates example techniques for navigating an extended reality experience according to some embodiments. Figure 8 Is a flowchart of a method for navigating an extended reality experience according to various embodiments. Figure 9 Is a flowchart of a method for navigating an extended reality experience according to some embodiments. Figures 7A to 7K The user interface in Figure 8 and Figure 9 illustrates the process in Figures 10A to 10G Illustrates example techniques for providing suggestions related to an extended reality experience according to some embodiments. Figure 11 Is a flowchart of a method for providing suggestions related to an extended reality experience according to various embodiments. Figures 10A to 10G The user interface of Figure 11 illustrates the process in Figures 12A to 12K Illustrates example techniques for gaze-based interaction according to some embodiments. Figure 13 Is a flowchart of a method for gaze-based interaction according to various embodiments. Figures 12A to 12K The user interface in Figure 13 illustrates the process in Figures 14A to 14L Illustrates example techniques for interacting with virtual content according to some embodiments. Figure 15 Is a flowchart of a method for interacting with virtual content according to various embodiments. Figures 14A to 14L The user interface in Figure 15 illustrates the process in
[0066] The processes described below enhance the operability of a device and make the user-device interface more efficient through various techniques (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device), including by providing the user with improved visual feedback, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation when a set of conditions has been met without further user input, improving privacy and / or security, providing a richer, more detailed, and / or more realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more effectively. Saving battery power, and thus weight, improves the ergonomics of the device. These techniques also enable real-time communication, allow for the use of fewer and / or less precise sensors, resulting in a more compact, lighter, and cheaper device, and enable the device to be used under various lighting conditions. These techniques reduce energy usage, thereby reducing the heat emitted by the device, which is particularly important for wearable devices where it can become uncomfortable for the user to wear the device if the device generates too much heat while operating within the operational parameters of the device components.
[0067] In addition, in a method where one or more of the steps described herein depend on one or more conditions being met, it should be understood that the method can be repeated in multiple iterations such that, during the course of the repetition, all of the conditions that determine the steps in the method are met in different iterations of the method. For example, if a method requires performing a first step (if a condition is met) and a second step (if the condition is not met), a person of ordinary skill in the art will know to repeat the stated steps until both the condition is met and the condition is not met (in no particular order). Thus, a method described as having one or more steps that depend on one or more conditions being met can be rewritten as a method that repeats until each of the conditions described in the method is met. However, this does not require the system or computer-readable medium to state that the system or computer-readable medium includes instructions for performing conditional operations based on the satisfaction of the corresponding one or more conditions and is thus capable of determining whether the possible conditions have been met without explicitly repeating the steps of the method until all of the conditions that determine the steps in the method are met. A person of ordinary skill in the art will also understand that, similar to a method with conditional steps, a system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all of the conditional steps have been performed.
[0068] In some embodiments, as Figure 1AAs shown, an XR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a haptic sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., household appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).
[0069] When describing an XR experience, various terms are used to distinctively refer to several related but different environments that a user can sense and / or with which a user can interact (e.g., interact using inputs detected by the computer system 101 that generates the XR experience, where these inputs cause the computer system that generates the XR experience to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:
[0070] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the help of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.
[0071] Extended Reality: In contrast, an extended reality (XR) environment is a fully or partially simulated environment in which people sense and / or interact via an electronic system. In XR, a subset of a person's physical movements or representations thereof are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that complies with at least one physical law. For example, an XR system can detect a person's head rotation, and in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the characteristics of virtual objects in the XR environment can be made in response to a representation of a physical movement (e.g., a voice command). A person can use any of their senses to sense and / or interact with XR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of point audio sources in 3D space. Additionally, an audio object can enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, a person can sense and / or interact only with audio objects.
[0072] Examples of XR include virtual reality and mixed reality.
[0073] Virtual Reality: A virtual reality (VR) environment is a simulated environment that is designed to be based entirely on computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects with which a person can sense and / or interact. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with the virtual objects in the VR environment by way of a simulation of the person's presence within the computer-generated environment and / or by way of a simulation of a subset of the person's physical movements within the computer-generated environment.
[0074] Mixed Reality: Compared to a VR environment that is designed to be based entirely on computer-generated sensory input, a mixed reality (MR) environment is an analog environment that is designed to include, in addition to computer-generated sensory input (e.g., virtual objects), sensory input from the physical environment or a representation thereof. On the virtual continuum, an MR environment is any condition between a fully physical environment at one end and a virtual reality environment at the other end, but excluding these two ends. In some MR environments, the computer-generated sensory input can respond to changes in the sensory input from the physical environment. Additionally, some electronic systems for presenting an MR environment can track the position and / or orientation relative to the physical environment so that virtual objects can interact with real objects (i.e., physical items from the physical environment or a representation thereof). For example, the system can cause movement such that a virtual tree appears stationary relative to the physical ground.
[0075] Examples of mixed reality include augmented reality and augmented virtuality.
[0076] Augmented Reality: An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are superimposed over a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or translucent display such that the person, using the system, perceives the virtual objects superimposed over the physical environment. Alternatively, the system can have an opaque display and one or more imaging sensors that capture images or video of the physical environment, which are representations of the physical environment. The system combines the images or video with the virtual objects and presents the combination on the opaque display. The person, using the system, indirectly views the physical environment via the images or video of the physical environment and perceives the virtual objects superimposed over the physical environment. As used herein, the video of the physical environment displayed on the opaque display is referred to as “passthrough video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system can have a projection system that projects virtual objects into the physical environment, such as as a hologram or on a physical surface, such that the person, using the system, perceives the virtual objects superimposed over the physical environment. An augmented reality environment is also a simulated environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, in providing passthrough video, the system can transform one or more sensor images to impose an alternative perspective (e.g., viewpoint) that is different from the perspective captured by the imaging sensors. As another example, a representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions thereof such that the modified portions can be a representative but not a true version of the originally captured image. As yet another example, a representation of the physical environment can be transformed by graphically removing portions thereof or blurring portions thereof.
[0077] Augmented Virtuality: An augmented virtuality (AV) environment is a simulated environment in which a virtual environment or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but the face of a person is a realistic reproduction of an image of a physical person. As another example, a virtual object can adopt the shape or color of a physical item imaged by one or more imaging sensors. As yet another example, a virtual object can adopt a shadow that conforms to the positioning of the sun in the physical environment.
[0078] In an augmented reality, mixed reality, or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of the three-dimensional environment is typically visible to the user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), the virtual viewport having a viewport boundary that defines the extent of the three-dimensional environment that is visible to the user via the one or more display generation components. In some embodiments, the region defined by the viewport boundary is smaller than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). In some embodiments, the region defined by the viewport boundary is larger than the user's visual field in one or more dimensions (e.g., based on the user's visual field, the size, optical properties, or other physical characteristics of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). The viewport and the viewport boundary typically move as the one or more display generation components move (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves). The user's viewpoint determines the content visible in the viewport, the viewpoint typically specifying a position and orientation relative to the three-dimensional environment, and as the viewpoint moves, the view of the three-dimensional environment will also move in the viewport. For a head-mounted device, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a perceptually accurate view of the three-dimensional environment that provides an immersive experience when the user is using the head-mounted device. For a handheld or stationary device, the viewpoint moves as the handheld or stationary device moves and / or as the user's positioning relative to the handheld or stationary device changes (e.g., the user moves toward, away from, up, down, right, and / or left). For a device that includes a display generation component with virtual passthrough, the portions of the physical environment that are visible (e.g., displayed and / or projected) via the one or more display generation components are based on the fields of view of one or more cameras that communicate with the display generation component, the one or more cameras typically moving as the display generation component moves (e.g., for a head-mounted device as the user's head moves, or for a handheld device such as a tablet or smartphone as the user's hand moves), because the user's viewpoint moves as the fields of view of the one or more cameras move (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual object are updated based on the movement of the user's viewpoint)).For a display generation component with optical passthrough, portions of the physical environment that are visible via one or more display generation components (e.g., optically visible through one or more partial or fully transparent portions of the display generation component) are based on the user's field of view through the partial or fully transparent portion of the display generation component (e.g., for a head-mounted device, moving as the user's head moves, or for a handheld device such as a tablet or smartphone, moving as the user's hand moves), because the user's viewpoint moves as the user's field of view through the partial or fully transparent portion of the display generation component moves (and the appearance of one or more virtual objects is updated based on the user's viewpoint).
[0079] In some embodiments, the representation of the physical environment (e.g., via virtual passthrough or optical passthrough display) may be partially or fully occluded by the virtual environment. In some embodiments, the amount of the virtual environment displayed (e.g., the amount of the physical environment not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally causes more of the virtual environment to be displayed, replacing and / or occluding more of the physical environment, and decreasing the immersion level optionally causes less of the virtual environment to be displayed, thereby revealing portions of the physical environment that were previously not displayed and / or occluded. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) more than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the immersion level includes the associated degree to which virtual content (e.g., virtual environment and / or virtual content) displayed by a computer system occludes background content (e.g., content other than the virtual environment and / or virtual content) around / behind the virtual environment, optionally including the number of items of the displayed background content and / or the displayed visual characteristics (e.g., color, contrast, and / or opacity) of the background content, the angular range of the virtual content displayed by a display generation component (e.g., 60 degrees of content displayed at low immersion, 120 degrees of content displayed at medium immersion, or 180 degrees of content displayed at high immersion), and / or the proportion of the field of view displayed by the display generation component occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content in the representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface generated by a computer system corresponding to an application), virtual objects not associated with and / or not included in the virtual environment and / or virtual content (e.g., files or representations of other users generated by a computer system, etc.), and / or real objects (e.g., passthrough objects representing real objects in the physical environment around the user, which are visible such that they are displayed by the display generation component and / or visible via a transparent or translucent component of the display generation component because the computer system does not occlude / hinder their visibility through the display generation component). In some embodiments, at a low immersion level (e.g., a first immersion level), the background, virtual, and / or real objects are displayed in an unoccluded manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with background content, which is optionally displayed at full brightness, color, and / or semi-transparency.In some embodiments, at a higher immersion level (e.g., a second immersion level higher than the first immersion level), background, virtual, and / or real objects are displayed in an occluded manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full-screen or fully immersive mode). As another example, a virtual environment displayed at a medium immersion level is displayed simultaneously with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary among the background objects. For example, at a particular immersion level, one or more first background objects are more visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, zero immersion or a zero immersion level corresponds to a virtual environment that ceases to be displayed, and instead a representation of the physical environment is displayed (optionally with one or more virtual objects, such as applications, windows, or virtual three-dimensional objects), and the representation of the physical environment is not occluded by the virtual environment. Adjusting the immersion level using physical input elements provides a quick and efficient way to adjust immersion, which enhances the operability of the computer system and makes the user-device interface more efficient.
[0080] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same position and / or orientation in the user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user looks straight ahead, the user's viewpoint is at least a portion of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments in which the computer system has a display generation component (e.g., a display screen) that is repositionable relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the computer system's display generation component. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or orientation at which a viewpoint-locked virtual object is displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments in which the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head such that the virtual object is also referred to as a "head-locked virtual object".
[0081] Environment-Locked Visual Objects: When a computer system displays a virtual object at a location and / or orientation within a user's field of view, the virtual object is environment-locked (alternatively, "world-locked"), where the location and / or orientation is based on a location and / or object within a three-dimensional environment (e.g., a physical environment or a virtual environment) (e.g., selected and / or anchored to the location and / or object with reference to the location and / or object). As the user's field of view moves, the location and / or object within the environment relative to the user's field of view changes, which causes the environment-locked virtual object to be displayed at a different location and / or orientation within the user's field of view. For example, an environment-locked virtual object locked to a tree directly in front of the user is displayed at the center of the user's field of view. When the user's field of view shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center within the user's field of view (e.g., the location of the tree within the user's field of view is offset), the environment-locked virtual object locked to the tree is displayed to the left of center within the user's field of view. In other words, the location and / or orientation at which the environment-locked virtual object is displayed within the user's field of view depends on the location and / or orientation of the location and / or object to which the virtual object is locked within the environment. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system anchored to a fixed location and / or object within a physical environment) in order to determine the orientation at which the environment-locked virtual object is displayed within the user's field of view. The environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object), or can be locked to a movable portion of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body such as the user's hand, wrist, arm, or foot that moves independently of the user's field of view) such that the virtual object moves as the field of view or that portion of the environment moves to maintain a fixed relationship between the virtual object and that portion of the environment.
[0082] In some embodiments, an environment-locked or view-locked virtual object exhibits a lazy follow behavior, which reduces or delays the movement of the environment-locked or view-locked virtual object relative to the movement of a reference point that the virtual object follows. In some embodiments, when exhibiting the lazy follow behavior, when the computer system detects movement of a reference point (e.g., a portion of the environment, a view point, or a point fixed relative to the view point, such as a point between 5 cm and 300 cm from the view point) that the virtual object is following, the computer system intentionally delays the movement of the virtual object. For example, when the reference point (e.g., a portion of the environment or a view point) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point but moves at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which time the virtual object begins to catch up to the reference point). In some embodiments, when the virtual object exhibits the lazy follow behavior, the device ignores small movements of the reference point (e.g., ignores movements of the reference point below a threshold movement amount, such as moving 0 to 5 degrees or moving 0 to 50 cm). For example, when the reference point (e.g., a portion of the environment to which the virtual object is locked or a view point) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed to remain fixed or substantially fixed relative to a different view point or portion of the environment than the reference point to which the virtual object is locked), and when the reference point (e.g., a portion of the environment to which the virtual object is locked or a view point) moves a second amount that is greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed to remain fixed or substantially fixed relative to a different view point or portion of the environment than the reference point to which the virtual object is locked), and then decreases when the movement amount of the reference point increases above a threshold (e.g., a "lazy follow" threshold) because the virtual object is moved by the computer system to remain fixed or substantially fixed relative to the reference point. In some embodiments, the virtual object remaining substantially fixed in position relative to the reference point includes the virtual object being displayed within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the positioning of the reference point).
[0083] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smartphones, tablet devices, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system for providing audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of a physical environment and / or one or more microphones for capturing audio of a physical environment. Instead of an opaque display, a head-mounted system may have a transparent or translucent display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium may be an optical waveguide, a hologram medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display may be configured to selectively become opaque. A projection-based system may employ retinal projection technology that projects graphic images onto a person's retina. The projection system may also be configured to project virtual objects into a physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. The following description with respect to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is local or remote relative to scene 105 (e.g., a physical environment). For example, controller 110 is a local server located within scene 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., a physical enclosure) of one or more of display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), one or more input devices of input device 125, one or more output devices of output device 155, one or more sensors of sensor 190, and / or one or more peripheral devices of peripheral device 195, or shares the same physical housing or support structure with one or more of the foregoing devices.
[0084] In some embodiments, display generation component 120 is configured to provide an XR experience (e.g., at least the visual component of the XR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. More details regarding Figure 3 display generation component 120 are described below. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.
[0085] According to some embodiments, when a user is virtually and / or physically present within scene 105, display generation component 120 provides an XR experience to the user.
[0086] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). Thus, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or a tablet device) configured to present XR content, and the user holds the device with a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR room, housing, or chamber configured to present XR content, where the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with XR content triggered based on an interaction occurring in the space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interaction occurs in the space in front of the HMD and the response to the XR content is displayed via the HMD. Similarly, a user interface showing an interaction with XR content triggered based on the movement of a handheld device or a tripod-mounted device relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by the movement of the HMD relative to the physical environment (e.g., scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).
[0087] Although relevant features of the operating environment 100 are shown in Figure 1A For the sake of brevity and to not obscure more relevant aspects of the example embodiments disclosed herein, various other features are not illustrated, although will be recognized by those of ordinary skill in the art from this disclosure.
[0088] Figures 1A to 1PIllustrates various examples of computer systems for performing methods and providing audio, visual, and / or tactile feedback as part of the user interfaces described herein. In some embodiments, the computer system includes one or more display generation components (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b) for displaying to a user of the computer system a representation of virtual elements and / or a physical environment optionally generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216, which are optionally removably attached to one or more of the optical modules such that the user interface is more easily viewable by users who would otherwise use glasses or contact lenses to correct their vision. Although many of the user interfaces shown herein show a single view of the user interface, the user interface in the HMD is optionally displayed using two optical modules (e.g., first display assembly 1-120a and second display assembly 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b), one optical module for the user's right eye and a different optical module for the user's left eye, and presenting slightly different images to the two different eyes to create an illusion of stereoscopic depth, the single view of the user interface is typically a right-eye view or a left-eye view, and the depth effect is explained in the text or using other schematic diagrams or views. In some embodiments, the computer system includes one or more external displays (e.g., display assembly 1-108) for displaying to a user of the computer system (when the computer system is not being worn) and / or to others in the vicinity of the computer system status information for the computer system, the status information optionally being generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronics component 1-112) for generating audio feedback, the audio feedback optionally being generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device (e.g., one or more sensors in sensor assembly 1-356, and / or Figure 1I )), which can be used (optionally in combination with one or more illuminators, such as Figure 1IThe illuminator described in [ID] generates a digital pass-through image, captures visual media corresponding to the physical environment (e.g., photos and / or videos), or determines the pose (e.g., position and / or orientation) of physical objects and / or surfaces in the physical environment, such that virtual objects can be placed based on the detected pose of the physical objects and / or surfaces. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand position and / or movement (e.g., sensor assemblies 1-356 and / or Figure 1I one or more of the sensors in [ID]), which can be used (optionally in combination with one or more illuminators, such as Figure 1I the illuminator 6-124 described in [ID]) to determine when one or more air gestures have been performed. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I the eye tracking and gaze tracking sensors in [ID]), which can be used (optionally in combination with one or more lights, such as Figure 1OThe lights in (11.3.2-110) determine the attention or fixation position and / or fixation movement, which can optionally be used to detect fixation-only input based on fixation movement and / or dwell. Combinations of the various sensors described above can be used to determine the user's facial expression and / or hand movement for generating a user avatar or representation, such as an anthropomorphic avatar or representation for a real-time communication session, where the avatar has facial expressions, hand movements, and / or body movements based on or similar to the detected facial expressions, hand movements, and / or body movements of the user of the device. The fixation and / or attention information is optionally combined with hand tracking information to determine the interaction between the user and one or more user interfaces based on direct and / or indirect input, such as an air gesture or input using one or more hardware input devices, such as one or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328), a knob (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328), a digital crown (e.g., a first button 1-128, button 11.1.1-114, and / or dial or button 1-328 that can be pressed and twisted or rotated), a touchpad, a touch screen, a keyboard, a mouse, and / or other input devices. One or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328) are optionally used to perform system operations, such as re-centering the content in the three-dimensional environment visible to the user of the device, displaying the main user interface for launching an application, starting a real-time communication session, or initiating the display of a virtual three-dimensional background. A knob or digital crown (e.g., a first button 1-128, button 11.1.1-114, and / or dial or button 1-328 that can be pressed and twisted or rotated) is optionally rotatable to adjust parameters of the visual content, such as the immersion level of the virtual three-dimensional environment (e.g., the extent to which the virtual content occupies the user's viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and the virtual content displayed via an optical module (e.g., first display component 1-120a and second display component 1-120b and / or first optical module 11.1.1-104a and second optical module 11.1.1-104b).
[0089] Figure 1BFront views, top views, and perspective views of examples of head - mountable display (HMD) devices 1 - 100 configured to be worn by a user and provide virtual and augmented reality (VR / AR) experiences are illustrated. The HMD 1 - 100 may include a display unit 1 - 102 or assembly, an electronic strip assembly 1 - 104 connected to and extending from the display unit 1 - 102, and a strap assembly 1 - 106 fixed to the electronic strip assembly 1 - 104 at either end. The electronic strip assembly 1 - 104 and the strap 1 - 106 may be part of a retention assembly configured to wrap around a user's head to hold the display unit 1 - 102 against the user's face.
[0090] In at least one example, the strap assembly 1 - 106 may include a first strap 1 - 116 configured to wrap around the back of the user's head and a second strap 1 - 117 configured to extend over the top of the user's head. As shown, the second strap may extend between a first electronic strip 1 - 105a and a second electronic strip 1 - 105b of the electronic strip assembly 1 - 104. The strip assembly 1 - 104 and the strap assembly 1 - 106 may be part of a fixation mechanism that extends rearward from the display unit 1 - 102 and is configured to hold the display unit 1 - 102 against the user's face.
[0091] In at least one example, the fixation mechanism includes a first electronic strip 1 - 105a that includes a first proximal end 1 - 134 coupled to the display unit 1 - 102 (e.g., the housing 1 - 150 of the display unit 1 - 102) and a first distal end 1 - 136 opposite the first proximal end 1 - 134. The fixation mechanism may also include a second electronic strip 1 - 105b that includes a second proximal end 1 - 138 coupled to the housing 1 - 150 of the display unit 1 - 102 and a second distal end 1 - 140 opposite the second proximal end 1 - 138. The fixation mechanism may also include a first strap 1 - 116 and a second strap 1 - 117, the first strap including a first end 1 - 142 coupled to the first distal end 1 - 136 and a second end 1 - 144 coupled to the second distal end 1 - 140, and the second strap extending between the first electronic strip 1 - 105a and the second electronic strip 1 - 105b. The strips 1 - 105a - b and the strap 1 - 116 may be coupled via a connection mechanism or assembly 1 - 114. In at least one example, the second strap 1 - 117 includes a first end 1 - 146 coupled to the first electronic strip 1 - 105a between the first proximal end 1 - 134 and the first distal end 1 - 136 and a second end 1 - 148 coupled to the second electronic strip 1 - 105b between the second proximal end 1 - 138 and the second distal end 1 - 140.
[0092] In at least one example, the first and second electronic strips 1-105a-b comprise plastic, metal, or other structural materials forming the shape of a substantially rigid strip 1-105a-b. In at least one example, the first strip 1-116 and the second strip 1-117 are formed of an elastic flexible material including woven textiles, rubber, etc. The first strip 1-116 and the second strip 1-117 may be flexible to conform to the shape of the user's head when wearing the HMD 1-100.
[0093] In at least one example, one or more of the first and second electronic strips 1-105a-b may define an internal strip volume and include one or more electronic components disposed within the internal strip volume. In one example, as Figure 1B shown, the first electronic strip 1-105a may include an electronic component 1-112. In one example, the electronic component 1-112 may include a speaker. In one example, the electronic component 1-112 may include a computing component, such as a processor.
[0094] In at least one example, the housing 1-150 defines a first front opening 1-152. The front opening is Figure 1B marked as 1-152 in dashed lines therein because the display assembly 1-108 is arranged to occlude the first opening 1-152 from view when assembling the HMD 1-100. The housing 1-150 may also define a second rear opening 1-154. The housing 1-150 also defines an internal volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover and a display screen (shown in other figures) disposed in or across the front opening 1-152 to occlude the front opening 1-152. In at least one example, the display screen of the display assembly 1-108 and generally the display assembly 1-108 have a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 may be curved as shown to complement the user's facial features and the overall curvature from one side of the face to the other, e.g., from left to right and / or from top to bottom, where the display unit 1-102 is pressed.
[0095] In at least one example, the housing 1-150 may define a first aperture 1-126 between a first opening 1-152 and a second opening 1-154, and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may further include a first button 1-126 disposed in the first aperture 1-128, and a second button 1-132 disposed in the second aperture 1-130. The first button 1-128 and the second button 1-132 can be pressed through the respective apertures 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 can be a twist dial as well as a push button. In at least one example, the first button 1-128 is a pushable and twistable dial button, and the second button 1-132 is a push button.
[0096] Figure 1C A rear perspective view of the HMD 1-100 is illustrated. The HMD 1-100 may include a light seal 1-110 that extends rearwardly around a perimeter of the housing 1-150 from the housing 1-150 of the display assembly 1-108, as shown. The light seal 1-110 may be configured to extend from the housing 1-150 to a user's face and around the user's eyes to block external light from being visible. In one example, the HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b that are disposed at or within a second opening 1-154 defined by the housing 1-150 that faces rearward and / or disposed within an interior volume of the housing 1-150 and configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include a respective display screen 1-122a, 1-122b that is configured to project light through the second opening 1-154 in a rearward direction toward the user's eyes.
[0097] In at least one example, referring Figure 1B and Figure 1C both, the display assembly 1-108 can be a front forward display assembly that includes a display screen configured to project light in a first forward direction, and the rear display screens 1-122a-b can be configured to project light in a second rearward direction that is opposite the first direction. As described above, the light seal 1-110 can be configured to block light external to the HMD 1-100 from reaching the user's eyes, including light projected by Figure 1B the front forward display screen of the display assembly 1-108 shown in the front perspective view of. In at least one example, the HMD 1-100 may further include a curtain 1-124 that occludes the second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a-b. In at least one example, the curtain 1-124 can be elastic or at least partially elastic.
[0098] Figure 1B and Figure 1C Any of the feature portions, components, and / or parts shown (including their arrangements and configurations) may be included, either individually or in any combination, in Figures 1D to 1F any example of the devices, feature portions, components, and parts shown and described herein. Similarly, reference Figures 1D to 1F to any of the feature portions, components, and / or parts shown or described (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1B and Figure 1C the examples of the devices, feature portions, components, and parts shown.
[0099] Figure 1D FIG. Figure 1D illustrates an exploded view of an example of HMD 1-200, which includes various parts or components separated according to the modularity and selective coupling of these parts. For example, HMD 1-200 may include a strap 1-216, which may be selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strip 1-205a may include a first electronic component 1-212a, and the second fixed strip 1-205b may include a second electronic component 1-212b. In at least one example, the first and second strips 1-205a-b can be removably coupled to the display unit 1-202.
[0100] In addition, HMD 1-200 may include a light seal 1-210 configured to be removably coupled to the display unit 1-202. HMD 1-200 may also include a lens 1-218, which may be removably coupled to the display unit 1-202, for example, on a first component and a second display component including a display screen. The lens 1-218 may include a custom prescription lens configured to correct vision. As noted, each of the parts shown in the exploded view of Figure 1D and described above can be removably coupled, attached, reattached, and replaced to update the parts or swap out parts for different users. For example, straps such as strap 1-216, light seals such as light seal 1-210, lenses such as lens 1-218, and electronic strips such as electronic strips 1-205a-b can be swapped out according to the user, such that these parts are customized to fit and correspond to a single user of HMD 1-200.
[0101] Figure 1D Any of the feature portions, components, and / or parts shown (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1B , Figure 1C and Figures 1E to 1F any example of the devices, feature portions, components, and parts shown and described herein. Similarly, referenceFigure 1B , Figure 1C and Figures 1E to 1F Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination in Figure 1D the examples of the devices, features, components, and parts shown.
[0102] Figure 1E An exploded view of an example of the display unit 1-306 of the HMD is illustrated. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320 that includes a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.
[0103] In at least one example, the display unit 1-306 may also include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the position of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, and each display screen 1-322a-b has at least one motor such that the motors can translate the display screens 1-322a-b to match the pupil spacing of the user's eyes.
[0104] In at least one example, the display unit 1-306 may include a dial or button 1-328 that can be pressed relative to the frame 1-350 and accessed by a user external to the frame 1-350. The button 1-328 may be electrically connected to the motor assembly 1-362 via a controller such that the button 1-328 can be manipulated by the user to cause the motors of the motor assembly 1-362 to adjust the position of the display screens 1-322a-b.
[0105] Figure 1E Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination in Figures 1B to 1D and Figure 1F any of the other examples of the devices, features, components, and parts shown and described herein. Similarly, with reference to Figures 1B to 1D and Figure 1F Any of the features, components, and / or parts shown and described (including their arrangement and configuration) may be included individually or in any combination in Figure 1EIn the examples of the devices, features, components, and parts shown.
[0106] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the positions of a first display sub-assembly 1-420a and a second display sub-assembly 1-420b of the rear display assembly 1-421, including a first corresponding display screen and a second corresponding display screen for inter-pupillary adjustment, as described above.
[0107] Figure 1F The various parts, systems, and components shown in the exploded view are described in more detail herein with reference to Figures 1B to 1E and the subsequent figures referenced in this disclosure. Figure 1F The display unit 1-406 shown may be assembled and integrated with Figures 1B to 1E the shown fixing mechanism, which includes an electronic strip, a belt, and other components including a light seal, a connection assembly, etc.
[0108] Figure 1F Any of the features, components, and / or parts shown (including their arrangements and configurations) may be included, either individually or in any combination, in Figures 1B to 1E any example of the other examples of the devices, features, components, and parts shown and described herein. Similarly, with reference to Figures 1B to 1E any of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1F the examples of the devices, features, components, and parts shown.
[0109] Figure 1G An exploded perspective view of a front cover assembly 3-100 of an HMD device described herein is illustrated, such as Figure 1G the front cover assembly 3-1 of the HMD 3-100 shown or any other HMD device shown and described herein. Figure 1GThe front cover assembly 3-100 shown may include a transparent or translucent cover 3-102, a shield 3-104 (or "cover"), an adhesive layer 3-106, a display assembly 3-108 including a bi-convex lens panel or array 3-110, and a structural trim 3-112. The adhesive layer 3-106 may secure the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the trim 3-112. The trim 3-112 may secure the various components of the front cover assembly 3-100 to the frame or base of the HMD device.
[0110] In at least one example, as Figure 1G shown, the transparent cover 3-102, the shield 3-104, and the display assembly 3-108 including the bi-convex lens array 3-110 may be curved to conform to the curvature of the user's face. The transparent cover 3-102 and the shield 3-104 may be curved in two or three dimensions, e.g., vertically in the Z direction inside and outside the Z-X plane, and horizontally in the X direction inside and outside the Z-X plane. In at least one example, the display assembly 3-108 may include the bi-convex lens array 3-110 and a display panel having pixels configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 may be curved in at least one direction (e.g., the horizontal direction) to conform to the curvature of the user's face from one side (e.g., the left side) to the other side (e.g., the right side) of the face. In at least one example, each layer or component of the display assembly 3-108 (which will be shown and described in more detail in subsequent figures, but which may include the bi-convex lens array 3-110 and the display layer) may be similarly or concentrically curved in the horizontal direction to conform to the curvature of the user's face.
[0111] In at least one example, the shield 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the shield 3-104 may include one or more opaque portions, such as an opaque ink printed portion or other opaque film portion on the back surface of the shield 3-104. When the HMD device is worn, the back surface may be the surface of the shield 3-104 facing the user's eyes. In at least one example, the opaque portion may be on the front surface of the shield 3-104 opposite the back surface. In at least one example, one or more opaque portions of the shield 3-104 may include a peripheral portion that visually hides any components around the outer perimeter of the display screen of the display assembly 3-108. In this way, the opaque portion of the shield hides any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shield 3-104, including electronic components, structural components, etc.
[0112] In at least one example, the shroud 3-104 may define one or more apertured transparent portions 3-120 through which the sensor may transmit and receive signals. In one example, portion 3-120 is an aperture through which the sensor may extend or through which the sensor may transmit and receive signals. In one example, portion 3-120 is a transparent portion, or a portion that is more transparent than the surrounding translucent or opaque portion of the shroud, through which the sensor may transmit and receive signals through the shroud and through the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.
[0113] Figure 1G Any one of the illustrated features, components, and / or parts (including their arrangement and configuration) may be included, either alone or in any combination, in any example of the other devices, features, components, and parts described herein. Similarly, any one of the features, components, and / or parts (including their arrangement and configuration) shown and described herein may be included, either alone or in any combination, in Figure 1G the examples of the devices, features, components, and parts shown.
[0114] Figure 1H An exploded view of an example of an HMD device 6-100 is illustrated. The HMD device 6-100 may include a sensor array or system 6-102 that includes one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 may be secured / fastened.
[0115] Figure 1I A portion of an HMD device 6-100 including a front transparent cover 6-104 and a sensor system 6-102 is illustrated. The sensor system 6-102 may include a plurality of different sensors, transmitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is shown in front of the sensor system 6-102 to illustrate the relative positions of the various sensors and transmitters and the orientation of each sensor / transmitter of the system 6-102. As referred to herein, "lateral", "side", "transverse", "horizontal", and other like terms refer to the orientation or direction indicated by the X-axis as Figure 1J shown. Terms such as "vertical", "upward", "downward", and like terms refer to the orientation or direction indicated by the Z-axis as Figure 1J shown. Terms such as "forward", "backward", "frontward", "backward", and like terms refer to the orientation or direction indicated by the Y-axis as Figure 1J shown.
[0116] In at least one example, a transparent cover 6-104 may define a front outer surface of the HMD device 6-100, and a sensor system 6-102 including various sensors and their components may be disposed behind the cover 6-104 in the Y-axis / direction. The cover 6-104 may be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted therefrom.
[0117] As described elsewhere herein, the HMD device 6-100 may include one or more controllers that include processors for electrically coupling the various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as display screens. Additionally, as will be shown in more detail with reference to other figures below, the various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to Figure 1I various structural frame members, brackets, etc. of the HMD device 6-100 not shown. For clarity, Figure 1I the components of the sensor system 6-102 are shown unattached and unelectrically coupled to other components.
[0118] In at least one example, the device may include one or more controllers having processors configured to execute instructions stored on a memory component electrically coupled to the processor. The instructions may include or cause the processor to execute one or more algorithms for self-correcting the angles and positions of the various cameras described herein over time as the initial position, angle, or orientation of the camera collides or deforms due to an accidental drop event or other event.
[0119] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. The system 6-102 may include two scene cameras 6-102 disposed on either side of the bridge or arch structure of the HMD device 6-100 such that each of the two cameras 6-106 generally corresponds to the position of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y-direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and provide images and content for MR video passthrough to a display screen facing the user's eyes when using the HMD device 6-100. The scene cameras 6-106 may also be used for environment and object reconstruction.
[0120] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 that typically points forward in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction and for tracking the user's hands and body. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 that is centered along the width of the HMD device 6-100 (e.g., along the X axis). For example, the second depth sensor 6-110 may be disposed above the central nose bridge or on an adapter structure above the nose when the user wears the HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction and for hand and body tracking. In at least one example, the second depth sensor may include a LIDAR sensor.
[0121] In at least one example, the sensor system 6-102 may include a depth projector 6-112 that is typically oriented forward to project electromagnetic waves (e.g., in the form of a pre-determined pattern of light points) into the field of view or within the field of view of the user and / or the scene camera 6-106, or into a field of view that includes and extends beyond the field of view of the user and / or the scene camera 6-106. In at least one example, the depth projector is capable of projecting electromagnetic waves of light in the form of a pattern of light points that are reflected from an object and back into the aforementioned depth sensors, including depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 may be used for environment and object reconstruction and for hand and body tracking.
[0122] In at least one example, the sensor system 6-102 may include a downward-facing camera 6-114 whose field of view typically points downward relative to the HDM device 6-100 along the Z axis. In at least one example, the downward camera 6-114 may be disposed on the left and right sides of the HMD device 6-100 as shown and is used for hand and body tracking, head-mounted headset tracking, and facial avatar detection and creation for displaying a user avatar on the forward display screen of the HMD device 6-100 described elsewhere herein. For example, the downward camera 6-114 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the cheeks, mouth, and chin.
[0123] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw camera 6-116 may be disposed on the left and right sides of the HMD device 6-100 as shown and is used for hand and body tracking, head-mounted headset tracking, and face avatar detection and creation for displaying a user avatar on the forward display of the HMD device 6-100 described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of the user's face below the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. For hand and body tracking, head-mounted headset tracking, and face avatar
[0124] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right views in the X-axis or with respect to the direction of the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, head-mounted headset tracking, and face avatar detection and recreation.
[0125] In at least one example, the sensor system 6-102 may include a plurality of eye tracking and gaze tracking sensors for determining the identity, status, and gaze direction of the user's eyes during and / or before use. In at least one example, the eye / gaze tracking sensors may include a nose-eye camera 6-120 that is disposed on either side of the user's nose and is adjacent to the user's nose when the HMD device 6-100 is worn. The eye / gaze sensors may also include a bottom eye camera 6-122 that is disposed below the respective user's eye for capturing an image of the eye for face avatar detection and creation, gaze tracking, and iris identification functions.
[0126] In at least one example, the sensor system 6-102 may include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 may detect the top light refresh rate to avoid display flicker. In one example, the infrared illuminator 6-124 may include a light-emitting diode and may be particularly used in low-light environments to illuminate the user's hand and other objects in low light for detection by the infrared sensors of the sensor system 6-102.
[0127] In at least one example, multiple sensors (including scene cameras 6-106, downward cameras 6-114, jaw cameras 6-116, side cameras 6-118, depth projectors 6-112, and depth sensors 6-108, 6-110) can be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for sizing, in order to better perform hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, the downward cameras 6-114, jaw cameras 6-116, and side cameras 6-118 described above and shown in Figure 1I can be wide-angle cameras capable of operating in the visible and infrared spectra. In at least one example, these cameras 6-114, 6-116, 6-118 can operate only in black-and-white light detection to simplify image processing and obtain sensitivity.
[0128] Figure 1I Any of the features, components, and / or parts shown (including their arrangements and configurations) can be included, either individually or in any combination, in Figures 1J to 1L any example of the devices, features, components, and parts shown and described herein. Similarly, any of the features, components, and / or parts shown and described with reference to Figures 1J to 1L can be included, either individually or in any combination, in Figure 1I the examples of the devices, features, components, and parts shown.
[0129] Figure 1J A bottom perspective view of an example of an HMD 6-200 including a cover or shroud 6-204 fixed to a frame 6-230 is illustrated. In at least one example, sensors 6-203 of the sensor system 6-202 can be disposed around the perimeter of the HDM 6-200 such that the sensors 6-203 are disposed outwardly around the perimeter of the display area or display zone 6-232 so as not to obstruct the viewing of the displayed light. In at least one example, the sensors can be disposed behind the shroud 6-204 and aligned with the transparent portion of the shroud, thereby allowing the sensors and projectors to allow light to pass back and forth through the shroud 6-204. In at least one example, an opaque ink or other opaque material or film / layer can be disposed on the shroud 6-204 around the display area 6-232 to hide the components of the HMD 6-200 outside the display area 6-232 rather than the transparent portion defined by the opaque portion through which the sensors and projectors transmit and receive light and electromagnetic signals during operation. In at least one example, the shroud 6-204 allows light to pass through from the display (e.g., within the display area 6-232), but does not allow light to pass radially outward from the display area around the perimeter of the display and the shroud 6-204.
[0130] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent regions 6-209 through which the sensors 6-203 of the sensor system 6-202 may send and receive signals. In the illustrated example, the sensors 6-203 of the sensor system 6-202 send and receive signals through the shield 6-204, or more specifically through the transparent regions 6-209 (or defined thereby) of the opaque portion 6-207 of the shield 6-204. The sensors may include the same or similar sensors as those shown in the example of Figure 1I such as depth sensors 6-108 and 6-110, depth projectors 6-112, first and second scene cameras 6-106, first and second downward cameras 6-114, first and second side cameras 6-118, and first and second infrared illuminators 6-124. These sensors are also shown in Figure 1K and Figure 1L Examples. Other sensors, sensor types, sensor quantities, and their relative positions may be included in one or more other examples of the HMD.
[0131] Figure 1J Any of the features, components, and / or parts shown (including their arrangements and configurations) may be included, either alone or in any combination, in any of the examples of the devices, features, components, and parts shown in Figure 1I and Figures 1K to 1L shown and described herein. Similarly, any of the features, components, and / or parts shown or described with reference to Figure 1I and Figures 1K to 1L shown (including their arrangements and configurations) may be included, either alone or in any combination, in the examples of the devices, features, components, and parts shown in Figure 1J shown.
[0132] Figure 1K Illustrates a front view of a portion of an example of the HMD device 6-300, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. Figure 1K The example shown does not include a front cover or shield to illustrate the brackets 6-336, 6-338. For example, Figure 1J the shield 6-204 shown includes an opaque portion 6-207 that will visually cover / block the viewing of anything external (e.g., radially / peripherally external) to the display / display area 6-334, including the sensor 6-303 and the bracket 6-338.
[0133] In at least one example, the various sensors of the sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances on the angles relative to each other. For example, the tolerance on the mounting angle between two scene cameras 6-306 can be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, the scene cameras 6-306 can be mounted to bracket 6-338 instead of the shroud. The bracket can include a cantilever on which the scene cameras 6-306 and other sensors of the sensor system 6-302 can be mounted to maintain their position and orientation unchanged in the event of a drop event that causes any deformation of the other brackets 6-226, the housing 6-330, and / or the shroud by the user.
[0134] Figure 1K Any of the illustrated features, components, and / or parts (including their arrangement and configuration) can be included, either alone or in any combination, in Figures 1I to 1J and Figure 1L any of the other examples of the devices, features, components, and parts illustrated and described herein. Similarly, reference Figures 1I to 1J and Figure 1L any of the illustrated or described features, components, and / or parts (including their arrangement and configuration) can be included, either alone or in any combination, in Figure 1K the examples of the devices, features, components, and parts illustrated.
[0135] Figure 1L Illustrated is a bottom view of an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402. The sensor system 6-402 can be similar to other sensor systems described above and elsewhere herein, including reference Figures 1I to 1K described. In at least one example, the chin camera 6-416 can face downward to capture images of the user's lower facial features. In one example, the chin camera 6-416 can be directly coupled to the frame or housing 6-430 or one or more internal brackets that are directly coupled to the illustrated frame or housing 6-430. The frame or housing 6-430 can include one or more holes / openings 6-415 through which the chin camera 6-416 can send and receive signals.
[0136] Figure 1L Any of the illustrated features, components, and / or parts (including their arrangement and configuration) can be included, either alone or in any combination, in Figures 1I to 1K any of the other examples of the devices, features, components, and parts illustrated and described herein. Similarly, reference Figures 1I to 1KAny of the features, components, and / or parts shown and described (including their arrangements and configurations) may be included, either individually or in any combination, in Figure 1L the examples of the devices, features, components, and parts shown.
[0137] Figure 1M A rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 is illustrated. The IPD adjustment system includes first and second optical modules 11.1.1-104a-b that are slidably engaged / coupled to respective guide rods 11.1.1-108a-b and motors 11.1.1-110a-b of left and right adjustment subsystems 11.1.1-106a-b. The IPD adjustment system 11.1.1-102 may be coupled to a bracket 11.1.1-112 and includes buttons 11.1.1-114 that are in electrical communication with the motors 11.1.1-110a-b. In at least one example, the buttons 11.1.1-114 may be in electrical communication with the first and second motors 11.1.1-110a-b via a processor or other circuit components such that the first and second motors 11.1.1-110a-b are activated and cause the first and second optical modules 11.1.1-104a-b to change positions relative to each other.
[0138] In at least one example, the first and second optical modules 11.1.1-104a-b may include respective display screens that are configured to project light toward a user's eyes when the HMD 11.1.1-100 is worn. In at least one example, a user may manipulate (e.g., press and / or rotate) the buttons 11.1.1-114 to activate position adjustment of the optical modules 11.1.1-104a-b to match the user's interpupillary distance. The optical modules 11.1.1-104a-b may also include one or more cameras or other sensor / sensor systems for imaging and measuring the user's IPD such that the optical modules 11.1.1-104a-b may be adjusted to match the IPD.
[0139] In one example, the user may manipulate button 11.1.1-114 to cause an automatic position adjustment of the first and second optical modules 11.1.1-104a-b. In one example, the user may manipulate button 11.1.1-114 to cause a manual adjustment such that the optical modules 11.1.1-104a-b move further or closer (e.g., when the user rotates button 11.1.1-114 in one way or another) until the user visually matches her / his own IPD. In one example, the manual adjustment is communicated electronically via one or more circuits, and the power for moving the optical modules 11.1.1-104a-b via motors 11.1.1-110a-b is provided by a power source. In one example, the adjustment and movement of the optical modules 11.1.1-104a-b via manipulation of button 11.1.1-114 is mechanically actuated via movement of button 11.1.1-114.
[0140] Figure 1M Any of the illustrated features, components, and / or parts (including their arrangements and configurations) may be included, either alone or in any combination, in any example of the other devices, features, components, and parts shown in any other figures and described herein. Similarly, any of the features, components, and / or parts (including their arrangements and configurations) shown or described with reference to any other figure may be included, either alone or in any combination, in Figure 1M the example of the devices, features, components, and parts shown.
[0141] Figure 1N A front perspective view of a portion of the HMD 11.1.2-100 is illustrated, including an outer structural frame 11.1.2-102 and an inner or intermediate structural frame 11.1.2-104 that define a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. The apertures 11.1.2-106a-b are shown Figure 1N in dashed lines because the view of the apertures 11.1.2-106a-b may be blocked by one or more other components of the HMD 11.1.2-100 that are coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 that is coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second apertures 11.1.2-106a-b.
[0142] The mounting bracket 11.1.2-108 may include an intermediate or central portion 11.1.2-109 coupled to the internal frame 11.1.2-104. In some examples, the intermediate or central portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Instead, the intermediate / central portion 11.1.2-109 may be disposed between a first cantilevered extension arm and a second cantilevered extension arm that extend away from the intermediate portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilever 11.1.2-112 and a second cantilever 11.1.2-114 that extend away from the intermediate portion 11.1.2-109 of the mounting bracket 11.1.2-108 that is coupled to the internal frame 11.1.2-104.
[0143] As Figure 1N shown, the outer frame 11.1.2-102 may define a curved geometry on its lower side to accommodate the user's nose when the user wears the HMD 11.1.2-100. The curved geometry may be referred to as the nose bridge 11.1.2-111 and is centered on the lower side of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 may be connected to the internal frame 11.1.2-104 between holes 11.1.2-106a-b such that the cantilevers 11.1.2-112, 11.1.2-114 extend downward and laterally outward away from the intermediate portion 11.1.2-109 to be geometrically complementary to the nose bridge 11.1.2-111 geometry of the outer frame 11.1.2-102. In this way, the mounting bracket 11.1.2-108 is configured to accommodate the user's nose, as described above. The geometry of the nose bridge 11.1.2-111 accommodates the nose because the nose bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, over, and around.
[0144] The first cantilever 11.1.2-112 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a first direction, and the second cantilever 11.1.2-114 can extend away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 in a second direction opposite to the first direction. The first cantilever 11.1.2-112 and the second cantilever 11.1.2-114 are referred to as "cantilevered" or "cantilever" arms because each arm 11.1.2-112, 11.1.2-114 respectively includes free distal ends 11.1.2-116, 11.1.2-118 that are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 overhang from the middle portion 11.1.2-109, which can be connected to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are not attached.
[0145] In at least one example, the HMD 11.1.2-100 can include one or more components coupled to the mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each of the plurality of sensors 11.1.2-110a-f can include various types of sensors, including cameras, IR sensors, etc. In some examples, one or more of the sensors 11.1.2-110a-f can be used for object recognition in three-dimensional space, such that it is important to maintain the precise relative positions of two or more of the plurality of sensors 11.1.2-110a-f. The cantilevered nature of the mounting bracket 11.1.2-108 can protect the sensors 11.1.2-110a-f from damage and displacement in the event of an accidental drop by the user. Since the sensors 11.1.2-110a-f are cantilevered on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, the stress and deformation of the inner frame and / or the outer frame 11.1.2-104, 11.1.2-102 are not transmitted to the cantilevers 11.1.2-112, 11.1.2-114 and thus do not affect the relative positions of the sensors 11.1.2-110a-f coupled / installed to the mounting bracket 11.1.2-108.
[0146] Figure 1NAny of the illustrated features, components, and / or parts (including their arrangement and configuration) may be included, either alone or in any combination, in any of the other examples of the devices, features, components described herein. Similarly, any of the features, components, and / or parts (including their arrangement and configuration) shown and described herein may be included, either alone or in any combination, in Figure 1N the examples of the devices, features, components, and parts shown.
[0147] Figure 1O An example of an optical module 11.3.2-100 for use in an electronic device (such as an HMD, including the HDM devices described herein) is illustrated. As shown in one or more other examples described herein, the optical module 11.3.2-100 may be one of two optical modules within the HMD, where each optical module is aligned to project light towards the user's eyes. In this manner, the first optical module may project light towards the user's first eye via a display screen, and the second optical module of the same device may project light towards the user's second eye via another display screen.
[0148] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a barrel or an optical module barrel. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light towards the user's eyes when wearing the HMD to which the display module 11.3.2-100 belongs during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.
[0149] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to a housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to a display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eyes during use. In at least one example, the optical module 11.3.2-100 may further include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the cameras 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light emitting diodes (LEDs) or other lights configured to project light toward a user's eyes when wearing the HMD. Each of the lights 11.3.2-110 in the light strip 11.3.2-108 may be spaced apart around the light strip 11.3.2-108 and thus may be spaced apart evenly or unevenly around the display 11.3.2-104 at various locations on the light strip 11.3.2-108 and around the display 11.3.2-104.
[0150] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user may view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light through the viewing opening 11.3.2-101 onto the user's eyes. In one example, the cameras 11.3.2-106 are configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.
[0151] As described above, Figure 1O each of the components and features of the illustrated optical module 11.3.2-100 may be replicated in another (e.g., second) optical module provided with the HMD to interact with the user's other eye (e.g., project light and capture images).
[0152] Figure 1O Any of the illustrated features, components, and / or parts (including their arrangement and configuration) may be included, alone or in any combination, in Figure 1P any example of other devices, features, components, and parts shown or otherwise described herein. Similarly, reference Figure 1P to or any of the features, components, and / or parts shown or otherwise described herein (including their arrangement and configuration) may be included, alone or in any combination, in Figure 1OIn the examples of the devices, features, components, and parts shown.
[0153] Figure 1P An example cross-sectional view of an example of an optical module 11.3.2-200 is illustrated, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first hole or passage 11.3.2-212 and a second hole or passage 11.3.2-214. The passages 11.3.2-212, 11.3.2-214 may be configured to slidably engage corresponding tracks or guide rods of the HMD device to allow the optical module 11.3.2-200 to adjust its position relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 is capable of slidably engaging the guide rods to fix the optical module 11.3.2-200 in place within the HMD.
[0154] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and disposed between the display assembly 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above a light bar 11.3.2-208 and one or more eye tracking cameras 11.3.2-206 such that the cameras 11.3.2-206 are configured to capture images of the user's eyes through the lens 11.3.2-216, and the light bar 11.3.2-208 includes lights configured to project light through the lens 11.3.2-216 onto the user's eyes during use.
[0155] Figure 1P Any of the features, components, and / or parts shown (including their arrangement and configuration) may be included, either individually or in any combination, in any of the other examples of the devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including their arrangement and configuration) may be included, either individually or in any combination, in Figure 1P the examples of the devices, features, components, and parts shown.
[0156] Figure 2FIG. 0 is a block diagram of an example of a controller 110 in accordance with some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features have not been shown for the sake of brevity and in order not to obscure more relevant aspects of the embodiments disclosed herein. To that end, by way of non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these components and various other components.
[0157] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.
[0158] The memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, the memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. The memory 220 includes non-transitory computer-readable storage medium. In some embodiments, the memory 220 or the non-transitory computer-readable storage medium of the memory 220 stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and an XR experience module 240.
[0159] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for respective groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0160] In some embodiments, the data acquisition unit 241 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least Figure 1A the display generation component 120, and optionally from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0161] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track the positioning / location of at least the display generation component 120 relative to Figure 1A the scene 105, and optionally the positioning / location relative to one or more of the tracking input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the positioning / location of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to Figure 1A the scene 105, relative to the display generation component 120, and / or relative to a coordinate system (which is defined relative to the user's hand). The hand tracking unit 244 is described in more detail below with respect to Figure 4 . In some embodiments, the eye tracking unit 243 is configured to track the positioning or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to the XR content displayed via the display generation component 120. The eye tracking unit 243 is described in more detail below with respect to Figure 5 .
[0162] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120 and, optionally, by one or more of the output device 155 and / or the peripheral device 195. For this purpose, in various embodiments, the coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for the heuristics.
[0163] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120 and, optionally, to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For this purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic for instructions, as well as heuristics and metadata for the heuristics.
[0164] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 may be located in separate computing devices.
[0165] In addition, Figure 2 This is more of a functional description of the various features that may be present in a particular implementation, as opposed to the structural schematic of the embodiments described herein. As will be recognized by those of ordinary skill in the art, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions, as well as how the features are allocated therein, will vary depending on the particular implementation and, in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for the particular implementation.
[0166] Figure 3FIG. 0 is a block diagram of an example of a display generation component 120 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that, for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein, various other features are not shown. For that purpose, by way of non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, an ASIC, an FPGA, a GPU, a CPU, a processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal and / or external image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these components and various other components.
[0167] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), and the like.
[0168] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to a user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffractive, reflective, polarization, holographic, and other waveguide displays. For example, the display generation component 120 (e.g., an HMD) includes a single XR display. In another example, the display generation component 120 includes XR displays for each eye of the user. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting MR or VR content.
[0169] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to a scene that the user would see in the absence of the display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.
[0170] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage media. In some embodiments, memory 320 or the non-transitory computer-readable storage media of memory 320 stores the following programs, modules, and data structures or subsets thereof, including optionally operating system 330 and XR rendering module 340.
[0171] Operating system 330 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, XR rendering module 340 is configured to present XR content to a user via one or more XR displays 312. To that end, in various embodiments, XR rendering module 340 includes data acquisition unit 342, XR rendering unit 344, XR mapping generation unit 346, and data transmission unit 348.
[0172] In some embodiments, data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) at least from Figure 1A controller 110. For that purpose, in various embodiments, data acquisition unit 342 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0173] In some embodiments, XR rendering unit 344 is configured to present XR content via one or more XR displays 312. To that end, in various embodiments, XR rendering unit 344 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0174] In some embodiments, XR mapping generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate extended reality) based on media content data. To that end, in various embodiments, XR mapping generation unit 346 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0175] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0176] Although the data acquisition unit 342, XR presentation unit 344, XR mapping generation unit 346, and data transmission unit 348 are shown as residing on a single device (e.g., Figure 1A the display generation component 120), it should be understood that in other embodiments, any combination of the data acquisition unit 342, XR presentation unit 344, XR mapping generation unit 346, and data transmission unit 348 may be located in separate computing devices.
[0177] Furthermore, Figure 3 this is more of a functional description of the various features that may be present in a particular implementation, as opposed to a schematic diagram of the structure of the embodiments described herein. As will be appreciated by those of ordinary skill in the art, items shown separately may be combined and some items may be separated. For example, Figure 3 some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific partitioning of functions and how the features are allocated therein will vary depending on the implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for the particular implementation.
[0178] Figure 4 is a schematic illustration of an example embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the positioning / position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to the scene 105 of FIG. 1 (e.g., relative to a portion of the physical environment around the user, relative to the display generation component 120, or relative to a part of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0179] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information including at least the hand 406 of a human user. The image sensor 404 captures hand images at a sufficient resolution such that the fingers and their corresponding positions can be distinguished. The image sensor 404 typically captures images of other parts of the user's body, and may also or possibly capture images of all parts of the body, and may have a zoom capability or a dedicated sensor with increased magnification to capture images of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in combination with other image sensors to capture the physical environment of the scene 105, or serves as the image sensor for capturing the physical environment of the scene 105. In some embodiments, the image sensor 404 is positioned relative to the user or the user's environment in such a way that the field of view of the image sensor 404 or a portion thereof is used to define an interaction space, in which hand movements captured by the image sensor are considered inputs to the controller 110.
[0180] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D map data (and in addition, possibly color image data) to the controller 110, which extracts high-level information from the map data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), and the application drives the display generation component 120 accordingly. For example, a user can interact with software running on the controller 110 by moving his hand 406 and changing his hand pose.
[0181] In some embodiments, the image sensor 404 projects a speckle pattern onto the scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offset of the speckles in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. This method gives the depth coordinates of points in the scene relative to a pre-determined reference plane at a specific distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis, such that the depth coordinates of points in the scene correspond to the z-component measured by the image sensor. Alternatively, the image sensor 404 (e.g., the hand tracking device) may use other 3D mapping methods, such as stereoscopy or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.
[0182] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps of the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in the image sensor 404 and / or the controller 110 processes the 3D map data to extract image patch descriptors of the hand in these depth maps. The software can match these descriptors with image patch descriptors stored in the database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose generally includes the 3D positions of the user's hand joints and finger tips.
[0183] The software can also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify gestures. The pose estimation function described herein can alternate with the motion tracking function such that image patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to find changes in the pose that occur in the remaining frames. Pose, motion, and gesture information are provided to an application running on the controller 110 via the aforementioned API. The program can, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.
[0184] In some embodiments, gestures include air gestures. An air gesture is detected when the user does not touch an input element (or independent of an input element that is part of a device (e.g., the computer system 101, one or more input devices 125, and / or the hand tracking device 140)) and is based on the detected movement of a part of the user's body (e.g., the head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the other user's hand, and / or the movement of the user's finger relative to another finger or part of the hand), and / or the absolute movement of a part of the user's body (e.g., a tap gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture that includes a predetermined speed or amount of rotation of a part of the user's body)).
[0185] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures performed by the movement of a user's finger relative to other fingers (or a part of the user's hand) for interacting with an XR environment (e.g., a virtual or mixed reality environment). In some embodiments, the air gesture is detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and is based on the detected movement of a part of the user's body through the air (including the movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), the movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one hand of the user relative to the other hand of the user, and / or the movement of the user's finger relative to another finger or a part of the hand of the user), and / or the absolute movement of a part of the user's body (e.g., a tap gesture including the hand moving a predetermined amount and / or speed in a predetermined pose, or a shake gesture including a predetermined speed or rotation amount of a part of the user's body)).
[0186] In some embodiments where the input gesture is an air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or touchpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in a specific implementation involving an air gesture, for example, the input gesture is detected in combination with (e.g., simultaneously) the movement of the user's finger and / or hand towards the user interface element along with the attention (e.g., gaze) to perform a pinch and / or tap input, as described below.
[0187] In some embodiments, an input gesture directed to a user interface object is performed with direct or indirect reference to the user interface object. For example, user input is performed directly on the user interface object by performing the input at a location corresponding to the positioning of the user's hand relative to the positioning of the user interface object in a three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, when attention (e.g., gaze) of the user to the user interface object is detected, the input gesture is performed indirectly on the user interface object in accordance with the positioning of the user's hand not being at the location corresponding to the positioning of the user interface object in the three-dimensional environment while the user performs the input gesture. For example, for a direct input gesture, the user can direct the user's input to the user interface object by initiating a gesture at or near a location corresponding to the display positioning of the user interface object (e.g., within a distance of 0.5 cm, 1 cm, 5 cm, or between 0 and 5 cm measured from the outer edge of the option or the central portion of the option). For an indirect input gesture, the user can direct the user's input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location not corresponding to the display positioning of the user interface object).
[0188] In some embodiments, according to some embodiments, input gestures (e.g., air gestures) used in the various examples and embodiments described herein include pinch inputs and tap inputs for interacting with a virtual or mixed reality environment. For example, the pinch inputs and tap inputs described below are performed as air gestures.
[0189] In some embodiments, a pinch input is part of an air gesture that includes one or more of the following: a pinch gesture, a long pinch gesture, a pinch-and-drag gesture, or a double-pinch gesture. For example, a pinch gesture as an air gesture includes the movement of two or more fingers of a hand to contact each other, i.e., optionally followed by an immediate (e.g., within 0 seconds to 1 second) interruption of contact with each other. A long pinch gesture as an air gesture includes the movement of two or more fingers of a hand contacting each other for at least a threshold amount of time (e.g., at least 1 second) before detecting an interruption of contact with each other. For example, a long pinch gesture includes a user maintaining a pinch gesture (e.g., where two or more fingers are in contact), and the long pinch gesture continues until an interruption of contact between the two or more fingers is detected. In some embodiments, a double-pinch gesture as an air gesture includes two (e.g., or more) pinch inputs (e.g., performed by the same hand) detected consecutively and immediately (e.g., within a predefined time period) with each other. For example, a user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.
[0190] In some embodiments, a pinching and dragging gesture as an air gesture includes a pinching gesture (e.g., a pinching gesture or a long - pinching gesture) performed in combination with (e.g., following) a dragging input that changes the positioning of the user's hand from a first positioning (e.g., the starting positioning of the drag) to a second positioning (e.g., the ending positioning of the drag). In some embodiments, the user maintains the pinching gesture while performing the dragging input and releases the pinching gesture (e.g., opens two or more of their fingers) to end the dragging gesture (e.g., at the second positioning). In some embodiments, the pinching input and the dragging input are performed by the same hand (e.g., the user pinches two or more fingers together and moves the same hand to a second positioning in the air using the dragging gesture). In some embodiments, the pinching input is performed by the user's first hand and the dragging input is performed by the user's second hand (e.g., while the user continues the pinching input with the user's first hand, the user's second hand moves from a first positioning to a second positioning in the air). In some embodiments, an input gesture as an air gesture includes an input performed using both of the user's hands (e.g., a pinching and / or tapping input). For example, the input gesture includes two (e.g., or more) pinching inputs performed in combination with each other (e.g., concurrently or within a predefined time period). For example, a first pinching gesture (e.g., a pinching input, a long - pinching input, or a pinching and dragging input) is performed using the user's first hand, and in combination with performing the pinching input with the first hand, a second pinching input is performed using the other hand (e.g., the second of the user's two hands). In some embodiments, there is movement between the two hands of the user (e.g., increasing and / or decreasing the distance or relative orientation between the two hands of the user).
[0191] In some embodiments, a tapping input performed as an air gesture (e.g., pointing to a user interface element) includes the movement of the user's finger towards the user interface element, the movement of the user's hand towards the user interface element (optionally, the user's finger extends towards the user interface element), the downward movement of the user's finger (e.g., mimicking a mouse click movement or a tap on a touch screen), or other predefined movements of the user's hand. In some embodiments, the tapping input performed as an air gesture is detected based on the movement characteristics of the finger or hand performing the tapping gesture movement, which is the movement of the finger or hand away from the user's viewpoint and / or towards an object that is the target of the tapping input, followed by the end of the movement. In some embodiments, the end of the movement is detected based on a change in the movement characteristics of the finger or hand performing the tapping gesture (e.g., the end of the movement away from the user's viewpoint and / or towards the object that is the target of the tapping input, the reversal of the movement direction of the finger or hand, and / or the reversal of the acceleration direction of the movement of the finger or hand).
[0192] In some embodiments, the determination of the user's attention being directed to a portion of a three-dimensional environment is based on the detection of a gaze directed to the portion of the three-dimensional environment (optionally, without the need for other conditions). In some embodiments, the determination of the user's attention being directed to a portion of a three-dimensional environment is based on the detection of a gaze directed to the portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed to the portion of the three-dimensional environment for at least a threshold duration (e.g., dwell duration) and / or requiring the gaze to be directed to the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, such that the device determines that the user's attention is directed to the portion of the three-dimensional environment, where if one of these additional conditions is not met, the device determines that the attention is not directed to the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).
[0193] In some embodiments, the detection of the readiness state configuration of the user or a part of the user is detected by a computer system. The detection of the readiness state configuration of the hand is used by the computer system as an indication that the user may be about to use one or more air gesture inputs (e.g., pinch, tap, pinch and drag, double pinch, long pinch, or other air gestures described herein) performed by the hand to interact with the computer system. For example, based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape where the thumb and one or more fingers are extended and spaced apart to prepare for a pinch or grab gesture, or a pre-tap where one or more fingers are extended and the palm is facing away from the user), based on whether the hand is in a predetermined positioning relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., moving towards an area in front of the user that is above the user's waist and below the user's head or moving away from the user's body or legs) to determine the readiness state of the hand. In some embodiments, the readiness state is used to determine whether interactive elements of the user interface respond to attention (e.g., gaze) input.
[0194] In a scenario where input is described with reference to air gestures, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect similar gestures, where optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units can be used to track the positioning of the hardware input device in space, and the positioning and / or movement of the hardware input device is used to replace the positioning and / or movement of one or more hands in the corresponding air gesture. In a scenario where input is described with reference to air gestures, it should be understood that a hardware input device attached to one or more of the user's hands or held by one or more of the user's hands can be used to detect similar gestures, and user input can be detected using the controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the positioning or change in positioning of parts of the hand and / or fingers relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, where user input using the controls contained in the hardware input device is used to replace hand and / or finger gestures such as an air tap or an air pinch in the corresponding air gesture. For example, a selection input described as being performed using an air tap or an air pinch input can alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. As another example, a movement input described as being performed using an air pinch and drag can alternatively be detected based on an interaction with a hardware input control (such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input after the movement of the hardware input device (e.g., along with the hand associated with the hardware input device) through space). Similarly, a two-handed input that includes the movement of the hands relative to each other can be performed using an air gesture and a hardware input device in a hand that is not performing an air gesture, two hardware input devices held in different hands, or two air gestures performed using various combinations of an air gesture and / or input detected by one or more of the aforementioned hardware input devices.
[0195] In some embodiments, the software can be downloaded electronically, for example, over a network, to the controller 110, or can alternatively be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is similarly stored in the memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer can be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although in Figure 4The controller 110 is shown, but for example, some or all of the processing functions of the controller, as a unit separate from the image sensor 404, may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other devices associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, a handheld device, or a head-mounted device) or integrated with any other suitable computerized device (such as a game console or a media player). The sensing function of the image sensor 404 may likewise be integrated into a computer or other computerized device to be controlled by the sensor output.
[0196] Figure 4 Also illustrated is a schematic diagram of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels having corresponding depth values. The pixels 412 corresponding to the hand 406 have been segmented from the background and the wrist in the figure. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z distance from the image sensor 404), where the gray shading becomes darker as the depth increases. The controller 110 processes these depth values to identify and segment the components of the image having human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, the overall size, shape, and movement from frame to frame in a sequence of depth maps.
[0197] Figure 4 Also schematically illustrated is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406. In Figure 4 , the hand skeleton 414 is superimposed on the hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points on the hand and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, palm center, the end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the positions and movements of these key feature points across multiple image frames to determine, according to some embodiments, the gesture performed by the hand or the current state of the hand.
[0198] Figure 5 An example embodiment of an eye tracking device 130 ( Figure 1A ) is illustrated. In some embodiments, the eye tracking device 130 is composed of an eye tracking unit 243 ( Figure 2)Control is used to track the positioning and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as, a head-mounted headset, helmet, goggles, or glasses) or a hand-held device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a hand-held device or an XR room, the eye tracking device 130 is optionally a device separate from the hand-held device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.
[0199] In some embodiments, the display generation component 120 uses display mechanisms (such as, a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, a head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display, and virtual objects are displayed on the transparent or translucent display through which the user can directly view the physical environment. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may be projected, for example, onto a physical surface or projected as a hologram such that an individual uses the system to observe the virtual objects superimposed over the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.
[0200] As Figure 5As shown, in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be directed at the user's eyes to receive IR or NIR light directly reflected from the eyes by the light source, or alternatively can be directed at a "hot" mirror located between the user's eyes and the display panel, which reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60 frames per second - 120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by corresponding eye tracking cameras and illumination sources.
[0201] In some embodiments, a device-specific calibration process is used to calibrate the eye tracking device 130 to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, cameras, hot mirrors (if present), eye lenses, and display screens. The device-specific calibration process can be performed at the factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process can include an estimation of the eye parameters of a particular user, such as pupil position, fovea position, optical axis, visual axis, interpupillary distance, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's fixation point relative to the display.
[0202] As Figure 5As shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is to be performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 may be directed at a mirror 550 (which reflects IR or NIR light from the eye 592 while allowing visible light to pass through) (e.g., as shown in the top portion of Figure 5 ), or alternatively may be directed at the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the bottom portion of Figure 5 ).
[0203] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for a left display panel and a right display panel) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash assist method or other suitable method. The gaze point estimated from the gaze tracking input 542 is optionally used to determine the direction in which the user is currently looking.
[0204] The following describes several possible use cases of the current gaze direction of a user and is not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal region determined according to the current gaze direction of the user than in the peripheral region. As another example, the controller may position or move virtual content in the view at least in part based on the current gaze direction of the user. As yet another example, the controller may display specific virtual content in the view at least in part based on the current gaze direction of the user. As another example use case in an AR application, the controller 110 may direct an external camera for capturing the physical environment of the XR experience to focus in the determined direction. Then, the autofocus mechanism of the external camera may focus on an object or surface in the environment that the user is currently looking at on the display 510. As another example use case, the eye lens 520 may be a focusable lens, and the controller uses the gaze tracking information to adjust the focus of the eye lens 520 such that the virtual object that the user is currently looking at has an appropriate vergence to match the convergence of the user's eyes 592. The controller 110 may utilize the gaze tracking information to direct the eye lens 520 to adjust the focus such that a nearby object that the user is looking at appears at the correct distance.
[0205] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510) mounted in a wearable housing, two eye lenses (e.g., eye lenses 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) towards the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each of the lenses, as Figure 5 shown. In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.
[0206] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and thus does not introduce noise in the gaze tracking system. Note that the position and angle of the eye tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0207] As Figure 5 shown, embodiments of the gaze tracking system may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a computer-generated reality, virtual reality, augmented reality, and / or augmented virtual experience to a user.
[0208] Figure 6 Illustrates a flash-assisted gaze tracking pipeline according to some embodiments. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., the eye tracking device 130 as Figure 1A and Figure 5 shown). The flash-assisted gaze tracking system may maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, when analyzing the current frame to track the pupil contour and flash in the current frame, the flash-assisted gaze tracking system uses previous information from the previous frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues with the next frame in the tracking state.
[0209] As Figure 6 shown, the gaze tracking camera may capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system may continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images may be input into the pipeline for processing. However, in some embodiments or under some conditions, not all of the captured frames are processed by the pipeline.
[0210] At 610, for the currently captured image, if the tracking state is yes, the method proceeds to element 640. At 610, if the tracking state is no, then as indicated at 620, the image is analyzed to detect the user's pupil and flash in the image. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.
[0211] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and flash based in part on previous information from a previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and flash detected in the current frame. The processing result at element 640 is checked to verify that the result of the tracking or detection can be trusted. For example, the result can be checked to determine whether the pupil and a sufficient number of flashes for performing gaze estimation are successfully tracked or detected in the current frame. At 650, if the result cannot be trusted, then at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trusted, the method proceeds to element 670. At 670, the tracking state is set to yes (if it is not already yes), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.
[0212] Figure 6 It is intended to be used as an example of an eye tracking technique that can be used for a particular specific implementation. As will be appreciated by those of ordinary skill in the art, according to various embodiments, in the computer system 101 for providing an XR experience to a user, other eye tracking techniques that currently exist or are developed in the future can be used to replace the flash-assisted eye tracking technique described herein or used in combination with the flash-assisted eye tracking technique.
[0213] In this disclosure, various input methods are described with respect to interaction with a computer system. When one input device or input method is used to provide an example and another input device or input method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. When one output device or output method is used to provide an example and another output device or output method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment via a computer system. When interaction with a virtual environment is used to provide an example and a mixed reality environment is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the methods described with respect to the other example. Accordingly, this disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each example embodiment.
[0214] User interface and associated processes
[0215] Attention is now focused on embodiments of a user interface (“UI”) and associated processes that can be implemented on a computer system that communicates with a display generation component, one or more input devices, and optionally one or more physical controls, such as a portable multifunctional device or a head-mounted device.
[0216] Figures 7A to 7K Examples of techniques for navigating an extended reality experience are illustrated. Figure 8 is a flowchart of an exemplary method 800 for navigating an extended reality experience. Figure 9 is a flowchart of an exemplary method 900 for navigating an extended reality experience. Figures 7A to 7K The user interface in is used to illustrate the processes described below, including Figure 8 and Figure 9 the processes in.
[0217] Figure 7Adepicts an electronic device 700, which is a smart phone including a touch-sensitive display 702, buttons 704a - 704c, and one or more input sensors 706 (e.g., one or more cameras, an eye gaze tracker, a hand movement tracker, and / or a head movement tracker). In some embodiments described below, the electronic device 700 is a smart phone. In some embodiments, the electronic device 700 is a tablet computer, a wearable device, a wearable smart watch device, a head-mounted system (e.g., a head-mounted headset), or other computer system including one or more display devices (e.g., a display screen, a projection device, etc.) and / or communicating with the one or more display devices. The electronic device 700 is a computer system (e.g., Figure 1A the computer system 101 in
[0218] At Figure 7A , the electronic device 700 is in a low-power, inactive, or sleep state, where content is not displayed via the display 702. At Figure 7A , the electronic device 700 detects a user input 708. In the depicted embodiment, the user input 708 is a button press input via the button 704c. However, in some embodiments, the user input 708 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 708 includes, for example, the user placing the electronic device 700 on his or her head, performing a gesture while wearing the electronic device 700, pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a specific manner), and / or any combination of the foregoing.
[0219] At Figure 7B , in response to the user input 708, the electronic device 700 transitions from the low-power, inactive, or sleep state to an active state, where the electronic device 700 displays a three-dimensional environment 712 and an extended reality experience 714 (e.g., an augmented reality experience and / or a virtual reality experience) via the display 702. In the depicted scenario, the three-dimensional environment 712 includes a chair, a table, and a place setting (e.g., a napkin, a fork, a knife, and a cup) placed on the table. In some embodiments, the three-dimensional environment 712 is displayed by the display (as Figure 7B depicted). In some embodiments, the three-dimensional environment 712 includes a virtual environment or is captured by one or more cameras (e.g., one or more cameras as part of the input sensors 706 and / or Figure 7BImages (or videos) of the physical environment captured by one or more cameras (not shown in the figure). In some embodiments, the three-dimensional environment 712 is visible to the user behind the extended reality experience 714 but is not displayed by the display. For example, in some embodiments, the three-dimensional environment 712 is a physical environment that is visible to the user behind the extended reality experience 712 (e.g., through a transparent display) but is not displayed by the display.
[0220] In Figure 7B , the extended reality experience 714 is a camera extended reality experience, as indicated by the identifier 716a, which includes the logo of the camera and the name of the extended reality experience. The camera extended reality experience 714 includes one or more selectable objects 716b-e that can be selected by the user to capture photos and / or video content via one or more cameras (e.g., one or more cameras that are part of the input sensor 706 and / or Figure 7B one or more cameras not shown in the figure). Object 716b is a shutter button that can be selected to capture photos and / or videos. Object 716c can be selected to enable the slow-motion capture mode. Option 716d can be selected to enable the photo capture mode. Option 716e can be selected to enable the video capture mode. In Figure 7B , the electronic device 700 detects that the user is looking to the right of the display 702, as indicated by the gaze indication 710. The gaze indication 710 is provided to better understand the described technology and is optionally not part of the user interface of the described device (e.g., not displayed by the electronic device 700). In Figure 7B , the electronic device 700 detects a user input 718. In the depicted embodiment, the user input 718 is a button press input via the button 704c. However, in some embodiments, the user input 718 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 718 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing the electronic device 700, pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a specific manner), and / or any combination of the foregoing.
[0221] In Figure 7C , in response to the user input 718, the electronic device 700 displays an animation in which the extended reality experience 714 appears to move away from the user. In Figure 7CIn [description], the electronic device 700 displays a representation 720, which represents a camera augmented reality experience 714. The representation 720 includes objects 722a - 722e, which represent objects 716a - 716e (e.g., these objects are smaller non - interactive versions of objects 716a - 716e) that are superimposed on a background portion 722f and surrounded by a boundary 719. The representation 720 appears to move away from the user, for example, by gradually getting smaller over time. In some embodiments, in response to user input 718 and / or when displaying an animation, the three - dimensional environment 712 is visually blurred (as indicated by the dashed lines in Figure 7C ), for example, displayed with reduced focus, reduced sharpness, reduced color saturation, and / or greater opacity) in order to attract the user's attention and gaze at the representation 720. As discussed above, in some embodiments, the three - dimensional environment 712 is a "see - through" environment that the user sees through a transparent display and is not displayed by the display. In some such embodiments, the three - dimensional environment 712 is visually de - emphasized by applying masking or other techniques to an area of the display (e.g., display 702) through which the user can view the three - dimensional environment 712.
[0222] In Figure 7D1At this point, the animation representing 720 (and / or the extended reality experience 714) that appears to move away from the user is complete, and the representation 720 is now shown at the top of the stack of representations 721, 724, 726. The representations 721, 724, and / or 726 represent other extended reality experiences that can be selected by the user and / or displayed by the electronic device 700. For example, as discussed above, the representation 720 represents a camera extended reality experience (e.g., the camera extended reality experience 714). In some embodiments, the representation 724 represents a music extended reality experience (e.g., which includes one or more selectable options for playing music), the representation 726 represents a translation extended reality experience (e.g., which includes one or more selectable options for translating content (e.g., content captured by one or more cameras and / or within the user's field of view and / or the field of view of the electronic device 700)), and the representation 721 includes, for example, a representation of a reading extended reality experience, a representation of a photo gallery extended reality experience, a representation of a video messaging extended reality experience, a representation of a navigation extended reality experience, and / or a representation of a fitness extended reality experience. As will be shown in subsequent figures, the user can scroll through the stack of representations 720, 721, 724, and / or 726 to select which extended reality experience the user wants to display. In some embodiments, each extended reality experience corresponds to a different color, and the representation corresponding to the extended reality experience is displayed in the corresponding color corresponding to the extended reality experience. For example, in some embodiments, the camera extended reality experience 714 corresponds to a first color, and the representation 720 is displayed in the first color (e.g., the background 722f is displayed in the first color, the border 720 is displayed in the first color, and / or the object 722a (e.g., logo and / or name) is displayed in the first color); and the music extended reality experience corresponds to a second color, such that the representation 724 is displayed in the second color (e.g., the background portion of the representation 724, the border of the representation 724, and / or the identifier of the representation 724 is displayed in the second color). In this way, the user can quickly identify the order of the extended reality experience stack based on the colors of the representations 720, 721, 724, and / or 726.
[0223] At Figure 7D1 this point, the representation 720 is shown at a first display position (e.g., at the top of the stack), indicating that a selection input (e.g., a press of the button 704c or other selection input) will result in the extended reality experience corresponding to the representation 720 being displayed (e.g., will result in the camera extended reality experience 714 being displayed). At Figure 7D1 this point, the electronic device 700 detects the user input 727. At Figure 7D1In this case, the user input 727 is a button press of button 704a. In some embodiments, a button press of button 704a indicates a request to navigate and / or scroll in a first direction (e.g., rotate the stack forward), and a button press of button 704b indicates a request to navigate and / or scroll in a second direction (e.g., rotate the stack backward). Additionally, in some embodiments, the user input 727 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 727 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing the electronic device 700, pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular manner), and / or any combination of the foregoing. For example, in some embodiments, a rotation of the rotatable input mechanism in a first direction (e.g., a clockwise rotation) (in some embodiments, a rotation of the rotatable input mechanism while looking at the stack) indicates a request to navigate and / or scroll in a second direction (e.g., a counterclockwise rotation), and a rotation of the rotatable input mechanism in a third direction (in some embodiments, a rotation of the rotatable input mechanism while looking at the stack) indicates a request to navigate and / or scroll in a fourth direction (e.g., rotate the stack backward).
[0224] In some embodiments, Figures 7A to 7K the techniques and user interfaces described in Figures 1A to 1P are provided by one or more of the devices described in Figures 7D2 to 7D4 For example, Figures 7B to 7D1 illustrates an embodiment in which the transition animation described in
[0225] In Figure 7D2 is displayed on the display module X702 of the head-mounted device (HMD) X700. In some embodiments, the device X700 includes a pair of display modules that provide stereoscopic content to different eyes of the same user. For example, the HMD X700 includes a display module X702 that provides content to the user's left eye and a second display module that provides content to the user's right eye. In some embodiments, the second display module displays an image that is slightly different from the display module X702 to create an illusion of stereoscopic depth.
[0225] In Figure 7D2 the extended reality experience 714 is a camera extended reality experience, as indicated by the identifier 716a, which includes a logo of the camera and the name of the extended reality experience. The camera extended reality experience 714 includes options that can be selected by the user to be captured via one or more cameras (e.g., one or more cameras that are part of the input sensor X706 and / or Figure 7D2One or more optional objects 716b-e for capturing photo and / or video content by one or more cameras (not shown in the figure). Object 716b is a shutter button that can be selected to capture a photo and / or video. Object 716c can be selected to enable the slow motion capture mode. Option 716d can be selected to enable the photo capture mode. Option 716e can be selected to enable the video capture mode. In Figure 7D2 In, the HMD X700 detects that the user is looking to the right of the display module X702, as indicated by the gaze indication 710. The gaze indication 710 is provided for better understanding of the described technology and is optionally not part of the user interface of the described device (e.g., not displayed by the HMD X700). In Figure 7D2 At, the HMD X700 detects a user input 718. In the depicted embodiment, the user input 718 is a button press input via the button X704c. However, in some embodiments, the user input 718 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, the user input 718 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing the HMD X700, pressing a button while wearing the HMD X700, rotating a rotatable input mechanism while wearing the HMD X700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a specific manner), and / or any combination of the foregoing.
[0226] In Figure 7D3 In response to the user input 718, the HMD X700 displays an animation in which the extended reality experience 714 appears to move away from the user. In Figure 7D3 In, the HMD X700 displays a representation 720 that represents the camera extended reality experience 714. The representation 720 includes objects 722a-722e that represent the objects 716a-716e (e.g., these objects are smaller non-interactive versions of the objects 716a-716e) superimposed on a background portion 722f and surrounded by a boundary 719. The representation 720 appears to move away from the user, for example, by gradually getting smaller over time. In some embodiments, in response to the user input 718 and / or while the animation is being displayed, the three-dimensional environment 712 is visually blurred (as Figure 7D3(indicated by the dashed line in) (e.g., displayed with reduced focus, reduced sharpness, reduced color saturation, and / or greater opacity) to attract the user's attention and gaze at the representation 720. As discussed above, in some embodiments, the three-dimensional environment 712 is a "see-through" environment that the user sees through the transparent display and is not displayed by the display. In some such embodiments, the three-dimensional environment 712 is visually de-emphasized by applying masking or other techniques to an area of the display (e.g., the display module X702) through which the user can view the three-dimensional environment 712).
[0227] At Figure 7D4 the animation of the representation 720 (and / or the extended reality experience 714) that appears to move away from the user is completed, and the representation 720 is now displayed on top of the stack of representations 721, 724, 726. The representations 721, 724, and / or 726 represent other extended reality experiences that can be selected by the user and / or displayed by the electronic device 700. For example, as discussed above, the representation 720 represents a camera extended reality experience (e.g., the camera extended reality experience 714). In some embodiments, the representation 724 represents a music extended reality experience (e.g., which includes one or more selectable options for playing music), the representation 726 represents a translation extended reality experience (e.g., which includes one or more selectable options for translating content (e.g., content captured by one or more cameras and / or content within the user's field of view and / or the field of view of the electronic device 700)), and the representation 721 includes, for example, a representation of a reading extended reality experience, a representation of a photo gallery extended reality experience, a representation of a video messaging extended reality experience, a representation of a navigation extended reality experience, and / or a representation of a fitness extended reality experience. As will be shown in subsequent figures, the user can scroll through the stack of representations 720, 721, 724, and / or 726 to select which extended reality experience the user wants to display. In some embodiments, each extended reality experience corresponds to a different color, and the representation corresponding to the extended reality experience is displayed in the corresponding color of the extended reality experience. For example, in some embodiments, the camera extended reality experience 714 corresponds to a first color, and the representation 720 is displayed in the first color (e.g., the background 722f is displayed in the first color, the border 720 is displayed in the first color, and / or the object 722a (e.g., logo and / or name) is displayed in the first color); and the music extended reality experience corresponds to a second color, such that the representation 724 is displayed in the second color (e.g., the background portion of the representation 724, the border of the representation 724, and / or the identifier of the representation 724 is displayed in the second color). In this way, the user can quickly identify the order of the extended reality experience stack based on the colors of the representations 720, 721, 724, and / or 726.
[0228] AtFigure 7D4 In, the representation 720 is shown at the first display position (e.g., at the top of the stack), indicating that a selection input (e.g., a press of button 704c or other selection input) will result in an extended reality experience corresponding to the representation 720 being displayed (e.g., will result in the camera extended reality experience 714 being displayed). In Figure 7D4 , the HMD X700 detects the user input 727. In Figure 7D4 , the user input 727 is a button press of button X704a. In some embodiments, a button press of button X704a indicates a request to navigate and / or scroll in a first direction (e.g., rotate the stack forward), and a button press of button X704b indicates a request to navigate and / or scroll in a second direction (e.g., rotate the stack backward). Additionally, in some embodiments, the user input 727 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, the user input 727 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing the HMD X700, pressing a button while wearing the HMD X700, rotating a rotatable input mechanism while wearing the HMD X700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular manner), and / or any combination of the foregoing. For example, in some embodiments, a rotation of the rotatable input mechanism in a first direction (e.g., a clockwise rotation) (in some embodiments, a rotation of the rotatable input mechanism while looking at the stack) indicates a request to navigate and / or scroll in a second direction (e.g., a counterclockwise rotation), and a rotation of the rotatable input mechanism in a third direction (in some embodiments, a rotation of the rotatable input mechanism while looking at the stack) indicates a request to navigate and / or scroll in a fourth direction (e.g., rotate the stack backward).
[0229] Figures 1B to 1PAny of the illustrated features, components, and / or parts (including their arrangements and configurations) may be included in the HMD X700 either individually or in any combination. For example, in some embodiments, the HMD X700 includes either individually or in any combination any of the features, components, and / or parts of HMDs 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100. In some embodiments, the display module X702 includes either individually or in any combination any of the features, components, and / or parts of display unit 1-102, display unit 1-202, display unit 1-306, display unit 1-406, display generation component 120, display screens 1-122a-b, first rear display screen 1-322a, and second rear display screen 1-322b, display 11.3.2-104, first display assembly 1-120a and second display assembly 1-120b, display assembly 1-320, display assembly 1-421, first display sub-assembly 1-420a and second display sub-assembly 1-420b, display assembly 3-108, display assembly 11.3.2-204, first optical module 11.1.1-104a and second optical module 11.1.1-104b, optical module 11.3.2-100, optical module 11.3.2-200, bi-convex lens array 3-110, display area or display region 6-232, and / or any of the features, components, and / or parts of display / display region 6-334. In some embodiments, the HMD X700 includes sensors, which include either individually or in any combination any of the features, components, and / or parts of sensors 190, 306, image sensor 314, image sensor 404, sensor assembly 1-356, sensor assembly 1-456, sensor system 6-102, sensor system 6-202, sensor 6-203, sensor system 6-302, sensor 6-303, sensor system 6-402, and / or sensors 11.1.2-110a-f. In some embodiments, the HMD X700 includes one or more input devices, which include either individually or in any combination any of the features, components, and / or parts of first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328. In some embodiments, the HMD X700 includes one or more audio output components (e.g., electronic component 1-112) for generating audio feedback (e.g., audio output X714-3), which is optionally generated based on detected events and / or user input detected by the HMD X700.
[0230] In Figure 7EAt , in response to user input 727, electronic device 700 stops displaying representation 720 at the top of the stack and now displays representation 724 (which is the second in the stack in ) at the top of the stack. Representation 724 represents a music extended reality experience and displays objects 728a - 728d superimposed on background 728e and surrounded by border 723. Objects 728a - 728d represent the objects that would be displayed in the music extended reality experience if the user selects the music extended reality experience for display. Thus, representation 724 provides the user with a preview of what the music extended reality experience will look like. At Figure 7D1 in, user input 729 is a button press of button 704a. As discussed above, in some embodiments, user input 729 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, electronic device 700 is a head - mounted system, and user input 729 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing electronic device 700, pressing a button while wearing electronic device 700, rotating a rotatable input mechanism while wearing electronic device 700, providing a gaze - based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing. Figure 7E At, electronic device 700 detects user input 729. At Figure 7E in, user input 729 is a button press of button 704a. As discussed above, in some embodiments, user input 729 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, electronic device 700 is a head - mounted system, and user input 729 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing electronic device 700, pressing a button while wearing electronic device 700, rotating a rotatable input mechanism while wearing electronic device 700, providing a gaze - based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0231] At Figure 7F in response to user input 729, electronic device 700 stops displaying representation 724 at the top of the stack and now displays representation 726 (which is the second in the stack in ) at the top of the stack. In some embodiments, if user input 729 was already a request to rotate the stack in the opposite direction (e.g., a button press of button 704b), then electronic device 700 will redisplay representation 720 (and the second representation 724 in the stack, as shown in ) at the top of the stack. At Figure 7E in, representation 726 represents a translation extended reality experience and includes objects 730a - 730d superimposed on background 730 and surrounded by border 725. In some embodiments, object 730a is an identifier that identifies the translation extended reality experience (e.g., via a logo and / or name), and objects 730b - 730d represent optional objects that would be displayed in the translation extended reality experience. In some embodiments, objects 730b - 730d represent optional objects (as will be described below with reference to Figure 7D1 ), but they are not individually selectable to perform any function themselves. At Figure 7F in, representation 726 represents a translation extended reality experience and includes objects 730a - 730d superimposed on background 730 and surrounded by border 725. In some embodiments, object 730a is an identifier that identifies the translation extended reality experience (e.g., via a logo and / or name), and objects 730b - 730d represent optional objects that would be displayed in the translation extended reality experience. In some embodiments, objects 730b - 730d represent optional objects (as will be described below with reference to Figure 7J ), but they are not individually selectable to perform any function themselves. At Figure 7F At, electronic device 700 detects user input 732. At Figure 7FIn [description], the user input 732 is a touchscreen swipe gesture with a downward direction. However, in some embodiments, the user input 732 is a different type of user input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 732 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing the electronic device 700, pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular manner), and / or any combination of the foregoing.
[0232] At Figure 7G , in response to the user input 732, the electronic device 700 displays a system control user interface 734 that includes selectable objects 736a - 736h. Object 736a can be selected to selectively engage or disengage the "Do Not Disturb" state or the sleep focus state, in which notifications received by the electronic device 700 are suppressed. Object 736b can be selected to selectively turn on or off WiFi. Object 736c can be selected to selectively engage or disengage airplane mode. Object 736d can be selected to selectively turn on or off the flashlight. Object 736e can be selected to initiate a process for streaming audio and / or video content to an external device. Object 736f can be selected to selectively engage or disengage silent mode. Object 736g can be selected to modify the volume settings of the electronic device 700. Object 736h can be selected to modify the brightness of the electronic device 700. In some embodiments, the electronic device 700 is a head-mounted system, and option 736h is selectable to modify the passthrough brightness setting and / or the passthrough opacity setting of the electronic device 700. At Figure 7G , the electronic device 700 detects the user input 736. At Figure 7G In [description], the user input 736 is a tap input on the touch-sensitive display 702. However, in some embodiments, the user input 736 is a different type of user input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 736 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing the electronic device 700, pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular manner), and / or any combination of the foregoing.
[0233] At Figure 7H , in response to the user input 736, the electronic device 700 stops displaying the system control user input 734. At Figure 7HAt , the electronic device 700 displays representation 726 at the top of the displayed stack, and when representation 736 is displayed at the top of the displayed stack, the electronic device 700 detects user input 740 (e.g., a selection input). At Figure 7H , the user input 740 is a button press input of button 704c. However, in some embodiments, the user input 740 is a different type of user input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 740 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing the electronic device 700, pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular manner), and / or any combination of the foregoing.
[0234] At Figure 7I the electronic device 700 stops displaying representation 721 (e.g., stops displaying the stack of representations) in response to the user input 740 and displays an animation in which representation 726 appears to move toward the user. For example, in Figure 7I representation 726 (including objects 730a - 730b and boundary 725) becomes larger. Additionally, in Figure 7I the background 730e changes from opaque to transparent to reveal the three-dimensional environment 712 behind representation 736.
[0235] At Figure 7J the electronic device 700 completes the animation of representation 721 becoming larger and now replaces the display of representation 721 with a display of a translated augmented reality experience 742. In some embodiments, when transitioning from the display of representation 721 to the display of the translated augmented reality experience 742, the electronic device 700 displays a cross-fade of objects 730a - 730d with corresponding objects 744a - 744d. Additionally, in Figure 7J the three-dimensional environment 712 is no longer visually de-emphasized (as indicated by the transition from the dashed line in Figure 7I to the solid line in Figure 7J ).
[0236] The translation augmented reality experience 742 includes an object 744a that identifies the augmented reality experience (e.g., using a logo and / or name), and objects 744b - 744d that can be selected to perform various tasks. For example, in some embodiments, object 744b can be selected to engage a microphone such that the user can provide spoken and / or verbal input for transitioning to a different language; object 744c can be selected to translate visual content captured by one or more cameras (e.g., input sensor 706); and object 744d can be selected to cause the electronic device 700 to read the translation aloud (e.g., play audio content that reads the translation aloud). In Figure 7J , the translation augmented reality experience 742 further includes an object 746 that indicates that the electronic device 700 has detected visual content that can be translated. In Figure 7J , a menu has moved into the view of the electronic device 700 (e.g., into the view of one or more cameras), and object 746 indicates that the menu includes text that can be translated into a different language. In Figure 7J , the electronic device 700 detects a user input 748 while also detecting that the user is looking at object 746 (e.g., as indicated by the gaze indication 710). In Figure 7J , the user input 748 is a tap input via the touch-sensitive display 702. However, in some embodiments, the user input 748 is a different type of user input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 748 includes, for example, the user performing a gesture (e.g., an air gesture) while wearing the electronic device 700, pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular manner), and / or any combination of the foregoing.
[0237] In Figure 7KAt 0, in response to user input 748 (e.g., in response to user input 748 when the user is gazing at object 746), electronic device 700 displays translations 750a - 750e. Translations 750a - 750e are displayed superimposed on three - dimensional environment 712. In some embodiments, objects 744a - 744d are viewpoint - locked objects such that even when the user changes the viewpoint of electronic device 700 (e.g., by moving and / or rotating electronic device 700), objects 744a - 744d do not move around display 702; and translations 750a - 750e are environment - locked (or world - locked) objects such that when the user changes the viewpoint of electronic device 700, translations 750a - 750e move around (and / or away from) display 702 based on how things move in three - dimensional environment 712. For example, translation 750a is "locked" to the word "MENU" and moves around display 702 with the word "MENU", and translation 750b is "locked" to the word "GARDEN SALAD" and moves around display 702 with the word "GARDEN SALAD".
[0238] The following reference provides additional description with respect to Figures 7A to 7K the methods 800 and 900 described with respect to Figures 7A to 7K ...
[0239] Figure 8 is a flowchart of an exemplary method 800 for navigating an extended reality experience according to some embodiments. In some embodiments, method 800 is executed at a computer system (e.g., Figure 1A computer system 101 in)(e.g., 700 and / or X700)(e.g., a smart phone, a smart watch, a tablet, a wearable device, and / or a head - mounted device), the computer system communicating with one or more display - generating components (e.g., 702 and / or X702)(e.g., a visual output device, a 3D display, a display having at least a portion of a transparent or translucent surface on which an image can be projected (e.g., a see - through display), a projector, a heads - up display, and / or a display controller) and one or more input devices (e.g., a touch - sensitive surface (e.g., a touch - sensitive display); a mouse; a keyboard; a remote control; a visual input device (e.g., one or more cameras (e.g., an infrared camera, a depth camera, a visible - light camera)); an audio input device; and / or a biometric sensor (e.g., a fingerprint sensor, a face - identification sensor, and / or an iris - identification sensor)). In some embodiments, method 800 is stored in a non - transitory (or transitory) computer - readable storage medium and executed by one or more processors of the computer system (such as one or more processors 202 of computer system 101)(e.g., Figure 1Ais managed by instructions executed by the control 110 in []. Some operations in method 800 are optionally combined, and / or the order of some operations is optionally changed.
[0240] In some embodiments, a computer system (e.g., 700 and / or X700) simultaneously displays (802) representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) (e.g., augmented reality user interfaces and / or augmented reality applications) (e.g., displays representations of multiple augmented reality experiences superimposed on a three-dimensional environment and / or displays representations of multiple augmented reality experiences simultaneously with the three-dimensional environment) in a three-dimensional environment (e.g., 712) (e.g., a virtual three-dimensional environment, a virtual passthrough three-dimensional environment, and / or an optical passthrough three-dimensional environment) via one or more display generation components (e.g., 702 and / or X702), including: a first representation (804) of a first augmented reality experience (e.g., 720, 721, 724, and / or 726); and a second representation (806) of a second augmented reality experience different from the first augmented reality experience (e.g., 720, 721, 724, and / or 726), where the second representation is different from the first representation. When simultaneously displaying (808) representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) in a three-dimensional environment (e.g., 712), the computer system receives (810) a first user input (e.g., 727, 729, and / or 740) (e.g., one or more user inputs and / or a first set of user inputs) (e.g., one or more mechanical inputs (e.g., button presses and / or rotations of physical input mechanisms), one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs) via one or more input devices (e.g., 702, 704a - 704c, and / or 706). In response to receiving the first user input (812), the computer system stops displaying (814) the representation(s) of one or more augmented reality experiences among the multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726); and based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience (816), the computer system displays (818) the first augmented reality experience (e.g., 714 and / or 742) in the three-dimensional environment via one or more display generation components (and in some embodiments, does not display the second augmented reality experience) (e.g., displays the first augmented reality experience applied to the three-dimensional environment, displays the first augmented reality experience superimposed on the three-dimensional environment, and / or displays the first augmented reality experience simultaneously with the three-dimensional environment).
[0241] In some embodiments, in response to receiving a first user input (e.g., 727, 729, and / or 740), and based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience (e.g., 720, 721, 724, and / or 726), the computer system displays the second augmented reality experience (e.g., 714 and / or 742) in a three-dimensional environment (e.g., 712) (e.g., applied to and / or concurrently with the three-dimensional environment) via one or more display generation components (and, in some embodiments, does not display the first augmented reality experience). In some embodiments, the computer system (e.g., 700 and / or X700) is a head-mounted system. In some embodiments, the three-dimensional environment (e.g., 712) is an optically see-through environment (e.g., a physical real-world environment) that is visible to the user via a transparent display generation component (e.g., a transparent optical lens display) that displays representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726), the first augmented reality experience (e.g., 714 and / or 742), and / or the second augmented reality experience (e.g., 714 and / or 742) thereon. In some embodiments, the three-dimensional environment (e.g., 712) is a virtual three-dimensional environment displayed by one or more display generation components (e.g., 702). In some embodiments, the three-dimensional environment is a virtual see-through environment (e.g., a virtual see-through environment that is a virtual representation of the user's physical real-world environment (e.g., as captured by one or more cameras communicating with the computer system)) displayed by one or more display generation components (e.g., 702 and / or X702). Displaying representations of multiple augmented reality experiences concurrently allows the user to switch between different augmented reality experiences with less user input, thereby reducing the amount of user input required to perform an operation. Displaying the first augmented reality experience based on determining that the first user input corresponds to a selection of a first representation of the first augmented reality experience provides the user with visual feedback regarding the state of the system (e.g., the system has detected the first user input corresponding to the selection of the first representation of the first augmented reality experience), thereby providing the user with improved visual feedback.
[0242] In some embodiments, representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) are displayed on one or more light field displays (e.g., perspective displays that display one or more elements while the real-world background behind the displayed elements is visible to the user); and a three-dimensional environment (e.g., 712) is an optically see-through environment (e.g., a physical real environment) that is visible to the user through one or more light field displays (e.g., behind and / or through the displayed representations of the multiple augmented reality experiences). Displaying representations of multiple augmented reality experiences simultaneously allows the user to switch between different augmented reality experiences with less user input, thereby reducing the amount of user input required to perform an operation.
[0243] In some embodiments, in response to receiving a first user input (e.g., 727, 729, and / or 740), and based on determining that the first user input corresponds to a selection of a second representation (e.g., 720, 721, 724, and / or 726) of a second augmented reality experience, the computer system displays the second augmented reality experience (e.g., 714 and / or 742) in a three-dimensional environment (e.g., 712) via one or more display generation components. In some embodiments, displaying the first augmented reality experience (e.g., 714 and / or 742) includes displaying a first set of interactive elements (e.g., objects 716a - 716e corresponding to augmented reality experience 714, and objects 744a - 744d corresponding to augmented reality experience 742) (e.g., one or more interactive elements) (e.g., one or more selectable options, selectable buttons, and / or affordance representations) (in some embodiments, displaying the first set of interactive elements superimposed on the three-dimensional environment); displaying the second augmented reality experience (e.g., 714 and / or 742) includes displaying a second set of interactive elements different from the first set of interactive elements (e.g., objects 716a - 716e corresponding to augmented reality experience 714, and objects 744a - 744d corresponding to augmented reality experience 742) (e.g., one or more interactive elements) (e.g., one or more selectable options, selectable buttons, and / or affordance representations) (e.g., not displaying the first set of interactive elements) (in some embodiments, displaying the second set of interactive elements superimposed on the three-dimensional environment); the first representation (e.g., 720 and / or 726) of the first augmented reality experience includes a representation of the first set of interactive elements (e.g., representation 720 includes objects 722a - 722e representing objects 716a - 716e; and representation 726 includes objects 730a - 730d representing objects 744a - 744d) (in some embodiments, a representation of the first set of interactive elements superimposed on a first representative background (e.g., a visual area and / or a displayed area representing the three-dimensional environment (e.g., a see-through environment, an optical see-through environment, and / or a virtual see-through environment)))); and the second representation (e.g., 720 and / or 726) of the second augmented reality experience includes a representation of the second set of interactive elements (e.g., representation 720 includes objects 722a - 722e representing objects 716a - 716e; and representation 726 includes objects 730a - 730d representing objects 744a - 744d) (in some embodiments, a representation of the second set of interactive elements superimposed on a second representation background), the representation of the second set of interactive elements being different from the representation of the first set of interactive elements. Displaying a representation of an augmented reality experience that provides a simplified preview of the augmented reality experience to the user enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0244] In some embodiments, displaying a first augmented reality experience (e.g., 714 and / or 742) includes displaying a first set of interactive elements (e.g., 716a - 716e and / or 744a - 744d) superimposed on a see-through environment (e.g., 712) (e.g., an optical see-through environment and / or a virtual see-through environment); displaying a second augmented reality experience (e.g., 714 and / or 742) includes displaying a second set of interactive elements (e.g., 716a - 7164 and / or 744a - 744d) superimposed on the see-through environment (e.g., 712); a first representation of the first augmented reality experience (e.g., 720 and / or 726) includes a first placeholder background content (e.g., 722f, 728e, and / or 730e) representing the see-through environment (e.g., placeholder content that is not the see-through environment and / or is a representation of the see-through environment) (e.g., an image, a virtual three-dimensional environment, a solid color, and / or a visual pattern) (in some embodiments, the representation of the first set of interactive elements is superimposed on the first placeholder background content); and a second representation of the second augmented reality experience (e.g., 720 and / or 726) includes a second placeholder background content (e.g., 722f, 728e, and / or 730e) representing the see-through environment (e.g., placeholder content that is not the see-through environment and / or is a representation of the see-through environment) (e.g., an image, a virtual three-dimensional environment, a solid color, and / or a visual pattern) (in some embodiments, the representation of the second set of interactive elements is superimposed on the second placeholder background content) (e.g., a second placeholder background content that is different from or the same as the first placeholder background content). Displaying a representation of an augmented reality experience that provides a simplified preview of the augmented reality experience to the user enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0245] In some embodiments, the representation of the first set of interactive elements (e.g., 722a - 722e, 728a - 728d, and / or 730a - 730d) is non - interactive (e.g., cannot be individually selected by the user and / or otherwise individually interacted with) (in some embodiments, one or more of the interactive elements of the first set of interactive elements (e.g., 716a - 716e and / or 744a - 744d) can be selected to perform corresponding actions (e.g., the first interactive element of the first set of interactive elements can be selected to perform a first action, and the second interactive element of the first set of interactive elements can be selected to perform a second action), and the representation of the first set of interactive elements (e.g., 722a - 722e, 728a - 728d, and / or 730a - 730d) cannot be selected (e.g., cannot be individually selected) to perform corresponding actions (e.g., the representation of the first set of interactive elements cannot be selected to perform the first action, and the representation of the second interactive element cannot be selected to perform the second action, and / or the representations of the first and second interactive elements cannot be individually selected); and / or the computer system is configured to distinguish between the selection of a first interactive element (e.g., 716a - 716e and / or 744a - 744d) and a second interactive element (e.g., 716a - 716e and / or 744a - 744d) of the first set of interactive elements, but is not configured to distinguish between the selection of the representation of a first interactive element (e.g., 722a - 722e, 728a - 728d, and / or 730a - 730d) and the representation of a second interactive element (e.g., 722a - 722e, 728a - 728d, and / or 730a - 730d)); and the representation of the second set of interactive elements (e.g., 722a - 722e, 728a - 728d, and / or 730a - 730d) is non - interactive. The representation of the augmented reality experience that provides a simplified preview of the augmented reality experience to the user enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0246] In some embodiments, at least a portion of a first representation of a first augmented reality experience (e.g., 720, 721, 724, and / or 726) is displayed in a first color corresponding to the first augmented reality experience (e.g., 714 and / or 742) (e.g., uniquely corresponding to the first augmented reality experience and / or corresponding to the first augmented reality experience and not corresponding to a second augmented reality experience); and at least a portion of a second representation of a second augmented reality experience (e.g., 720, 721, 724, and / or 726) is displayed in a second color corresponding to the second augmented reality experience (e.g., 714 and / or 742) (e.g., uniquely corresponding to the second augmented reality experience and / or corresponding to the second augmented reality experience and not corresponding to the first augmented reality experience), wherein the second color is different from the first color. In some embodiments, the first representation of the first augmented reality experience does not include the second color, and the second representation of the second augmented reality experience does not include the first color. Displaying representations of augmented reality experiences in different colors that uniquely correspond to different augmented reality experiences allows a user to more easily select a particular augmented reality experience, which enhances the operability of a computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0247] In some embodiments, a first representation of a first augmented reality experience (e.g., 720, 721, 724, and / or 726) includes a first identifier (e.g., 722a, 728a, and / or 730a) (e.g., a first icon, a first set of text (e.g., a name and / or other text identifier), and / or a first color) corresponding to the first augmented reality experience (e.g., uniquely corresponding to the first augmented reality experience and / or corresponding to the first augmented reality experience and not corresponding to a second augmented reality experience or other augmented reality experiences available via the computer system); a second representation of a second augmented reality experience (e.g., 720, 721, 724, and / or 726) includes a second identifier (e.g., 722a, 728, and / or 730a) (e.g., a second icon, a second set of text (e.g., a name and / or other text identifier), and / or a second color) different from the first identifier and corresponding to the second augmented reality experience (e.g., uniquely corresponding to the second augmented reality experience and / or corresponding to the second augmented reality experience and not corresponding to the first augmented reality experience); displaying the first augmented reality experience (e.g., 714 and / or 742) includes displaying the first identifier (e.g., 716a and / or 744a) as part of the first augmented reality experience (e.g., not displaying the second identifier); and displaying the second augmented reality experience (e.g., 714 and / or 742) includes displaying the second identifier (e.g., 716a and / or 744a) as part of the second augmented reality experience (e.g., not displaying the first identifier). Displaying representations of augmented reality experiences with different identifiers that uniquely correspond to different augmented reality experiences allows a user to more easily select a particular augmented reality experience, which enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0248] In some embodiments, in response to receiving a first input (e.g., 718, 727, 729, and / or 740): Based on determining that the first user input corresponds to a selection of a first representation of a first augmented reality experience (e.g., 720, 721, 724, and / or 726): Before displaying the first augmented reality experience (e.g., 714 and / or 742), the computer system displays, via one or more display generation components, a first animation in which the first representation of the first augmented reality experience moves toward the viewpoint of the user of the computer system (e.g., Figures 7H to 7J, indicating that 726 moves towards the user's viewpoint until the augmented reality experience 742 is displayed (e.g., where the first representation of the first augmented reality experience becomes larger and / or appears to move closer to the user's viewpoint). In some embodiments, in response to receiving a first user input and based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience: before displaying the second augmented reality experience, the computer system displays a second animation via one or more display generation components, in which the second representation of the second augmented reality experience moves towards the viewpoint of the user of the computer system. Displaying the animation in which the first representation of the first augmented reality experience moves towards the user's viewpoint provides the user with visual feedback regarding the state of the system (e.g., the system is transitioning to the first augmented reality experience), thereby providing the user with improved visual feedback.
[0249] In some embodiments, the first representation of the first augmented reality experience (e.g., 720, 721, 724, and / or 726) includes a first boundary around (e.g., partially and / or completely around) a first set of interactive elements (e.g., 722a - 722e, 728a - 728d, and / or 730a - 730d); the second representation of the second augmented reality experience (e.g., 720, 721, 724, and / or 726) includes a second boundary around (e.g., partially and / or completely around) a second set of interactive elements (e.g., 722a - 722e, 728a - 728d, and / or 730a - 730d) (e.g., different and / or separate from the first boundary); and displaying the first animation includes displaying the first boundary moving towards the viewpoint of the user of the computer system until the first boundary is no longer displayed (e.g., Figures 7H to 7J, move around the boundary representing 726 towards the user's viewing point until the augmented reality experience 742 is displayed (e.g., until the first boundary moves out of the display area of the computer system and / or out of the portion of the display area of the computer system visible to the user) (e.g., increase the size of the first boundary until the first boundary is no longer displayed by the computer system and / or is outside the portion of the display area of the computer system visible to the user). In some embodiments, in response to receiving a first user input and based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience: before displaying the second augmented reality experience, the computer system displays a second animation via one or more display generation components, in which the second representation of the second augmented reality experience moves towards the viewing point of the user of the computer system, where displaying the second animation includes displaying a second boundary that moves towards the viewing point of the user of the computer system until the second boundary is no longer displayed. Displaying an animation in which the first representation (including the boundary of the first representation) of the first augmented reality experience moves towards the user's viewing point provides the user with visual feedback about the state of the system (e.g., the system is transitioning to the first augmented reality experience), thus providing the user with improved visual feedback.
[0250] In some embodiments, in response to receiving a first input (e.g., 718, 727, 729, and / or 740): based on determining that the first user input corresponds to a selection of a first representation of a first augmented reality experience: the computer system displays, via one or more display generation components, a cross-fade of the representation of the first set of interactive elements (e.g., 730a - 730d) with the first set of interactive elements (e.g., 744a - 744d) (e.g., during the display of the first animation and / or after displaying the first animation). In some embodiments, in response to receiving a first input: based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience: the computer system displays, via one or more display generation components, a cross-fade of the representation of the second set of interactive elements with the second set of interactive elements. Displaying the cross-fade of the representation of the first set of interactive elements with the first set of interactive elements provides the user with visual feedback about the state of the system (e.g., the system is transitioning to the first augmented reality experience), thus providing the user with improved visual feedback.
[0251] In some embodiments, in response to receiving a first input (e.g., 718, 727, 729, and / or 740): Based on determining that the first user input corresponds to a selection of a first representation of a first augmented reality experience: The computer system stops displaying the representation of a first set of interactive elements (e.g., 730a - 7330d); and displays a first set of interactive elements (e.g., 744a - 744d) via one or more display generation components. Displaying the replacement of the representation of the first set of interactive elements with the first set of interactive elements provides visual feedback to the user regarding the state of the system (e.g., the system is transitioning to the first augmented reality experience), thereby providing improved visual feedback to the user.
[0252] In some embodiments, before receiving the first user input, and when representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) are simultaneously displayed in a three - dimensional environment (e.g., 712), the computer system displays a first representation of a first augmented reality experience (e.g., 720, 721, 724, and / or 726) at a first display position via one or more display generation components (e.g., Figure 7D1 representation 720 in Figure 7D1 ), and in some embodiments, simultaneously displays a second representation of a second augmented reality experience (e.g., Figures 7D1 to 7E representation 724 in Figures 7E to 7F ) at a second display position different from the first display position. (In some embodiments, the first display position represents the currently selected object and / or the currently focused object). When the first representation of the first augmented reality experience is displayed at the first display position, the computer system receives, via one or more input devices, a second user input (e.g., 727 and / or 729) corresponding to a request to navigate from the first representation of the first augmented reality experience (e.g., 720, 721, 724, and / or 726) to a second representation of a second augmented reality experience (e.g., 720, 721, 724, and / or 726) (e.g., one or more user inputs and / or a first set of user inputs) (e.g., one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs). In response to receiving the second user input (e.g., 727 and / or 729), the computer system stops displaying the first representation of the first augmented reality experience at the first display position (e.g., in Figures 7D1 to 7E , in response to user input 727, electronic device 700 stops displaying representation 720 at the front - most position of the stack; and in Figures 7E to 7F , in response to user input 729, electronic device 700 stops displaying representation 724 at the front - most position of the stack) (and in some embodiments, simultaneously maintains the display of at least a portion of the first representation of the first augmented reality experience); and displays the second representation of the second augmented reality experience at the first display position via one or more display generation components (e.g., inFigure 7E In, 724 is shown at the front position of the stack, and in Figure 7F In, 726 is shown at the front position of the stack). Displaying navigation from a first representation of a first augmented reality experience to a second representation of a second augmented reality experience in response to a second user input provides the user with visual feedback regarding the state of the system (e.g., the system has detected the second user input), thereby providing the user with improved visual feedback.
[0253] In some embodiments, displaying representations of multiple augmented reality experiences simultaneously (e.g., 720, 721, 724, and / or 726) includes displaying representations of multiple augmented reality experiences in a stack, where a first representation of a first augmented reality experience is stacked on top of a second representation of a second augmented reality experience (e.g., the first representation of the first augmented reality experience is located on top of the second representation of the second augmented reality experience and / or partially obscures the second representation of the second augmented reality experience). Displaying representations of augmented reality experiences in a stack in which the user can navigate allows the user to more easily select a particular augmented reality experience, which enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0254] In some embodiments, before receiving a first user input, and when representations of multiple augmented reality experiences are displayed simultaneously in a three-dimensional environment (including simultaneously displaying a first representation of a first augmented reality experience and a second representation of a second augmented reality experience), the computer system receives a third user input (e.g., 727 and / or 729) corresponding to a request to navigate among the representations of the multiple augmented reality experiences via one or more input devices (e.g., one or more user inputs and / or a first set of user inputs) (e.g., one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs). In response to receiving the third user input, the computer system stops displaying the first representation of the first augmented reality experience while maintaining the display of the second representation of the second augmented reality experience (e.g., in Figures 7D1 to 7E In, in response to user input 727, electronic device 700 stops displaying representation 720 while maintaining the display of representations 724 and / or 726; and / or in Figures 7E to 7F In, in response to user input 729, electronic device 700 stops displaying representation 724 while maintaining the display of representation 726). Displaying representations of augmented reality experiences in a stack in which the user can navigate allows the user to more easily select a particular augmented reality experience, which enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0255] In some embodiments, determining that a first user input corresponds to a selection of a first representation of a first augmented reality experience includes determining that the first user input is a selection input, the selection input including: a gaze input toward the first representation of the first augmented reality experience (e.g., in Figure 7H a gaze indication 710 indicates that the user is looking at a representation 726) (e.g., user gaze toward an optional object; user gaze toward a corresponding representation among representations of multiple augmented reality experiences and / or user gaze corresponding to and / or identifying a particular augmented reality experience); and a hardware press input detected when the gaze input is toward the first representation of the first augmented reality experience (e.g., 740) (e.g., a press of a hardware button and / or a press of a pressable input mechanism (e.g., a rotatable and pressable input mechanism)) (e.g., a hardware press input that occurs simultaneously with the gaze input). In some embodiments, in response to receiving the first user input and based on determining that the first user input is not a selection input (e.g., based on determining that the first user input does not include a gaze input toward the first representation of the first augmented reality experience and / or a hardware press input when the gaze input is toward the first representation of the first augmented reality experience), the computer system abandons displaying the first augmented reality experience. In some embodiments, in response to receiving the first user input and based on determining that the first user input is not a selection input, the computer system abandons stopping the display of a representation of one or more augmented reality experiences among multiple augmented reality experiences (e.g., the computer system maintains the display of a representation of one or more augmented reality experiences among multiple augmented reality experiences). In some embodiments, stopping the display of a representation of one or more augmented reality experiences among multiple augmented reality experiences is performed based on determining that the first user input is a selection input. In some embodiments, the first user input includes: a first gaze input (e.g., user gaze toward a corresponding representation among representations of multiple augmented reality experiences and / or user gaze corresponding to and / or identifying a particular augmented reality experience); and a hardware press input (e.g., a press of a hardware button and / or a press of a pressable input mechanism (e.g., a rotatable and pressable input mechanism)). In some embodiments, the first user input includes a first gaze input and a hardware press input that occur simultaneously (e.g., a hardware press input when the user is gazing at a particular object and / or a hardware press input when the user is gazing at a corresponding representation among representations of multiple augmented reality experiences). Allowing the user to select a particular augmented reality experience using gaze and hardware press inputs enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0256] In some embodiments, determining that a first user input corresponds to a selection of a first augmented reality experience representation includes determining that the first user input is a selection input, the selection input including: a voice input indicating a user request to select an optional object (e.g., a voice input identifying a specific optional object; and / or a voice input identifying a corresponding augmented reality experience among multiple augmented reality experiences and / or a corresponding representation among representations of multiple augmented reality experiences (e.g., in Figure 7D1 and (or in Figure 7B ), the user states "Apply the translation extended reality experience", and in response to the user voice input, the electronic device 700 and / or the HMD X700 display the translation extended reality experience, as Figures 7I to 7J shown). In some embodiments, in response to receiving the first user input, and based on determining that the first user input is not a selection input (e.g., based on determining that the first user input does not include a voice input indicating a user request to select an optional object), the computer system abandons displaying the first augmented reality experience (e.g., the computer system continues to display the stack of representations as Figure 7D1 shown). In some embodiments, in response to receiving the first user input, and based on determining that the first user input is not a selection input, the computer system abandons stopping the display of the representations of one or more augmented reality experiences among multiple augmented reality experiences (e.g., the computer system maintains the display of the representations of one or more augmented reality experiences among multiple augmented reality experiences). In some embodiments, stopping the display of the representations of one or more augmented reality experiences among multiple augmented reality experiences is performed based on determining that the first user input is a selection input. In some embodiments, the first user input includes: a first voice input (e.g., a voice input identifying a corresponding augmented reality experience among multiple augmented reality experiences and / or a corresponding representation among representations of multiple augmented reality experiences). In some embodiments, the first user input includes a first voice input and a first gaze input (e.g., a user gaze toward a corresponding representation among representations of multiple augmented reality experiences) (e.g., in Figure 7H , the user voice input of stating "Display the extended reality experience" while looking at the representation 726). In some embodiments, the first user input includes a first voice input that occurs simultaneously with the first gaze input (e.g., a voice input when the user gazes at a specific object and / or a voice input when the user gazes at a corresponding representation of the representations of multiple augmented reality experiences). Allowing the user to use voice input to select a specific augmented reality experience enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0257] In some embodiments, determining that a first user input corresponds to a selection of a first representation of a first augmented reality experience includes determining that the first user input is a selection input, the selection input including: a gaze input (e.g., Figure 7H the gaze indication 720 in) toward the first representation of the first augmented reality experience (e.g., 726) (e.g., a user gaze toward an optional object; a user gaze toward a respective representation among representations of multiple augmented reality experiences and / or a user gaze corresponding to and / or identifying a particular augmented reality experience, the user gaze meeting a first set of gaze duration criteria (e.g., a user gaze toward an optional object and maintained on the optional object for a threshold duration (e.g., without interruption and / or with less than a threshold amount of interruption); and / or a user gaze toward a respective representation among representations of multiple augmented reality experiences and maintained on the respective representation for a threshold duration (e.g., without interruption and / or with less than a threshold amount of interruption))).
[0258] In some embodiments, in response to receiving the first user input and based on determining that the first user input is not a selection input (e.g., based on determining that the first user input does not include a gaze input toward the first representation of the first augmented reality experience that meets the first set of gaze duration criteria, because the gaze input is not toward the first representation of the first augmented reality experience or because the gaze input moves away from the first representation of the first augmented reality experience before the first set of gaze duration criteria have been met), the computer system abandons displaying the first augmented reality experience (e.g., in some embodiments, in Figure 7H , if the user maintains his or her gaze on representation 726 for a threshold duration, the electronic device 700 and / or the HMDX700 displays the translated extended reality experience 742 as shown in Figures 7I to 7J , but if the user does not maintain his or her gaze on representation 726 for a threshold duration, the electronic device 700 maintains Figure 7H the display of representations 726, 721 in). In some embodiments, in response to receiving the first user input and based on determining that the first user input is not a selection input, the computer system abandons stopping the display of representations of one or more of the multiple augmented reality experiences (e.g., the computer system maintains the display of representations of one or more of the multiple augmented reality experiences) (e.g., maintains Figure 7H the display of representations 721, 726 in). In some embodiments, stopping the display of representations of one or more of the multiple augmented reality experiences is performed based on determining that the first user input is a selection input. In some embodiments, the first user input includes a first gaze input that meets the first set of gaze duration criteria (e.g., Figure 7Hin the 710) (e.g., a user gaze that is directed toward and maintained on a respective representation among a plurality of representations of augmented reality experiences for a threshold duration (e.g., without interruption and / or with less than a threshold amount of interruption)). In some embodiments, determining that a first user input corresponds to a selection of a first representation of a first augmented reality experience includes the user having gazed at the first representation of the first augmented reality experience for a threshold duration (e.g., without interruption and / or with less than a threshold amount of interruption) (e.g., in Figure 7H the user has gazed at representation 726 for a threshold duration). Allowing a user to select a particular augmented reality experience using gaze and dwell inputs enhances the operability of a computer system by helping the user provide appropriate inputs and reducing user errors when operating / interacting with the computer system.
[0259] In some embodiments, when representations of a plurality of augmented reality experiences are simultaneously displayed, the computer system displays, via one or more display generation components, one or more setting controls (e.g., 736a - 736h), including a first setting control corresponding to a first setting of the computer system (in some embodiments, the computer system simultaneously displays a second setting control corresponding to a second setting of the computer system that is different from the first setting). When one or more setting controls are displayed, the computer system receives, via one or more input devices, a first setting input corresponding to the first setting of the computer system (e.g., Figure 7GThe user input in selects one of the setting options 736a - 736h and / or modifies the setting). In response to receiving the first setting input, the computer system modifies the first setting from a first value to a second value different from the first value. When multiple representations of augmented reality experiences are displayed simultaneously and when the first setting is set to the second value, the computer system receives a third user input (e.g., 740) (e.g., one or more user inputs and / or a third set of user inputs) (e.g., one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs) via one or more input devices. In response to receiving the third user input: Based on determining that the third user input corresponds to a selection of a first representation (e.g., 720, 724, and / or 726) of a first augmented reality experience, the computer system displays the first augmented reality experience (e.g., 714 and / or 742) in a three - dimensional environment via one or more display generation components while maintaining the first setting at the second value; and based on determining that the first user input corresponds to a selection of a second representation (e.g., 720, 724, and / or 726) of a second augmented reality experience, the computer system displays the second augmented reality experience (e.g., 714 and / or 742) in a three - dimensional environment via one or more display generation components while maintaining the first setting at the second value. Displaying one or more setting controls to modify one or more device settings and maintaining these settings across different augmented reality experiences allows the user to modify the device settings with fewer user inputs, thereby reducing the number of user inputs required to perform an operation.
[0260] In some embodiments, the first setting is a passthrough coloring setting (e.g., option 736h) (e.g., a setting that controls how much masking and / or darkening is applied to the three - dimensional environment (e.g., passthrough background, optical passthrough background, and / or virtual passthrough background)); the first value corresponds to a first amount of coloring (e.g., a first amount of masking and / or darkening; and / or a first brightness) applied to the three - dimensional environment; and the second value corresponds to a second amount of coloring (e.g., a second amount of masking and / or darkening; and / or a second brightness) different from the first amount of coloring applied to the three - dimensional environment. Displaying setting controls to modify the passthrough coloring and maintaining the passthrough coloring setting across different augmented reality experiences allows the user to modify the passthrough coloring setting with fewer user inputs, thereby reducing the number of user inputs required to perform an operation.
[0261] In some embodiments, the first setting is a volume setting (e.g., option 736g); the first value corresponds to a first volume; and the second value corresponds to a second volume different from the first volume. Displaying setting controls to modify the volume and maintaining the volume setting across different augmented reality experiences allows the user to modify the volume setting with fewer user inputs, thereby reducing the number of user inputs required to perform an operation.
[0262] In some embodiments, when displaying representations of multiple augmented reality experiences and one or more settings controls simultaneously, the computer system displays, via one or more display generation components, device status information indicating the status of one or more characteristics of the computer system (e.g., Wi-Fi network name, Wi-Fi signal strength, computer system battery level, computer system location tracking indicator, microphone recording indicator, camera recording indicator, and / or volume slider) (e.g., a Wi-Fi level indicator and / or a battery level indicator in the upper right of the display 702 and / or display module X702 in Figures 7D1 to 7H . Displaying the device status information provides the user with visual feedback regarding the status of the system (e.g., information regarding the status of one or more characteristics of the computer system), thereby providing the user with improved visual feedback.
[0263] In some embodiments, representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) are view-locked objects that remain in corresponding regions of the computer system user's field of view when the user's viewpoint is shifted relative to the three-dimensional environment (e.g., when the user's viewpoint is shifted and the background three-dimensional environment 712 moves, the representations 720, 721, 724, and / or 726 do not move). Displaying the representations of multiple augmented reality experiences as view-locked objects enhances the operability of the computer system by keeping the representations of the multiple augmented reality experiences within the user's line of sight, thereby assisting the user in providing appropriate input and reducing user errors when operating / interacting with the computer system.
[0264] In some embodiments, simultaneously displaying representations of multiple augmented reality experiences includes simultaneously displaying the representations of the multiple augmented reality experiences in a first orientation in which the representations of the multiple augmented reality experiences are aligned with gravity (e.g., in Figure 7D1In [the figure], representations 720, 721, 724, and / or 726 are displayed in an orientation such that the bottom surface of the representations 720, 721, 724, and / or 726 faces the ground (e.g., each representation has a bottom portion and a top portion, and the bottom portion is displayed closer to the ground and / or the center of the Earth than the top portion). In some embodiments, when multiple representations of augmented reality experiences are displayed simultaneously, the computer system detects a change in the orientation of the user's viewpoint (e.g., rotation of the electronic device 700, which, for example, causes the representations 720, 721, 724, and / or 726 to no longer be aligned with gravity (e.g., the bottom of the representations 720, 721, 724, and / or 726 no longer faces the ground)) (e.g., detecting rotation and / or movement of the user's head and / or detecting rotation and / or movement of a headset and / or other wearable device (e.g., a wearable device worn on the user's head)). In response to detecting a change in the orientation of the user's viewpoint: The computer system rotates the multiple representations of augmented reality experiences (e.g., 720, 721, 724, and / or 726) from a first orientation to a second orientation (e.g., a second orientation different from the first orientation) based on the change in the orientation of the user's viewpoint to continue to align the multiple representations of augmented reality experiences with gravity (e.g., to display the multiple representations of augmented reality experiences in a manner that maintains alignment with gravity (e.g., each representation has a bottom portion and a top portion, and even when the user moves and / or rotates his or her field of view, the bottom portion remains closer to the ground and / or the center of the Earth than the top portion)). In some embodiments, the multiple representations of augmented reality experiences are aligned with gravity (e.g., to display the multiple representations of augmented reality experiences in a manner that maintains alignment with gravity (e.g., each representation has a bottom portion and a top portion, and even when the user moves and / or rotates his or her field of view, the bottom portion remains closer to the ground and / or the center of the Earth than the top portion)). In some embodiments, when the computer system detects rotation of the computer system, the computer system rotates the multiple representations of augmented reality experiences based on the rotation of the computer system such that the bottom portion of the representation remains closer to the ground and / or the center of the Earth than the top portion of the representation. Displaying the multiple representations of augmented reality experiences as gravity-aligned viewpoint-locked objects enhances the operability of the computer system by keeping the multiple representations of augmented reality experiences within the user's line of sight and maintaining consistent alignment (even when the user moves and / or the computer system moves), thereby helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0265] In some embodiments, rotating the representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) from a first orientation to a second orientation includes: at a first time after detecting a change in the orientation of the user's viewpoint, displaying the representations of the multiple augmented reality experiences in the first orientation via one or more display generation components, where at the first time, at least in part due to the change in the orientation of the user's viewpoint, the representations of the multiple augmented reality experiences are not aligned with gravity (e.g., displaying representations 720, 721, 724, and / or 726 where the bottom edge of the representation does not face the ground); and at a second time after the first time, displaying the representations of the multiple augmented reality experiences in the second orientation via one or more display generation components to align the representations of the multiple augmented reality experiences with gravity (e.g., representations 720, 721, 724, and / or 726 as shown Figure 7D1 ). In some embodiments, the computer system displays a gradual rotation of the representations of the multiple augmented reality experiences from the first orientation to the second orientation over time. In some embodiments, at a third time after the first time and before the second time, the computer system displays the representations of the multiple augmented reality experiences in a third orientation different from the first and second orientations, where the third orientation is between the first and second orientations (e.g., at an angle between the angle of the first orientation and the angle of the second orientation). In some embodiments, the representations of the multiple augmented reality experiences exhibit an inertia following behavior (e.g., a behavior of reducing or delaying the movement of the representations of the multiple augmented reality experiences relative to a detected physical movement of the user (e.g., relative to a detected physical movement of the user's head) and / or relative to a detected physical movement of the computer system). Displaying the representations of the multiple augmented reality experiences as a viewpoint-locked object exhibiting inertia following behavior provides the user with visual feedback about the state of the system (e.g., when the user's head moves, the system intentionally moves the representations of the multiple augmented reality experiences), thereby providing the user with improved visual feedback.
[0266] In some embodiments, displaying a first augmented reality experience (e.g., 742) includes simultaneously displaying a first set of objects (e.g., 744a - 744d, 750a - 750e), including a first object (e.g., 744a - 744d) and a second object (e.g., 750a - 750e), and wherein: the first object is a viewpoint - locked object (e.g., 744a - 744d are viewpoint - locked objects); and the second object is an environment - locked object (e.g., 750a - 750e are environment - locked objects). In some embodiments, the second augmented reality experience includes a second set of objects, including a third object and a fourth object, wherein the third object is a viewpoint - locked object and the fourth object is an environment - locked object. Displaying certain objects in an AR experience as viewpoint - locked objects and other objects as environment - locked objects enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0267] In some embodiments, a computer system displays a first augmented reality experience (e.g., Figure 7B 714 in) in a three - dimensional environment (e.g., 712) via one or more display - generating components. When displaying the first augmented reality experience (e.g., Figure 7B 714 in), the computer system receives, via one or more input devices, a first voice input (e.g., an input including the voice of the user and / or an input spoken by the user) indicating a user request to change from the first augmented reality experience to a second augmented reality experience (e.g., "switch to the next AR experience" and / or "switch to the camera AR view"). In response to receiving the first voice input, the computer system stops displaying the first augmented reality experience (e.g., stops displaying experience 714); and displays, via one or more input devices, a second augmented reality experience (e.g., Figure 7J 742 in) in the three - dimensional environment (e.g., 712). Allowing the user to use voice input to switch between different augmented reality experiences enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system. Allowing the user to use voice input to switch between different augmented reality experiences allows the user to switch between different augmented reality experiences with less user input, thereby reducing the amount of user input required to perform an operation.
[0268] In some embodiments, when the computer system is in a sleep state (e.g., Figure 7A)(e.g., off state, locked state, and / or sleep state), the computer system receives, via one or more input devices, a first wake input (e.g., 708) corresponding to a request to transition the computer system from a sleep state to a wake state (e.g., one or more user inputs and / or a first set of user inputs) (e.g., one or more mechanical inputs (e.g., button presses and / or rotations of a physical input mechanism), one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs). In response to receiving the first wake input (and in some embodiments, based on determining that the first wake input meets a first set of wake criteria (e.g., unlock criteria, user authentication criteria, and / or biometric authentication criteria)), the computer system displays, via one or more display generation components, a first augmented reality experience (e.g., 714 and / or 742) (e.g., without displaying a second augmented reality experience and / or representations of multiple augmented reality experiences). In some embodiments, the first augmented reality experience represents the default augmented reality experience displayed when the computer system transitions from a sleep state to a wake state. Automatically displaying the first augmented reality experience when the computer system transitions from a sleep state to a wake state allows a user to access the first augmented reality experience with fewer user inputs, thereby reducing the number of user inputs required to perform an operation.
[0269] In some embodiments, when the computer system is in a sleep state (e.g., Figure 7A )(e.g., off state, locked state, and / or sleep state), the computer system receives, via one or more input devices, a first wake input (e.g., 708) corresponding to a request to transition the computer system from a sleep state to a wake state (e.g., one or more user inputs and / or a first set of user inputs) (e.g., one or more mechanical inputs (e.g., button presses and / or rotations of a physical input mechanism), one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs). In response to receiving the first wake input (and in some embodiments, based on determining that the first wake input meets a first set of wake criteria (e.g., unlock criteria, user authentication criteria, and / or biometric authentication criteria)), the computer system displays, via one or more display generation components, representations of multiple augmented reality experiences (e.g., Figure 7D1720, 721, 724, and / or 726) therein (e.g., not displaying the first augmented reality experience and / or the second augmented reality experience). In some embodiments, an AR experience switcher user interface representing a plurality of augmented reality experiences represents a default user interface displayed when the computer system transitions from a sleep state to a wake state. Automatically displaying the representations of the plurality of augmented reality experiences when the computer system transitions from a sleep state to a wake state allows users to access the representations of the plurality of augmented reality experiences with less user input, thereby reducing the amount of user input required to perform an operation.
[0270] In some embodiments, the multiple augmented reality experiences include one or more of the following: a camera augmented reality experience (e.g., 714) (e.g., an augmented reality experience including a shutter button that can be selected to capture more photos and / or videos (e.g., one or more photos and / or videos of the environment around the computer system) using one or more cameras of the computer system) (e.g., an augmented reality experience in which a user can capture one or more photos and / or videos of the user's surrounding environment); a translation augmented reality experience (e.g., 742) (e.g., an augmented reality experience including one or more options that can be selected to translate text captured by one or more cameras of the computer system (e.g., translate text in the environment around the computer system)) (e.g., an augmented reality experience in which a user can translate content in the user's environment from a first language to a second language); a reading augmented reality experience (e.g., an augmented reality experience that displays books, articles, and / or other text content) (e.g., an augmented reality experience in which a user can read content (e.g., books and / or articles)); a music augmented reality experience (e.g., an augmented reality experience including one or more selectable options that can be selected to output audio content (e.g., music and / or other audio content)) (e.g., an augmented reality experience in which a user can listen to music and / or other audio content); a navigation augmented reality experience (e.g., an augmented reality experience that displays navigation instructions to a geographical location) (e.g., an augmented reality experience in which a user can receive navigation instructions to a geographical location); a photo augmented reality experience (e.g., an augmented reality experience that displays one or more selectable objects and / or user interfaces for navigating through photo and / or video content in a media library) (e.g., an augmented reality experience in which a user can navigate through photo and / or video content in a media library and / or view the photo and / or video content); a video messaging augmented reality experience (e.g., an augmented reality experience including one or more selectable options that can be selected to initiate and / or terminate a video conference and / or video call with one or more contacts) (e.g., an augmented reality experience in which a user can participate in a video conference and / or video call with one or more contacts); and / or a fitness augmented reality experience (e.g., an augmented reality experience that displays one or more fitness metrics and / or physical activity metrics corresponding to the user); and / or an augmented reality experience that displays fitness instructions (e.g., fitness videos and / or demonstrations) (e.g., an augmented reality experience in which a user can track fitness and / or physical activity metrics; and / or an augmented reality experience in which a user can view fitness instructions (e.g., fitness videos and / or demonstrations). Displaying representations of multiple augmented reality experiences allows a user to switch between different augmented reality experiences with less user input, thereby reducing the amount of user input required to perform an operation.
[0271] In some embodiments, aspects / operations of methods 800, 900, 1100, 1300, and / or 1500 may be interchanged, substituted, and / or added between these methods. For example, in some embodiments, the augmented reality experience in method 800 is an extended reality experien...
Claims
1. A method, the method comprising: At a computer system in communication with one or more display generation components and one or more input devices: Simultaneously display, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: A first representation of a first augmented reality experience; and A second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; When simultaneously displaying the representations of the multiple augmented reality experiences in the three-dimensional environment, receive a first user input via the one or more input devices; and In response to receiving the first user input: Stop displaying the representations of one or more of the multiple augmented reality experiences; and In accordance with determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience, display the first augmented reality experience in the three-dimensional environment via the one or more display generation components.
2. The method according to claim 1, wherein: The representations of the multiple augmented reality experiences are displayed on one or more additive light displays; and The three-dimensional environment is an optically see-through environment visible to a user through the one or more additive light displays.
3. The method according to any one of claims 1 to 2, the method further comprising: In response to receiving the first user input: In accordance with determining that the first user input corresponds to a selection of the second representation of the second augmented reality experience, display the second augmented reality experience in the three-dimensional environment via the one or more display generation components, wherein: Displaying the first augmented reality experience includes displaying a first set of interactive elements; Displaying the second augmented reality experience includes displaying a second set of interactive elements different from the first set of interactive elements; The first representation of the first augmented reality experience includes a representation of the first set of interactive elements; and The second representation of the second augmented reality experience includes a representation of the second set of interactive elements different from the representation of the first set of interactive elements.
4. The method according to claim 3, wherein: Displaying the first augmented reality experience includes displaying the first set of interactive elements superimposed on the see-through environment; Displaying the second augmented reality experience includes displaying the second set of interactive elements superimposed on the see-through environment; The first representation of the first augmented reality experience includes a first placeholder background content representing the see-through environment; and The second representation of the second augmented reality experience includes a second placeholder background content representing the see-through environment.
5. The method according to claim 3, wherein: The representation of the first set of interactive elements is non-interactive; and The representation of the second set of interactive elements is non-interactive.
6. The method according to claim 3, wherein: At least a portion of the first representation of the first augmented reality experience is displayed in a first color corresponding to the first augmented reality experience; and At least a portion of the second representation of the second augmented reality experience is displayed in a second color corresponding to the second augmented reality experience, wherein the second color is different from the first color.
7. The method according to claim 3, wherein: The first representation of the first augmented reality experience includes a first identifier corresponding to the first augmented reality experience; The second representation of the second augmented reality experience includes a second identifier that is different from the first identifier and corresponds to the second augmented reality experience; Displaying the first augmented reality experience includes displaying the first identifier as part of the first augmented reality experience; And Displaying the second augmented reality experience includes displaying the second identifier as part of the second augmented reality experience.
8. The method according to claim 3, the method further comprising: In response to receiving the first input: Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience: Before displaying the first augmented reality experience, display a first animation via the one or more display generation components, in which the first representation of the first augmented reality experience moves towards the viewpoint of the user of the computer system.
9. The method according to claim 8, wherein: The first representation of the first augmented reality experience includes a first boundary around the representation of the first set of interactive elements; The second representation of the second augmented reality experience includes a second boundary around the representation of the second set of interactive elements; and Displaying the first animation includes displaying the first boundary moving towards the viewpoint of the user of the computer system until the first boundary is no longer displayed.
10. The method according to claim 8, the method further comprising: In response to receiving the first input: Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience: Display a cross-fade between the representation of the first set of interactive elements and the first set of interactive elements via the one or more display generation components.
11. The method according to claim 3, the method further comprising: In response to receiving the first input: Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience: Stop displaying the representation of the first set of interactive elements; and Display the first set of interactive elements via the one or more display generation components.
12. The method according to any one of claims 1 to 2, the method further comprising: Before receiving the first user input, and when the representations of the plurality of augmented reality experiences are simultaneously displayed in the three-dimensional environment: Display the first representation of the first augmented reality experience at a first display position via the one or more display generation components; When the first representation of the first augmented reality experience is displayed at the first display position, receive a second user input corresponding to a request to navigate from the first representation of the first augmented reality experience to the second representation of the second augmented reality experience via the one or more input devices; And In response to receiving the second user input: Stop displaying the first representation of the first augmented reality experience at the first display position; And Display the second representation of the second augmented reality experience at the first display position via the one or more display generation components.
13. The method according to any one of claims 1 to 2, wherein simultaneously displaying the representations of the plurality of augmented reality experiences includes displaying the representations of the plurality of augmented reality experiences in a stacked form, wherein the first representation of the first augmented reality experience is stacked on top of the second representation of the second augmented reality experience.
14. The method according to any one of claims 1 to 2, the method further comprising: Before receiving the first user input and when displaying the representations of the plurality of augmented reality experiences simultaneously in the three-dimensional environment, including simultaneously displaying the first representation of the first augmented reality experience and the second representation of the second augmented reality experience, receiving a third user input via the one or more input devices corresponding to a request to navigate through the representations of the plurality of augmented reality experiences; And In response to receiving the third user input: Stop displaying the first representation of the first augmented reality experience while maintaining the display of the second representation of the second augmented reality experience.
15. The method according to any one of claims 1 to 2, wherein: Determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes determining that the first user input is a selection input, the selection input including: A gaze input towards the first representation of the first augmented reality experience; and A hardware press input detected when the gaze input is towards the first representation of the first augmented reality experience.
16. According to the method according to any one of claims 1 to 2, wherein: Determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes determining that the first user input is a selection input, the selection input including: A voice input indicating a user request to select an optional object.
17. According to the method according to any one of claims 1 to 2, wherein: Determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes determining that the first user input is a selection input, the selection input including: A gaze input towards the first representation of the first augmented reality experience, where the first augmented reality experience meets a first set of gaze duration criteria.
18. According to the method according to any one of claims 1 to 2, the method further comprising: When displaying the representations of the plurality of augmented reality experiences simultaneously, display one or more setting controls via the one or more display generation components, the one or more setting controls including a first setting control corresponding to a first setting of the computer system; When displaying the one or more setting controls, receive a first setting input corresponding to the first setting of the computer system via the one or more input devices; In response to receiving the first setting input, modify the first setting from a first value to a second value different from the first value; When displaying the representations of the plurality of augmented reality experiences simultaneously and when the first setting is set to the second value, receive a third user input via the one or more input devices; And In response to receiving the third user input: According to determining that the third user input corresponds to a selection of the first representation of the first augmented reality experience, display the first augmented reality experience in the three-dimensional environment via the one or more display generation components while maintaining the first setting at the second value; and According to determining that the first user input corresponds to a selection of the second representation of the second augmented reality experience, display the second augmented reality experience in the three-dimensional environment via the one or more display generation components while maintaining the first setting at the second value.
19. According to the method according to claim 18, wherein: The first setting is a passthrough shading setting; The first value corresponds to a first amount applied to the three-dimensional environment; and The second value corresponds to a second shading amount different from the first shading amount applied to the three-dimensional environment.
20. According to the method according to claim 18, wherein: The first setting is a volume setting; The first value corresponds to a first volume; and The second value corresponds to a second volume different from the first volume.
21. According to the method according to claim 18, the method further comprising: When the representations of the plurality of augmented reality experiences and the one or more setting controls are displayed simultaneously, device status information indicating the status of one or more characteristics of the computer system is displayed via the one or more display generation components.
22. According to the method according to any one of claims 1 to 2, wherein the representations of the plurality of augmented reality experiences are view-locked objects that remain in corresponding regions of the user's field of view when the viewpoint of the user of the computer system shifts relative to the three-dimensional environment.
23. According to the method according to claim 22, wherein: Simultaneously displaying the representations of the plurality of augmented reality experiences includes simultaneously displaying the representations of the plurality of augmented reality experiences in a first orientation, in which the representations of the plurality of augmented reality experiences are aligned with gravity; And The method further includes: When the representations of the plurality of augmented reality experiences are displayed simultaneously, detecting a change in the orientation of the user's viewpoint; And In response to detecting the change in the orientation of the user's viewpoint: Based on the change in the orientation of the user's viewpoint, rotating the representations of the plurality of augmented reality experiences from the first orientation to a second orientation to continue aligning the representations of the plurality of augmented reality experiences with gravity.
24. The method according to claim 23, wherein rotating the representations of the plurality of augmented reality experiences from the first orientation to the second orientation comprises: At a first time after detecting the change in the orientation of the user's viewpoint, the representations of the plurality of augmented reality experiences are displayed in the first orientation via the one or more display generation components, where at the first time, the representations of the plurality of augmented reality experiences are at least partially misaligned with gravity due to the change in the orientation of the user's viewpoint; And At a second time after the first time, the representations of the plurality of augmented reality experiences are displayed in the second orientation via the one or more display generation components to align the representations of the plurality of augmented reality experiences with gravity.
25. The method according to claim 22, wherein: Displaying the first augmented reality experience includes simultaneously displaying a first set of objects including a first object and a second object, and wherein: The first object is a viewpoint-locked object; and The second object is an environment-locked object.
26. The method according to any one of claims 1 to 2, the method further comprising: The first augmented reality experience is displayed in the three-dimensional environment via the one or more display generation components; When the first augmented reality experience is displayed, a first voice input indicating a user request to change from the first augmented reality experience to the second augmented reality experience is received via the one or more input devices; And In response to receiving the first voice input: Stop displaying the first augmented reality experience; And The second augmented reality experience is displayed in the three-dimensional environment via the one or more input devices.
27. The method according to any one of claims 1 to 2, the method further comprising: When the computer system is in a sleep state, a first wake-up input corresponding to a request to transition the computer system from the sleep state to the wake state is received via the one or more input devices; And In response to receiving the first wake-up input, the first augmented reality experience is displayed via the one or more display generation components.
28. The method according to any one of claims 1 to 2, the method further comprising: When the computer system is in a sleep state, receive a first wake-up input corresponding to a request to transition the computer system from the sleep state to a wake state via the one or more input devices; and In response to receiving the first wake-up input, display the representations of the plurality of augmented reality experiences via the one or more display generation components.
29. The method according to any one of claims 1 to 2, wherein the plurality of augmented reality experiences comprises one or more of the following: a camera augmented reality experience; a translation augmented reality experience; a reading augmented reality experience; a music augmented reality experience; a navigation augmented reality experience; a photo augmented reality experience; a video messaging augmented reality experience; and / or a fitness augmented reality experience.
30. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 29.
31. A computer system configured to communicate with one or more display generation components and one or more input devices, the computer system comprising: One or more processors; and A memory that stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 29.
32. A computer system configured to communicate with one or more display generation components and one or more input devices, the computer system comprising: Means for performing the method according to any one of claims 1 to 29.
33. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for performing the method according to any one of claims 1 to 29.