Devices, methods, and graphical user interfaces for interacting with augmented reality experience
By communicating with display generation components and input devices in a computer system, the representation of augmented reality experience is realized selectively displayed in a three-dimensional environment, and the problems of interaction complexity and inefficiency in the prior art are solved, and the efficiency and intuitiveness of user interaction are improved.
Patent Information
- Application Number
- CN202380067076.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-15
- Filing Date
- 2023-09-21
- Publication Date
- 2025-05-13
AI Technical Summary
The methods and interfaces for interacting with extended real-life experience in the prior art have problems such as insufficient feedback, complex operations, cumbersome and error-prone, resulting in a large cognitive burden on users and low interaction efficiency.
By communicating with the display generation components and input devices in a computer system, representations of multiple augmented reality experiences are realized simultaneously in a three-dimensional environment and selectively display or hide these representations according to user inputs to simplify the user interaction process.
This method reduces the number and complexity of user input, improves the efficiency and intuitiveness of the human-computer interface, reduces the power consumption of the battery-driven device, and extends the battery life of the device.
Smart Images

Figure CN119998762A_ABST
Abstract
Description
Technical Field
[0001] This application claims priority to U.S. patent application Ser. No. 18 / 369,075, filed Sep. 15, 2023, entitled “DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR INTERACTING WITH EXTENDED REALITY EXPERIENCES,” and U.S. Provisional Patent Application Ser. No. 63 / 538,453, filed Sep. 14, 2023, entitled “DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR INTERACTING WITH EXTENDED REALITY EXPERIENCES,” and U.S. Provisional Patent Application Ser. No. 63 / 538,453, filed Sep. 14, 2023, entitled “DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR INTERACTING WITH EXTENDED REALITY EXPERIENCES,” and U.S. Provisional Patent Application Ser. No. 63 / 538,453, filed Sep. 14, 2023, entitled “DEVICES, METHODS, AND GRAPHICAL USER INTERFACES FOR INTERACTING WITH EXTENDED REALITY The contents of each of these patent applications are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure generally relates to computer systems that provide computer-generated experiences in communication with one or more display generation components and one or more input devices, including but not limited to electronic devices that provide virtual reality experiences and mixed reality experiences via displays. Background Art
[0004] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices for computer systems and other electronic computing devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touchscreen displays) are used to interact with virtual / augmented reality environments. Example virtual elements include virtual objects such as digital images, videos, text, icons, and control elements (such as buttons and other graphics). Summary of the Invention
[0005] Some methods and interfaces for interacting with environments that include at least some virtual elements (e.g., applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve desired results in an augmented reality environment, and systems where virtual object manipulation is complex, cumbersome, and error-prone can place a significant cognitive burden on users and detract from the experience of the virtual / augmented reality environment. Furthermore, these methods take longer than necessary, wasting the computer system's energy. This latter consideration is particularly important in battery-powered devices.
[0006] Therefore, there is a need for computer systems with improved methods and interfaces for providing computer-generated experiences (such as, for example, extended reality experiences) to users, thereby making user interactions with computer systems more efficient and intuitive for the users. Such methods and interfaces optionally supplement or replace conventional methods for providing extended reality experiences to users. Such methods and interfaces reduce the amount, extent, and / or nature of inputs from the user by helping the user understand the connection between the inputs provided and the device's responses to those inputs, thereby forming a more effective human-computer interface.
[0007] The above-mentioned defects and other problems associated with the user interface of the computer system are reduced or eliminated by the disclosed system. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a notebook computer, a tablet computer, or a handheld device). In some embodiments, the computer system is a personal electronic device (e.g., a wearable electronic device, such as a watch or a head-mounted device). In some embodiments, the computer system has a touch pad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye tracking components. In some embodiments, the computer system has one or more hand tracking components. In some embodiments, in addition to the display generation component, the computer system also has one or more output devices, which include one or more tactile output generators and / or one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, a program or instruction set stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contacts and gestures of a stylus and / or fingers on a touch-sensitive surface, movement of the user's eyes and hands in space relative to the GUI (and / or computer system) or the user's body (as captured by a camera and other motion sensors), and / or voice input (as captured by one or more audio input devices). In some embodiments, the functions performed by interaction optionally include image editing, drawing, presentations, word processing, spreadsheet creation, playing games, making and receiving calls, video conferencing, sending and receiving emails, instant messaging, test support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. Executable instructions for performing these functions are optionally included in a transient and / or non-transient computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0008] There is a need for electronic devices with improved methods and interfaces for interacting with extended reality experiences. Such methods and interfaces can supplement or replace conventional methods for interacting with extended reality experiences. Such methods and interfaces reduce the amount, extent, and / or nature of input from the user and produce a more efficient human-computer interface. For battery-powered computing devices, such methods and interfaces conserve power and increase the time between battery charges.
[0009] According to some embodiments, a method is described. The method includes: at a computer system in communication with one or more display generation components and one or more input devices: simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; while simultaneously displaying the representations of the multiple augmented reality experiences in the three-dimensional environment, receiving, via the one or more input devices, a first user input; and in response to receiving the first user input: ceasing display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0010] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations comprising: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; while simultaneously displaying the representations of the multiple augmented reality experiences in the three-dimensional environment, receiving, via the one or more input devices, a first user input; and in response to receiving the first user input: ceasing display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0011] According to some embodiments, a transient computer-readable storage medium is described. In some embodiments, the transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations comprising: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; while simultaneously displaying the representations of the multiple augmented reality experiences in the three-dimensional environment, receiving, via the one or more input devices, a first user input; and in response to receiving the first user input: ceasing display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0012] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; while simultaneously displaying the representations of the multiple augmented reality experiences in the three-dimensional environment, receiving, via the one or more input devices, a first user input; and in response to receiving the first user input: ceasing display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0013] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: means for simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; means for receiving, via the one or more input devices, a first user input while simultaneously displaying the representations of the multiple augmented reality experiences in the three-dimensional environment; and means for, in response to receiving the first user input, performing the following operations: stopping display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences; and displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience.
[0014] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: simultaneously displaying, via the one or more display generation components, representations of multiple augmented reality experiences in a three-dimensional environment, the representations including: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; while simultaneously displaying the representations of the multiple augmented reality experiences in the three-dimensional environment, receiving, via the one or more input devices, a first user input; and in response to receiving the first user input: ceasing display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences; and based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience, displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment.
[0015] According to some embodiments, a method is described. The method includes: at a computer system in communication with one or more display generation components and one or more input devices: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience in the three-dimensional environment that is different from the first extended reality experience via the one or more display generation components.
[0016] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0017] According to some embodiments, a transient computer-readable storage medium is described. In some embodiments, the transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0018] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of the one or more user inputs: based on determining that the first sequence of the one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of the one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0019] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: a component for receiving a first sequence of one or more user inputs via a first physical control; and a component for performing the following operations in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0020] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: receiving a first sequence of one or more user inputs via a first physical control; and in response to receiving the first sequence of one or more user inputs: based on determining that the first sequence of one or more user inputs has a first magnitude: displaying a first extended reality experience in a three-dimensional environment via the one or more display generation components; and based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: displaying a second extended reality experience different from the first extended reality experience in the three-dimensional environment via the one or more display generation components.
[0021] According to some embodiments, a method is described that includes: at a computer system in communication with one or more display generation components and one or more input devices: detecting, via the one or more input devices, a first set of conditions in a three-dimensional environment in which the computer system is located, while a view of the three-dimensional environment is visible; and in response to detecting the first set of conditions in the three-dimensional environment: displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
[0022] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: detecting, when a view of a three-dimensional environment in which the computer system is located is visible, via the one or more input devices, a first set of conditions in the three-dimensional environment; and in response to detecting the first set of conditions in the three-dimensional environment, displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
[0023] According to some embodiments, a transient computer-readable storage medium is described. In some embodiments, the transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: detecting, when a view of a three-dimensional environment in which the computer system is located is visible, via the one or more input devices, a first set of conditions in the three-dimensional environment; and in response to detecting the first set of conditions in the three-dimensional environment, displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
[0024] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting, when a view of a three-dimensional environment in which the computer system is located is visible, via the one or more input devices, a first set of conditions in the three-dimensional environment; and in response to detecting the first set of conditions in the three-dimensional environment, displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
[0025] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: means for detecting, via the one or more input devices, a first set of conditions in a three-dimensional environment when a view of the three-dimensional environment in which the computer system is located is visible; and means for, in response to detecting the first set of conditions in the three-dimensional environment, displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
[0026] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: detecting, when a view of a three-dimensional environment in which the computer system is located is visible, via the one or more input devices, a first set of conditions in the three-dimensional environment; and responsive to detecting the first set of conditions in the three-dimensional environment, displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience, wherein the first augmented reality experience is selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
[0027] According to some embodiments, a method is described. The method includes: at a computer system in communication with one or more display generation components and one or more input devices: detecting, via the one or more input devices, a gaze of a user corresponding to a first display location of the one or more display generation components; in response to detecting the gaze of the user corresponding to the first display location of the one or more display generation components, displaying, via the one or more display generation components, a first object; while displaying the first object, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying, via the one or more display generation components, movement of the first object; and after displaying the movement of the first object: performing a first operation based on determining that the gaze of the user satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and forgoing performing the first operation based on determining that the gaze of the user does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0028] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for the following operations: detecting a user's gaze corresponding to a first display position of the one or more display generation components via the one or more input devices; in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components, displaying a first object via the one or more display generation components; when displaying the first object, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying movement of the first object via the one or more display generation components; and after displaying the movement of the first object: performing a first operation based on determining that the user's gaze satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and abandoning performing the first operation based on determining that the user's gaze does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0029] According to some embodiments, a transient computer-readable storage medium is described. In some embodiments, the transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for the following operations: detecting, via the one or more input devices, a user's gaze corresponding to a first display position of the one or more display generation components; in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components, displaying a first object via the one or more display generation components; when displaying the first object, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying movement of the first object via the one or more display generation components; and after displaying the movement of the first object: performing a first operation based on determining that the user's gaze satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and abandoning performing the first operation based on determining that the user's gaze does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0030] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: detecting, via the one or more input devices, a gaze of a user corresponding to a first display position of the one or more display generation components; in response to detecting the gaze of the user corresponding to the first display position of the one or more display generation components, displaying a first object via the one or more display generation components; when displaying the first object, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying movement of the first object via the one or more display generation components; and after displaying the movement of the first object: performing a first operation based on determining that the gaze of the user satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and abandoning performing the first operation based on determining that the gaze of the user does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0031] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: a component for detecting a user's gaze corresponding to a first display position of the one or more display generation components via the one or more input devices; a component for displaying a first object via the one or more display generation components in response to detecting the user's gaze corresponding to the first display position of the one or more display generation components; a component for detecting that a first set of criteria is satisfied when displaying the first object; a component for displaying movement of the first object via the one or more display generation components in response to detecting that the first set of criteria is satisfied; and a component for performing the following operations after displaying the movement of the first object: performing a first operation based on determining that the user's gaze satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and abandoning the first operation based on determining that the user's gaze does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0032] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: detecting, via the one or more input devices, a gaze of a user corresponding to a first display position of the one or more display generation components; in response to detecting the gaze of the user corresponding to the first display position of the one or more display generation components, displaying a first object via the one or more display generation components; when displaying the first object, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying movement of the first object via the one or more display generation components; and after displaying the movement of the first object: performing a first operation based on determining that the gaze of the user satisfies a second set of criteria indicating gaze tracking of the movement of the first object; and abandoning performing the first operation based on determining that the gaze of the user does not satisfy the second set of criteria indicating gaze tracking of the movement of the first object.
[0033] According to some embodiments, a method is described that includes: at a computer system in communication with one or more display generation components and one or more input devices: displaying virtual content via the one or more display generation components; while displaying the virtual content, detecting, via the one or more input devices, a first gesture in front of a face of a user of the computer system; and in response to detecting the first gesture: based on a determination that the first gesture in front of the face of the user satisfies a first set of criteria, ceasing display of at least a portion of the virtual content; and based on a determination that the first gesture in front of the face of the user does not satisfy the first set of criteria, maintaining display of the virtual content.
[0034] According to some embodiments, a non-transitory computer-readable storage medium is described. In some embodiments, the non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; detecting a first gesture in front of a face of a user of the computer system via the one or more input devices while displaying the virtual content; and in response to detecting the first gesture: based on determining that the first gesture in front of the face of the user meets a first set of criteria, ceasing display of at least a portion of the virtual content; and based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.
[0035] According to some embodiments, a transient computer-readable storage medium is described. In some embodiments, the transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; detecting a first gesture in front of a face of a user of the computer system via the one or more input devices while displaying the virtual content; and in response to detecting the first gesture: based on determining that the first gesture in front of the face of the user meets a first set of criteria, ceasing display of at least a portion of the virtual content; and based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.
[0036] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; detecting a first gesture in front of a face of a user of the computer system via the one or more input devices while displaying the virtual content; and in response to detecting the first gesture: based on determining that the first gesture in front of the face of the user meets a first set of criteria, ceasing display of at least a portion of the virtual content; and based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.
[0037] According to some embodiments, a computer system is described. In some embodiments, the computer system is configured to communicate with one or more display generation components and one or more input devices, and the computer system includes: means for displaying virtual content via the one or more display generation components; means for detecting, via the one or more input devices, a first gesture in front of a face of a user of the computer system while displaying the virtual content; and means for, in response to detecting the first gesture, performing the following operations: based on determining that the first gesture in front of the face of the user meets a first set of criteria, ceasing display of at least a portion of the virtual content; and based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.
[0038] According to some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generation components and one or more input devices, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; detecting a first gesture in front of a face of a user of the computer system via the one or more input devices while displaying the virtual content; and in response to detecting the first gesture: based on a determination that the first gesture in front of the face of the user satisfies a first set of criteria, ceasing display of at least a portion of the virtual content; and based on a determination that the first gesture in front of the face of the user does not satisfy the first set of criteria, maintaining display of the virtual content.
[0039] It should be noted that the various embodiments described above can be combined with any other embodiment described herein. The features and advantages described in this specification are not comprehensive. In particular, many additional features and advantages will be apparent to those skilled in the art from the drawings, the specification, and the claims. In addition, it should be noted that the language used in this specification has been selected in principle for readability and instructional purposes, and may not be selected to describe or define the subject matter of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] For a better understanding of the various described embodiments, reference should be made to the following detailed description taken in conjunction with the following drawings, wherein like reference numerals designate corresponding parts throughout the several views.
[0041] Figure 1A is a block diagram illustrating an operating environment for a computer system for providing an XR experience according to some embodiments.
[0042] Figure 1B to Figure 1P is used in Figure 1A An example of a computer system that provides an XR experience in an operating environment.
[0043] Figure 2 is a block diagram illustrating a controller of a computer system configured to manage and coordinate a user's XR experience according to some embodiments.
[0044] Figure 3 is a block diagram illustrating display generation components of a computer system configured to provide the visual component of an XR experience to a user according to some embodiments.
[0045] Figure 4 is a block diagram illustrating a hand tracking unit of a computer system configured to capture gesture input from a user according to some embodiments.
[0046] Figure 5 is a block diagram illustrating an eye tracking unit of a computer system configured to capture gaze input from a user according to some embodiments.
[0047] Figure 6 is a flow chart illustrating a flash-assisted gaze tracking pipeline according to some embodiments.
[0048] Figures 7A to 7K Example techniques for navigating an extended reality experience according to some embodiments are illustrated.
[0049] Figure 8 is a flowchart of a method of navigating an extended reality experience according to various embodiments.
[0050] Figure 9is a flowchart of a method of navigating an extended reality experience according to various embodiments.
[0051] Figures 10A to 10G Example techniques for providing recommendations related to extended reality experiences according to some embodiments are illustrated.
[0052] Figure 11 is a flow chart of a method of providing recommendations related to an extended reality experience according to various embodiments.
[0053] Figures 12A to 12K Example techniques for gaze-based interaction according to some embodiments are illustrated.
[0054] Figure 13 is a flowchart of a method for gaze-based interaction according to some embodiments.
[0055] Figures 14A to 14L Example techniques for interacting with virtual content according to various embodiments are illustrated.
[0056] Figure 15 is a flow chart of a method of interacting with virtual content according to some embodiments. DETAILED DESCRIPTION
[0057] According to some embodiments, the present disclosure relates to a user interface for providing an extended reality (XR) experience to a user.
[0058] The systems, methods, and GUIs described herein improve user interface interaction with augmented reality environments and other virtual content in a number of ways.
[0059] In some embodiments, a computer system simultaneously displays representations of multiple augmented reality experiences in a three-dimensional environment, the representations including a first representation of a first augmented reality experience and a second representation of a second augmented reality experience. While the representations of the multiple augmented reality experiences are displayed simultaneously, the computer system receives a first user input. In response to receiving the first user input, the computer system stops displaying the representations of one or more of the multiple augmented reality experiences and, based on the direction and / or magnitude of the first user input, displays either the first augmented reality experience or the second augmented reality experience. The computer system thereby provides the user with the ability to switch between different augmented reality experiences in an intuitive and efficient manner.
[0060] In some embodiments, the computer system receives a first sequence of one or more user inputs via a first physical control. In some embodiments, the first physical control is a rotatable and depressible physical control, enabling the user to provide rotational input and depressive input via the first physical control. In response to receiving the first sequence of one or more user inputs, the computer system displays a first extended reality experience or a second extended reality experience based on the direction and / or magnitude of the first sequence of one or more user inputs. The computer system thereby provides the user with the ability to switch between different augmented reality experiences in an intuitive and efficient manner.
[0061] In some embodiments, the computer system detects a first set of conditions in a three-dimensional environment in which the computer system is located. For example, in various embodiments, the computer system detects one or more visual objects, audio content, or other conditions in the physical environment in which the computer system is located. In response to detecting the first set of conditions in the three-dimensional environment, the computer system displays a first suggestion corresponding to a first augmented reality experience. The first augmented reality experience is selected from a plurality of augmented reality experiences that can be displayed by the computer system. For example, in some embodiments, the first augmented reality experience is selected based on the first set of conditions in the three-dimensional environment. The computer system thereby provides a suggestion of potentially relevant augmented reality experiences to the user based on the conditions detected by the computer system.
[0062] In some embodiments, a computer system detects a gaze of a user corresponding to a first display position of one or more display generating components. In response to detecting the gaze of the user corresponding to the first display position, the computer system displays a first object. For example, the computer system displays a gaze target that the user intends to track with his or her eyes. When displaying the first object, the computer system detects that a first set of criteria are met, and in response to detecting that the first set of criteria are met, the computer system displays the movement of the first object. If the user successfully tracks the movement of the first object with his or her gaze, the computer system performs a first operation, and if the user does not successfully track the movement of the first object with his or her gaze, the computer system does not perform the first operation. For example, in some embodiments, the user can unlock the computer system using a gaze input by tracking the movement of the first object. The computer system thus provides the user with an intuitive and efficient way to perform operations such as unlocking the computer system using a gaze input.
[0063] In some embodiments, a computer system displays virtual content. While the virtual content is displayed, the computer system detects a first gesture in front of the face of a user of the computer system. If the first gesture meets a first set of criteria, the computer system stops displaying at least a portion of the virtual content. In this manner, the computer system allows the user to quickly and easily clear some or all of the virtual content using a gesture.
[0064] Figures 1A to 6 A description of an example computer system for providing an XR experience to a user is provided. Figures 7A to 7K Example techniques for navigating an extended reality experience according to some embodiments are illustrated. Figure 8 is a flowchart of a method of navigating an extended reality experience according to various embodiments. Figure 9 is a flowchart of a method of navigating an extended reality experience according to some embodiments. Figures 7A to 7K The user interface in Figure 8 and Figure 9 in the process. Figures 10A to 10G Example techniques for providing recommendations related to extended reality experiences according to some embodiments are illustrated. Figure 11 is a flow chart of a method of providing recommendations related to an extended reality experience according to various embodiments. Figures 10A to 10G The user interface is used to illustrate Figure 11 in the process. Figures 12A to 12K Example techniques for gaze-based interaction according to some embodiments are illustrated. Figure 13 is a flowchart of a method of gaze-based interaction according to various embodiments. Figures 12A to 12K The user interface in Figure 13 in the process. Figures 14A to 14L Example techniques for interacting with virtual content according to some embodiments are illustrated. Figure 15 is a flow chart of a method of interacting with virtual content according to various embodiments. Figures 14A to 14L The user interface in Figure 15 in the process.
[0065] The processes described below enhance the operability of the device and make the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device) through various techniques, including by providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation without further user input when a set of conditions have been met, improving privacy and / or security, providing a richer, more detailed, and / or more realistic user experience while saving storage space, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device faster and more efficiently. This saves battery power and, therefore, weight, improving the ergonomics of the device. These techniques also enable real-time communication, allow the use of fewer and / or less precise sensors, resulting in a more compact, lighter, and less expensive device, and enable the device to be used in a variety of lighting conditions. These techniques reduce energy usage and thereby reduce the heat emitted by the device, which is particularly important for wearable devices where if the device generates too much heat well within the operating parameters of the device components, it may become uncomfortable for the user to wear the device.
[0066] In addition, in the method described herein where one or more steps depend on having met one or more conditions, it should be understood that the method can be repeated in multiple repetitions so that in the process of repetition, all conditions of the steps in the method of determining the method have been met in different repetitions of the method. For example, if the method needs to perform the first step (if the condition is met), and perform the second step (if the condition is not met), then those of ordinary skill will know that the steps stated are repeated until both the condition is met and the condition is not met (in no particular order). Therefore, the method described as having one or more steps depending on having met one or more conditions can be rewritten as a method of repeating until each condition described in the method is met. However, this does not require a system or computer-readable medium to declare that the system or computer-readable medium includes instructions for performing a contingent operation based on the satisfaction of the corresponding one or more conditions, and is therefore able to determine whether a possible situation has been met without explicitly repeating the steps of the method until all conditions of the steps in the method of determining the method have been met. Those of ordinary skill in the art will also understand that, similar to the method with a contingent step, a system or computer-readable storage medium can repeat the steps of the method as needed multiple times to ensure that all contingent steps have been performed.
[0067] In some embodiments, as Figure 1AAs shown, an XR experience is provided to a user via an operating environment 100 including a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., a household appliance, a wearable device, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).
[0068] When describing an XR experience, various terms are used to distinctly refer to several related but distinct environments that a user can sense and / or with which the user can interact (e.g., using inputs detected by the computer system 101 generating the XR experience, which inputs cause the computer system generating the XR experience to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:
[0069] Physical Environment: The physical environment refers to the physical world that people can sense and / or interact with without the aid of electronic systems. A physical environment, such as a physical park, includes physical objects, such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through sight, touch, hearing, taste, and smell.
[0070] Extended Reality: In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via electronic systems. In XR, a subset of a person's physical movements, or representations thereof, is tracked, and in response, one or more properties of one or more virtual objects simulated in the XR environment are adjusted in a manner consistent with at least one law of physics. For example, an XR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), adjustments to the properties of virtual objects in the XR environment can be made in response to representations of physical movement (e.g., voice commands). People can sense and / or interact with XR objects using any of their senses, including vision, hearing, touch, taste, and smell. For example, people can sense and / or interact with audio objects, which create a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. As another example, audio objects can enable audio transparency, which selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some XR environments, people can sense and / or interact only with audio objects.
[0071] Examples of XR include virtual reality and mixed reality.
[0072] Virtual Reality: A virtual reality (VR) environment is a simulated environment designed to be based entirely on computer-generated sensory input to one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing human avatars are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through the simulation of the person's presence within the computer-generated environment and / or through the simulation of a subset of the person's physical movement within the computer-generated environment.
[0073] Mixed Reality: In contrast to VR environments, which are designed to be based entirely on computer-generated sensory input, a mixed reality (MR) environment refers to a simulated environment that is designed to include sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). On the virtuality continuum, a mixed reality environment is anything between, but not including, a fully physical environment at one end and a virtual reality environment at the other. In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. In addition, some electronic systems used to render MR environments can track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical items from the physical environment, or representations thereof). For example, the system can cause motion so that virtual trees appear stationary relative to the physical ground.
[0074] Examples of mixed reality include augmented reality and augmented virtuality.
[0075] Augmented Reality: An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on a transparent or translucent display so that a person uses the system to perceive the virtual objects superimposed on the physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with the virtual objects and presents the combination on the opaque display. A person uses the system to indirectly view the physical environment via the images or videos of the physical environment and perceives the virtual objects superimposed on the physical environment. As used herein, a video of the physical environment displayed on an opaque display is referred to as "transparent video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system may have a projection system that projects virtual objects into a physical environment, such as as holograms or on a physical surface, so that a person using the system perceives virtual objects superimposed on the physical environment. An augmented reality environment also refers to a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing a pass-through video, the system may transform one or more sensor images to apply a selected perspective (e.g., a viewpoint) that is different from the perspective captured by the imaging sensor. For another example, the representation of the physical environment may be transformed by graphically modifying (e.g., enlarging) a portion thereof so that the modified portion may be a representative but not real version of the original captured image. For another example, the representation of the physical environment may be transformed by graphically eliminating a portion thereof or blurring a portion thereof.
[0076] Augmented Virtual: An augmented virtual (AV) environment is a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but people's faces are realistically reproduced from images taken of physical people. In another example, a virtual object can adopt the shape or color of a physical object imaged by one or more imaging sensors. In another example, a virtual object can adopt a shadow that conforms to the position of the sun in the physical environment.
[0077] In an augmented reality, mixed reality, or virtual reality environment, a view of a three-dimensional environment is visible to a user. The view of the three-dimensional environment is typically visible to the user through a virtual viewport via one or more display generation components (e.g., a display or a pair of display modules that provide stereoscopic content to different eyes of the same user), and the virtual viewport has a viewport boundary that defines the range of the three-dimensional environment visible to the user via the one or more display generation components. In some embodiments, the area defined by the viewport boundary is smaller than the user's visual range in one or more dimensions (e.g., based on the user's visual range, the size, optical properties, or other physical properties of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). In some embodiments, the area defined by the viewport boundary is larger than the user's visual range in one or more dimensions (e.g., based on the user's visual range, the size, optical properties, or other physical properties of the one or more display generation components, and / or the position and / or orientation of the one or more display generation components relative to the user's eyes). The viewport and viewport boundary typically move with the movement of one or more display generation components (e.g., with the user's head for a head-mounted device, or with the user's hand for a handheld device such as a tablet or smart phone). The user's viewpoint determines what is visible in the viewport. The viewpoint typically specifies a position and orientation relative to the three-dimensional environment, and as the viewpoint moves, the view of the three-dimensional environment will also move in the viewport. For head-mounted devices, the viewpoint is typically based on the position and orientation of the user's head, face, and / or eyes to provide a view of the three-dimensional environment that is perceptually accurate and provides an immersive experience when the user is using the head-mounted device. For handheld or fixed devices, the viewpoint moves as the handheld or fixed device moves and / or as the user's positioning relative to the handheld or fixed device changes (e.g., the user moves toward, away from, up, down, right, and / or left). For devices that include display generation components with virtual pass-through, portions of the physical environment that are visible (e.g., displayed and / or projected) via one or more display generation components are based on the field of view of one or more cameras in communication with the display generation components, which typically move with movement of the display generation components (e.g., with movement of the user's head for a head-mounted device, or with movement of the user's hands for a handheld device such as a tablet or smartphone) as the user's viewpoint moves with movement of the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via the one or more display generation components is updated based on the user's viewpoint (e.g., the display position and pose of the virtual objects are updated based on movement of the user's viewpoint)).For display generation components with optical transmittance, portions of the physical environment that are visible via one or more display generation components (e.g., optically visible through one or more partially or fully transparent portions of the display generation components) are based on the user's field of view through the partially or fully transparent portions of the display generation components (e.g., moves as the user's head moves for a head-mounted device, or moves as the user's hands move for a handheld device such as a tablet or smartphone) because the user's viewpoint moves as the user moves through the field of view of the partially or fully transparent portions of the display generation components (and the appearance of one or more virtual objects is updated based on the user's viewpoint).
[0078] In some embodiments, the representation of the physical environment (e.g., displayed via virtual see-through or optical see-through) may be partially or completely obscured by the virtual environment. In some embodiments, the amount of the virtual environment displayed (e.g., the amount of the physical environment that is not displayed) is based on the immersion level of the virtual environment (e.g., relative to the representation of the physical environment). For example, increasing the immersion level optionally causes more of the virtual environment to be displayed, replacing and / or obscuring more of the physical environment, and decreasing the immersion level optionally causes less of the virtual environment to be displayed, thereby revealing portions of the physical environment that were previously not displayed and / or obscured. In some embodiments, at a particular immersion level, one or more first background objects (e.g., in the representation of the physical environment) are visually de-emphasized (e.g., dimmed, blurred, displayed with increased transparency) more than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, the immersion level includes an associated degree to which virtual content displayed by the computer system (e.g., a virtual environment and / or virtual content) obscures background content (e.g., content other than the virtual environment and / or virtual content) surrounding / behind the virtual environment, optionally including the number of items of background content displayed and / or the displayed visual characteristics of the background content (e.g., color, contrast, and / or opacity), the angular range of the virtual content displayed via the display generation component (e.g., 60 degrees for content displayed at low immersion, 120 degrees for content displayed at medium immersion, or 180 degrees for content displayed at high immersion), and / or the proportion of the field of view displayed via the display generation component that is occupied by the virtual content (e.g., 33% of the field of view occupied by the virtual content at low immersion, 66% of the field of view occupied by the virtual content at medium immersion, or 100% of the field of view occupied by the virtual content at high immersion). In some embodiments, the background content is included in the background on which the virtual content is displayed (e.g., background content in a representation of the physical environment). In some embodiments, the background content includes a user interface (e.g., a user interface corresponding to an application generated by a computer system), virtual objects that are not associated with or included in the virtual environment and / or virtual content (e.g., files generated by a computer system or representations of other users, etc.), and / or real objects (e.g., see-through objects representing real objects in the physical environment surrounding the user, which are visible so that they are displayed via the display generation component and / or visible via transparent or translucent components of the display generation component because the computer system does not block / impede their visibility through the display generation component). In some embodiments, at a low immersion level (e.g., a first immersion level), the background, virtual and / or real objects are displayed in an unobstructed manner. For example, a virtual environment with a low immersion level is optionally displayed simultaneously with the background content, which is optionally displayed at full brightness, color and / or translucency.In some embodiments, at a higher immersion level (e.g., a second immersion level that is higher than the first immersion level), background, virtual and / or real objects are displayed in an obscured manner (e.g., dimmed, blurred, or removed from the display). For example, a corresponding virtual environment with a high immersion level is displayed without simultaneously displaying background content (e.g., in full screen or fully immersive mode). As another example, a virtual environment displayed at a medium immersion level is displayed simultaneously with background content that is dimmed, blurred, or otherwise de-emphasized. In some embodiments, the visual characteristics of background objects vary between background objects. For example, at a particular immersion level, one or more first background objects are visually de-emphasized (e.g., dimmed, blurred, and / or displayed with increased transparency) more than one or more second background objects, and one or more third background objects cease to be displayed. In some embodiments, zero immersion or zero immersion level corresponds to a virtual environment that ceases to be displayed, and instead displays a representation of the physical environment (optionally with one or more virtual objects, such as applications, windows, or virtual three-dimensional objects), without the representation of the physical environment being obscured by the virtual environment. Adjusting the immersion level using physical input elements provides a fast and efficient method of adjusting immersion, which enhances the operability of the computer system and makes the user-device interface more efficient.
[0079] Viewpoint-locked virtual objects: When a computer system displays a virtual object at the same position and / or location in a user's viewpoint, the virtual object is viewpoint-locked even if the user's viewpoint shifts (e.g., changes). In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the forward direction of the user's head (e.g., when the user is looking straight ahead, the user's viewpoint is at least a portion of the user's field of view); thus, without moving the user's head, the user's viewpoint remains fixed even when the user's gaze shifts. In embodiments where the computer system has a display generation component (e.g., a display screen) that is repositionable relative to the user's head, the user's viewpoint is the augmented reality view presented to the user on the display generation component of the computer system. For example, a viewpoint-locked virtual object that is displayed in the upper left corner of the user's viewpoint when the user's viewpoint is in a first orientation (e.g., the user's head is facing north) continues to be displayed in the upper left corner of the user's viewpoint even when the user's viewpoint changes to a second orientation (e.g., the user's head is facing west). In other words, the position and / or location of the viewpoint-locked virtual object displayed in the user's viewpoint is independent of the user's position and / or orientation in the physical environment. In embodiments where the computer system is a head-mounted device, the user's viewpoint is locked to the orientation of the user's head, such that the virtual object is also referred to as a "head-locked virtual object."
[0080] Environment-locked visual objects: A virtual object is environment-locked (alternatively, "world-locked") when a computer system displays it at a location and / or position in a user's viewpoint that is based on (e.g., selected with reference to and / or anchored to) a location and / or object in a three-dimensional environment (e.g., a physical environment or a virtual environment). As the user's viewpoint moves, the location and / or objects in the environment change relative to the user's viewpoint, which causes the environment-locked virtual object to be displayed at a different location and / or position in the user's viewpoint. For example, an environment-locked virtual object locked to a tree immediately in front of the user is displayed at the center of the user's viewpoint. When the user's viewpoint shifts to the right (e.g., the user's head turns to the right) such that the tree is now to the left of center in the user's viewpoint (e.g., the tree's position in the user's viewpoint shifts), the environment-locked virtual object locked to the tree is displayed to the left of center in the user's viewpoint. In other words, the position and / or location at which an environment-locked virtual object is displayed in the user's viewpoint depends on the position and / or orientation of the object in the environment to which the virtual object is locked. In some embodiments, the computer system uses a stationary reference frame (e.g., a coordinate system anchored to fixed locations and / or objects in the physical environment) to determine the location at which an environment-locked virtual object is displayed in the user's viewpoint. An environment-locked virtual object can be locked to a stationary portion of the environment (e.g., a floor, wall, table, or other stationary object), or can be locked to a movable portion of the environment (e.g., a vehicle, animal, person, or even a representation of a part of the user's body that moves independently of the user's viewpoint, such as a hand, wrist, arm, or foot of the user) so that the virtual object moves as the viewpoint or that portion of the environment moves to maintain a fixed relationship between the virtual object and that portion of the environment.
[0081] In some embodiments, an environment-locked or viewpoint-locked virtual object exhibits an inertial following behavior that reduces or delays the movement of the environment-locked or viewpoint-locked virtual object relative to the movement of a reference point that the virtual object follows. In some embodiments, when exhibiting inertial following behavior, the computer system intentionally delays the movement of the virtual object when movement of a reference point (e.g., a portion of the environment, a viewpoint, or a point fixed relative to the viewpoint, such as a point between 5 cm and 300 cm from the viewpoint) that the virtual object is following is detected. For example, when the reference point (e.g., a portion of the environment or a viewpoint) moves at a first speed, the virtual object is moved by the device to remain locked to the reference point, but at a second speed that is slower than the first speed (e.g., until the reference point stops moving or slows down, at which point the virtual object begins to catch up with the reference point). In some embodiments, when the virtual object exhibits inertial following behavior, the device ignores small amounts of movement of the reference point (e.g., ignoring movements of the reference point below a threshold movement amount, such as movement of 0 to 5 degrees or movement of 0 to 50 cm). For example, when a reference point (e.g., a portion or viewpoint of an environment to which a virtual object is locked) moves a first amount, the distance between the reference point and the virtual object increases (e.g., because the virtual object is being displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and when the reference point (e.g., the portion or viewpoint of the environment to which the virtual object is locked) moves a second amount greater than the first amount, the distance between the reference point and the virtual object first increases (e.g., because the virtual object is being displayed so as to maintain a fixed or substantially fixed position relative to a viewpoint or portion of the environment different from the reference point to which the virtual object is locked), and then decreases when the amount of movement of the reference point increases above a threshold (e.g., a “lazy follow” threshold) because the virtual object is moved by the computer system to maintain a fixed or substantially fixed position relative to the reference point. In some embodiments, maintaining a substantially fixed position of the virtual object relative to the reference point includes displaying the virtual object within a threshold distance (e.g., 1 cm, 2 cm, 3 cm, 5 cm, 15 cm, 20 cm, 50 cm) of the reference point in one or more dimensions (e.g., up / down, left / right, and / or forward / backward relative to the position of the reference point).
[0082] Hardware: There are many different types of electronic systems that enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablet devices, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system for providing audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, a laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to selectively become opaque. Projection-based systems may employ retinal projection technology that projects graphic images onto a person's retina. The projection system may also be configured to project virtual objects into a physical environment, such as as a hologram or on a physical surface. In some embodiments, the controller 110 is configured to manage and coordinate the user's XR experience. In some embodiments, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is located locally or remotely relative to scene 105 (e.g., physical environment). For example, controller 110 is a local server located within scene 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., HMD, display, projector, touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In another example, the controller 110 is included within a housing (e.g., a physical housing) of the display generating component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), one or more input devices of the input devices 125, one or more output devices of the output devices 155, one or more sensors of the sensors 190, and / or one or more peripheral devices 195, or shares the same physical housing or support structure with one or more of the above devices.
[0083] In some embodiments, the display generation component 120 is configured to provide an XR experience (e.g., at least the visual component of the XR experience) to the user. In some embodiments, the display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Figure 3 Display generation component 120 is described in further detail. In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.
[0084] According to some embodiments, display generation component 120 provides an XR experience to the user when the user is virtually and / or physically present within scene 105.
[0085] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). In this way, the display generation component 120 includes one or more XR displays provided for displaying XR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or tablet device) configured to present XR content, and the user holds a device with a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is an XR room, housing, or room configured to present XR content, wherein the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying XR content (e.g., an HMD or other wearable computing device). For example, a user interface showing interactions with XR content that are triggered based on interactions occurring in the space in front of a handheld device or a tripod-mounted device can similarly be implemented with an HMD, where the interactions occur in the space in front of the HMD and the responses to the XR content are displayed via the HMD. Similarly, a user interface showing interactions with XR content that are triggered based on movement of a handheld device or a tripod-mounted device relative to a physical environment (e.g., scene 105 or a part of a user's body (e.g., the user's eyes, head, or hands)) can similarly be implemented with an HMD, where the movement is caused by movement of the HMD relative to the physical environment (e.g., scene 105 or a part of a user's body (e.g., the user's eyes, head, or hands)).
[0086] Despite Figure 1A Relevant features of the operating environment 100 are shown in FIG, but those skilled in the art will recognize from this disclosure that various other features are not illustrated for the sake of brevity and so as not to obscure more relevant aspects of the example embodiments disclosed herein.
[0087] Figures 1A to 1PVarious examples of computer systems for performing the methods and providing audio, visual, and / or tactile feedback as part of the user interfaces described herein are illustrated. In some embodiments, the computer system includes one or more display generation components (e.g., a first display assembly 1-120a and a second display assembly 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b) for displaying to a user of the computer system a representation of a virtual element and / or a physical environment, optionally generated based on detected events and / or user input detected by the computer system. The user interface generated by the computer system is optionally corrected by one or more corrective lenses 11.3.2-216, which are optionally removably attached to one or more of the optical modules to make the user interface easier to view by users who would otherwise use glasses or contact lenses to correct their vision. While many of the user interfaces shown herein show a single view of the user interface, a user interface in an HMD is optionally displayed using two optical modules (e.g., a first display component 1-120a and a second display component 1-120b and / or a first optical module 11.1.1-104a and a second optical module 11.1.1-104b), one optical module for the user's right eye and a different optical module for the user's left eye, and presenting slightly different images to the two different eyes to create the illusion of stereoscopic depth, the single view of the user interface being typically either a right eye view or a left eye view, with the depth effect being explained in text or using other diagrams or views. In some embodiments, a computer system includes one or more external displays (e.g., display component 1-108) for displaying status information for the computer system to a user of the computer system (when the computer system is not being worn) and / or to other people near the computer system, the status information being optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback, which is optionally generated based on detected events and / or user input detected by the computer system. In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting information about the physical environment of the device (e.g., one or more sensors in sensor components 1-356, and / or Figure 1I ), which information can be used (optionally in conjunction with one or more luminaires, such as Figure 1IIn some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting hand position and / or movement (e.g., sensor assembly 1-356 and / or sensor assembly 1-357). Figure 1I One or more sensors in ), which can be used (optionally in combination with one or more illuminators, such as Figure 1I In some embodiments, the computer system includes one or more input devices for detecting input, such as one or more sensors for detecting eye movement (e.g., Figure 1I eye tracking and gaze tracking sensors in the , which can be used (optionally in conjunction with one or more lights, such as Figure 1O11.3.2-110) determine attention or gaze location and / or gaze movement, which can optionally be used to detect gaze-only input based on gaze movement and / or dwell. Combinations of the various sensors described above can be used to determine user facial expressions and / or hand movements for use in generating an avatar or representation of the user, such as an anthropomorphic avatar or representation for a real-time communication session, wherein the avatar has facial expressions, hand movements, and / or body movements that are based on or similar to the detected facial expressions, hand movements, and / or body movements of the user of the device. Gaze and / or attention information is optionally combined with hand tracking information to determine interaction between a user and one or more user interfaces based on direct and / or indirect input, such as air gestures or input using one or more hardware input devices, such as one or more buttons (e.g., first button 1-128, button 11.1.1-114, second button 1-132, and / or dial or button 1-328), knobs (e.g., first button 1-128, button 11.1.1-114, and / or dial or button 1-328), a digital crown (e.g., a pressable and twistable or rotatable first button 1-128, button 11.1.1-114, and / or dial or button 1-328), a touchpad, a touch screen, a keyboard, a mouse, and / or other input devices. One or more buttons (e.g., a first button 1-128, a button 11.1.1-114, a second button 1-132, and / or a dial or button 1-328) are optionally used to perform system operations, such as re-centering content in a three-dimensional environment visible to a user of the device, displaying a primary user interface for launching an application, starting a real-time communication session, or initiating display of a virtual three-dimensional background. A knob or digital crown (e.g., a pressable and twistable or rotatable first button 1-128, a button 11.1.1-114, and / or a dial or button 1-328) is optionally rotatable to adjust parameters of visual content, such as the immersion level of the virtual three-dimensional environment (e.g., the extent to which virtual content occupies the user's viewport in the three-dimensional environment) or other parameters associated with the three-dimensional environment and virtual content displayed via the optical modules (e.g., the first and second display components 1-120a, 1-120b, and / or the first and second optical modules 11.1.1-104a, 11.1.1-104b).
[0088] Figure 1BIllustrated are front, top, and perspective views of an example head-mountable display (HMD) device 1-100 configured to be worn by a user and to provide a virtual and altered / mixed reality (VR / AR) experience. The HMD 1-100 may include a display unit 1-102 or assembly, an electronic strap assembly 1-104 connected to and extending from the display unit 1-102, and a strap assembly 1-106 secured at either end to the electronic strap assembly 1-104. The electronic strap assembly 1-104 and the strap 1-106 may be part of a retaining assembly configured to wrap around a user's head to hold the display unit 1-102 against the user's face.
[0089] In at least one example, the strap assembly 1-106 can include a first strap 1-116 configured to wrap around the back of a user's head and a second strap 1-117 configured to extend over the top of the user's head. As shown, the second strap can extend between the first electronic strip 1-105a and the second electronic strip 1-105b of the electronic strip assembly 1-104. The strap assembly 1-104 and the strap assembly 1-106 can be part of a securing mechanism that extends rearward from the display unit 1-102 and is configured to hold the display unit 1-102 against the user's face.
[0090] In at least one example, the securing mechanism includes a first electronic strip 1-105a including a first proximal end 1-134 coupled to the display unit 1-102 (e.g., the housing 1-150 of the display unit 1-102) and a first distal end 1-136 opposite the first proximal end 1-134. The securing mechanism may also include a second electronic strip 1-105b including a second proximal end 1-138 coupled to the housing 1-150 of the display unit 1-102 and a second distal end 1-140 opposite the second proximal end 1-138. The securing mechanism may also include a first band 1-116 and a second band 1-117, the first band including a first end 1-142 coupled to the first distal end 1-136 and a second end 1-144 coupled to the second distal end 1-140, and the second band extending between the first electronic strip 1-105a and the second electronic strip 1-105b. The strips 1-105a-b and the strip 1-116 may be coupled via a connecting mechanism or assembly 1-114. In at least one example, the second strip 1-117 includes a first end 1-146 coupled to the first electronic strip 1-105a between the first proximal end 1-134 and the first distal end 1-136 and a second end 1-148 coupled to the second electronic strip 1-105b between the second proximal end 1-138 and the second distal end 1-140.
[0091] In at least one example, the first and second electronic strips 1-105a-b include plastic, metal, or other structural materials formed into the shape of a substantially rigid strip 1-105a-b. In at least one example, the first band 1-116 and the second band 1-117 are formed from a resilient, flexible material including a woven textile, rubber, or the like. The first band 1-116 and the second band 1-117 can be flexible to conform to the shape of the user's head when the HMD 1-100 is worn.
[0092] In at least one example, one or more of the first and second electronic strips 1-105a-b can define an interior strip volume and include one or more electronic components disposed within the interior strip volume. Figure 1B As shown, the first electronic strip 1-105a may include an electronic component 1-112. In one example, the electronic component 1-112 may include a speaker. In one example, the electronic component 1-112 may include a computing component, such as a processor.
[0093] In at least one example, the housing 1-150 defines a first front opening 1-152. Figure 1B 1-152 in dashed lines because the display assembly 1-108 is configured to obscure the first opening 1-152 from view when the HMD 1-100 is assembled. The housing 1-150 may also define a rear-mounted second opening 1-154. The housing 1-150 further defines an interior volume between the first opening 1-152 and the second opening 1-154. In at least one example, the HMD 1-100 includes a display assembly 1-108, which may include a front cover and a display screen (shown in other figures) disposed in or across the front opening to obscure the front opening 1-152. In at least one example, the display screen of the display assembly 1-108, and the display assembly 1-108 generally, has a curvature configured to follow the curvature of the user's face. The display screen of the display assembly 1-108 may be curved as shown to complement the user's facial features and the overall curvature from one side of the face to the other, such as from left to right and / or from top to bottom, where the display unit 1-102 is depressed.
[0094] In at least one example, the housing 1-150 may define a first aperture 1-126 between the first opening 1-152 and the second opening 1-154, and a second aperture 1-130 between the first opening 1-152 and the second opening 1-154. The HMD 1-100 may also include a first button 1-126 disposed in the first aperture 1-128, and a second button 1-132 disposed in the second aperture 1-130. The first button 1-128 and the second button 1-132 are capable of being pressed through the respective apertures 1-126, 1-130. In at least one example, the first button 1-126 and / or the second button 1-132 may be a twistable dial and a pressable button. In at least one example, the first button 1-128 is a pressable and twistable dial button, and the second button 1-132 is a pressable button.
[0095] Figure 1C A rear perspective view of an HMD 1-100 is illustrated. The HMD 1-100 may include a light seal 1-110 extending rearwardly from a housing 1-150 of a display assembly 1-108 around the perimeter of the housing 1-150, as shown. The light seal 1-110 may be configured to extend from the housing 1-150 to the user's face, surrounding the user's eyes, to block external light from being visible. In one example, the HMD 1-100 may include a first display assembly 1-120a and a second display assembly 1-120b, which are disposed at or within a rearward-facing second opening 1-154 defined by the housing 1-150 and / or within the interior volume of the housing 1-150 and are configured to project light through the second opening 1-154. In at least one example, each display assembly 1-120a-b may include a respective display screen 1-122a, 1-122b, which are configured to project light in a rearward direction through the second opening 1-154 toward the user's eyes.
[0096] In at least one example, reference Figure 1B and Figure 1C In both cases, the display assembly 1-108 may be a front-facing, forward-facing display assembly including a display screen configured to project light in a first, forward direction, and the rear-facing display screens 1-122a-b may be configured to project light in a second, rearward direction opposite the first direction. As described above, the light seal 1-110 may be configured to block light external to the HMD 1-100 from reaching the user's eyes, including by Figure 1B 1-108 is shown in a front perspective view of the HMD 1-100. In at least one example, the HMD 1-100 may further include a curtain 1-124 that obscures a second opening 1-154 between the housing 1-150 and the rear display assemblies 1-120a-b. In at least one example, the curtain 1-124 may be elastic or at least partially elastic.
[0097] Figure 1B and Figure 1C Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figures 1D to 1F any of the other examples of devices, features, components, and parts shown and described herein. Figures 1D to 1F Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1B and Figure 1C Examples of devices, features, components, and parts are shown.
[0098] Figure 1D An exploded view of an example of an HMD 1-200 is illustrated, the HMD including various parts or components that are separable according to the modularity and selective coupling of these components. For example, the HMD 1-200 may include a strap 1-216 that is selectively coupled to a first electronic strip 1-205a and a second electronic strip 1-205b. The first fixed strap 1-205a may include a first electronic component 1-212a, and the second fixed strap 1-205b may include a second electronic component 1-212b. In at least one example, the first and second straps 1-205a-b are removably coupled to the display unit 1-202.
[0099] Additionally, the HMD 1-200 may include an optical seal 1-210 configured to be removably coupled to the display unit 1-202. The HMD 1-200 may also include a lens 1-218 that may be removably coupled to the display unit 1-202, for example, on a first assembly and a second display assembly that include a display screen. The lens 1-218 may include a custom prescription lens configured to correct vision. As noted, in Figure 1D Each of the parts shown in the exploded view of the HMD 1-200 and described above can be removably coupled, attached, reattached, and replaced to update parts or swap out parts for different users. For example, bands such as the band 1-216, optical seals such as the optical seal 1-210, lenses such as the lens 1-218, and electronic strips such as the electronic strips 1-205a-b can be swapped out depending on the user so that these parts are customized to fit and correspond to an individual user of the HMD 1-200.
[0100] Figure 1D Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figure 1B 、 Figure 1C and Figures 1E to 1F any of the other examples of devices, features, components, and parts shown and described herein. Figure 1B 、 Figure 1C and Figures 1E to 1F Any of the features, components and / or parts shown and described, including their arrangement and configuration, may be included in the Figure 1D Examples of devices, features, components, and parts are shown.
[0101] Figure 1E An exploded view of an example of a display unit 1-306 of an HMD is illustrated. The display unit 1-306 may include a front display assembly 1-308, a frame / housing assembly 1-350, and a curtain assembly 1-324. The display unit 1-306 may also include a sensor assembly 1-356, a logic board assembly 1-358, and a cooling assembly 1-360 disposed between the frame assembly 1-350 and the front display assembly 1-308. In at least one example, the display unit 1-306 may also include a rear display assembly 1-320 including a first rear display screen 1-322a and a second rear display screen 1-322b disposed between the frame 1-350 and the curtain assembly 1-324.
[0102] In at least one example, the display unit 1-306 may further include a motor assembly 1-362 configured as an adjustment mechanism for adjusting the position of the display screens 1-322a-b of the display assembly 1-320 relative to the frame 1-350. In at least one example, the display assembly 1-320 is mechanically coupled to the motor assembly 1-362, with each display screen 1-322a-b having at least one motor, such that the motors are capable of translating the display screens 1-322a-b to match the interpupillary distance of the user's eyes.
[0103] In at least one example, the display unit 1-306 may include a dial or button 1-328 that is depressible relative to the frame 1-350 and accessible by a user external to the frame 1-350. The button 1-328 may be electrically connected to the motor assembly 1-362 via a controller such that the button 1-328 may be manipulated by a user to cause a motor of the motor assembly 1-362 to adjust the position of the display screens 1-322a-b.
[0104] Figure 1E Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figures 1B to 1D and Figure 1F any of the other examples of devices, features, components, and parts shown and described herein. Figures 1B to 1D and Figure 1F Any of the features, components and / or parts shown and described, including their arrangement and configuration, may be included in the Figure 1EExamples of devices, features, components, and parts are shown.
[0105] Figure 1F An exploded view of another example of a display unit 1-406 of an HMD device similar to other HMD devices described herein is illustrated. The display unit 1-406 may include a front display assembly 1-402, a sensor assembly 1-456, a logic board assembly 1-458, a cooling assembly 1-460, a frame assembly 1-450, a rear display assembly 1-421, and a curtain assembly 1-424. The display unit 1-406 may also include a motor assembly 1-462 for adjusting the position of a first display subassembly 1-420a and a second display subassembly 1-420b of the rear display assembly 1-421, including first and second corresponding display screens for interpupillary adjustment, as described above.
[0106] Figure 1F The various parts, systems and assemblies shown in exploded views herein are referenced Figures 1B to 1E and subsequent figures referenced throughout this disclosure describe in more detail. Figure 1F The display unit 1-406 shown can be used with Figures 1B to 1E The shown fixing mechanism is assembled and integrated, and includes the electronic strips, ribbons, and other components including optical seals, connection components, etc.
[0107] Figure 1F Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figures 1B to 1E any of the other examples of devices, features, components, and parts shown and described herein. Figures 1B to 1E Any of the features, components and / or parts shown and described, including their arrangement and configuration, may be included in the Figure 1F Examples of devices, features, components, and parts are shown.
[0108] Figure 1G An exploded perspective view of the front cover assembly 3-100 of the HMD device described herein is illustrated, for example Figure 1G The front cover assembly 3-1 of the illustrated HMD 3-100 or any other HMD device shown and described herein. Figure 1GThe illustrated front cover assembly 3-100 may include a transparent or translucent cover 3-102, a shield 3-104 (or "cover"), an adhesive layer 3-106, a display assembly 3-108 including a lenticular lens panel or array 3-110, and a structural decorative member 3-112. The adhesive layer 3-106 may secure the shield 3-104 and / or the transparent cover 3-102 to the display assembly 3-108 and / or the decorative member 3-112. The decorative member 3-112 may secure the various components of the front cover assembly 3-100 to the frame or base of the HMD device.
[0109] In at least one example, Figure 1G As shown, the transparent cover 3-102, the shield 3-104, and the display assembly 3-108 including the lenticular lens array 3-110 can be bent to accommodate the curvature of the user's face. The transparent cover 3-102 and the shield 3-104 can be bent in two or three dimensions, for example, vertically in the Z direction within and outside the ZX plane, and horizontally in the X direction within and outside the ZX plane. In at least one example, the display assembly 3-108 may include the lenticular lens array 3-110 and a display panel having pixels that are configured to project light through the shield 3-104 and the transparent cover 3-102. The display assembly 3-108 can be bent in at least one direction (e.g., horizontally) to accommodate the curvature of the user's face from one side of the face (e.g., the left side) to the other side (e.g., the right side). In at least one example, each layer or component of the display assembly 3-108 (which will be shown in subsequent figures and described in more detail, but which may include the lenticular lens array 3-110 and the display layer) may be curved similarly or concentrically in the horizontal direction to accommodate the curvature of the user's face.
[0110] In at least one example, the shield 3-104 may include a transparent or translucent material through which the display assembly 3-108 projects light. In one example, the shield 3-104 may include one or more opaque portions, such as an opaque ink-printed portion or other opaque film portion on the back of the shield 3-104. When the HMD device is worn, the back surface may be the surface of the shield 3-104 that faces the user's eyes. In at least one example, the opaque portion may be on the front surface of the shield 3-104, opposite the back surface. In at least one example, the one or more opaque portions of the shield 3-104 may include a peripheral portion that visually conceals any components surrounding the outer perimeter of the display screen of the display assembly 3-108. In this manner, the opaque portion of the shield conceals any other components of the HMD device that would otherwise be visible through the transparent or translucent cover 3-102 and / or the shield 3-104, including electronic components, structural components, etc.
[0111] In at least one example, the shield 3-104 may define one or more aperture transparent portions 3-120 through which sensors may transmit and receive signals. In one example, the portion 3-120 is an aperture through which a sensor may extend or transmit and receive signals. In one example, the portion 3-120 is a transparent portion, or a portion that is more transparent than surrounding translucent or opaque portions of the shield, through which sensors may transmit and receive signals through the shield and through the transparent cover 3-102. In one example, the sensor may include a camera, an IR sensor, a LUX sensor, or any other visual or non-visual environmental sensor of the HMD device.
[0112] Figure 1G Any of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any of the other examples of devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included, alone or in any combination, in any of the other examples of devices, features, components, and parts described herein. Figure 1G Examples of devices, features, components, and parts are shown.
[0113] Figure 1H An exploded view of an example of an HMD device 6-100 is illustrated. The HMD device 6-100 may include a sensor array or system 6-102 including one or more sensors, cameras, projectors, etc. mounted to one or more components of the HMD 6-100. In at least one example, the sensor system 6-102 may include a bracket 1-338 to which one or more sensors of the sensor system 6-102 may be secured / fastened.
[0114] Figure 1I A portion of an HMD device 6-100 is illustrated that includes a front transparent cover 6-104 and a sensor system 6-102. The sensor system 6-102 may include a plurality of different sensors, emitters, receivers, including cameras, IR sensors, projectors, etc. The transparent cover 6-104 is shown in front of the sensor system 6-102 to illustrate the relative positions of the various sensors and emitters and the orientation of each sensor / emitter of the system 6-102. As referred to herein, "lateral," "sideways," "horizontal," and other similar terms refer to the orientation of the sensor system 6-102. Figure 1J The orientation or direction indicated by the X-axis shown. Terms such as "vertical", "upward", "downward" and similar terms refer to Figure 1J The orientation or direction indicated by the Z-axis shown. Terms such as "forward," "backward," "forward," "backward" and similar terms refer to the orientation or direction indicated by the Z-axis shown. Figure 1J The Y-axis shown indicates the orientation or direction.
[0115] In at least one example, a transparent cover 6-104 may define a front exterior surface of the HMD device 6-100, and a sensor system 6-102, including various sensors and components thereof, may be disposed in the Y axis / direction behind the cover 6-104. The cover 6-104 may be transparent or translucent to allow light to pass through the cover 6-104, including both light detected by the sensor system 6-102 and light emitted thereby.
[0116] As described elsewhere herein, the HMD device 6-100 may include one or more controllers including processors for electrically coupling the various sensors and transmitters of the sensor system 6-102 to one or more motherboards, processing units, and other electronic devices such as display screens. Furthermore, as will be shown in greater detail below with reference to other figures, the various sensors, transmitters, and other components of the sensor system 6-102 may be coupled to Figure 1I For clarity, various structural frame members, brackets, etc. of the HMD device 6-100 are not shown. Figure 1I Components of the sensor system 6-102 are shown unattached and unelectrically coupled to other components.
[0117] In at least one example, the device may include one or more controllers having processors configured to execute instructions stored on a memory component electrically coupled to the processors. The instructions may include or cause the processors to execute one or more algorithms for self-correcting the angles and positions of the various cameras described herein over time as the initial position, angle, or orientation of the camera is bumped or deformed due to an accidental drop event or other event.
[0118] In at least one example, the sensor system 6-102 may include one or more scene cameras 6-106. The system 6-102 may include two scene cameras 6-102, one located on either side of the nose bridge or arch of the HMD device 6-100, such that each of the two cameras 6-106 roughly corresponds to the position of the user's left and right eyes behind the cover 6-103. In at least one example, the scene cameras 6-106 are generally oriented forward in the Y direction to capture images in front of the user during use of the HMD 6-100. In at least one example, the scene cameras are color cameras and provide images and content for MR video pass-through to the display screen facing the user's eyes when the HMD device 6-100 is in use. The scene cameras 6-106 may also be used for environment and object reconstruction.
[0119] In at least one example, the sensor system 6-102 may include a first depth sensor 6-108 pointing generally forward in the Y direction. In at least one example, the first depth sensor 6-108 may be used for environment and object reconstruction and hand and body tracking of the user. In at least one example, the sensor system 6-102 may include a second depth sensor 6-110 centrally located along the width of the HMD device 6-100 (e.g., along the X axis). For example, the second depth sensor 6-110 may be located above a central nose bridge or on an adaptable structure above the nose of the user when wearing the HMD 6-100. In at least one example, the second depth sensor 6-110 may be used for environment and object reconstruction and hand and body tracking. In at least one example, the second depth sensor may include a LIDAR sensor.
[0120] In at least one example, the sensor system 6-102 may include a depth projector 6-112 that faces generally forward to project electromagnetic waves (e.g., in a predetermined pattern of light dots) into or within the field of view of the user and / or scene camera 6-106, or into or within a field of view that includes and extends beyond the field of view of the user and / or scene camera 6-106. In at least one example, the depth projector may be capable of projecting electromagnetic waves of light in the form of a pattern of light dots that reflect off an object and return to the depth sensors described above, including the depth sensors 6-108, 6-110. In at least one example, the depth projector 6-112 may be used for environment and object reconstruction and hand and body tracking.
[0121] In at least one example, the sensor system 6-102 may include downward-facing cameras 6-114 whose fields of view are generally directed downward on the Z-axis relative to the HMD device 6-100. In at least one example, the downward-facing cameras 6-114 may be disposed on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for displaying a user avatar on a forward-facing display screen of the HMD device 6-100 as described elsewhere herein. For example, the downward-facing cameras 6-114 may be used to capture facial expressions and movements of a user's face, including cheeks, mouth, and chin, beneath the HMD device 6-100.
[0122] In at least one example, the sensor system 6-102 may include a jaw camera 6-116. In at least one example, the jaw cameras 6-116 may be positioned on the left and right sides of the HMD device 6-100 as shown and used for hand and body tracking, headset tracking, and facial avatar detection and creation for displaying a user avatar on a front-facing display screen of the HMD device 6-100 as described elsewhere herein. For example, the jaw camera 6-116 may be used to capture facial expressions and movements of a user's face beneath the HMD device 6-100, including the user's jaw, cheeks, mouth, and chin. Used for hand and body tracking, headset tracking, and facial avatar
[0123] In at least one example, the sensor system 6-102 may include a side camera 6-118. The side camera 6-118 may be oriented to capture left and right side views in an X-axis or direction relative to the HMD device 6-100. In at least one example, the side camera 6-118 may be used for hand and body tracking, headset tracking, and facial avatar detection and reconstruction.
[0124] In at least one example, the sensor system 6-102 may include a plurality of eye tracking and gaze tracking sensors for determining the identity, status, and gaze direction of a user's eyes during and / or prior to use. In at least one example, the eye / gaze tracking sensors may include nose-eye cameras 6-120 that are positioned on either side of the user's nose and adjacent to the user's nose when the HMD device 6-100 is worn. The eye / gaze sensors may also include bottom eye cameras 6-122 positioned below the respective user's eyes for capturing images of the eyes for use in facial avatar detection and creation, gaze tracking, and iris identification functionality.
[0125] In at least one example, the sensor system 6-102 may include an infrared illuminator 6-124 that points outward from the HMD device 6-100 to illuminate the external environment and any objects therein with IR light for IR detection using one or more IR sensors of the sensor system 6-102. In at least one example, the sensor system 6-102 may include a flicker sensor 6-126 and an ambient light sensor 6-128. In at least one example, the flicker sensor 6-126 may detect the refresh rate of overhead light to avoid display flicker. In one example, the infrared illuminator 6-124 may include a light emitting diode and may be particularly useful in low-light environments for illuminating a user's hands and other objects in low light for detection by the infrared sensors of the sensor system 6-102.
[0126] In at least one example, a plurality of sensors (including a scene camera 6-106, a downward camera 6-114, a jaw camera 6-116, a side camera 6-118, a depth projector 6-112, and depth sensors 6-108, 6-110) may be used in combination with an electrically coupled controller to combine depth data with camera data for hand tracking and for size determination to better perform hand tracking and object recognition and tracking functions of the HMD device 6-100. In at least one example, as described above and in Figure 1I The downward camera 6-114, the jaw camera 6-116, and the side camera 6-118 shown in the figure can be wide-angle cameras capable of operating in the visible and infrared spectrums. In at least one example, these cameras 6-114, 6-116, 6-118 can operate only in black and white light detection to simplify image processing and gain sensitivity.
[0127] Figure 1I Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figures 1J to 1L any of the other examples of devices, features, components, and parts shown and described herein. Figures 1J to 1L Any of the features, components and / or parts shown and described, including their arrangement and configuration, may be included in the Figure 1I Examples of devices, features, components, and parts are shown.
[0128] Figure 1J A lower perspective view of an example of an HMD 6-200 including a cover or shroud 6-204 secured to a frame 6-230 is illustrated. In at least one example, sensors 6-203 of a sensor system 6-202 may be disposed around the perimeter of the HMD 6-200 such that the sensors 6-203 are disposed outwardly around the perimeter of a display area or display region 6-232 so as to not obstruct a view of displayed light. In at least one example, the sensors may be disposed behind the shroud 6-204 and aligned with a transparent portion of the shroud, thereby allowing the sensors and projector to pass light back and forth through the shroud 6-204. In at least one example, opaque ink or other opaque material or film / layer may be disposed on the shroud 6-204 around the display area 6-232 to conceal components of the HMD 6-200 outside of the display area 6-232 rather than the transparent portion defined by the opaque portion through which the sensors and projector transmit and receive light and electromagnetic signals during operation. In at least one example, the shield 6-204 allows light to pass from the display (eg, within the display area 6-232), but does not allow light to pass radially outward from the display area around the display and the perimeter of the shield 6-204.
[0129] In some examples, the shield 6-204 includes a transparent portion 6-205 and an opaque portion 6-207, as described above and elsewhere herein. In at least one example, the opaque portion 6-207 of the shield 6-204 may define one or more transparent areas 6-209 through which the sensor 6-203 of the sensor system 6-202 may send and receive signals. In the illustrated example, the sensor 6-203 of the sensor system 6-202 sends and receives signals through the shield 6-204, or more specifically, through (or defined by) the transparent areas 6-209 of the opaque portion 6-207 of the shield 6-204, which may include a plurality of transparent regions 6-209. Figure 1I The same or similar sensors as those shown in the example of FIG, such as the depth sensors 6-108 and 6-110, the depth projector 6-112, the first and second scene cameras 6-106, the first and second downward cameras 6-114, the first and second side cameras 6-118, and the first and second infrared illuminators 6-124. These sensors are also Figure 1K and Figure 1L Other sensors, sensor types, number of sensors, and their relative positions may be included in one or more other examples of an HMD.
[0130] Figure 1J Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figure 1I and Figures 1K to 1L any of the other examples of devices, features, components, and parts shown and described herein. Figure 1I and Figures 1K to 1L Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1J Examples of devices, features, components, and parts are shown.
[0131] Figure 1K Illustrated is a front view of a portion of an example of an HMD device 6-300, including a display 6-334, brackets 6-336, 6-338, and a frame or housing 6-330. Figure 1K The example shown does not include a front cover or shield in order to illustrate the brackets 6-336, 6-338. For example, Figure 1J The illustrated shield 6-204 includes an opaque portion 6-207 that would visually cover / block viewing of anything external to (e.g., radially / peripherally external to) the display / display area 6-334, including the sensor 6-303 and bracket 6-338.
[0132] In at least one example, the various sensors of the sensor system 6-302 are coupled to brackets 6-336, 6-338. In at least one example, the scene cameras 6-306 include tight tolerances on angles relative to each other. For example, the tolerance on mounting angles between two scene cameras 6-306 may be 0.5 degrees or less, such as 0.3 degrees or less. To achieve and maintain such tight tolerances, in one example, the scene cameras 6-306 may be mounted to the bracket 6-338 instead of the shield. The bracket may include a cantilever on which the scene camera 6-306 and other sensors of the sensor system 6-302 may be mounted to maintain position and orientation in the event of a drop by a user that causes any deformation of the other bracket 6-226, the housing 6-330, and / or the shield.
[0133] Figure 1K Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figures 1I to 1J and Figure 1L any of the other examples of devices, features, components, and parts shown and described herein. Figures 1I to 1J and Figure 1L Any of the features, components and / or parts shown or described (including their arrangement and configuration) may be included alone or in any combination in the Figure 1K Examples of devices, features, components, and parts are shown.
[0134] Figure 1L A bottom view of an example of an HMD 6-400 including a front display / cover assembly 6-404 and a sensor system 6-402 is illustrated. The sensor system 6-402 may be similar to other sensor systems described above and elsewhere herein, including with reference to Figures 1I to 1K As described. In at least one example, the jaw camera 6-416 can face downward to capture images of the user's lower facial features. In one example, the jaw camera 6-416 can be directly coupled to the frame or housing 6-430 or one or more internal brackets that are directly coupled to the frame or housing 6-430 as shown. The frame or housing 6-430 can include one or more holes / openings 6-415 through which the jaw camera 6-416 can send and receive signals.
[0135] Figure 1L Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figures 1I to 1K any of the other examples of devices, features, components, and parts shown and described herein. Figures 1I to 1KAny of the features, components and / or parts shown and described, including their arrangement and configuration, may be included in the Figure 1L Examples of devices, features, components, and parts are shown.
[0136] Figure 1M Illustrated is a rear perspective view of an interpupillary distance (IPD) adjustment system 11.1.1-102 comprising first and second optical modules 11.1.1-104a-b slidably engaged / coupled to respective guide rods 11.1.1-108a-b and motors 11.1.1-110a-b of left and right adjustment subsystems 11.1.1-106a-b. The IPD adjustment system 11.1.1-102 may be coupled to a bracket 11.1.1-112 and include a button 11.1.1-114 in electrical communication with the motors 11.1.1-110a-b. In at least one example, the button 11.1.1-114 may be in electrical communication with the first and second motors 11.1.1-110a-b via a processor or other circuit components to cause the first and second motors 11.1.1-110a-b to activate and respectively cause the first and second optical modules 11.1.1-104a-b to change position relative to each other.
[0137] In at least one example, the first and second optical modules 11.1.1-104a-b may include respective display screens configured to project light toward the user's eyes when the HMD 11.1.1-100 is worn. In at least one example, the user may manipulate (e.g., press and / or rotate) a button 11.1.1-114 to activate positional adjustment of the optical modules 11.1.1-104a-b to match the interpupillary distance of the user's eyes. The optical modules 11.1.1-104a-b may also include one or more cameras or other sensors / sensor systems for imaging and measuring the user's IPD so that the optical modules 11.1.1-104a-b can be adjusted to match the IPD.
[0138] In one example, a user can manipulate button 11.1.1-114 to cause automatic position adjustment of the first and second optical modules 11.1.1-104a-b. In one example, a user can manipulate button 11.1.1-114 to cause manual adjustment, causing the optical modules 11.1.1-104a-b to move further or closer (e.g., when the user rotates button 11.1.1-114 one way or another) until the user visually matches their IPD. In one example, the manual adjustment is communicated electronically via one or more circuits, and power for moving the optical modules 11.1.1-104a-b via motors 11.1.1-110a-b is provided by a power source. In one example, adjustment and movement of the optical modules 11.1.1-104a-b via manipulation button 11.1.1-114 is mechanically actuated via movement button 11.1.1-114.
[0139] Figure 1M Any of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any of the other examples of devices, features, components, and parts shown in any other drawing and described herein. Similarly, any of the features, components, and / or parts shown or described with reference to any other drawing (including arrangements and configurations thereof) may be included, alone or in any combination, in any of the other examples of devices, features, components, and parts shown in any other drawing and described herein. Figure 1M Examples of devices, features, components, and parts are shown.
[0140] Figure 1N Illustrated is a front perspective view of a portion of an HMD 11.1.2-100 including an outer structural frame 11.1.2-102 and an inner or intermediate structural frame 11.1.2-104 defining a first aperture 11.1.2-106a and a second aperture 11.1.2-106b. Figure 1N 2-106a-b may be blocked by one or more other components of the HMD 11.1.2-100 coupled to the inner frame 11.1.2-104 and / or the outer frame 11.1.2-102, as shown. In at least one example, the HMD 11.1.2-100 may include a first mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104. In at least one example, the mounting bracket 11.1.2-108 is coupled to the inner frame 11.1.2-104 between the first and second apertures 11.1.2-106a-b.
[0141] The mounting bracket 11.1.2-108 can include a middle or center portion 11.1.2-109 coupled to the inner frame 11.1.2-104. In some examples, the middle or center portion 11.1.2-109 may not be the geometric middle or center of the bracket 11.1.2-108. Instead, the middle / center portion 11.1.2-109 can be disposed between first and second cantilevered extension arms extending away from the middle portion 11.1.2-109. In at least one example, the mounting bracket 108 includes a first cantilevered arm 11.1.2-112 and a second cantilevered arm 11.1.2-114 extending away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108 coupled to the inner frame 11.1.2-104.
[0142] like Figure 1N As shown, the outer frame 11.1.2-102 can define a curved geometry on its underside to accommodate the user's nose when the user wears the HMD 11.1.2-100. The curved geometry can be referred to as a nose bridge 11.1.2-111 and is centrally located on the underside of the HMD 11.1.2-100 as shown. In at least one example, the mounting bracket 11.1.2-108 can be connected to the inner frame 11.1.2-104 between the holes 11.1.2-106a-b so that the cantilevered arms 11.1.2-112, 11.1.2-114 extend downwardly and laterally outwardly away from the middle portion 11.1.2-109 to complement the nose bridge 11.1.2-111 geometry of the outer frame 11.1.2-102. In this manner, the mounting bracket 11.1.2-108 is configured to accommodate the user's nose, as described above. The geometry of the nose bridge 11.1.2-111 adapts to the nose in that the nose bridge 11.1.2-111 provides a curvature that conforms to the shape of the user's nose, providing a comfortable fit from above, over, and around.
[0143] The first cantilever arm 11.1.2-112 can extend in a first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108, and the second cantilever arm 11.1.2-114 can extend in a second direction opposite to the first direction away from the middle portion 11.1.2-109 of the mounting bracket 11.1.2-108. The first cantilever arm 11.1.2-112 and the second cantilever arm 11.1.2-114 are referred to as "cantilevered" or "cantilever" arms because each arm 11.1.2-112, 11.1.2-114 includes a free distal end 11.1.2-116, 11.1.2-118, respectively, which are not attached to the inner frame 11.1.2-102 and the outer frame 11.1.2-104. In this way, the arms 11.1.2-112, 11.1.2-114 depend from the middle portion 11.1.2-109, which is connectable to the inner frame 11.1.2-104, while the distal ends 11.1.2-102, 11.1.2-104 are unattached.
[0144] In at least one example, the HMD 11.1.2-100 may include one or more components coupled to a mounting bracket 11.1.2-108. In one example, the components include a plurality of sensors 11.1.2-110a-f. Each of the plurality of sensors 11.1.2-110a-f may include various types of sensors, including cameras, IR sensors, and the like. In some examples, one or more of the sensors 11.1.2-110a-f may be used for object recognition in three-dimensional space, making it important to maintain the precise relative position of two or more of the plurality of sensors 11.1.2-110a-f. The cantilevered nature of the mounting bracket 11.1.2-108 may protect the sensors 11.1.2-110a-f from damage and change of position if accidentally dropped by a user. Because the sensors 11.1.2-110a-f are cantilevered on the arms 11.1.2-112, 11.1.2-114 of the mounting bracket 11.1.2-108, stresses and deformations of the inner and / or outer frames 11.1.2-104, 11.1.2-102 are not transferred to the cantilevered arms 11.1.2-112, 11.1.2-114 and therefore do not affect the relative positions of the sensors 11.1.2-110a-f coupled / mounted to the mounting bracket 11.1.2-108.
[0145] Figure 1NAny of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any of the other examples of devices, features, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included, alone or in any combination, in any of the other examples of devices, features, and parts described herein. Figure 1N Examples of devices, features, components, and parts are shown.
[0146] Figure 1O An example of an optical module 11.3.2-100 for use in an electronic device (such as an HMD, including the HMD devices described herein) is illustrated. As shown in one or more other examples described herein, the optical module 11.3.2-100 can be one of two optical modules within the HMD, where each optical module is aligned to project light toward an eye of a user. In this manner, a first optical module can project light toward a first eye of a user via a display screen, and a second optical module of the same device can project light toward a second eye of the user via another display screen.
[0147] In at least one example, the optical module 11.3.2-100 may include an optical frame or housing 11.3.2-102, which may also be referred to as a barrel or optical module barrel. The optical module 11.3.2-100 may also include a display 11.3.2-104 coupled to the housing 11.3.2-102, the display including one or more display screens. The display 11.3.2-104 may be coupled to the housing 11.3.2-102 such that the display 11.3.2-104 is configured to project light toward the eyes of a user when the HMD to which the display module 11.3.2-100 belongs is worn during use. In at least one example, the housing 11.3.2-102 may surround the display 11.3.2-104 and provide connection features for coupling other components of the optical module described herein.
[0148] In one example, the optical module 11.3.2-100 may include one or more cameras 11.3.2-106 coupled to the housing 11.3.2-102. The cameras 11.3.2-106 may be positioned relative to the display 11.3.2-104 and the housing 11.3.2-102 such that the cameras 11.3.2-106 are configured to capture one or more images of a user's eyes during use. In at least one example, the optical module 11.3.2-100 may also include a light strip 11.3.2-108 surrounding the display 11.3.2-104. In one example, the light strip 11.3.2-108 is disposed between the display 11.3.2-104 and the cameras 11.3.2-106. The light strip 11.3.2-108 may include a plurality of lights 11.3.2-110. The plurality of lights may include one or more light emitting diodes (LEDs) or other lights configured to project light toward the eyes of the user when the HMD is worn. The individual lights 11.3.2-110 in the light strip 11.3.2-108 may be spaced apart around the light strip 11.3.2-108 and thus evenly or unevenly spaced around the display 11.3.2-104 at various locations on the light strip 11.3.2-108 and around the display 11.3.2-104.
[0149] In at least one example, the housing 11.3.2-102 defines a viewing opening 11.3.2-101 through which a user can view the display 11.3.2-104 when wearing the HMD device. In at least one example, the LEDs are configured and arranged to emit light through the viewing opening 11.3.2-101 toward the user's eyes. In one example, the camera 11.3.2-106 is configured to capture one or more images of the user's eyes through the viewing opening 11.3.2-101.
[0150] As mentioned above, Figure 1O Each of the components and features of the illustrated optical module 11.3.2-100 may be replicated in another (eg, second) optical module provided with the HMD to interact with (eg, project light and capture images) the user's other eye.
[0151] Figure 1O Any of the features, components and / or parts shown (including their arrangement and configuration) may be included alone or in any combination in the Figure 1P any of the other examples of devices, features, components, and parts shown or otherwise described herein. Figure 1P Any of the features, components and / or parts shown or described herein (including their arrangement and configuration) may be included alone or in any combination. Figure 1OExamples of devices, features, components, and parts are shown.
[0152] Figure 1P A cross-sectional view of an example of an optical module 11.3.2-200 is illustrated, including a housing 11.3.2-202, a display assembly 11.3.2-204 coupled to the housing 11.3.2-202, and a lens 11.3.2-216 coupled to the housing 11.3.2-202. In at least one example, the housing 11.3.2-202 defines a first aperture or channel 11.3.2-212 and a second aperture or channel 11.3.2-214. The channels 11.3.2-212, 11.3.2-214 can be configured to slidably engage corresponding tracks or guides of an HMD device to allow the optical module 11.3.2-200 to be adjusted relative to the user's eyes to match the user's interpupillary distance (IPD). The housing 11.3.2-202 can slidably engage the guides to secure the optical module 11.3.2-200 in place within the HMD.
[0153] In at least one example, the optical module 11.3.2-200 may further include a lens 11.3.2-216 coupled to the housing 11.3.2-202 and positioned between the display assembly 11.3.2-204 and the user's eyes when the HMD is worn. The lens 11.3.2-216 may be configured to direct light from the display assembly 11.3.2-204 to the user's eyes. In at least one example, the lens 11.3.2-216 may be part of a lens assembly including a corrective lens that is removably attached to the optical module 11.3.2-200. In at least one example, the lens 11.3.2-216 is disposed above the light strip 11.3.2-208 and the one or more eye tracking cameras 11.3.2-206 such that the camera 11.3.2-206 is configured to capture an image of the user's eyes through the lens 11.3.2-216, and the light strip 11.3.2-208 includes lights configured to project light into the user's eyes through the lens 11.3.2-216 during use.
[0154] Figure 1P Any of the features, components, and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any of the other examples of devices, features, components, and parts described herein. Similarly, any of the features, components, and / or parts shown and described herein (including arrangements and configurations thereof) may be included, alone or in any combination, in any of the other examples of devices, features, components, and parts described herein. Figure 1P Examples of devices, features, components, and parts are shown.
[0155] Figure 2is a block diagram of an example of a controller 110 according to some embodiments. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein. To this end, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., a universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, a memory 220, and one or more communication buses 204 for interconnecting these components and various other components.
[0156] In some embodiments, the one or more communication buses 204 include circuits that interconnect and control communications between system components. In some embodiments, the one or more I / O devices 206 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and the like.
[0157] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located away from one or more processing units 202. Memory 220 includes non-transitory computer-readable storage media. In some embodiments, memory 220 or a non-transitory computer-readable storage medium of memory 220 stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 230 and an XR experience module 240.
[0158] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR experience module 240 is configured to manage and coordinate single or multiple XR experiences for one or more users (e.g., a single XR experience for one or more users, or multiple XR experiences for corresponding groups of one or more users). To this end, in various embodiments, the XR experience module 240 includes a data acquisition unit 241, a tracking unit 242, a coordination unit 246, and a data transmission unit 248.
[0159] In some embodiments, the data acquisition unit 241 is configured to at least Figure 1A 1 and / or peripherals 195. The display generation component 120 of the embodiment of the present invention may also be used to obtain data (e.g., presentation data, interaction data, sensor data, position data, etc.) from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the data acquisition unit 241 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.
[0160] In some embodiments, the tracking unit 242 is configured to map the scene 105 and track at least the display generation component 120 relative to the scene 105. Figure 1A 105, and optionally relative to the position / location of one or more of the tracked input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. To this end, in various embodiments, the tracking unit 242 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics. In some embodiments, the tracking unit 242 includes a hand tracking unit 244 and / or an eye tracking unit 243. In some embodiments, the hand tracking unit 244 is configured to track the position / location of one or more parts of the user's hand and / or the position / location of one or more parts of the user's hand relative to the user's Figure 1A The movement of the scene 105 relative to the display generation component 120 and / or relative to the coordinate system (the coordinate system is defined relative to the user's hand). Figure 4 The hand tracking unit 244 is described in more detail. In some embodiments, the eye tracking unit 243 is configured to track the position or movement of the user's gaze (or more broadly, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hands)) or relative to the XR content displayed via the display generation component 120. Figure 5 The eye tracking unit 243 is described in more detail.
[0161] In some embodiments, the coordination unit 246 is configured to manage and coordinate the XR experience presented to the user by the display generation component 120, and optionally by one or more of the output device 155 and / or peripheral devices 195. To this end, in various embodiments, the coordination unit 246 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0162] In some embodiments, data transmission unit 248 is configured to transmit data (e.g., presentation data, position data, etc.) to at least display generation component 120, and optionally to one or more of input device 125, output device 155, sensor 190, and / or peripheral device 195. To this end, in various embodiments, data transmission unit 248 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.
[0163] Although the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 241, the tracking unit 242 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 may be located in separate computing devices.
[0164] also, Figure 2 It serves more as a functional description of various features that may be present in a particular implementation, rather than as a structural diagram of the embodiments described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of specific functions and how features are distributed among them will vary depending on the specific implementation and, in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.
[0165] Figure 3is a block diagram of an example of a display generation component 120 according to some embodiments. While some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the embodiments disclosed herein. For this purpose, as a non-limiting example, in some embodiments, the display generation component 120 (e.g., an HMD) includes one or more processing units 302 (e.g., a microprocessor, an ASIC, an FPGA, a GPU, a CPU, a processing core, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more XR displays 312, one or more optional internal-facing and / or external-facing image sensors 314, memory 320, and one or more communication buses 304 for interconnecting these and various other components.
[0166] In some embodiments, the one or more communication buses 304 include circuits for interconnecting and controlling communications between various system components. In some embodiments, the one or more I / O devices and sensors 306 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.
[0167] In some embodiments, one or more XR displays 312 are configured to provide an XR experience to the user. In some embodiments, one or more XR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS) and / or similar display types. In some embodiments, one or more XR displays 312 correspond to diffraction, reflection, polarization, holographic and other waveguide displays. For example, the display generation component 120 (e.g., HMD) includes a single XR display. In another example, the display generation component 120 includes an XR display for each eye of the user. In some embodiments, one or more XR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more XR displays 312 are capable of presenting MR or VR content.
[0168] In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, the one or more image sensors 314 are configured to acquire image data corresponding to the user's hands and, optionally, at least a portion of the user's arms (and may be referred to as a hand-tracking camera). In some embodiments, the one or more image sensors 314 are configured to face forward so as to acquire image data corresponding to the scene that the user would see in the absence of the display generation component 120 (e.g., an HMD) (and may be referred to as a scene camera). The one or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, among others.
[0169] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located away from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage media. In some embodiments, memory 320 or a non-transitory computer-readable storage medium of memory 320 stores the following programs, modules, and data structures, or a subset thereof, including an optional operating system 330 and an XR rendering module 340.
[0170] The operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some embodiments, the XR rendering module 340 is configured to present XR content to the user via one or more XR displays 312. To this end, in various embodiments, the XR rendering module 340 includes a data acquisition unit 342, an XR rendering unit 344, an XR map generation unit 346, and a data transmission unit 348.
[0171] In some embodiments, the data acquisition unit 342 is configured to at least Figure 1A The controller 110 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). For this purpose, in various embodiments, the data acquisition unit 342 includes instructions and / or logic for instructions and heuristics and metadata for the heuristics.
[0172] In some embodiments, the XR rendering unit 344 is configured to render XR content via one or more XR displays 312. For such purposes, in various embodiments, the XR rendering unit 344 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0173] In some embodiments, the XR map generation unit 346 is configured to generate an XR map (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate an extended reality) based on the media content data. For this purpose, in various embodiments, the XR map generation unit 346 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0174] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, position data, etc.) to at least the controller 110, and optionally one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For such purposes, in various embodiments, the data transmission unit 348 includes instructions and / or logic for the instructions and heuristics and metadata for the heuristics.
[0175] Although the data acquisition unit 342, the XR rendering unit 344, the XR map generation unit 346, and the data transmission unit 348 are shown as residing on a single device (e.g., Figure 1A , but it should be understood that in other embodiments, any combination of the data acquisition unit 342, the XR rendering unit 344, the XR map generation unit 346, and the data transmission unit 348 may be located in a separate computing device.
[0176] also, Figure 3 It serves more as a functional description of various features that may be present in a particular implementation, rather than as a structural diagram of the embodiments described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 3 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the division of specific functions and how features are distributed among them will vary depending on the specific implementation and, in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.
[0177] Figure 4 is a schematic illustration of an example embodiment of a hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1A ) is controlled by the hand tracking unit 244 ( Figure 2 ) to track the position / location of one or more parts of the user's hand, and / or the movement of one or more parts of the user's hand relative to the scene 105 of Figure 1 (e.g., relative to a portion of the physical environment surrounding the user, relative to the display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system (which is defined relative to the user's hand)). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to the head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).
[0178] In some embodiments, the hand tracking device 140 includes an image sensor 404 (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.) that captures three-dimensional scene information, including at least a human user's hand 406. The image sensor 404 captures hand images at a sufficient resolution to enable the fingers and their respective positioning to be distinguished. The image sensor 404 typically captures images of other parts of the user's body, or may also capture images of all parts of the body, and may have zoom capabilities or specialized sensors with increased magnification to capture images of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in conjunction with other image sensors to capture the physical environment of the scene 105, or serves as an image sensor for capturing the physical environment of the scene 105. In some embodiments, the image sensor is positioned relative to the user or the user's environment in such a way that the field of view of the image sensor 404, or a portion thereof, is used to define an interaction space in which hand movements captured by the image sensor are treated as input to the controller 110.
[0179] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D image data (and possibly color image data) to the controller 110, which extracts high-level information from the image data. This high-level information is typically provided via an application program interface (API) to an application running on the controller, which in turn drives the display generation component 120. For example, a user can interact with the software running on the controller 110 by moving his hand 406 and changing his hand pose.
[0180] In some embodiments, the image sensor 404 projects a speckled pattern onto a scene containing the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offsets of the spots in the pattern. This approach is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. The method gives the depth coordinates of a point in the scene relative to a predetermined reference plane at a specific distance from the image sensor 404. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x-axis, y-axis, and z-axis such that the depth coordinates of a point in the scene correspond to the z component measured by the image sensor. Alternatively, the image sensor 404 (e.g., a hand tracking device) may use other 3D mapping methods based on a single or multiple cameras or other types of sensors, such as stereo imaging or time-of-flight measurement.
[0181] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps containing the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on the image sensor 404 and / or the processor in the controller 110 processes the 3D map data to extract image patch descriptors of the hand in these depth maps. The software can match these descriptors with image patch descriptors stored in the database 408 based on a previous learning process to estimate the pose of the hand in each frame. The pose typically includes the 3D positions of the user's hand joints and fingertips.
[0182] The software can also analyze the trajectory of the hand and / or finger over multiple frames in the sequence to identify gestures. The pose estimation function described herein can be alternated with the motion tracking function so that the image block-based pose estimation is performed only once every two (or more) frames, and tracking is used to find changes in pose that occur on the remaining frames. The pose, motion, and gesture information is provided to the application running on the controller 110 via the above-mentioned API. The program can, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.
[0183] In some embodiments, gestures include air gestures. An air gesture is a gesture that is detected without the user touching an input element that is part of a device (e.g., computer system 101, one or more input devices 125, and / or hand tracking device 140) (or independent of an input element that is part of the device) and is based on detected movement of a part of the user's body (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs) through air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture in which the hand moves a predetermined amount and / or speed in a predetermined posture, or a shake gesture including a predetermined speed or amount of rotation of a part of the user's body).
[0184] In some embodiments, according to some embodiments, the input gestures used in the various examples and embodiments described herein include air gestures for interacting with an XR environment (e.g., a virtual or mixed reality environment) performed by movement of a user's fingers relative to other fingers (or parts of the user's hands). In some embodiments, an air gesture is a gesture detected without the user touching an input element that is part of the device (or independent of an input element that is part of the device) and based on detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of a user's finger relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tap gesture in which the hand moves a predetermined amount and / or speed in a predetermined posture, or a shake gesture in which a part of the user's body is rotated at a predetermined speed or amount)).
[0185] In some embodiments where the input gesture is an in-air gesture (e.g., in the absence of physical contact with an input device that provides information to the computer system about which user interface element is the target of the user input, such as contact with a user interface element displayed on a touch screen, or contact with a mouse or trackpad to move a cursor to a user interface element), the gesture takes into account the user's attention (e.g., gaze) to determine the target of the user input (e.g., for direct input, as described below). Thus, in specific implementations involving in-air gestures, for example, the input gesture is combined (e.g., simultaneously) with movement of the user's fingers and / or hand to detect attention (e.g., gaze) toward a user interface element to perform a pinch and / or tap input, as described below.
[0186] In some embodiments, an input gesture directed to a user interface object is performed directly or indirectly with reference to the user interface object. For example, user input is performed directly on the user interface object based on performing input with the user's hand at a location corresponding to the location of the user interface object in the three-dimensional environment (e.g., as determined based on the user's current viewpoint). In some embodiments, upon detecting user attention (e.g., gaze) to the user interface object, an input gesture is performed indirectly on the user interface object based on the user's hand being located not at the location corresponding to the location of the user interface object in the three-dimensional environment while the user performs the input gesture. For example, for a direct input gesture, the user can direct the user's input to the user interface object by initiating a gesture at or near a location corresponding to the displayed location of the user interface object (e.g., within 0.5 cm, 1 cm, 5 cm, or a distance between 0 and 5 cm measured from the outer edge of the option or the center portion of the option). For an indirect input gesture, the user can direct the user's input to the user interface object by focusing on the user interface object (e.g., by gazing at the user interface object), and while focusing on the option, the user initiates an input gesture (e.g., at any location detectable by the computer system) (e.g., at a location that does not correspond to the displayed location of the user interface object).
[0187] In some embodiments, according to some embodiments, input gestures (e.g., air gestures) used in various examples and embodiments described herein include pinch input and tap input for interacting with a virtual or mixed reality environment. For example, the pinch input and tap input described below are performed as air gestures.
[0188] In some embodiments, a pinch input is part of an air gesture that includes one or more of: a pinch gesture, a long pinch gesture, a pinch and drag gesture, or a double pinch gesture. For example, a pinch gesture as an air gesture includes movement of two or more fingers of a hand to contact each other, i.e., optionally followed by a break in contact with each other immediately (e.g., within 0 seconds to 1 second). A long pinch gesture as an air gesture includes movement of two or more fingers of a hand in contact with each other for at least a threshold amount of time (e.g., at least 1 second) before a break in contact with each other is detected. For example, a long pinch gesture includes the user maintaining a pinch gesture (e.g., in which the two or more fingers are in contact), and the long pinch gesture continues until a break in contact between the two or more fingers is detected. In some embodiments, a double pinch gesture as an air gesture includes two (e.g., more) pinch inputs (e.g., performed by the same hand) that are detected consecutively immediately (e.g., within a predefined time period) with respect to each other. For example, the user performs a first pinch input (e.g., a pinch input or a long pinch input), releases the first pinch input (e.g., interrupts contact between two or more fingers), and performs a second pinch input within a predefined time period (e.g., within 1 second or within 2 seconds) after releasing the first pinch input.
[0189] In some embodiments, a pinch and drag gesture as an air gesture includes a pinch gesture (e.g., a pinch gesture or a long pinch gesture) performed in conjunction with (e.g., following) a drag input that changes the position of the user's hand from a first position (e.g., the starting position of the drag) to a second position (e.g., the ending position of the drag). In some embodiments, the user maintains the pinch gesture while performing the drag input, and releases the pinch gesture (e.g., opens their two or more fingers) to end the drag gesture (e.g., at the second position). In some embodiments, the pinch input and the drag input are performed by the same hand (e.g., the user pinches two or more fingers to contact each other and moves the same hand to a second position in the air using a drag gesture). In some embodiments, the pinch input is performed by the user's first hand, and the drag input is performed by the user's second hand (e.g., the user's second hand moves in the air from the first position to the second position while the user continues the pinch input with the user's first hand). In some embodiments, an input gesture as an air gesture includes input performed using both hands of the user (e.g., a pinch and / or tap input). For example, the input gesture includes two (e.g., more) pinch inputs performed in conjunction with each other (e.g., concurrently or within a predefined time period). For example, a first pinch gesture (e.g., a pinch input, a long pinch input, or a pinch and drag input) is performed using a first hand of a user, and a second pinch input is performed using another hand (e.g., a second hand of the user) in conjunction with the pinch input performed using the first hand. In some embodiments, movement between the user's two hands (e.g., increasing and / or decreasing the distance or relative orientation between the user's two hands) occurs.
[0190] In some embodiments, a tap input performed as an air gesture (e.g., pointing to a user interface element) includes movement of a user's finger toward the user interface element, movement of the user's hand toward the user interface element (optionally, extension of the user's finger toward the user interface element), a downward motion of the user's finger (e.g., mimicking a mouse click motion or a tap on a touch screen), or other predefined movement of the user's hand. In some embodiments, a tap input performed as an air gesture is detected based on movement characteristics of the finger or hand performing the tap gesture movement of the finger or hand, which is a movement of the finger or hand away from the user's viewpoint and / or toward an object that is the target of the tap input, followed by an end of the movement. In some embodiments, the end of the movement is detected based on a change in movement characteristics of the finger or hand performing the tap gesture (e.g., an end of movement away from the user's viewpoint and / or toward an object that is the target of the tap input, a reversal of the direction of movement of the finger or hand, and / or a reversal of the acceleration direction of the movement of the finger or hand).
[0191] In some embodiments, the user's attention is determined to be directed toward a portion of the three-dimensional environment based on detection of a gaze directed toward the portion of the three-dimensional environment (optionally, no other conditions are required). In some embodiments, the user's attention is determined to be directed toward a portion of the three-dimensional environment based on detection of a gaze directed toward the portion of the three-dimensional environment using one or more additional conditions, such as requiring the gaze to be directed toward the portion of the three-dimensional environment for at least a threshold duration (e.g., a dwell duration) and / or requiring the gaze to be directed toward the portion of the three-dimensional environment when the user's viewpoint is within a distance threshold from the portion of the three-dimensional environment, so that the device determines that the user's attention is directed toward the portion of the three-dimensional environment, wherein if one of these additional conditions is not met, the device determines that the attention is not directed toward the portion of the three-dimensional environment to which the gaze is directed (e.g., until the one or more additional conditions are met).
[0192] In some embodiments, the detection of a ready state configuration of a user or a portion of a user is detected by a computer system. The detection of a ready state configuration of a hand is used by the computer system as an indication that the user may be preparing to interact with the computer system using one or more air gesture inputs performed by the hand (e.g., a pinch, a tap, a pinch and drag, a double pinch, a long pinch, or other air gestures described herein). For example, the ready state of a hand is determined based on whether the hand has a predetermined hand shape (e.g., a pre-pinch shape with the thumb and one or more fingers extended and spaced apart in preparation for a pinch or grab gesture, or a pre-tap with one or more fingers extended and the palm facing away from the user), based on whether the hand is in a predetermined position relative to the user's viewpoint (e.g., below the user's head and above the user's waist and extending at least 15 cm, 20 cm, 25 cm, 30 cm, or 50 cm from the body), and / or based on whether the hand has moved in a particular manner (e.g., toward an area in front of the user above the user's waist and below the user's head, or away from the user's body or legs). In some embodiments, the ready state is used to determine whether an interactive element of a user interface responds to attention (e.g., gaze) input.
[0193] In scenarios where input is described with reference to in-air gestures, it should be understood that similar gestures can be detected using a hardware input device attached to or held by one or more hands of a user, where the positioning of the hardware input device in space can be tracked using optical tracking, one or more accelerometers, one or more gyroscopes, one or more magnetometers, and / or one or more inertial measurement units, and the positioning and / or movement of the hardware input device is used instead of the positioning and / or movement of the one or more hands in the corresponding in-air gesture. In scenarios where input is described with reference to in-air gestures, it should be understood that similar gestures can be detected using a hardware input device attached to or held by one or more hands of a user, and user input can be detected using controls contained in the hardware input device, such as one or more touch-sensitive input elements, one or more pressure-sensitive input elements, one or more buttons, one or more knobs, one or more dials, one or more joysticks, one or more hand or finger overlays that can detect the position or change in position of parts of a hand and / or finger relative to each other, relative to the user's body, and / or relative to the user's physical environment, and / or other hardware input device controls, wherein user input performed using controls contained in the hardware input device is used in place of hand and / or finger gestures such as an air tap or air pinch in the corresponding in-air gesture. For example, a selection input described as being performed using an air tap or air pinch input can alternatively be detected using a button press, a tap on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input. For another example, movement input described as being performed using a mid-air pinch and drag may alternatively be detected based on interaction with a hardware input control, such as a button press and hold, a touch on a touch-sensitive surface, a press on a pressure-sensitive surface, or other hardware input following movement of a hardware input device (e.g., along with a hand associated with the hardware input device) through space. Similarly, two-handed input involving movement of hands relative to each other may be performed using one mid-air gesture and one hardware input device in a hand that is not performing the mid-air gesture, two hardware input devices held in different hands, or two mid-air gestures performed by different hands using various combinations of mid-air gestures and / or input detected by one or more of the aforementioned hardware input devices.
[0194] In some embodiments, the software may be downloaded to the controller 110 in electronic form, for example, over a network, or may alternatively be provided on tangible, non-transitory media such as optical, magnetic, or electronic memory media. In some embodiments, the database 408 is also stored in memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although in Figure 4, but some or all of the processing functions of the controller may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of the image sensor 404 (e.g., a hand tracking device) or other device associated with the image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with the display generation component 120 (e.g., in a television receiver, handheld device, or head-mounted device) or integrated with any other suitable computerized device (such as a game console or media player). The sensing functions of the image sensor 404 may also be integrated into a computer or other computerized device to be controlled by the sensor output.
[0195] Figure 4 Also included is a schematic diagram of a depth map 410 captured by the image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels with corresponding depth values. Pixels 412 corresponding to the hand 406 have been segmented from the background and wrist in the figure. The brightness of each pixel within the depth map 410 is inversely proportional to its depth value (i.e., the measured z distance from the image sensor 404), where shades of gray become darker with increasing depth. The controller 110 processes these depth values in order to identify and segment components of the image (i.e., a group of adjacent pixels) that have characteristics of a human hand. These characteristics may include, for example, overall size, shape, and motion from frame to frame in the depth map sequence.
[0196] Figure 4 Also schematically illustrated is a hand skeleton 414 that the controller 110 ultimately extracts from the depth map 410 of the hand 406 according to some embodiments. Figure 4 , a hand skeleton 414 is superimposed on a hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points of the hand and, optionally, on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, the center of the palm, the end of the hand connected to the wrist, etc.) are identified and located on the hand skeleton 414. In some embodiments, the controller 110 uses the position and movement of these key feature points over multiple image frames to determine a gesture performed by the hand or the current state of the hand according to some embodiments.
[0197] Figure 5 The eye tracking device 130 ( Figure 1A ). In some embodiments, the eye tracking device 130 is composed of an eye tracking unit 243 ( Figure 2) controls to track the position and movement of the user's gaze relative to the scene 105 or relative to the XR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as a headset, helmet, goggles, or glasses) or a handheld device placed in a wearable frame, the head-mounted device includes both components for generating XR content for the user to view and components for tracking the user's gaze relative to the XR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a handheld device or an XR room, the eye tracking device 130 is optionally a device separate from the handheld device or the XR room. In some embodiments, the eye tracking device 130 is a head-mounted device or a part of the head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in conjunction with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in conjunction with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.
[0198] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, the head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, the head-mounted display generation component may have a transparent or translucent display and display virtual objects on the transparent or translucent display, through which the user can directly view the physical environment. In some embodiments, the display generation component projects the virtual objects into the physical environment. The virtual objects may, for example, be projected onto a physical surface or projected as a hologram, so that the individual using the system observes the virtual objects superimposed on the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.
[0199] like Figure 5As shown in , in some embodiments, the eye tracking device 130 (e.g., a gaze tracking device) includes at least one eye tracking camera (e.g., an infrared (IR) or near infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera can be pointed at the user's eyes to receive IR or NIR light that the light source reflects directly from the eyes, or alternatively can be pointed at "hot" mirrors located between the user's eyes and the display panel, which reflect IR or NIR light from the eyes toward the eye tracking camera while allowing visible light to pass through. The eye tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames per second (fps)), analyzes these images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, both eyes of the user are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one eye of the user is tracked by corresponding eye tracking camera and illumination source.
[0200] In some embodiments, the eye tracking device 130 is calibrated using a device-specific calibration process to determine the parameters of the eye tracking device for a specific operating environment 100, such as the 3D geometry and parameters of the LED, camera, thermal mirror (if present), eye lens, and display screen. The device-specific calibration process can be performed at a factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process can include an estimation of eye parameters of a specific user, such as pupil position, fovea position, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's gaze point relative to the display.
[0201] like Figure 5As shown in FIG, an eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near infrared (NIR) camera) positioned on the side of the user's face on which eye tracking is performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 may be directed toward a mirror 550 (which reflects the IR or NIR light from the eye 592 while allowing visible light to pass) located between the user's eye 592 and a display 510 (e.g., a left display panel or a right display panel of a head-mounted display, or a display of a handheld device, a projector, etc.). Figure 5 ), or alternatively may be directed toward the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the top portion of Figure 5 (as shown in the bottom portion of the ).
[0202] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for left and right display panels) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained from the eye tracking camera 540 using a flash-assisted method or other suitable method. The gaze point estimated from the gaze tracking input 542 is optionally used to determine the direction the user is currently looking.
[0203] The following describes several possible use cases for the user's current gaze direction and is not intended to be limiting. As an example use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal region determined based on the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content within a view based at least in part on the user's current gaze direction. As another example, the controller may display specific virtual content within a view based at least in part on the user's current gaze direction. As another example use case in an AR application, the controller 110 may direct an external camera used to capture the physical environment of an XR experience to focus in a determined direction. The external camera's autofocus mechanism may then focus on an object or surface in the environment on the display 510 that the user is currently looking at. As another example use case, the eye lens 520 may be a focusable lens, and the controller may use gaze tracking information to adjust the focus of the eye lens 520 so that the virtual object the user is currently looking at has the appropriate vergence to match the convergence of the user's eye 592. The controller 110 can use the gaze tracking information to guide the eye lens 520 to adjust the focus so that nearby objects that the user is looking at appear at the correct distance.
[0204] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510), two eye lenses (e.g., eye lenses 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) toward the user's eyes 592. In some embodiments, the light sources can be arranged in a ring or circle around each of the lenses, such as Figure 5 In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 can be used, and other arrangements and positions of the light sources 530 can be used.
[0205] In some embodiments, the display 510 emits light in the visible range and does not emit light in the IR or NIR range, and therefore does not introduce noise into the gaze tracking system. It should be noted that the positions and angles of the eye tracking cameras 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, a camera 540 with a wider field of view (FOV) and a camera 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, a camera 540 operating at one wavelength (e.g., 850 nm) and a camera 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.
[0206] like Figure 5 The illustrated embodiments of the gaze tracking system may be used, for example, in computer-generated reality, virtual reality, and / or mixed reality applications to provide a user with a computer-generated reality, virtual reality, augmented reality, and / or enhanced virtual experience.
[0207] Figure 6 A flash-assisted gaze tracking pipeline according to some embodiments is illustrated. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., Figure 1A and Figure 5 The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no." When in the tracking state, the flash-assisted gaze tracking system uses previous information from previous frames when analyzing the current frame to track the pupil outline and glint in the current frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and glint in the current frame, and if successful, initializes the tracking state to "yes" and continues to the next frame in the tracking state.
[0208] like Figure 6 As shown in , the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input to the gaze tracking pipeline for processing starting at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes at a rate of, for example, 60 to 120 frames per second. In some embodiments, each set of captured images can be input to the pipeline for processing. However, in some embodiments or under some conditions, not all captured frames are processed by the pipeline.
[0209] At 610, for the currently captured image, if the tracking status is yes, the method proceeds to element 640. At 610, if the tracking status is no, the image is analyzed to detect the user's pupil and glint in the image, as indicated at 620. At 630, if the pupil and glint are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.
[0210] At 640, if proceeding from element 610, the current frame is analyzed to track the pupil and glint based in part on previous information from the previous frame. At 640, if proceeding from element 630, the tracking state is initialized based on the pupil and glint detected in the current frame. The processing result at element 640 is checked to verify that the tracking or detection result can be trusted. For example, the result can be checked to determine whether the pupil and a sufficient number of glints were successfully tracked or detected in the current frame to perform gaze estimation. At 650, if the result is not likely to be trusted, at element 660, the tracking state is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trustworthy, the method proceeds to element 670. At 670, the tracking state is set to yes (if not already yes), and the pupil and glint information is passed to element 680 to estimate the user's gaze point.
[0211] Figure 6 This is intended to be used as an example of an eye tracking technology that may be used for a particular implementation. As one of ordinary skill in the art will appreciate, according to various embodiments, other eye tracking technologies currently existing or developed in the future may be used in place of or in combination with the flash-assisted eye tracking technology described herein in the computer system 101 for providing an XR experience to a user.
[0212] In this disclosure, various input methods are described with respect to interaction with a computer system. When an example is provided using one input device or input method, and another example is provided using another input device or input method, it should be understood that each example is compatible with and optionally utilizes the input device or input method described with respect to the other example. Similarly, various output methods are described with respect to interaction with a computer system. When an example is provided using one output device or output method, and another example is provided using another output device or output method, it should be understood that each example is compatible with and optionally utilizes the output device or output method described with respect to the other example. Similarly, various methods are described with respect to interaction with a virtual environment or a mixed reality environment through a computer system. When an example is provided using interaction with a virtual environment, and another example is provided using a mixed reality environment, it should be understood that each example is compatible with and optionally utilizes the methods described with respect to the other example. Therefore, this disclosure discloses embodiments that are combinations of features from multiple examples, without necessarily listing all features of the embodiments in detail in the description of each example embodiment.
[0213] User interface and associated processes
[0214] Attention is now focused on embodiments of a user interface ("UI") and associated processes that may be implemented on a computer system, such as a portable multifunction device or a head-mounted device, in communication with a display generating component, one or more input devices, and optionally one or more physical controls.
[0215] Figures 7A to 7K Examples of techniques for navigating an extended reality experience are illustrated. Figure 8 is a flow chart of an exemplary method 800 for navigating an extended reality experience. Figure 9 is a flow chart of an exemplary method 900 for navigating an extended reality experience. Figures 7A to 7K The user interface in the example is used to illustrate the process described below, including Figure 8 and Figure 9 in the process.
[0216] Figure 7AAn electronic device 700 is depicted that is a smartphone that includes a touch-sensitive display 702, buttons 704a-704c, and one or more input sensors 706 (e.g., one or more cameras, eye gaze trackers, hand movement trackers, and / or head movement trackers). In some embodiments described below, the electronic device 700 is a smartphone. In some embodiments, the electronic device 700 is a tablet computer, a wearable device, a wearable smartwatch device, a head-mounted system (e.g., a headset), or other computer system that includes one or more display devices (e.g., a display screen, a projection device, etc.) and / or communicates with the one or more display devices. The electronic device 700 is a computer system (e.g., Figure 1A Computer system 101 in).
[0217] exist Figure 7A At , the electronic device 700 is in a low-power, inactive, or dormant state, where content is not displayed via the display 702. Figure 7A At , the electronic device 700 detects user input 708. In the depicted embodiment, the user input 708 is a button press input via button 704c. However, in some embodiments, the user input 708 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 708 includes, for example, the user placing the electronic device 700 on his or her head, performing a gesture while wearing the electronic device 700, pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0218] exist Figure 7B At , in response to user input 708, the electronic device 700 transitions from a low-power, inactive, or sleep state to an active state, wherein the electronic device 700 displays a three-dimensional environment 712 and an extended reality experience 714 (e.g., an augmented reality experience and / or a virtual reality experience) via the display 702. In the depicted scene, the three-dimensional environment 712 includes a chair, a table, and a place setting (e.g., a napkin, a fork, a knife, and a cup) placed on the table. In some embodiments, the three-dimensional environment 712 is displayed by the display (e.g., Figure 7B In some embodiments, the three-dimensional environment 712 includes a virtual environment or a scene captured by one or more cameras (e.g., one or more cameras as part of the input sensor 706 and / or Figure 7BIn some embodiments, three-dimensional environment 712 is visible to the user behind extended reality experience 714, but is not displayed by a display. For example, in some embodiments, three-dimensional environment 712 is a physical environment that is visible to the user behind extended reality experience 712 (e.g., through a transparent display) but is not displayed by a display.
[0219] exist Figure 7B , the extended reality experience 714 is a camera extended reality experience, as indicated by identifier 716a, which includes a camera logo and the name of the extended reality experience. The camera extended reality experience 714 includes a camera that can be selected by the user to be viewed via one or more cameras (e.g., one or more cameras as part of input sensor 706 and / or Figure 7B One or more selectable objects 716b-e for capturing photos and / or video content (one or more cameras not shown). Object 716b is a shutter button that can be selected to capture photos and / or video. Object 716c can be selected to enable slow motion capture mode. Option 716d can be selected to enable photo capture mode. Option 716e can be selected to enable video capture mode. Figure 7B , the electronic device 700 detects that the user is looking toward the right side of the display 702, as indicated by the gaze indication 710. The gaze indication 710 is provided for better understanding of the described technology and is optionally not part of the user interface of the described device (e.g., not displayed by the electronic device 700). Figure 7B At , the electronic device 700 detects user input 718. In the depicted embodiment, the user input 718 is a button press input via button 704c. However, in some embodiments, the user input 718 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 718 includes, for example, the user performing a gesture while wearing the electronic device 700 (e.g., an air gesture), pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0220] exist Figure 7C In response to user input 718, electronic device 700 displays an animation in which extended reality experience 714 appears to move away from the user. Figure 7C, the electronic device 700 displays a representation 720 representing a camera extended reality experience 714. The representation 720 includes objects 722a-722e representing objects 716a-716e superimposed on a background portion 722f and surrounded by a border 719 (e.g., these objects are smaller, non-interactive versions of the objects 716a-716e). The representation 720 appears to move away from the user by, for example, gradually becoming smaller over time. In some embodiments, in response to user input 718 and / or when displaying animations, the three-dimensional environment 712 is visually blurred (e.g., Figure 7C ) (e.g., displayed with reduced focus, reduced sharpness, reduced color saturation, and / or greater opacity) in order to draw the user's attention and fixation on representation 720. As discussed above, in some embodiments, three-dimensional environment 712 is a "see-through" environment that the user sees through a transparent display and is not displayed by the display. In some such embodiments, three-dimensional environment 712 is visually de-emphasized by applying masking or other techniques to areas of a display (e.g., display 702) through which the user can view three-dimensional environment 712.
[0221] exist Figure 7D1At , the animation of representation 720 (and / or extended reality experience 714) appearing to move away from the user completes, and representation 720 is now displayed at the top of the stack of representations 721, 724, and 726. Representations 721, 724, and / or 726 represent other extended reality experiences that can be selected by the user and / or displayed by electronic device 700. For example, as discussed above, representation 720 represents a camera extended reality experience (e.g., camera extended reality experience 714). In some embodiments, representation 724 represents a music extended reality experience (e.g., it includes one or more selectable options for playing music), representation 726 represents a translation extended reality experience (e.g., it includes one or more selectable options for translating content (e.g., content captured by one or more cameras and / or content within the user's field of view and / or the field of view of electronic device 700)), and representation 721 includes, for example, a representation of a reading extended reality experience, a representation of a photo gallery extended reality experience, a representation of a video messaging extended reality experience, a representation of a navigation extended reality experience, and / or a representation of a fitness extended reality experience. As will be shown in subsequent figures, a user can scroll through the stack of representations 720, 721, 724, and / or 726 to select which extended reality experience the user wants to display. In some embodiments, each extended reality experience corresponds to a different color, and the representation corresponding to the extended reality experience is displayed in the corresponding color corresponding to the extended reality experience. For example, in some embodiments, the camera extended reality experience 714 corresponds to a first color, and the representation 720 is displayed in the first color (e.g., the background 722f is displayed in the first color, the border 720 is displayed in the first color, and / or the object 722a (e.g., the logo and / or name) is displayed in the first color); and the music extended reality experience corresponds to a second color, so that the representation 724 is displayed in the second color (e.g., the background portion of the representation 724, the border of the representation 724, and / or the identifier of the representation 724 are displayed in the second color). In this way, the user can quickly identify the order of the extended reality experience stack based on the colors of the representations 720, 721, 724, and / or 726.
[0222] exist Figure 7D1 , representation 720 is displayed at a first display position (e.g., at the top of the stack), indicating that a selection input (e.g., a press of button 704c or other selection input) will result in an extended reality experience corresponding to representation 720 being displayed (e.g., will result in camera extended reality experience 714 being displayed). Figure 7D1 At , the electronic device 700 detects user input 727. Figure 7D1, user input 727 is a button press of button 704a. In some embodiments, the button press of button 704a indicates a request to navigate and / or scroll in a first direction (e.g., rotate the stack forward), and the button press of button 704b indicates a request to navigate and / or scroll in a second direction (e.g., rotate the stack backward). In addition, in some embodiments, user input 727 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, the electronic device 700 is a head-mounted system, and the user input 727 includes, for example, the user performing a gesture while wearing the electronic device 700 (e.g., an air gesture), pressing a button while wearing the electronic device 700, rotating a rotatable input mechanism while wearing the electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing. For example, in some embodiments, rotation of the rotatable input mechanism in a first direction (e.g., rotation in a clockwise direction) (in some embodiments, rotation of the rotatable input mechanism while looking toward the stack) indicates a request to navigate and / or scroll in a second direction (e.g., rotation in a counterclockwise direction), and rotation of the rotatable input mechanism in a third direction (in some embodiments, rotation of the rotatable input mechanism while looking toward the stack) indicates a request to navigate and / or scroll in a fourth direction (e.g., to rotate the stack backward).
[0223] In some embodiments, Figures 7A to 7K The techniques and user interfaces described in Figures 1A to 1P For example, Figures 7D2 to 7D4 Illustrated in which Figures 7B to 7D1 Implementations of the transition animation described in
[0015] are displayed on a display module X702 of a head-mounted device (HMD) X700. In some implementations, the device X700 includes a pair of display modules that provide stereoscopic content to different eyes of the same user. For example, the HMD X700 includes a display module X702 that provides content to the user's left eye and a second display module that provides content to the user's right eye. In some implementations, the second display module displays a slightly different image than the display module X702 to create the illusion of stereoscopic depth.
[0224] exist Figure 7D2 , the extended reality experience 714 is a camera extended reality experience, as indicated by the identifier 716a, which includes a camera logo and the name of the extended reality experience. The camera extended reality experience 714 includes one or more cameras (e.g., one or more cameras as part of the input sensor X706 and / or Figure 7D2One or more selectable objects 716b-e for capturing photos and / or video content (one or more cameras not shown). Object 716b is a shutter button that can be selected to capture photos and / or video. Object 716c can be selected to enable slow motion capture mode. Option 716d can be selected to enable photo capture mode. Option 716e can be selected to enable video capture mode. Figure 7D2 , HMD X700 detects that the user is looking to the right of display module X702, as indicated by gaze indication 710. Gaze indication 710 is provided for better understanding of the described technology and is optionally not part of the user interface of the described device (e.g., not displayed by HMD X700). Figure 7D2 At , HMD X700 detects user input 718. In the depicted embodiment, user input 718 is a button press input via button X704c. However, in some embodiments, user input 718 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, user input 718 includes, for example, the user performing a gesture while wearing the HMD X700 (e.g., an air gesture), pressing a button while wearing the HMD X700, rotating a rotatable input mechanism while wearing the HMD X700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0225] exist Figure 7D3 In response to user input 718, HMD X700 displays an animation in which the extended reality experience 714 appears to move away from the user. Figure 7D3 , HMD X700 displays a representation 720 representing a camera extended reality experience 714. Representation 720 includes objects 722a-722e representing objects 716a-716e superimposed on a background portion 722f and surrounded by a border 719 (e.g., these objects are smaller, non-interactive versions of objects 716a-716e). Representation 720 appears to move away from the user by, for example, gradually becoming smaller over time. In some embodiments, in response to user input 718 and / or when displaying animations, three-dimensional environment 712 is visually blurred (e.g., Figure 7D3) (e.g., displayed with reduced focus, reduced sharpness, reduced color saturation, and / or greater opacity) in order to draw the user's attention and fixation on representation 720. As discussed above, in some embodiments, three-dimensional environment 712 is a "see-through" environment that the user sees through a transparent display and is not displayed by the display. In some such embodiments, three-dimensional environment 712 is visually de-emphasized by applying masking or other techniques to areas of the display (e.g., display module X 702) through which the user can view three-dimensional environment 712.
[0226] exist Figure 7D4 At , the animation of representation 720 (and / or extended reality experience 714) appearing to move away from the user completes, and representation 720 is now displayed at the top of the stack of representations 721, 724, and 726. Representations 721, 724, and / or 726 represent other extended reality experiences that can be selected by the user and / or displayed by electronic device 700. For example, as discussed above, representation 720 represents a camera extended reality experience (e.g., camera extended reality experience 714). In some embodiments, representation 724 represents a music extended reality experience (e.g., it includes one or more selectable options for playing music), representation 726 represents a translation extended reality experience (e.g., it includes one or more selectable options for translating content (e.g., content captured by one or more cameras and / or content within the user's field of view and / or the field of view of electronic device 700)), and representation 721 includes, for example, a representation of a reading extended reality experience, a representation of a photo gallery extended reality experience, a representation of a video messaging extended reality experience, a representation of a navigation extended reality experience, and / or a representation of a fitness extended reality experience. As will be shown in subsequent figures, a user can scroll through the stack of representations 720, 721, 724, and / or 726 to select which extended reality experience the user wants to display. In some embodiments, each extended reality experience corresponds to a different color, and the representation corresponding to the extended reality experience is displayed in the corresponding color corresponding to the extended reality experience. For example, in some embodiments, the camera extended reality experience 714 corresponds to a first color, and the representation 720 is displayed in the first color (e.g., the background 722f is displayed in the first color, the border 720 is displayed in the first color, and / or the object 722a (e.g., the logo and / or name) is displayed in the first color); and the music extended reality experience corresponds to a second color, so that the representation 724 is displayed in the second color (e.g., the background portion of the representation 724, the border of the representation 724, and / or the identifier of the representation 724 are displayed in the second color). In this way, the user can quickly identify the order of the extended reality experience stack based on the colors of the representations 720, 721, 724, and / or 726.
[0227] exist Figure 7D4 , representation 720 is displayed at a first display position (e.g., at the top of the stack), indicating that a selection input (e.g., a press of button 704c or other selection input) will result in an extended reality experience corresponding to representation 720 being displayed (e.g., will result in camera extended reality experience 714 being displayed). Figure 7D4 At , HMD X700 detects user input 727. Figure 7D4 , user input 727 is a button press of button X704a. In some embodiments, the button press of button X704a indicates a request to navigate and / or scroll in a first direction (e.g., rotate the stack forward), and the button press of button X704b indicates a request to navigate and / or scroll in a second direction (e.g., rotate the stack backward). In addition, in some embodiments, user input 727 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, user input 727 includes, for example, the user performing a gesture while wearing HMD X700 (e.g., an air gesture), pressing a button while wearing HMD X700, rotating a rotatable input mechanism while wearing HMD X700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing. For example, in some embodiments, rotation of the rotatable input mechanism in a first direction (e.g., rotation in a clockwise direction) (in some embodiments, rotation of the rotatable input mechanism while looking toward the stack) indicates a request to navigate and / or scroll in a second direction (e.g., rotation in a counterclockwise direction), and rotation of the rotatable input mechanism in a third direction (in some embodiments, rotation of the rotatable input mechanism while looking toward the stack) indicates a request to navigate and / or scroll in a fourth direction (e.g., to rotate the stack backward).
[0228] Figure 1B to Figure 1PAny of the features, components, and / or parts shown (including their arrangements and configurations) may be included, alone or in any combination, in HMD X700. For example, in some embodiments, HMD X700 includes any of the features, components, and / or parts of HMDs 1-100, 1-200, 3-100, 6-100, 6-200, 6-300, 6-400, 11.1.1-100, and / or 11.1.2-100, alone or in any combination. In some embodiments, the display module X702 includes, alone or in any combination, a display unit 1-102, a display unit 1-202, a display unit 1-306, a display unit 1-406, a display generation component 120, display screens 1-122a-b, a first rear display screen 1-322a and a second rear display screen 1-322b, a display 11.3.2-104, a first display component 1-120a and a second display component 1-120b, a display component 1-320, and a display component 1-421. , the first display subassembly 1-420a and the second display subassembly 1-420b, the display component 3-108, the display component 11.3.2-204, the first optical module 11.1.1-104a and the second optical module 11.1.1-104b, the optical module 11.3.2-100, the optical module 11.3.2-200, the double convex lens array 3-110, the display area or display area 6-232 and / or any of the features, components and / or parts of the display / display area 6-334. In some embodiments, the HMD X700 includes sensors including, individually or in any combination, any of the features, components, and / or parts of sensor 190, sensor 306, image sensor 314, image sensor 404, sensor assembly 1-356, sensor assembly 1-456, sensor system 6-102, sensor system 6-202, sensor 6-203, sensor system 6-302, sensor 6-303, sensor system 6-402, and / or any of sensors 11.1.2-110a-f. In some embodiments, the HMD X700 includes one or more input devices including, individually or in any combination, any of the features, components, and / or parts of first button 1-128, button 11.1.1-114, second button 1-132, and / or any of dial or button 1-328. In some embodiments, HMD X700 includes one or more audio output components (e.g., electronic components 1-112) for generating audio feedback (e.g., audio output X714-3), which is optionally generated based on detected events and / or user input detected by HMD X700.
[0229] exist Figure 7EIn response to user input 727, electronic device 700 stops displaying representation 720 at the top of the stack and now displays representation 724 (which is Figure 7D1 724 represents the music extended reality experience and displays objects 728a-728d superimposed on a background 728e and surrounded by a border 723. Objects 728a-728d represent objects that would be displayed in the music extended reality experience if the user selected the music extended reality experience for display. Thus, representation 724 provides the user with a preview of what the music extended reality experience will look like. Figure 7E At , the electronic device 700 detects user input 729. Figure 7E , user input 729 is a button press of button 704a. As discussed above, in some embodiments, user input 729 is a different type of input, such as a gesture or other action taken by the user. For example, in some embodiments, electronic device 700 is a head-mounted system, and user input 729 includes, for example, the user performing a gesture while wearing electronic device 700 (e.g., an air gesture), pressing a button while wearing electronic device 700, rotating a rotatable input mechanism while wearing electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0230] exist Figure 7F In response to user input 729, electronic device 700 stops displaying representation 724 at the top of the stack and now displays representation 726 (which is Figure 7E In some embodiments, if the user input 729 was a request to rotate the stack in the opposite direction (e.g., a button press of button 704b), the electronic device 700 will redisplay representation 720 at the top of the stack (and the second representation 724 in the stack, as shown in FIG. Figure 7D1 ). Figure 7F , representation 726 represents a translation extended reality experience and includes objects 730a-730d superimposed on a background 730 and surrounded by a border 725. In some embodiments, object 730a is an identifier that identifies the translation extended reality experience (e.g., via a logo and / or name), and objects 730b-730d represent selectable objects that will be displayed in the translation extended reality experience. In some embodiments, objects 730b-730d represent selectable objects (as will be described below with reference to Figure 7J described), but they are not individually selectable by themselves to perform any function. Figure 7F At , the electronic device 700 detects user input 732. Figure 7FIn some embodiments, user input 732 is a touch screen swipe gesture with a downward direction. However, in some embodiments, user input 732 is a different type of user input, such as a gesture or other action taken by the user. For example, in some embodiments, electronic device 700 is a head-mounted system, and user input 732 includes, for example, the user performing a gesture while wearing electronic device 700 (e.g., an air gesture), pressing a button while wearing electronic device 700, rotating a rotatable input mechanism while wearing electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0231] exist Figure 7G At 732, in response to user input 732, the electronic device 700 displays a system control user interface 734 including selectable objects 736a-736h. Object 736a can be selected to selectively engage or disengage a "Do Not Disturb" state or a dormant focus state, in which notifications received by the electronic device 700 are suppressed. Object 736b can be selected to selectively turn WiFi on or off. Object 736c can be selected to selectively engage or disengage airplane mode. Object 736d can be selected to selectively turn the flashlight on or off. Object 736e can be selected to initiate a process for streaming audio and / or video content to an external device. Object 736f can be selected to selectively engage or disengage silent mode. Object 736g can be selected to modify the volume setting of the electronic device 700. Object 736h can be selected to modify the brightness of the electronic device 700. In some embodiments, the electronic device 700 is a head-mounted system, and option 736h is selectable to modify the pass-through brightness setting and / or pass-through opacity setting of the electronic device 700. exist Figure 7G At , the electronic device 700 detects user input 736. Figure 7G , user input 736 is a tap input on touch-sensitive display 702. However, in some embodiments, user input 736 is a different type of user input, such as a gesture or other action taken by the user. For example, in some embodiments, electronic device 700 is a head-mounted system, and user input 736 includes, for example, the user performing a gesture while wearing electronic device 700 (e.g., an air gesture), pressing a button while wearing electronic device 700, rotating a rotatable input mechanism while wearing electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0232] exist Figure 7H At , in response to user input 736, electronic device 700 stops the display of system controlled user input 734. Figure 7HAt , the electronic device 700 displays representation 726 at the top of the stack of representations, and when representation 736 is displayed at the top of the stack of representations, the electronic device 700 detects user input 740 (e.g., a selection input). Figure 7H , user input 740 is a button press input of button 704c. However, in some embodiments, user input 740 is a different type of user input, such as a gesture or other action taken by the user. For example, in some embodiments, electronic device 700 is a head-mounted system, and user input 740 includes, for example, the user performing a gesture while wearing electronic device 700 (e.g., an air gesture), pressing a button while wearing electronic device 700, rotating a rotatable input mechanism while wearing electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0233] exist Figure 7I At , in response to user input 740, electronic device 700 stops display of representation 721 (e.g., stops display of the stack of representations) and displays an animation in which representation 726 appears to move toward the user. Figure 7I , representation 726 (including objects 730a-730b and boundary 725) becomes larger. Figure 7I , the background 730e changes from opaque to transparent to show the three-dimensional environment 712 behind the representation 736.
[0234] exist Figure 7J At , electronic device 700 completes the animation of representation 721 becoming larger and now replaces the display of representation 721 with the display of translated extended reality experience 742. In some embodiments, when transitioning from the display of representation 721 to the display of translated extended reality experience 742, electronic device 700 displays a cross-fade of objects 730a-730d and corresponding objects 744a-744d. Figure 7J In the example, the three-dimensional environment 712 is no longer visually de-emphasized (e.g., from Figure 7I The dotted line to Figure 7J (indicated by the solid line transition in ).
[0235] The translation extended reality experience 742 includes an object 744a that identifies the extended reality experience (e.g., using a logo and / or name) and objects 744b-744d that can be selected to perform various tasks. For example, in some embodiments, object 744b can be selected to engage a microphone, enabling a user to provide spoken and / or verbal input for transitioning to a different language; object 744c can be selected to translate visual content captured by one or more cameras (e.g., input sensor 706); and object 744d can be selected to cause the electronic device 700 to read the translation aloud (e.g., play audio content reading the translation aloud). Figure 7J In the example, the translation extended reality experience 742 also includes an object 746 that indicates that the electronic device 700 has detected visual content that can be translated. Figure 7J , the menu has been moved into the view of the electronic device 700 (e.g., into the view of one or more cameras), and object 746 indicates that the menu includes text that can be translated into different languages. Figure 7J At , the electronic device 700 detects user input 748 while also detecting that the user is looking at the object 746 (e.g., as indicated by the gaze indication 710). Figure 7J , user input 748 is a tap input via touch-sensitive display 702. However, in some embodiments, user input 748 is a different type of user input, such as a gesture or other action taken by the user. For example, in some embodiments, electronic device 700 is a head-mounted system, and user input 748 includes, for example, the user performing a gesture while wearing electronic device 700 (e.g., an air gesture), pressing a button while wearing electronic device 700, rotating a rotatable input mechanism while wearing electronic device 700, providing a gaze-based gesture (e.g., looking at an object and / or moving his or her gaze in a particular way), and / or any combination of the foregoing.
[0236] exist Figure 7KAt 746, in response to user input 748 (e.g., in response to user input 748 when the user is looking at object 746), the electronic device 700 displays translations 750a-750e. The translations 750a-750e are displayed superimposed on the three-dimensional environment 712. In some embodiments, the objects 744a-744d are viewpoint-locked objects such that even when the user changes the viewpoint of the electronic device 700 (e.g., by moving and / or rotating the electronic device 700), the objects 744a-744d do not move around the display 702; and the translations 750a-750e are environment-locked (or world-locked) objects such that when the user changes the viewpoint of the electronic device 700, the translations 750a-750e move around the display 702 (and / or move away from the display 702) based on how things in the three-dimensional environment 712 move. For example, translation 750a is "locked" to the word "MENU" and moves around display 702 with the word "MENU," and translation 750b is "locked" to the word "GARDEN SALAD" and moves around display 702 with the word "GARDEN SALAD."
[0237] The following references are relative to Figures 7A to 7K The described methods 800 and 900 provide information on Figures 7A to 7K Additional description of .
[0238] Figure 8 is a flow chart of an exemplary method 800 for navigating an extended reality experience according to some embodiments. In some embodiments, the method 800 is performed on a computer system (e.g., Figure 1A , a computer system 101 in (e.g., 700 and / or X700) (e.g., a smartphone, a smartwatch, a tablet, a wearable device, and / or a head-mounted device) that communicates with one or more display generating components (e.g., 702 and / or X702) (e.g., a visual output device, a 3D display, a display having at least a portion that is transparent or translucent onto which an image can be projected (e.g., a see-through display), a projector, a head-up display, and / or a display controller) and one or more input devices (e.g., a touch-sensitive surface (e.g., a touch-sensitive display); a mouse; a keyboard; a remote control; a visual input device (e.g., one or more cameras (e.g., an infrared camera, a depth camera, a visible light camera)); an audio input device; and / or a biometric sensor (e.g., a fingerprint sensor, a facial identification sensor, and / or an iris identification sensor)). In some embodiments, method 800 is performed by storing in a non-transitory (or transient) computer-readable storage medium and executed by one or more processors of a computer system (such as one or more processors 202 of computer system 101) (e.g., Figure 1ASome operations in method 800 may be optionally combined, and / or the order of some operations may be optionally changed.
[0239] In some embodiments, a computer system (e.g., 700 and / or X700) displays (802) via one or more display generation components (e.g., 702 and / or X702) multiple representations of augmented reality experiences (e.g., 720, 721, 724, and / or 726) (e.g., augmented reality user interfaces and / or augmented reality applications) in a three-dimensional environment (e.g., 712) (e.g., a virtual three-dimensional environment, a virtual see-through three-dimensional environment, and / or an optical see-through three-dimensional environment) (e.g., displaying representations of multiple augmented reality experiences superimposed on the three-dimensional environment and / or displaying representations of multiple augmented reality experiences simultaneously with the three-dimensional environment), including: a first representation (804) (e.g., 720, 721, 724, and / or 726) of a first augmented reality experience; and a second representation (806) (e.g., 720, 721, 724, and / or 726) of a second augmented reality experience that is different from the first augmented reality experience, wherein the second representation is different from the first representation. While simultaneously displaying (808) multiple representations of augmented reality experiences (e.g., 720, 721, 724, and / or 726) in a three-dimensional environment (e.g., 712), the computer system receives (810) a first user input (e.g., 727, 729, and / or 740) (e.g., one or more user inputs and / or a first set of user inputs) (e.g., one or more mechanical inputs (e.g., button presses and / or rotations of a physical input mechanism), one or more touch inputs, one or more gestures, one or more mid-air gestures, and / or one or more gaze inputs) via one or more input devices (e.g., 702, 704a-704c, and / or 706). In response to receiving the first user input (812), the computer system stops displaying (814) representations of one or more augmented reality experiences (e.g., 720, 721, 724, and / or 726) of the plurality of augmented reality experiences; and based on determining that the first user input corresponds to a selection of a first representation of the first augmented reality experience (816), the computer system displays (818) the first augmented reality experience (e.g., 714 and / or 742) in the three-dimensional environment (and in some embodiments, does not display the second augmented reality experience) via one or more display generation components (e.g., displays the first augmented reality experience applied to the three-dimensional environment, displays the first augmented reality experience superimposed on the three-dimensional environment, and / or displays the first augmented reality experience simultaneously with the three-dimensional environment).
[0240] In some embodiments, in response to receiving the first user input (e.g., 727, 729, and / or 740), and based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience (e.g., 720, 721, 724, and / or 726), the computer system displays the second augmented reality experience (e.g., 714 and / or 742) in (e.g., applied to and / or simultaneously with) a three-dimensional environment (e.g., 712) via one or more display generation components (and, in some embodiments, without displaying the first augmented reality experience). In some embodiments, the computer system (e.g., 700 and / or X700) is a head-mounted system. In some embodiments, the three-dimensional environment (e.g., 712) is an optically transparent environment (e.g., a physical, real environment) that is visible to the user via a transparent display generating component (e.g., a transparent optical lens display) on which representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726), a first augmented reality experience (e.g., 714 and / or 742), and / or a second augmented reality experience (e.g., 714 and / or 742) are displayed. In some embodiments, the three-dimensional environment (e.g., 712) is a virtual three-dimensional environment displayed by one or more display generating components (e.g., 702). In some embodiments, the three-dimensional environment is a virtual transparent environment (e.g., a virtual transparent environment that is a virtual representation of the user's physical, real-world environment (e.g., as captured by one or more cameras in communication with the computer system)) displayed by one or more display generating components (e.g., 702 and / or X702). Simultaneously displaying representations of multiple augmented reality experiences allows the user to switch between different augmented reality experiences with less user input, thereby reducing the amount of user input required to perform operations. Displaying the first augmented reality experience based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience provides visual feedback to the user regarding the state of the system (e.g., the system has detected the first user input corresponding to the selection of the first representation of the first augmented reality experience), thereby providing improved visual feedback to the user.
[0241] In some embodiments, representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) are displayed on one or more additive light displays (e.g., a see-through display that displays one or more elements while a real-world background is visible to the user behind the displayed elements); and a three-dimensional environment (e.g., 712) is an optical see-through environment (e.g., a physical, real environment) visible to the user through the one or more additive light displays (e.g., behind and / or through the displayed representations of the multiple augmented reality experiences). Simultaneously displaying representations of multiple augmented reality experiences allows the user to switch between different augmented reality experiences with less user input, thereby reducing the amount of user input required to perform operations.
[0242] In some embodiments, in response to receiving a first user input (e.g., 727, 729, and / or 740), and based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience (e.g., 720, 721, 724, and / or 726), the computer system displays a second augmented reality experience (e.g., 714 and / or 742) in a three-dimensional environment (e.g., 712) via one or more display generation components. In some embodiments, displaying a first augmented reality experience (e.g., 714 and / or 742) includes displaying a first set of interactive elements (e.g., objects 716a-716e correspond to augmented reality experience 714, and objects 744a-744d correspond to augmented reality experience 742) (e.g., one or more interactive elements) (e.g., one or more selectable options, selectable buttons, and / or affordances) (in some embodiments, displaying the first set of interactive elements superimposed on a three-dimensional environment); displaying a second augmented reality experience (e.g., 714 and / or 742) includes displaying a second set of interactive elements that is different from the first set of interactive elements (e.g., objects 716a-716e correspond to augmented reality experience 714, and objects 744a-744d correspond to augmented reality experience 742) (e.g., one or more interactive elements) (e.g., one or more selectable options, selectable buttons, and / or affordances) (e.g., not displaying the first set of interactive elements) (in some embodiments, displaying the second set of interactive elements superimposed on a three-dimensional environment); The first representation of the experience (e.g., 720 and / or 726) includes representations of a first set of interactive elements (e.g., representation 720 includes objects 722a-722e representing objects 716a-716e; and representation 726 includes objects 730a-730d representing objects 744a-744d) (in some embodiments, the first set of interactive elements is superimposed on a first representative background (e.g., a visual area and / or a displayed area representing a three-dimensional environment (e.g., representing a see-through environment, an optical see-through environment, and / or a virtual see-through environment)). In some embodiments, the second representation of the second augmented reality experience includes representations of a second set of interactive elements (e.g., representations of objects 716a-716e; and representation 726 includes representations of objects 730a-730d representing objects 744a-744d) (in some embodiments, representations of the second set of interactive elements superimposed on a background of the second representation), the representations of the second set of interactive elements being different from the representations of the first set of interactive elements. Displaying a representation of an augmented reality experience that provides a simplified preview of the augmented reality experience to the user enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0243] In some embodiments, displaying a first augmented reality experience (e.g., 714 and / or 742) includes displaying a first set of interactive elements (e.g., 716a-716e and / or 744a-744d) superimposed on a see-through environment (e.g., 712) (e.g., an optical see-through environment and / or a virtual see-through environment); displaying a second augmented reality experience (e.g., 714 and / or 742) includes displaying a second set of interactive elements (e.g., 716a-716e and / or 744a-744d) superimposed on the see-through environment (e.g., 712); a first representation (e.g., 720 and / or 726) of the first augmented reality experience includes representing the see-through environment (e.g., not the see-through environment and / or placeholder content for a representation of the see-through environment) (e.g., an image, a virtual three-dimensional environment, a solid color image); and / or visual patterns) (in some embodiments, a representation of the first set of interactive elements is overlaid on the first placeholder background content); and a second representation of a second augmented reality experience (e.g., 720 and / or 726) includes second placeholder background content (e.g., 722f, 728e, and / or 730e) representing a see-through environment (e.g., placeholder content that is not a see-through environment and / or is a representation of a see-through environment) (e.g., an image, a virtual three-dimensional environment, a solid color, and / or a visual pattern) (in some embodiments, a representation of the second set of interactive elements is overlaid on the second placeholder background content) (e.g., a second placeholder background content that is different from or the same as the first placeholder background content). Displaying a representation of an augmented reality experience that provides a simplified preview of the augmented reality experience to the user enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0244] In some embodiments, representations of the first group of interactive elements (e.g., 722a-722e, 728a-728d, and / or 730a-730d) are non-interactive (e.g., cannot be individually selected by a user and / or otherwise individually interacted with) (in some embodiments, one or more interactive elements of the first group of interactive elements (e.g., 716a-716e and / or 744a-744d) are selectable to perform a corresponding action (e.g., a first interactive element of the first group of interactive elements is selectable to perform a first action, and a second interactive element of the first group of interactive elements is selectable to perform a second action), and representations of the first group of interactive elements (e.g., 722a-722e, 728a-728d, and / or 730a-730d) are not selectable (e.g., cannot be individually selected) to perform a corresponding action (e.g., the representations of the first group of interactive elements are not selectable to perform a first action, and the representations of the second interactive elements are not selectable to perform a second action). representations are not selectable to perform the second action, and / or the representations of the first and second interactive elements are not individually selectable); and / or the computer system is configured to distinguish between selections of a first interactive element (e.g., 716a-716e and / or 744a-744d) and a second interactive element (e.g., 716a-716e and / or 744a-744d) of the first set of interactive elements (e.g., 716a-716e and / or 744a-744d) , but is not configured to distinguish between selections of representations of the first interactive elements (e.g., 722a-722e, 728a-728d, and / or 730a-730d) and representations of the second interactive elements (e.g., 722a-722e, 728a-728d, and / or 730a-730d); and the representations of the second set of interactive elements (e.g., 722a-722e, 728a-728d, and / or 730a-730d) are non-interactive. Displaying a representation of an augmented reality experience that provides a simplified preview of the augmented reality experience to the user enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0245] In some embodiments, at least a portion of a first representation (e.g., 720, 721, 724, and / or 726) of a first augmented reality experience is displayed in a first color corresponding to the first augmented reality experience (e.g., 714 and / or 742) (e.g., corresponding uniquely to the first augmented reality experience and / or corresponding to the first augmented reality experience and not to the second augmented reality experience); and at least a portion of a second representation (e.g., 720, 721, 724, and / or 726) of a second augmented reality experience is displayed in a second color corresponding to the second augmented reality experience (e.g., 714 and / or 742) (e.g., corresponding uniquely to the second augmented reality experience and / or corresponding to the second augmented reality experience and not to the first augmented reality experience), where the second color is different from the first color. In some embodiments, the first representation of the first augmented reality experience does not include the second color, and the second representation of the second augmented reality experience does not include the first color. Displaying representations of augmented reality experiences in different colors that uniquely correspond to different augmented reality experiences allows a user to more easily select a particular augmented reality experience, which enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0246] In some embodiments, the first representation of the first augmented reality experience (e.g., 720, 721, 724, and / or 726) includes a first identifier (e.g., 722a, 728a, and / or 730a) (e.g., a first icon, a first set of text (e.g., a name and / or other text identifier), and / or a first color) that corresponds to the first augmented reality experience (e.g., uniquely corresponds to the first augmented reality experience and / or corresponds to the first augmented reality experience and not to the second augmented reality experience or other augmented reality experiences available via the computer system); the second representation of the second augmented reality experience (e.g., 720, 721, 724, and / or 726) includes a first identifier that is different from the first identifier and corresponds to the second augmented reality experience (e.g., uniquely corresponds to the first augmented reality experience and / or corresponds to the first augmented reality experience and not to the second augmented reality experience or other augmented reality experiences available via the computer system). In some embodiments, the present invention provides a second augmented reality experience and / or a second identifier (e.g., 722a, 728, and / or 730a) corresponding to the second augmented reality experience and / or corresponding to the second augmented reality experience and not the first augmented reality experience (e.g., a second icon, a second set of text (e.g., a name and / or other textual identifier), and / or a second color); displaying the first augmented reality experience (e.g., 714 and / or 742) includes displaying the first identifier (e.g., 716a and / or 744a) as part of the first augmented reality experience (e.g., without displaying the second identifier); and displaying the second augmented reality experience (e.g., 714 and / or 742) includes displaying the second identifier (e.g., 716a and / or 744a) as part of the second augmented reality experience (e.g., without displaying the first identifier). Displaying representations of augmented reality experiences with different identifiers that uniquely correspond to different augmented reality experiences allows a user to more easily select a particular augmented reality experience, which enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0247] In some embodiments, in response to receiving a first input (e.g., 718, 727, 729, and / or 740): based on determining that the first user input corresponds to a selection of a first representation of a first augmented reality experience (e.g., 720, 721, 724, and / or 726): before displaying the first augmented reality experience (e.g., 714 and / or 742), the computer system displays, via one or more display generation components, a first animation in which the first representation of the first augmented reality experience moves toward a viewpoint of a user of the computer system (e.g., Figures 7H to 7J, representation 726 moves toward the user's viewpoint until augmented reality experience 742 is displayed) (e.g., wherein the first representation of the first augmented reality experience becomes larger and / or appears to move closer to the user's viewpoint). In some embodiments, in response to receiving the first user input, and based on determining that the first user input corresponds to a selection of a second representation of the second augmented reality experience: before displaying the second augmented reality experience, the computer system displays, via one or more display generation components, a second animation in which the second representation of the second augmented reality experience moves toward the viewpoint of the user of the computer system. Displaying the animation in which the first representation of the first augmented reality experience moves toward the user's viewpoint provides visual feedback to the user about the state of the system (e.g., the system is transitioning to the first augmented reality experience), thereby providing improved visual feedback to the user.
[0248] In some embodiments, the first representation (e.g., 720, 721, 724, and / or 726) of the first augmented reality experience includes a first boundary surrounding (e.g., partially and / or completely surrounding) the representation of the first set of interactive elements (e.g., 722a-722e, 728a-728d, and / or 730a-730d); the second representation (e.g., 720, 721, 724, and / or 726) of the second augmented reality experience includes a second boundary (e.g., different from and / or separate from the first boundary) surrounding (e.g., partially and / or completely surrounding) the representation of the second set of interactive elements (e.g., 722a-722e, 728a-728d, and / or 730a-730d); and displaying the first animation includes displaying the first boundary moving toward the viewpoint of the user of the computer system until the first boundary is no longer displayed (e.g., Figures 7H to 7J, moving the border of representation 726 toward the user's viewpoint until augmented reality experience 742 is displayed) (e.g., until the first border moves out of the display area of the computer system and / or out of the portion of the display area of the computer system that is visible to the user) (e.g., increasing the size of the first border until the first border is no longer displayed by the computer system and / or is outside the portion of the display area of the computer system that is visible to the user). In some embodiments, in response to receiving the first user input, and based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience: before displaying the second augmented reality experience, the computer system, via one or more display generation components, displays a second animation in which the second representation of the second augmented reality experience moves toward the viewpoint of the user of the computer system, wherein displaying the second animation includes displaying the second border moving toward the viewpoint of the user of the computer system until the second border is no longer displayed. Displaying the animation in which the first representation of the first augmented reality experience (including the border of the first representation) moves toward the user's viewpoint provides visual feedback to the user about the state of the system (e.g., the system is transitioning to the first augmented reality experience), thereby providing improved visual feedback to the user.
[0249] In some embodiments, in response to receiving a first input (e.g., 718, 727, 729, and / or 740): based on determining that the first user input corresponds to a selection of a first representation of a first augmented reality experience: the computer system displays, via one or more display generation components, representations of a first set of interactive elements (e.g., 730a-730d) and a cross-fade of the first set of interactive elements (e.g., 744a-744d) (e.g., during display of the first animation and / or after display of the first animation). In some embodiments, in response to receiving a first input: based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience: the computer system displays, via one or more display generation components, representations of a second set of interactive elements and a cross-fade of the second set of interactive elements. Displaying the cross-fade of the representations of the first set of interactive elements and the first set of interactive elements provides visual feedback to the user regarding the state of the system (e.g., that the system is transitioning to the first augmented reality experience), thereby providing improved visual feedback to the user.
[0250] In some embodiments, in response to receiving a first input (e.g., 718, 727, 729, and / or 740): based on determining that the first user input corresponds to a selection of a first representation of a first augmented reality experience: the computer system stops displaying representations of the first set of interactive elements (e.g., 730a-7330d); and displays the first set of interactive elements (e.g., 744a-744d) via one or more display generation components. Displaying the first set of interactive elements instead of the representations of the first set of interactive elements provides visual feedback to the user about the state of the system (e.g., that the system is transitioning to the first augmented reality experience), thereby providing improved visual feedback to the user.
[0251] In some embodiments, prior to receiving the first user input, and while simultaneously displaying representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) in a three-dimensional environment (e.g., 712), the computer system displays, via one or more display generation components, a first representation of a first augmented reality experience (e.g., 720, 721, 724, and / or 726) at a first display location (e.g., Figure 7D1 720) (and in some embodiments, simultaneously displaying a second representation of a second augmented reality experience at a second display location different from the first display location (e.g., Figure 7D1 724) in the representation 724) (in some embodiments, the first display position represents the currently selected object and / or the currently focused object). While the first representation of the first augmented reality experience is displayed at the first display position, the computer system receives, via one or more input devices, a second user input (e.g., 727 and / or 729) (e.g., one or more user inputs and / or a first group of user inputs) (e.g., one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs) corresponding to a request to navigate from the first representation of the first augmented reality experience (e.g., 720, 721, 724, and / or 726) to the second representation of the second augmented reality experience (e.g., 720, 721, 724, and / or 726). In response to receiving the second user input (e.g., 727 and / or 729), the computer system stops displaying the first representation of the first augmented reality experience at the first display position (e.g., in Figures 7D1 to 7E In response to user input 727, the electronic device 700 stops displaying the representation 720 at the front position of the stack; and Figures 7E to 7F In response to user input 729, electronic device 700 stops display of representation 724 at the front position of the stack (and in some embodiments, while maintaining display of at least a portion of the first representation of the first augmented reality experience); and displays, via one or more display generation components, a second representation of the second augmented reality experience at the first display position (e.g., at Figure 7E , representation 724 is shown at the front of the stack, and in Figure 7F , representation 726 is displayed at the front position of the stack). Displaying navigation from the first representation of the first augmented reality experience to the second representation of the second augmented reality experience in response to the second user input provides visual feedback to the user regarding the state of the system (e.g., that the system has detected the second user input), thereby providing improved visual feedback to the user.
[0252] In some embodiments, simultaneously displaying representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) includes displaying representations of the multiple augmented reality experiences in a stack, where a first representation of a first augmented reality experience is stacked on top of a second representation of a second augmented reality experience (e.g., the first representation of the first augmented reality experience is on top of and / or partially obscures the second representation of the second augmented reality experience). Displaying representations of augmented reality experiences in a stack that a user can navigate allows the user to more easily select a particular augmented reality experience, which enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0253] In some embodiments, before receiving the first user input, and while representations of multiple augmented reality experiences are displayed simultaneously in the three-dimensional environment (including simultaneously displaying a first representation of a first augmented reality experience and a second representation of a second augmented reality experience), the computer system receives a third user input (e.g., 727 and / or 729) (e.g., one or more user inputs and / or a first group of user inputs) (e.g., one or more touch inputs, one or more gestures, one or more mid-air gestures, and / or one or more gaze inputs) via one or more input devices corresponding to a request to navigate among representations of the multiple augmented reality experiences. In response to receiving the third user input, the computer system stops display of the first representation of the first augmented reality experience while maintaining display of the second representation of the second augmented reality experience (e.g., in Figures 7D1 to 7E In response to user input 727, electronic device 700 stops display of representation 720 while maintaining display of representations 724 and / or 726; and / or Figures 7E to 7F , in response to user input 729, electronic device 700 ceases display of representation 724 while maintaining display of representation 726. Displaying representations of augmented reality experiences in a stack that the user can navigate allows the user to more easily select a particular augmented reality experience, which enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0254] In some embodiments, determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes determining that the first user input is a selection input, the selection input comprising: a gaze input (e.g., in the Figure 7H , gaze indication 710 indicates that the user is looking at representation 726) (e.g., user gaze toward a selectable object; user gaze toward a corresponding representation in representations of multiple augmented reality experiences and / or user gaze corresponding to and / or identifying a particular augmented reality experience); and a hardware press input (e.g., 740) detected when the gaze input is toward the first representation of the first augmented reality experience (e.g., a press of a hardware button and / or a press of a pressable input mechanism (e.g., a rotatable and pressable input mechanism)) (e.g., a hardware press input occurring simultaneously with the gaze input). In some embodiments, in response to receiving the first user input, and based on determining that the first user input is not a selection input (e.g., based on determining that the first user input does not include a gaze input toward the first representation of the first augmented reality experience and / or a hardware press input when the gaze input is toward the first representation of the first augmented reality experience), the computer system forgoes displaying the first augmented reality experience. In some embodiments, in response to receiving a first user input, and based on determining that the first user input is not a selection input, the computer system abandons stopping the display of the representation of one or more augmented reality experiences in the multiple augmented reality experiences (e.g., the computer system maintains the display of the representation of one or more augmented reality experiences in the multiple augmented reality experiences). In some embodiments, stopping the display of the representation of one or more augmented reality experiences in the multiple augmented reality experiences is performed based on determining that the first user input is a selection input. In some embodiments, the first user input includes: a first gaze input (e.g., a user gaze toward a corresponding representation in the multiple augmented reality experiences and / or a user gaze corresponding to and / or identifying a specific augmented reality experience); and a hardware press input (e.g., a press on a hardware button and / or a press on a pressable input mechanism (e.g., a rotatable and pressable input mechanism)). In some embodiments, the first user input includes a first gaze input and a hardware press input that occur simultaneously (e.g., a hardware press input when the user gazes at a specific object and / or a hardware press input when the user gazes at a corresponding representation in the multiple augmented reality experiences). Allowing a user to select a particular augmented reality experience using gaze and hardware press inputs enhances the operability of a computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0255] In some embodiments, determining that the first user input corresponds to a selection of a first representation of a first augmented reality experience includes determining that the first user input is a selection input, the selection input including: voice input indicating a user request to select a selectable object (e.g., voice input identifying a particular selectable object; and / or voice input identifying a corresponding augmented reality experience in a plurality of augmented reality experiences and / or a corresponding representation in a plurality of representations of augmented reality experiences (e.g., in Figure 7D1 in (and / or in Figure 7B ), the user states “apply translation extended reality experience”, and in response to the user voice input, the electronic device 700 and / or HMD X700 displays the translation extended reality experience, such as Figures 7I to 7J In some embodiments, in response to receiving the first user input, and based on determining that the first user input is not a selection input (e.g., based on determining that the first user input does not include a voice input indicating a user request to select a selectable object), the computer system abandons displaying the first augmented reality experience (e.g., the computer system continues to display the Figure 7D1 In some embodiments, in response to receiving a first user input, and based on determining that the first user input is not a selection input, the computer system abandons stopping the display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences (e.g., the computer system maintains the display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences). In some embodiments, stopping the display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences is performed based on determining that the first user input is a selection input. In some embodiments, the first user input includes: a first voice input (e.g., a voice input identifying a corresponding augmented reality experience in the multiple augmented reality experiences and / or a corresponding representation in the representations of the multiple augmented reality experiences). In some embodiments, the first user input includes a first voice input and a first gaze input (e.g., a user gaze toward a corresponding representation in the representations of the multiple augmented reality experiences) (e.g., in Figure 7H , stating "show this extended reality experience" while looking at representation 726). In some embodiments, the first user input includes a first voice input that occurs simultaneously with the first gaze input (e.g., voice input when the user gazes at a particular object and / or voice input when the user gazes at a respective representation of a plurality of augmented reality experiences). Allowing a user to select a particular augmented reality experience using voice input enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0256] In some embodiments, determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes determining that the first user input is a selection input, the selection input comprising: a gaze input (e.g., 726) toward the first representation of the first augmented reality experience; Figure 7H 720 in the present invention) (e.g., a user gaze toward a selectable object; a user gaze toward a corresponding representation among representations of a plurality of augmented reality experiences and / or a user gaze corresponding to and / or identifying a particular augmented reality experience, the user gaze satisfying a first set of gaze duration criteria (e.g., a user gaze toward a selectable object and maintained on the selectable object for a threshold duration (e.g., without interruption and / or with less than a threshold amount of interruption); and / or a user gaze toward a corresponding representation among representations of a plurality of augmented reality experiences and maintained on the corresponding representation for a threshold duration (e.g., without interruption and / or with less than a threshold amount of interruption)).
[0257] In some embodiments, in response to receiving the first user input, and based on determining that the first user input is not a selection input (e.g., based on determining that the first user input does not include a gaze input toward a first representation of the first augmented reality experience that satisfies a first set of gaze duration criteria, because the gaze input is not toward the first representation of the first augmented reality experience, or because the gaze input moves away from the first representation of the first augmented reality experience before the first set of gaze duration criteria have been met), the computer system forgoes displaying the first augmented reality experience (e.g., in some embodiments, Figure 7H In FIG. 7 , if the user maintains his or her gaze on the representation 726 for a threshold duration, the electronic device 700 and / or HMDX 700 displays the following Figures 7I to 7J The translated extended reality experience 742 is shown, but if the user does not maintain his or her gaze on the representation 726 for a threshold duration, the electronic device 700 remains Figure 7H In some embodiments, in response to receiving the first user input, and based on determining that the first user input is not a selection input, the computer system abandons stopping the display of the representations of one or more augmented reality experiences in the multiple augmented reality experiences (e.g., the computer system maintains the display of the representations of the one or more augmented reality experiences in the multiple augmented reality experiences) (e.g., maintains Figure 7H In some embodiments, stopping display of representations 721, 726 in the plurality of augmented reality experiences is performed based on determining that the first user input is a selection input. In some embodiments, the first user input includes a first gaze input that meets a first set of gaze duration criteria (e.g., Figure 7HIn some embodiments, determining that the first user input corresponds to a selection of a first representation of the first augmented reality experience includes the user having gazed at the first representation of the first augmented reality experience for a threshold duration (e.g., without interruption and / or with less than a threshold amount of interruption) (e.g., at 710 in the example embodiment). In some embodiments, determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes the user having gazed at the first representation of the first augmented reality experience for a threshold duration (e.g., without interruption and / or with less than a threshold amount of interruption) (e.g., at Figure 7H , the user has gazed at representation 726 for a threshold duration). Allowing a user to utilize gaze and dwell input to select a particular augmented reality experience enhances the operability of a computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0258] In some embodiments, when simultaneously displaying representations of multiple augmented reality experiences, the computer system displays, via one or more display generation components, one or more setting controls (e.g., 736a-736h), including a first setting control corresponding to a first setting of the computer system (in some embodiments, the computer system simultaneously displays a second setting control corresponding to a second setting of the computer system that is different from the first setting). While displaying the one or more setting controls, the computer system receives, via one or more input devices, a first setting input corresponding to the first setting of the computer system (e.g., Figure 7GIn response to receiving the first setting input, the computer system modifies the first setting from the first value to a second value different from the first value. When representations of multiple augmented reality experiences are displayed simultaneously and when the first setting is set to the second value, the computer system receives a third user input (e.g., 740) (e.g., one or more user inputs and / or a third group of user inputs) (e.g., one or more touch inputs, one or more gestures, one or more mid-air gestures, and / or one or more gaze inputs) via one or more input devices. In response to receiving a third user input: based on determining that the third user input corresponds to a selection of a first representation of a first augmented reality experience (e.g., 720, 724, and / or 726), the computer system displays the first augmented reality experience (e.g., 714 and / or 742) in the three-dimensional environment via one or more display generation components while maintaining the first setting at a second value; and based on determining that the first user input corresponds to a selection of a second representation of a second augmented reality experience (e.g., 720, 724, and / or 726), the computer system displays the second augmented reality experience (e.g., 714 and / or 742) in the three-dimensional environment via one or more display generation components while maintaining the first setting at the second value. Displaying one or more setting controls to modify one or more device settings and maintaining these settings between different augmented reality experiences allows the user to modify the device settings with less user input, thereby reducing the amount of user input required to perform an operation.
[0259] In some embodiments, the first setting is a see-through shading setting (e.g., option 736h) (e.g., a setting that controls how much masking and / or shading is applied to a three-dimensional environment (e.g., a see-through background, an optical see-through background, and / or a virtual see-through background); the first value corresponds to a first amount of shading applied to the three-dimensional environment (e.g., a first amount of masking and / or shading; and / or a first brightness); and the second value corresponds to a second amount of shading applied to the three-dimensional environment that is different from the first amount of shading (e.g., a second amount of masking and / or shading; and / or a second brightness). Displaying a setting control to modify the see-through shading, and maintaining the see-through shading setting between different augmented reality experiences allows the user to modify the see-through shading setting with less user input, thereby reducing the amount of user input required to perform the operation.
[0260] In some embodiments, the first setting is a volume setting (e.g., option 736g); the first value corresponds to a first volume; and the second value corresponds to a second volume different from the first volume. Displaying a setting control to modify the volume and persisting the volume setting between different augmented reality experiences allows the user to modify the volume setting with less user input, thereby reducing the amount of user input required to perform the operation.
[0261] In some embodiments, when representations of multiple augmented reality experiences and one or more settings controls are displayed simultaneously, the computer system displays, via one or more display generation components, device status information (e.g., in the display) indicating the status of one or more characteristics of the computer system (e.g., Wi-Fi network name, Wi-Fi signal strength, computer system battery level, computer system location tracking indicator, microphone recording indicator, camera recording indicator, and / or volume slider). Figures 7D1 to 7H Displaying device status information provides visual feedback to the user regarding the status of the system (e.g., information regarding the status of one or more characteristics of the computer system), thereby providing improved visual feedback to the user.
[0262] In some embodiments, representations of the multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) are viewpoint-locked objects that stay in corresponding areas of the user's field of view of the computer system as the user's viewpoint shifts relative to the three-dimensional environment (e.g., representations 720, 721, 724, and / or 726 do not move as the user's viewpoint shifts and the background three-dimensional environment 712 moves). Displaying the representations of the multiple augmented reality experiences as viewpoint-locked objects enhances the operability of the computer system by keeping the representations of the multiple augmented reality experiences within the user's line of sight, thereby helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0263] In some embodiments, simultaneously displaying representations of multiple augmented reality experiences includes simultaneously displaying representations of the multiple augmented reality experiences in a first orientation in which the representations of the multiple augmented reality experiences are aligned with gravity (e.g., in a Figure 7D1, representations 720, 721, 724, and / or 726 are displayed in an orientation such that a bottom surface of representations 720, 721, 724, and / or 726 faces the ground (e.g., each representation has a bottom portion and a top portion, and the bottom portion is displayed closer to the ground and / or the center of the earth than the top portion). In some embodiments, when representations of multiple augmented reality experiences are displayed simultaneously, the computer system detects a change in the orientation of the user's viewpoint (e.g., a rotation of the electronic device 700, which, for example, causes representations 720, 721, 724, and / or 726 to no longer be aligned with gravity (e.g., the bottom of representations 720, 721, 724, and / or 726 is no longer facing the ground)) (e.g., detecting rotation and / or movement of the user's head and / or detecting rotation and / or movement of headphones and / or other wearable devices (e.g., wearable devices worn on the user's head)). In response to detecting a change in the orientation of the user's viewpoint: the computer system rotates representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) from a first orientation to a second orientation (e.g., a second orientation different from the first orientation) based on the change in the orientation of the user's viewpoint to continue to align the representations of the multiple augmented reality experiences with gravity (e.g., displaying the representations of the multiple augmented reality experiences in a manner such that the representations of the multiple augmented reality experiences remain aligned with gravity (e.g., each representation has a bottom portion and a top portion, and even when the user moves and / or rotates his or her field of view, the bottom portion remains closer to the ground and / or the center of the earth than the top portion)). In some embodiments, the representations of the multiple augmented reality experiences are aligned with gravity (e.g., displaying the representations of the multiple augmented reality experiences in a manner such that the representations of the multiple augmented reality experiences remain aligned with gravity (e.g., each representation has a bottom portion and a top portion, and even when the user moves and / or rotates his or her field of view, the bottom portion remains closer to the ground and / or the center of the earth than the top portion)). In some embodiments, when the computer system detects a rotation of the computer system, the computer system rotates the representations of the multiple augmented reality experiences based on the rotation of the computer system such that bottom portions of the representations remain closer to the ground and / or the center of the earth than top portions of the representations. Displaying the representations of the multiple augmented reality experiences as gravity-aligned, viewpoint-locked objects enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system by keeping the representations of the multiple augmented reality experiences within the user's line of sight and in a consistent alignment even as the user moves and / or the computer system moves.
[0264] In some embodiments, rotating representations of multiple augmented reality experiences (e.g., 720, 721, 724, and / or 726) from a first orientation to a second orientation includes: at a first time after detecting a change in the orientation of the user's viewpoint, displaying, via one or more display generation components, representations of the multiple augmented reality experiences in the first orientation, wherein at the first time, due at least in part to the change in the orientation of the user's viewpoint, the representations of the multiple augmented reality experiences are not aligned with gravity (e.g., displaying representations 720, 721, 724, and / or 726 with bottom edges of the representations not facing the ground); and at a second time after the first time, displaying, via the one or more display generation components, representations of the multiple augmented reality experiences in the second orientation to align the representations of the multiple augmented reality experiences with gravity (e.g., as shown in FIG. 2 ). Figure 7D1 Representations 720, 721, 724, and / or 726 shown). In some embodiments, the computer system displays a gradual rotation of the representations of the multiple augmented reality experiences from a first orientation to a second orientation over time. In some embodiments, at a third time after the first time and before the second time, the computer system displays the representations of the multiple augmented reality experiences at a third orientation different from the first orientation and the second orientation, via one or more display generation components, where the third orientation is between the first orientation and the second orientation (e.g., at an angle between the angle of the first orientation and the angle of the second orientation). In some embodiments, the representations of the multiple augmented reality experiences exhibit inertial following behavior (e.g., behavior that reduces or delays the movement of the representations of the multiple augmented reality experiences relative to detected physical movement of the user (e.g., relative to detected physical movement of the user's head) and / or relative to detected physical movement of the computer system). Displaying the representations of the multiple augmented reality experiences as viewpoint-locked objects that exhibit inertial following behavior provides visual feedback to the user about the state of the system (e.g., the system intentionally moves the representations of the multiple augmented reality experiences when the user's head moves), thereby providing improved visual feedback to the user.
[0265] In some embodiments, displaying the first augmented reality experience (e.g., 742) includes simultaneously displaying a first group of objects (e.g., 744a-744d, 750a-750e), including a first object (e.g., 744a-744d) and a second object (e.g., 750a-750e), and wherein: the first object is a viewpoint-locked object (e.g., 744a-744d are viewpoint-locked objects); and the second object is an environment-locked object (e.g., 750a-750e are environment-locked objects). In some embodiments, the second augmented reality experience includes a second group of objects, including a third object and a fourth object, wherein the third object is a viewpoint-locked object and the fourth object is an environment-locked object. Displaying certain objects in the AR experience as viewpoint-locked objects and displaying other objects as environment-locked objects enhances the operability of the computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system.
[0266] In some embodiments, the computer system displays a first augmented reality experience (e.g., 712) in a three-dimensional environment (e.g., 712) via one or more display generation components. Figure 7B 714 in ). When the first augmented reality experience is displayed (e.g., Figure 7B In response to receiving the first voice input, the computer system stops display of the first augmented reality experience (e.g., stops display of experience 714); and displays the second augmented reality experience in the three-dimensional environment (e.g., 712) via the one or more input devices (e.g., Figure 7J 742 of ). Allowing a user to use voice input to switch between different augmented reality experiences enhances the operability of a computer system by helping the user provide appropriate input and reducing user errors when operating / interacting with the computer system. Allowing a user to use voice input to switch between different augmented reality experiences allows the user to switch between different augmented reality experiences with less user input, thereby reducing the amount of user input required to perform an operation.
[0267] In some embodiments, when the computer system is in a sleep state (e.g., Figure 7A) (e.g., an off state, a locked state, and / or a sleep state), the computer system receives, via one or more input devices, a first wake-up input (e.g., 708) (e.g., one or more user inputs and / or a first group of user inputs) (e.g., one or more mechanical inputs (e.g., button presses and / or rotations of a physical input mechanism), one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs) corresponding to a request to transition the computer system from the sleep state to the wake state. In response to receiving the first wake-up input (and in some embodiments, based on determining that the first wake-up input satisfies a first set of wake-up criteria (e.g., unlock criteria, user authentication criteria, and / or biometric authentication criteria)), the computer system displays a first augmented reality experience (e.g., 714 and / or 742) via one or more display generation components (e.g., without displaying a second augmented reality experience and / or representations of multiple augmented reality experiences). In some embodiments, the first augmented reality experience represents a default augmented reality experience displayed when the computer system transitions from the sleep state to the wake state. Automatically displaying the first augmented reality experience when the computer system transitions from a dormant state to an awake state allows the user to access the first augmented reality experience with less user input, thereby reducing the amount of user input required to perform an operation.
[0268] In some embodiments, when the computer system is in a sleep state (e.g., Figure 7A ) (e.g., an off state, a locked state, and / or a sleep state), the computer system receives, via one or more input devices, a first wake-up input (e.g., 708) (e.g., one or more user inputs and / or a first set of user inputs) (e.g., one or more mechanical inputs (e.g., button presses and / or rotations of a physical input mechanism), one or more touch inputs, one or more gestures, one or more air gestures, and / or one or more gaze inputs) corresponding to a request to transition the computer system from the sleep state to the wake-up state. In response to receiving the first wake-up input (and in some embodiments, based on determining that the first wake-up input satisfies a first set of wake-up criteria (e.g., unlock criteria, user authentication criteria, and / or biometric authentication criteria)), the computer system displays, via one or more display generation components, representations of multiple augmented reality experiences (e.g., Figure 7D1720, 721, 724, and / or 726 in the figure) (e.g., the first augmented reality experience and / or the second augmented reality experience are not displayed). In some embodiments, the AR experience switcher user interface including representations of multiple augmented reality experiences represents a default user interface that is displayed when the computer system transitions from a sleep state to a wake state. Automatically displaying representations of multiple augmented reality experiences when the computer system transitions from a sleep state to a wake state allows the user to access representations of multiple augmented reality experiences with less user input, thereby reducing the amount of user input required to perform operations.
[0269] In some embodiments, the plurality of augmented reality experiences includes one or more of the following: a camera augmented reality experience (e.g., 714) (e.g., an augmented reality experience including a shutter button that can be selected to capture more photos and / or videos (e.g., one or more photos and / or videos of the environment surrounding the computer system) using one or more cameras of the computer system) (e.g., an augmented reality experience in which a user is able to capture one or more photos and / or videos of the user's surroundings); a translation augmented reality experience (e.g., 742) (e.g., an augmented reality experience including a shutter button that can be selected to translate text captured by one or more cameras of the computer system (e.g., translate text in the environment surrounding the computer system)); an augmented reality experience that includes one or more selectable options that can be selected to output audio content (e.g., music and / or other audio content) (e.g., an augmented reality experience in which a user is able to listen to music and / or other audio content); a navigation augmented reality experience (e.g., displaying navigation instructions to a geographic location); an augmented reality experience (e.g., an augmented reality experience in which a user can receive navigation instructions for a geographic location); a photo augmented reality experience (e.g., an augmented reality experience that displays one or more selectable objects and / or user interfaces for navigating through photo and / or video content in a media library) (e.g., an augmented reality experience in which a user can navigate through and / or view photo and / or video content in a media library); a video messaging augmented reality experience (e.g., an augmented reality experience that includes one or more selectable options that can be selected to initiate and / or terminate a video conference and / or video call with one or more contacts) (e.g., an augmented reality experience in which a user can participate in a video chat with one or more contacts); Simultaneously displaying representations of multiple augmented reality experiences allows a user to switch between different augmented reality experiences with less user input, thereby reducing the amount of user input required to perform operations.
[0270] In some embodiments, aspects / operations of methods 800, 900, 1100, 1300, and / or 1500 may be interchanged, replaced, and / or added between these methods. For example, in some embodiments, the augmented reality experience in method 800 is the extended reality experience in methods 900 and / or 1100. For another example, in some embodiments, the virtual content in method 1500 includes virtual content related to the augmented reality experience in method 800 and / or the extended reality experience in methods 900 and / or 1100. For another example, in some embodiments, the computer system in method 1300 is the computer system in any of methods 800, 900, 1100, and / or 1500. For the sake of brevity, these details are not repeated here.
[0271] Figure 9 is a flow chart of an exemplary method 900 for navigating an extended reality experience according to some embodiments. In some embodiments, the method 900 is performed on a computer system (e.g., Figure 1A , a computer system 101; 700; and / or HMDX 700 in a computer system (e.g., a smartphone, a smartwatch, a tablet, a wearable device, and / or a head-mounted device) that communicates with one or more display generating components (e.g., a visual output device, a 3D display, a display having at least a portion that is transparent or translucent onto which an image can be projected (e.g., a see-through display), a projector, a heads-up display, and / or a display controller) and one or more input devices (e.g., a touch-sensitive surface (e.g., a touch-sensitive display); a mouse; a keyboard; a remote control; a visual input device (e.g., one or more cameras (e.g., an infrared camera, a depth camera, a visible light camera)); an audio input device; and / or a biometric sensor (e.g., a fingerprint sensor, a facial identification sensor, and / or an iris identification sensor)). In some embodiments, method 900 is performed by storing in a non-transitory (or transitory) computer-readable storage medium and executed by one or more processors of a computer system (such as one or more processors 202 of computer system 101) (e.g., Figure 1A Some operations in method 900 may be optionally combined, and / or the order of some operations may be optionally changed.
[0272] In some embodiments, the computer system (e.g., 700 and / or HMD X700) receives (900) a first sequence of one or more user inputs (e.g., 718, 727, 729, and / or 740) via a first physical control (e.g., 704a-704c and / or X704a-X704c) (e.g., a physical button, a rotatable input mechanism, a pressable input mechanism, and / or a rotatable and pressable input mechanism) (e.g., a first physical control of one or more input devices) (e.g., one or more presses of a pressable input mechanism, one or more rotations of a rotatable input mechanism, and / or one or more presses of and / or rotations of a rotatable and pressable input mechanism). In response to receiving a first sequence of one or more user inputs (904): based on determining that the first sequence of one or more user inputs has a first magnitude (906) (e.g., an amount of movement, a speed of movement, and / or a duration of the input / movement), the computer system displays (908) a first augmented reality experience (e.g., 714 and / or 742) in a three-dimensional environment (e.g., 712) via one or more display generation components (e.g., 702 and / or X702) (e.g., based on Figure 7D1 , if the first sequence of one or more user inputs includes only one press of button 704a and / or button X 704a prior to user input 740, the computer system 700 and / or HMD X 700 displays an extended reality experience corresponding to representation 724 (e.g., a music extended reality experience); and if the first sequence of one or more user inputs includes two presses of button 704a and / or button X 704a prior to user input 740 (e.g., Figures 7D1 to 7J ), the computer system 700 and / or the HMDX 700 displays an extended reality experience 742 (e.g., a translated extended reality experience)) (e.g., a first extended reality user interface and / or a first extended reality application) (e.g., displays the first extended reality experience applied to the three-dimensional environment, displays the first extended reality experience overlaid on the three-dimensional environment, and / or displays the first extended reality experience simultaneously with the three-dimensional environment); and based on determining that the first sequence of one or more user inputs has a second magnitude (910) different from the first magnitude (e.g., an amount of movement, a speed of movement, and / or a duration of the input / movement), the computer system displays (912) in the three-dimensional environment via the one or more display generation components a second extended reality experience (e.g., 714 and / or 742) different from the first extended reality experience (e.g., based on Figure 7D1, if the first sequence of one or more user inputs includes only one press of button 704a and / or button X 704a prior to user input 740, the computer system 700 and / or HMD X 700 displays an extended reality experience corresponding to representation 724 (e.g., a music extended reality experience); and if the first sequence of one or more user inputs includes two presses of button 704a and / or button X 704a prior to user input 740 (e.g., Figures 7D1 to 7J ), the computer system 700 and / or HMD X700 displays the extended reality experience 742 (e.g., translates the extended reality experience).
[0273] In some embodiments, in response to receiving a first sequence of one or more user inputs: based on determining that the first user input has a second magnitude and a first direction (e.g., button 704a and / or X704a corresponds to the first direction and button 704b and / or X704b corresponds to the second direction), the computer system displays, via one or more display generation components, a third extended reality experience in the three-dimensional environment that is different from the first extended re...
Claims
1. A method, comprising: At a computer system in communication with one or more display generating components and one or more input devices: Simultaneously displaying, via the one or more display generation components, representations of a plurality of augmented reality experiences in a three-dimensional environment, the representations comprising: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receiving a first user input via the one or more input devices while the representations of the plurality of augmented reality experiences are simultaneously displayed in the three-dimensional environment; and In response to receiving the first user input: ceasing display of the representation of one or more augmented reality experiences in the plurality of augmented reality experiences; and Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience, displaying the first augmented reality experience in the three-dimensional environment via the one or more display generation components.
2. The method according to claim 1, wherein: The representations of the plurality of augmented reality experiences are displayed on one or more additive light displays; and The three-dimensional environment is an optically see-through environment that is visible to a user through the one or more added-light displays.
3. The method according to any one of claims 1 to 2, further comprising: In response to receiving the first user input: displaying, via the one or more display generating components, the second augmented reality experience in the three-dimensional environment based on determining that the first user input corresponds to a selection of the second representation of the second augmented reality experience, in: Displaying the first augmented reality experience includes displaying a first set of interactive elements; Displaying the second augmented reality experience includes displaying a second set of interactive elements different from the first set of interactive elements; The first representation of the first augmented reality experience includes a representation of the first set of interactive elements; and The second representation of the second augmented reality experience includes a representation of the second set of interactive elements that is different from the representation of the first set of interactive elements.
4. The method according to claim 3, wherein: Displaying the first augmented reality experience includes displaying the first set of interactive elements superimposed on the transparent environment; Displaying the second augmented reality experience includes displaying the second set of interactive elements superimposed on the pass-through environment; The first representation of the first augmented reality experience includes first placeholder background content representing the pass-through environment; and The second representation of the second augmented reality experience includes second placeholder background content representing the pass-through environment.
5. The method according to any one of claims 3 to 4, wherein: said representation of said first set of interactive elements is non-interactive; and The representation of the second set of interactive elements is non-interactive.
6. The method according to any one of claims 3 to 5, wherein: At least a portion of the first representation of the first augmented reality experience is displayed in a first color corresponding to the first augmented reality experience; and At least a portion of the second representation of the second augmented reality experience is displayed in a second color corresponding to the second augmented reality experience, wherein the second color is different from the first color.
7. The method according to any one of claims 3 to 6, wherein: the first representation of the first augmented reality experience comprising a first identifier corresponding to the first augmented reality experience; the second representation of the second augmented reality experience comprising a second identifier different from the first identifier and corresponding to the second augmented reality experience; Displaying the first augmented reality experience includes displaying the first identifier as part of the first augmented reality experience; and Displaying the second augmented reality experience includes displaying the second identifier as part of the second augmented reality experience.
8. The method according to any one of claims 3 to 7, further comprising: In response to receiving the first input: Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience: Prior to displaying the first augmented reality experience, a first animation is displayed via the one or more display generation components in which the first representation of the first augmented reality experience moves toward a viewpoint of a user of the computer system.
9. The method according to claim 8, wherein: the first representation of the first augmented reality experience comprising a first border surrounding the representation of the first set of interactive elements; The second representation of the second augmented reality experience includes a second border surrounding the representation of the second set of interactive elements; and Displaying the first animation includes displaying the first border moving toward the viewpoint of the user of the computer system until the first border is no longer displayed.
10. The method according to any one of claims 8 to 9, further comprising: In response to receiving the first input: Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience: The representation of the first set of interactive elements is displayed via the one or more display generating components, cross-fading with the first set of interactive elements.
11. The method according to any one of claims 3 to 10, further comprising: In response to receiving the first input: Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience: ceasing display of said representations of said first set of interactive elements; and The first set of interactive elements is displayed via the one or more display generating components.
12. The method according to any one of claims 1 to 11, further comprising: Prior to receiving the first user input, and while the representations of the plurality of augmented reality experiences are displayed simultaneously in the three-dimensional environment: displaying, via the one or more display generating components, the first representation of the first augmented reality experience at a first display location; receiving, via the one or more input devices, a second user input corresponding to a request to navigate from the first representation of the first augmented reality experience to the second representation of the second augmented reality experience while displaying the first representation of the first augmented reality experience at the first display location; as well as In response to receiving the second user input: ceasing display of the first representation of the first augmented reality experience at the first display location; as well as The second representation of the second augmented reality experience is displayed at the first display location via the one or more display generating components.
13. The method of any one of claims 1 to 12, wherein simultaneously displaying the representations of the multiple augmented reality experiences comprises displaying the representations of the multiple augmented reality experiences in a stacked form, wherein the first representation of the first augmented reality experience is stacked on top of the second representation of the second augmented reality experience.
14. The method according to any one of claims 1 to 13, further comprising: prior to receiving the first user input, and while simultaneously displaying the representations of the plurality of augmented reality experiences in the three-dimensional environment, including simultaneously displaying the first representation of the first augmented reality experience and the second representation of the second augmented reality experience, receiving, via the one or more input devices, a third user input corresponding to a request to navigate through the representations of the plurality of augmented reality experiences; as well as In response to receiving the third user input: Display of the first representation of the first augmented reality experience is ceased while display of the second representation of the second augmented reality experience is maintained.
15. The method according to any one of claims 1 to 14, wherein: Determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes determining that the first user input is a selection input, the selection input comprising: gaze input toward the first representation of the first augmented reality experience; and A hardware press input is detected when the gaze input is toward the first representation of the first augmented reality experience.
16. The method according to any one of claims 1 to 15, wherein: Determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes determining that the first user input is a selection input, the selection input comprising: Indicates user-requested voice input to select a selectable object.
17. The method according to any one of claims 1 to 16, wherein: Determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience includes determining that the first user input is a selection input, the selection input comprising: A gaze input is directed toward the first representation of the first augmented reality experience, the first augmented reality experience satisfying a first set of gaze duration criteria.
18. The method according to any one of claims 1 to 17, further comprising: while simultaneously displaying the representations of the plurality of augmented reality experiences, displaying, via the one or more display generating components, one or more settings controls, the one or more settings controls including a first settings control corresponding to a first setting of the computer system; receiving, via the one or more input devices, a first setting input corresponding to the first setting of the computer system while the one or more setting controls are displayed; in response to receiving the first setting input, modifying the first setting from a first value to a second value different from the first value; while the representations of the plurality of augmented reality experiences are displayed simultaneously, and when the first setting is set to the second value, receiving a third user input via the one or more input devices; as well as In response to receiving the third user input: Based on determining that the third user input corresponds to a selection of the first representation of the first augmented reality experience, displaying the first augmented reality experience in the three-dimensional environment via the one or more display generation components while maintaining the first setting at the second value; and Based on determining that the first user input corresponds to a selection of the second representation of the second augmented reality experience, displaying the second augmented reality experience in the three-dimensional environment via the one or more display generation components while maintaining the first setting at the second value.
19. The method according to claim 18, wherein: The first setting is a transparent shading setting; The first value corresponds to a first quantity applied to the three-dimensional environment; and The second value corresponds to a second amount of shading applied to the three-dimensional environment that is different from the first amount of shading.
20. The method of claim 18, wherein: The first setting is a volume setting; The first value corresponds to a first volume; and The second value corresponds to a second volume different from the first volume.
21. The method according to any one of claims 18 to 20, further comprising: When the representations of the plurality of augmented reality experiences and the one or more settings controls are concurrently displayed, device state information indicating a state of one or more characteristics of the computer system is displayed via the one or more display generating components.
22. A method according to any one of claims 1 to 21, wherein the representations of the multiple augmented reality experiences are viewpoint-locked objects, and when the viewpoint of a user of the computer system shifts relative to the three-dimensional environment, the viewpoint-locked objects remain in a corresponding area of the user's field of view.
23. The method of claim 22, wherein: concurrently displaying the representations of the plurality of augmented reality experiences comprises concurrently displaying the representations of the plurality of augmented reality experiences in a first orientation in which the representations of the plurality of augmented reality experiences are aligned with gravity; and The method further comprises: detecting a change in orientation of the viewpoint of the user while simultaneously displaying the representations of the plurality of augmented reality experiences; as well as In response to detecting the change in the orientation of the viewpoint of the user: The representations of the plurality of augmented reality experiences are rotated from the first orientation to a second orientation based on the change in orientation of the viewpoint of the user to continue to align the representations of the plurality of augmented reality experiences with gravity.
24. The method of claim 23, wherein rotating the representations of the plurality of augmented reality experiences from the first orientation to the second orientation comprises: at a first time after detecting the change in orientation of the viewpoint of the user, displaying, via the one or more display generating components, the representations of the plurality of augmented reality experiences in the first orientation, wherein at the first time, the representations of the plurality of augmented reality experiences are not aligned with gravity due, at least in part, to the change in orientation of the viewpoint of the user; as well as At a second time after the first time, the representations of the plurality of augmented reality experiences are displayed, via the one or more display generating components, in the second orientation to align the representations of the plurality of augmented reality experiences with gravity.
25. A method according to any one of claims 22 to 24, wherein: Displaying the first augmented reality experience includes simultaneously displaying a first set of objects including a first object and a second object, and wherein: The first object is a viewpoint-locked object; and The second object is an environment-locked object.
26. The method according to any one of claims 1 to 25, further comprising: displaying, via the one or more display generation components, the first augmented reality experience in the three-dimensional environment; while displaying the first augmented reality experience, receiving, via the one or more input devices, a first voice input indicating a user request to change from the first augmented reality experience to the second augmented reality experience; as well as In response to receiving the first voice input: ceasing display of the first augmented reality experience; as well as The second augmented reality experience is displayed in the three-dimensional environment via the one or more input devices.
27. The method according to any one of claims 1 to 26, further comprising: When the computer system is in a sleep state, receiving, via the one or more input devices, a first wake-up input corresponding to a request to transition the computer system from the sleep state to a wake-up state; as well as In response to receiving the first wake-up input, the first augmented reality experience is displayed via the one or more display generation components.
28. The method according to any one of claims 1 to 26, further comprising: When the computer system is in a sleep state, receiving, via the one or more input devices, a first wake-up input corresponding to a request to transition the computer system from the sleep state to a wake-up state; as well as In response to receiving the first wake-up input, the representations of the plurality of augmented reality experiences are displayed via the one or more display generating components.
29. A method according to any one of claims 1 to 28, wherein the multiple augmented reality experiences include one or more of the following: a camera augmented reality experience; a translation augmented reality experience; a reading augmented reality experience; a music augmented reality experience; a navigation augmented reality experience; a photo augmented reality experience; a video messaging augmented reality experience; and / or a fitness augmented reality experience.
30. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system communicating with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing the method according to any one of claims 1 to 29.
31. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 1 to 29.
32. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 1 to 29.
33. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 29.
34. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: Simultaneously displaying, via the one or more display generation components, representations of a plurality of augmented reality experiences in a three-dimensional environment, the representations comprising: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receiving, via the one or more input devices, a first user input while simultaneously displaying the representations of the plurality of augmented reality experiences in the three-dimensional environment; as well as In response to receiving the first user input: ceasing display of the representation of one or more augmented reality experiences in the plurality of augmented reality experiences; as well as Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience, displaying the first augmented reality experience in the three-dimensional environment via the one or more display generation components.
35. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: Simultaneously displaying, via the one or more display generation components, representations of a plurality of augmented reality experiences in a three-dimensional environment, the representations comprising: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receiving a first user input via the one or more input devices while the representations of the plurality of augmented reality experiences are simultaneously displayed in the three-dimensional environment; and In response to receiving the first user input: ceasing display of the representation of one or more augmented reality experiences in the plurality of augmented reality experiences; and Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience, displaying the first augmented reality experience in the three-dimensional environment via the one or more display generation components.
36. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: Means for simultaneously displaying, via the one or more display generation components, representations of a plurality of augmented reality experiences in a three-dimensional environment, the representations comprising: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; means for receiving a first user input via the one or more input devices while simultaneously displaying the representations of the plurality of augmented reality experiences in the three-dimensional environment; and means for, in response to receiving the first user input: ceasing display of the representation of one or more augmented reality experiences in the plurality of augmented reality experiences; and Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience, displaying the first augmented reality experience in the three-dimensional environment via the one or more display generation components.
37. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: Simultaneously displaying, via the one or more display generation components, representations of a plurality of augmented reality experiences in a three-dimensional environment, the representations comprising: a first representation of a first augmented reality experience; and a second representation of a second augmented reality experience different from the first augmented reality experience, wherein the second representation is different from the first representation; receiving, via the one or more input devices, a first user input while simultaneously displaying the representations of the plurality of augmented reality experiences in the three-dimensional environment; as well as In response to receiving the first user input: ceasing display of the representation of one or more augmented reality experiences in the plurality of augmented reality experiences; as well as Based on determining that the first user input corresponds to a selection of the first representation of the first augmented reality experience, displaying the first augmented reality experience in the three-dimensional environment via the one or more display generation components.
38. A method comprising: At a computer system in communication with one or more display generating components and one or more input devices: receiving a first sequence of one or more user inputs via a first physical control; as well as In response to receiving the first sequence of one or more user inputs: Based on determining that the first sequence of one or more user inputs has a first magnitude: displaying, via the one or more display generation components, a first extended reality experience in a three-dimensional environment; as well as Based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: A second extended reality experience different from the first extended reality experience is displayed in the three-dimensional environment via the one or more display generation components.
39. The method of claim 38, further comprising: After receiving the first sequence of one or more user inputs: receiving a second sequence of one or more user inputs via the first physical control; as well as In response to receiving the second sequence of one or more user inputs: Based on determining that the second sequence of one or more user inputs has a first direction: displaying, via the one or more display generation components, a third extended reality experience in the three-dimensional environment; as well as Based on determining that the second sequence of one or more user inputs has a second direction different from the first direction: A fourth extended reality experience different from the third extended reality experience is displayed in the three-dimensional environment via the one or more display generation components.
40. The method according to any one of claims 38 to 39, further comprising: displaying, via the one or more display generating components, a representation of the first extended reality experience in a first manner, the first manner indicating that selection input will cause the first extended reality experience to be displayed; receiving, via the first physical control, a second sequence of one or more user inputs while displaying the representation of the first extended reality experience in the first manner; as well as In response to receiving the second sequence of one or more user inputs: ceasing display of the representation of the first extended reality experience in the first manner; as well as A representation of the second extended reality experience is displayed in the first manner via the one or more display generation components.
41. The method of claim 40, further comprising: receiving, via the first physical control, a third sequence of one or more user inputs while displaying the representation of the second extended reality experience in the first manner; as well as In response to receiving a third sequence of the one or more user inputs: scrolling the representations of the plurality of extended reality experiences based on determining that the third sequence of the one or more user inputs has a third direction; as well as Based on determining that the third sequence of the one or more user inputs has a fifth direction different from the third direction, scrolling the representations of the plurality of extended reality experiences in a sixth direction different from the fourth direction.
42. The method according to any one of claims 40 to 41, further comprising: receiving, via the first physical control, a fourth sequence of one or more user inputs while displaying the representation of the second extended reality experience in the first manner; as well as In response to receiving a fourth sequence of the one or more user inputs: scrolling the representations of the plurality of extended reality experiences by a first amount based on determining that the fourth sequence of the one or more user inputs has a third magnitude; as well as Based on determining that the fourth sequence of the one or more user inputs has a fourth magnitude different from the third magnitude, scrolling the representations of the plurality of extended reality experiences a second amount different from the first amount.
43. A method according to any one of claims 38 to 42, wherein: The first sequence of receiving the one or more user inputs via the first physical control comprises: receiving, via the first physical control, a first input corresponding to a request to exit a currently displayed extended reality experience; and A second input is received via the first physical control corresponding to a request to select a next extended reality experience for display.
44. The method of claim 43, further comprising: In response to receiving the first input corresponding to a request to exit a currently displayed extended reality experience, displaying, via the one or more display generating components, a first animation in which a representation of the currently displayed extended reality experience moves away from a viewpoint of a user of the computer system.
45. A method according to any one of claims 43 to 44, wherein: The first sequence of receiving the one or more user inputs via the first physical control further includes: After receiving the first input and before receiving the second input, receiving a navigation input via the first physical control; and The first input includes a press input on the first physical control.
46. The method of claim 45, wherein the navigation input comprises rotation of the first physical control.
47. A method according to any one of claims 38 to 46, wherein: Displaying the first extended reality experience includes simultaneously displaying a first set of objects including a first object and a second object, and wherein: The first object is a viewpoint-locked object; and The second object is an environment-locked object.
48. The method according to any one of claims 38 to 47, further comprising: receiving, when the computer system is in a low-power state, via the one or more input devices, a first wake-up input corresponding to a request to transition the computer system from the low-power state to a higher-power state (e.g., a wake-up state); as well as In response to receiving the first wake-up input, the first extended reality experience is displayed via the one or more display generation components.
49. The method according to any one of claims 38 to 47, further comprising: receiving, when the computer system is in a low power state, via the one or more input devices, a first wake-up input corresponding to a request to transition the computer system from the low power state to a high power state; as well as In response to receiving the first wake-up input, simultaneously displaying, via the one or more display generation components, representations of a plurality of extended reality experiences, including simultaneously displaying: a representation of the first extended reality experience; and A representation of the second extended reality experience separate from the representation of the first extended reality experience.
50. The method according to any one of claims 38 to 49, further comprising: receiving a fourth sequence of one or more user inputs via the first physical control; as well as In response to receiving a fourth sequence of the one or more user inputs, a volume setting of the computer system is modified.
51. The method according to any one of claims 38 to 50, further comprising: receiving a fifth sequence of one or more user inputs via the first physical control; as well as In response to receiving the fifth sequence of the one or more user inputs, a pass-through shading setting of the computer system is modified.
52. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 38 to 51.
53. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 38 to 51.
54. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 38 to 51.
55. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 38 to 51.
56. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: receiving a first sequence of one or more user inputs via a first physical control; and In response to receiving the first sequence of one or more user inputs: Based on determining that the first sequence of one or more user inputs has a first magnitude: displaying, via the one or more display generation components, a first extended reality experience in a three-dimensional environment; and Based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: A second extended reality experience different from the first extended reality experience is displayed in the three-dimensional environment via the one or more display generation components.
57. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: receiving a first sequence of one or more user inputs via a first physical control; and In response to receiving the first sequence of one or more user inputs: Based on determining that the first sequence of one or more user inputs has a first magnitude: displaying, via the one or more display generation components, a first extended reality experience in a three-dimensional environment; as well as Based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: A second extended reality experience different from the first extended reality experience is displayed in the three-dimensional environment via the one or more display generation components.
58. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: means for receiving a first sequence of one or more user inputs via a first physical control; and means for, in response to receiving the first sequence of one or more user inputs, performing the following operations: Based on determining that the first sequence of one or more user inputs has a first magnitude: displaying, via the one or more display generation components, a first extended reality experience in a three-dimensional environment; as well as Based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: A second extended reality experience different from the first extended reality experience is displayed in the three-dimensional environment via the one or more display generation components.
59. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: receiving a first sequence of one or more user inputs via a first physical control; and In response to receiving the first sequence of one or more user inputs: Based on determining that the first sequence of one or more user inputs has a first magnitude: displaying, via the one or more display generation components, a first extended reality experience in a three-dimensional environment; and Based on determining that the first sequence of one or more user inputs has a second magnitude different from the first magnitude: A second extended reality experience different from the first extended reality experience is displayed in the three-dimensional environment via the one or more display generation components.
60. A method comprising: At a computer system in communication with one or more display generating components and one or more input devices: detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment while a view of the three-dimensional environment in which the computer system is located is visible; as well as In response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: Displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
61. The method of claim 60, further comprising: detecting, via the one or more input devices, one or more objects in a second three-dimensional environment while the view of the second three-dimensional environment is visible; as well as In response to detecting the first group of objects in the second three-dimensional environment: displaying, via the one or more display generation components and concurrently with at least a portion of the view of the second three-dimensional environment, a second suggestion corresponding to a second augmented reality experience based on determining that the one or more objects in the second three-dimensional environment include a first group of objects; as well as Based on determining that the one or more objects in the second three-dimensional environment include a second group of objects different from the first group of objects, a third suggestion corresponding to a third augmented reality experience different from the second augmented reality experience is displayed via the one or more display generation components and simultaneously with at least a portion of the view of the second three-dimensional environment.
62. The method of any one of claims 60 to 61, wherein the first augmented reality experience is selected from the plurality of augmented reality experiences displayable by the computer system based on audio content received by the computer system.
63. The method of any one of claims 60 to 62, wherein: The three-dimensional environment is a transparent environment; and The first suggestion corresponding to the first augmented reality experience is superimposed on the transparent environment.
64. The method according to any one of claims 60 to 63, further comprising: receiving, via the one or more input devices, accepting user input while displaying the first suggestion corresponding to the first augmented reality experience; as well as In response to receiving the accept user input, the first augmented reality experience is displayed via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system.
65. The method of claim 64, wherein: The accepting user input includes a first gaze input corresponding to the first suggestion.
66. A method according to any one of claims 64 to 65, wherein: The accepting user input includes a first hand input corresponding to the first suggestion.
67. A method according to any one of claims 64 to 66, wherein: The accepting user input includes physical control input via a physical input mechanism.
68. The method according to any one of claims 64 to 67, further comprising: receiving, via a first physical control, a first sequence of one or more user inputs while displaying the first augmented reality experience; as well as In response to receiving the first sequence of one or more user inputs: ceasing display of the first augmented reality experience; as well as A second augmented reality experience different from the first augmented reality experience is displayed via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system.
69. The method according to any one of claims 60 to 68, further comprising: while displaying the first suggestion, determining that a first set of exclusion criteria are satisfied, wherein the first set of exclusion criteria includes a first criterion that is satisfied when the first suggestion has been displayed for a threshold duration without a user of the computer system providing user input accepting the first suggestion; as well as In response to determining that the first set of exclusion criteria is satisfied, display of the first suggestion is ceased.
70. The method of any one of claims 60 to 69, wherein: detecting the first set of conditions while displaying a second augmented reality experience different from the first augmented reality experience; and Displaying the first suggestion corresponding to the first augmented reality experience includes: The first suggestion corresponding to the first augmented reality experience is displayed while maintaining display of the second augmented reality experience.
71. The method of claim 70, further comprising: receiving, via the one or more input devices, a second accepting user input while displaying the first suggestion corresponding to the first augmented reality experience; as well as In response to receiving the second accepted user input: ceasing display of the second augmented reality experience; as well as The first augmented reality experience is displayed via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system.
72. The method of claim 70, further comprising: receiving, via the one or more input devices, a third accepting user input while displaying the first suggestion corresponding to the first augmented reality experience; as well as In response to receiving the third accepted user input: The first augmented reality experience is displayed via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system while maintaining display of the second augmented reality experience.
73. The method of any one of claims 60 to 72, wherein: Displaying the first suggestion includes displaying the first suggestion in a first display area of the one or more display generating components; and The method further comprises: while displaying the first suggestion in the first display area of the one or more display generating components, detecting a change in the viewpoint of the user from pointing in a first direction to pointing in a second direction different from the first direction; as well as After detecting the change of the viewpoint of the user from pointing in the first direction to pointing in the second direction: When the viewpoint of the user is directed to the second direction, the first suggestion is displayed in the first display area of the one or more display generating components via the one or more display generating components.
74. The method of claim 73, wherein: Displaying the first suggestion includes displaying the first suggestion in a first orientation, wherein the first suggestion is aligned with gravity; and The method further comprises: detecting a change in orientation of the viewpoint of the user while the first suggestion is displayed in the first orientation; as well as In response to detecting the change in the orientation of the viewpoint of the user: The first suggestion is rotated from the first orientation to a second orientation based on the change in orientation of the viewpoint of the user to continue to align the first suggestion with gravity.
75. The method of claim 74, wherein rotating the first suggestion from the first orientation to the second orientation comprises: displaying, via the one or more display generating components, the first suggestion in the first orientation at a first time after detecting the change in orientation of the viewpoint of the user, wherein at the first time, the first suggestion is not aligned with gravity due, at least in part, to the change in orientation of the viewpoint of the user; as well as At a second time after the first time, the first suggestion is displayed, via the one or more display generating components, in the second orientation to align the first suggestion with gravity.
76. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 60 to 75.
77. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 60 to 75.
78. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 60 to 75.
79. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for performing a method according to any one of claims 60 to 75.
80. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: When a view of the three-dimensional environment in which the computer system is located is visible, detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment; and In response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: Displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
81. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment while a view of the three-dimensional environment in which the computer system is located is visible; as well as In response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: Displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
82. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: means for detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment when a view of the three-dimensional environment in which the computer system is located is visible; and means for, in response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: Displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
83. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: When a view of the three-dimensional environment in which the computer system is located is visible, detecting, via the one or more input devices, a first set of conditions in the three-dimensional environment; and In response to detecting the first set of conditions in the three-dimensional environment, performing the following operations: Displaying, via the one or more display generation components and concurrently with at least a portion of the view of the three-dimensional environment of the computer system, a first suggestion corresponding to a first augmented reality experience selected from a plurality of augmented reality experiences capable of being displayed by the computer system.
84. A method comprising: At a computer system in communication with one or more display generating components and one or more input devices: detecting, via the one or more input devices, a gaze of a user corresponding to a first display position of the one or more display generating components; in response to detecting the gaze of the user corresponding to the first display position of the one or more display generating components, displaying a first object via the one or more display generating components; When the first object is displayed, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying, via the one or more display generating components, movement of the first object; as well as After showing the movement of the first object: performing a first operation based on determining that the gaze of the user satisfies a second set of criteria indicative of gaze tracking of the movement of the first object; as well as Based on determining that the gaze of the user does not satisfy the second set of criteria indicative of gaze tracking of the movement of the first object, performing the first operation is abandoned.
85. The method of claim 84, further comprising: Prior to detecting the gaze of the user at the first display location corresponding to the one or more display generating components, an initial object different from the first object is displayed at the first display location via the one or more display generating components.
86. The method of claim 85, further comprising: In response to detecting the gaze of the user corresponding to the first display location of the one or more display generating components, the appearance of the initial object is changed.
87. A method according to any one of claims 85 to 86, wherein the initial object is persistently displayed without displaying the first object.
88. The method of any one of claims 85 to 86, wherein: The initial object is displayed at the first display location in response to detecting a gaze of the user in a first display area corresponding to the one or more display generating components, wherein the first display area is larger than and includes the first display location.
89. The method of any one of claims 84 to 88, further comprising: In response to detecting the gaze of the user corresponding to the first display location of the one or more display generating components, displaying, via the one or more display generating components and concurrently with the first object, a first instruction instructing the user to look at the first object.
90. The method of any one of claims 84 to 89, wherein the first set of criteria includes a first criterion that is satisfied when the computer system detects a user gaze corresponding to the first display location for a threshold duration.
91. A method according to any one of claims 84 to 90, wherein the first set of criteria includes a second criterion that is satisfied when the computer system detects a user gaze corresponding to the first object.
92. The method of any one of claims 84 to 91, further comprising: When displaying the first object, detecting that the gaze of the user is directed to a display area of the one or more display generating components that does not correspond to the first display position or the first object; as well as In response to detecting that the gaze of the user is directed to a display area of the one or more display generating components that does not correspond to the first display position or the first object, display of the first object is stopped.
93. The method of any one of claims 84 to 92, wherein displaying the movement of the first object comprises displaying the movement of the first object at a first predetermined rate of motion.
94. The method of any one of claims 84 to 93, wherein displaying the movement of the first object comprises displaying the movement of the first object at a first movement rate, wherein the first movement rate is determined based on the gaze of the user.
95. A method according to any one of claims 84 to 94, wherein the second set of criteria includes a second criterion that is met when the movement of the gaze of the user meets a similarity criterion relative to the movement of the first object.
96. The method of claim 95, wherein the second set of criteria comprises a smooth movement criterion that is satisfied when the movement of the gaze of the user satisfies a smoothness criterion that indicates the smoothness of the movement of the gaze of the user.
97. A method according to any one of claims 95 to 96, wherein determining whether the movement of the user's gaze meets the smooth movement standard excludes glances.
98. The method of any one of claims 84 to 97, wherein: Displaying the movement of the first object includes displaying the movement of the first object from an initial position to a destination position; and The second set of criteria includes a third criterion that is satisfied when the gaze of the user moves to the destination location.
99. The method of claim 98, further comprising: When displaying the movement of the first object, a destination indication indicating the destination location is displayed via the one or more display generating components and concurrently with the movement of the first object.
100. The method of any one of claims 98 to 99, further comprising: while displaying the movement of the first object and before the first object reaches the destination location, detecting, via the one or more input devices, a gaze of the user corresponding to the destination location; as well as In response to detecting the gaze of the user corresponding to the destination location, the first operation is performed.
101. A method according to any one of claims 98 to 100, wherein the second set of criteria includes a fourth criterion that is satisfied when the gaze of the user moves to the destination location and remains at the destination location for a threshold duration.
102. The method of any one of claims 84 to 101, wherein: Displaying the movement of the first object includes displaying the movement of the first object from an initial position to a destination position; and The method further comprises: Based on determining that the gaze of the user does not satisfy the second set of criteria indicative of gaze tracking of the movement of the first object, the first object is displayed at the initial position.
103. The method of any one of claims 84 to 102, wherein performing the first operation comprises transitioning the computer system from a locked state to an unlocked state.
104. The method of any one of claims 84 to 103, further comprising: When showing the movement of the first object: Based on determining that the gaze of the user satisfies progress criteria indicating progress towards satisfying the second set of criteria, providing a first audio output.
105. The method of claim 104, further comprising: After showing the movement of the first object: Based on determining that the gaze of the user satisfies a second set of criteria indicative of gaze tracking of the movement of the first object, a second audio output is provided indicative of the gaze of the user satisfying the second set of criteria.
106. The method according to any one of claims 104 to 105, further comprising: While displaying movement of the first object and after providing the first audio output: Based on determining that the gaze of the user no longer satisfies the progress criteria indicating progress towards satisfying the second set of criteria, output of the first audio output is ceased.
107. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 84 to 106.
108. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 84 to 106.
109. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 84 to 106.
110. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 84 to 106.
111. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: detecting, via the one or more input devices, a gaze of a user corresponding to a first display position of the one or more display generating components; in response to detecting the gaze of the user corresponding to the first display position of the one or more display generating components, displaying a first object via the one or more display generating components; When the first object is displayed, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying, via the one or more display generating components, movement of the first object; as well as After showing the movement of the first object: performing a first operation based on determining that the gaze of the user satisfies a second set of criteria indicative of gaze tracking of the movement of the first object; as well as Based on determining that the gaze of the user does not satisfy the second set of criteria indicative of gaze tracking of the movement of the first object, performing the first operation is abandoned.
112. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: detecting, via the one or more input devices, a gaze of a user corresponding to a first display position of the one or more display generating components; in response to detecting the gaze of the user corresponding to the first display position of the one or more display generating components, displaying a first object via the one or more display generating components; When the first object is displayed, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying, via the one or more display generating components, movement of the first object; as well as After showing the movement of the first object: performing a first operation based on determining that the gaze of the user satisfies a second set of criteria indicative of gaze tracking of the movement of the first object; as well as Based on determining that the gaze of the user does not satisfy the second set of criteria indicative of gaze tracking of the movement of the first object, performing the first operation is abandoned.
113. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: means for detecting, via the one or more input devices, a gaze of a user corresponding to a first display position of the one or more display generating components; means for displaying a first object via the one or more display generating components in response to detecting the gaze of the user corresponding to the first display position of the one or more display generating components; means for detecting that a first set of criteria is satisfied when displaying the first object; means for displaying, via the one or more display generating components, movement of the first object in response to detecting that the first set of criteria is satisfied; and Means for performing the following operations after displaying the movement of the first object: performing a first operation based on determining that the gaze of the user satisfies a second set of criteria indicative of gaze tracking of the movement of the first object; as well as Based on determining that the gaze of the user does not satisfy the second set of criteria indicative of gaze tracking of the movement of the first object, performing the first operation is abandoned.
114. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: detecting, via the one or more input devices, a gaze of a user corresponding to a first display position of the one or more display generating components; in response to detecting the gaze of the user corresponding to the first display position of the one or more display generating components, displaying a first object via the one or more display generating components; When the first object is displayed, detecting that a first set of criteria is satisfied; in response to detecting that the first set of criteria is satisfied, displaying, via the one or more display generating components, movement of the first object; as well as After showing the movement of the first object: performing a first operation based on determining that the gaze of the user satisfies a second set of criteria indicative of gaze tracking of the movement of the first object; as well as Based on determining that the gaze of the user does not satisfy the second set of criteria indicative of gaze tracking of the movement of the first object, performing the first operation is abandoned.
115. A method comprising: At a computer system in communication with one or more display generating components and one or more input devices: displaying virtual content via the one or more display generation components; detecting, via the one or more input devices, a first gesture in front of a face of a user of the computer system while displaying the virtual content; as well as In response to detecting the first gesture: Based on determining that the first gesture in front of the face of the user satisfies a first set of criteria, ceasing display of at least a portion of the virtual content; and Based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.
116. The method of claim 115, wherein determining that the first gesture in front of the face of the user satisfies the first set of criteria comprises determining that a speed of the first gesture satisfies a speed criterion.
117. A method according to any one of claims 115 to 116, wherein determining that the first gesture in front of the face of the user satisfies the first set of criteria includes determining that the distance between the user's hand and the user's face when performing the first gesture satisfies a distance criterion.
118. The method of claim 117, wherein the first set of criteria includes a first criterion that is satisfied when the distance between the hand of the user and the face of the user is greater than a minimum distance threshold when performing the first gesture.
119. A method according to any one of claims 117 to 118, wherein the first set of criteria includes a second criterion that is satisfied when the distance between the user's hand and the user's face is less than a maximum distance threshold when performing the first gesture.
120. The method of any one of claims 115 to 119, wherein determining that the first gesture in front of the face of the user satisfies the first set of criteria comprises determining that the location of the first gesture satisfies location criteria.
121. The method of claim 120, wherein determining that the first gesture in front of the face of the user does not satisfy the first set of criteria comprises determining that the first gesture is performed within a first predetermined portion of the field of view of the computer system.
122. A method according to any one of claims 115 to 121, wherein determining that the first gesture in front of the face of the user satisfies the first set of criteria includes determining that a direction of the first gesture satisfies an orientation criterion.
123. The method of any one of claims 115 to 122, wherein the virtual content comprises a virtual environment.
124. A method according to any one of claims 115 to 123, wherein the virtual content includes virtual content superimposed on a three-dimensional augmented reality environment, and the three-dimensional augmented reality environment includes one or more elements representing a three-dimensional environment in which the computer system is located.
125. The method of any one of claims 115 to 124, wherein ceasing display of at least a portion of the virtual content comprises ceasing display of the virtual content.
126. The method of any one of claims 115 to 124, wherein ceasing display of at least a portion of the virtual content comprises ceasing display of a first portion of the virtual content while maintaining display of a second portion of the virtual content.
127. The method of claim 126, wherein ceasing display of the first portion of the virtual content while maintaining display of the second portion of the virtual content comprises: ceasing display of the first portion of the virtual content based on determining that the user's gaze is directed toward the first portion of the virtual content when the first gesture is detected; as well as Based on determining that the gaze of the user is not directed toward the second portion of the virtual content when the first gesture is detected, maintaining display of the second portion of the virtual content.
128. The method of claim 126, wherein: The first portion of the virtual content corresponds to a first application; The second portion of the virtual content corresponds to a system user interface; and Stopping display of the first portion of the virtual content while maintaining display of the second portion of the virtual content includes: Based on determining that the first portion of the virtual content corresponds to the first application, stopping display of the first portion of the virtual content; as well as Based on determining that the second portion of the virtual content corresponds to the system user interface, the display of the second portion of the virtual content is maintained.
129. The method of claim 126, wherein: The first portion of the virtual content corresponds to foreground content; The second portion of the virtual content corresponds to background content; and Stopping display of the first portion of the virtual content while maintaining display of the second portion of the virtual content includes: ceasing display of the first portion of the virtual content based on determining that the first portion of the virtual content corresponds to foreground content; as well as Based on determining that the second portion of the virtual content corresponds to background content, display of the second portion of the virtual content is maintained.
130. The method of any one of claims 115 to 129, further comprising: while displaying the virtual content, detecting, via the one or more input devices, a second gesture in front of the face of the user; In response to the first portion of the second gesture, stopping display of a third portion of the virtual content while maintaining display of a fourth portion of the virtual content; as well as In response to a second portion of the second gesture being a continuation of the first portion of the second gesture, display of the fourth portion of the virtual content is ceased.
131. The method of claim 130, wherein: ceasing display of the fourth portion of the virtual content is performed based on determining that the second gesture satisfies the first set of criteria; and The method further comprises: In response to the second part of the second gesture: Based on determining that the second gesture does not satisfy the first set of criteria: maintaining display of the fourth portion of the virtual content; and The third portion of the virtual content is redisplayed.
132. The method of any one of claims 115 to 131, further comprising: After ceasing display of the at least a portion of the virtual content and while the at least a portion of the virtual content is not being displayed, detecting, via the one or more input devices, a third gesture in front of a face of the user; as well as In response to detecting the third gesture: Based on determining that the third gesture in front of the face of the user satisfies a second set of criteria, redisplaying the at least a portion of the virtual content.
133. The method of claim 132, further comprising: In response to detecting the third gesture: Based on determining that the third gesture in front of the face of the user does not satisfy the second set of criteria, redisplaying the at least a portion of the virtual content is abandoned.
134. The method of any one of claims 132 to 133, wherein: The first gesture comprises movement in a third direction; and The second set of criteria includes a third criterion that is satisfied when the third gesture includes movement in a fourth direction different from the third direction.
135. The method of any one of claims 132 to 133, wherein: The first gesture comprises movement in a fifth direction; and The second set of criteria includes a fourth criterion that is satisfied when the third gesture includes movement in the fifth direction.
136. The method of any one of claims 132 to 135, wherein: The second set of criteria includes a fifth criterion that is satisfied when a duration elapsed between the first gesture and the third gesture is less than a threshold duration.
137. The method of any one of claims 115 to 136, further comprising: detecting, via the one or more input devices, a first air gesture input at a first elapsed time after ceasing display of the at least a portion of the virtual content and when the at least a portion of the virtual content is not displayed; as well as In response to detecting the first air gesture input: Redisplaying the at least a portion of the virtual content based on determining that the first elapsed time is less than a first threshold duration; and Based on determining that the first elapsed time is greater than the first threshold duration, redisplaying the at least a portion of the virtual content is foregone.
138. The method of claim 137, further comprising: detecting, via the one or more input devices, a first mechanical hardware input at a second elapsed time after ceasing display of the at least a portion of the virtual content and while the at least a portion of the virtual content is not displayed, the second elapsed time being greater than the first elapsed time; and In response to detecting the first mechanical hardware input: The at least a portion of the virtual content is redisplayed.
139. The method of any one of claims 115 to 138, further comprising: After stopping the display of the at least a portion of the virtual content, and while the at least a portion of the virtual content is not displayed: displaying, via the one or more display generation components, a redisplay indication based on determining that the at least a portion of the virtual content is available for redisplay; as well as Based on determining that the at least a portion of the virtual content is not available for redisplay, display of the redisplay indication is abandoned.
140. The method of claim 139, further comprising: When the redisplay indication is displayed, first audio content indicating virtual content that can be used for redisplay is output.
141. The method of any one of claims 139 to 140, further comprising: When the redisplay indication is displayed, determining that a content removal criterion has been met; as well as In response to determining that the content removal criteria have been met: An output is provided indicating second audio content that satisfies the content removal criteria.
142. The method of claim 141, further comprising: In response to determining that the content removal criteria have been met: The display of the redisplay indication is stopped.
143. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system that communicates with one or more display generating components and one or more input devices, the one or more programs comprising instructions for executing a method according to any one of claims 115 to 142.
144. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 115 to 142.
145. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: A component for carrying out the method according to any one of claims 115 to 142.
146. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for performing a method according to any one of claims 115 to 142.
147. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: displaying virtual content via the one or more display generation components; detecting, via the one or more input devices, a first gesture in front of a face of a user of the computer system while displaying the virtual content; as well as In response to detecting the first gesture: Based on determining that the first gesture in front of the face of the user satisfies a first set of criteria, ceasing display of at least a portion of the virtual content; and Based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.
148. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: one or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for: displaying virtual content via the one or more display generation components; detecting, via the one or more input devices, a first gesture in front of a face of a user of the computer system while displaying the virtual content; as well as In response to detecting the first gesture: Based on determining that the first gesture in front of the face of the user satisfies a first set of criteria, ceasing display of at least a portion of the virtual content; and Based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.
149. A computer system configured to communicate with one or more display generating components and one or more input devices, the computer system comprising: means for displaying virtual content via the one or more display generation components; means for detecting, via the one or more input devices, a first gesture in front of a face of a user of the computer system while displaying the virtual content; and means for, in response to detecting the first gesture, performing the following operations: Based on determining that the first gesture in front of the face of the user satisfies a first set of criteria, ceasing display of at least a portion of the virtual content; and Based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.
150. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more display generating components and one or more input devices, the one or more programs comprising instructions for: displaying virtual content via the one or more display generation components; detecting, via the one or more input devices, a first gesture in front of a face of a user of the computer system while displaying the virtual content; as well as In response to detecting the first gesture: Based on determining that the first gesture in front of the face of the user satisfies a first set of criteria, ceasing display of at least a portion of the virtual content; and Based on determining that the first gesture in front of the face of the user does not meet the first set of criteria, maintaining display of the virtual content.