Devices, methods, and graphical user interfaces for interacting with a three-dimensional environment

By using a camera in a computer system to detect the micro gestures of the user's hand and the ready state configuration of the hand, combined with the gaze input, the interaction with the virtual/augmented reality environment is simplified, the problems of low efficiency and high complexity in the prior art are solved, and a more efficient and intuitive user experience is achieved.

CN114637376BActive Publication Date: 2025-07-04APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210375296.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-23
Filing Date
2020-09-25
Publication Date
2025-07-04
Estimated Expiration
2040-09-25

AI Technical Summary

Technical Problem

In the prior art, the methods and interfaces used to interact with virtual/augmented reality environments are inefficient, complex operations and error-prone, resulting in a large cognitive burden on users and taking a long time, especially in battery-driven devices.

Method used

By using a camera in a computer system to detect tiny gesture movements of a user's hands, combining hand-ready state configuration and gaze input, simplify interaction with the three-dimensional environment, reduce the number and nature of inputs, and provide intuitive operational context indications and feedback.

Benefits of technology

It improves the efficiency and security of user interaction, reduces cognitive burden, provides a more intuitive and efficient human-computer interface, suitable for a variety of computer systems and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114637376B_ABST
    Figure CN114637376B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a device, a method, and a graphical user interface for interacting with a three-dimensional environment. When displaying a three-dimensional environment, a computer system detects a hand at a first location corresponding to a portion of the three-dimensional environment. In response to detecting the hand at the first location: based on determining that the hand is held in a first predefined configuration, the computer system displays a visual indication of a first operational context for a gesture input using a gesture in the three-dimensional environment; and based on determining that the hand is not held in the first predefined configuration, the computer system abandons displaying the visual indication.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application for invention with the application date of September 25, 2020, application number 202080064537.8, and invention title "Devices, Methods, and Graphical User Interfaces for Interacting with a Three-Dimensional Environment".

[0002] Related patent applications

[0003] This patent application claims the priority of U.S. Provisional Patent Application No. 62 / 907,480 filed on September 27, 2019 and U.S. Patent Application No. 17 / 030,200 filed on September 23, 2020, and is a continuation of U.S. Patent Application No. 17 / 030,200 filed on September 23, 2020. Technical Field

[0004] The present disclosure generally relates to computer systems that provide computer-generated experiences and have a display generation component and one or more input devices, including but not limited to electronic devices that provide virtual reality and mixed reality experiences via a display. Background Art

[0005] In recent years, the development of computer systems for augmented reality has increased significantly. Example augmented reality environments include at least some virtual elements that replace or enhance the physical world. Input devices (such as cameras, controllers, joysticks, touch-sensitive surfaces, and touch-sensitive displays) for computer systems and other electronic computing devices are used to interact with virtual / augmented reality environments. Exemplary virtual elements include virtual objects (including digital images, videos, text, icons, control elements (such as buttons), and other graphics).

[0006] However, the methods and interfaces for interacting with environments that include at least some virtual elements (such as applications, augmented reality environments, mixed reality environments, and virtual reality environments) are cumbersome, inefficient, and limited. For example, systems that provide insufficient feedback for performing actions associated with virtual objects, systems that require a series of inputs to achieve a desired result in an augmented reality environment, and systems where virtual object manipulation is complex, cumbersome, and error-prone impose a significant cognitive burden on users and detract from the experience of virtual / augmented reality environments. In addition, these methods take longer than necessary, thereby wasting energy. This latter consideration is particularly important in battery-powered devices. Summary of the Invention

[0007] Accordingly, there is a need for computer systems with improved methods and interfaces to provide computer-generated experiences to users, making the interaction between the user and the computer system more efficient and intuitive for the user. Such methods and interfaces optionally supplement or replace conventional methods for providing computer-generated reality experiences to users. Such methods and interfaces form a more effective human-machine interface by helping the user understand the connection between the provided inputs and the device's response to these inputs, reducing the amount, degree, and / or nature of the inputs from the user.

[0008] The disclosed systems reduce or eliminate the above-mentioned deficiencies and other problems associated with the user interfaces for computer systems having a display generation component and one or more input devices. In some embodiments, the computer system is a desktop computer with an associated display. In some embodiments, the computer system is a portable device (e.g., a laptop computer, a tablet computer, or a handheld device). In some embodiments, the computer system includes a personal electronic device (e.g., a wearable electronic device such as a watch or a head-mounted device). In some embodiments, the computer system has a touchpad. In some embodiments, the computer system has one or more cameras. In some embodiments, the computer system has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display"). In some embodiments, the computer system has one or more eye-tracking components. In some embodiments, the computer system has one or more hand-tracking components. In some embodiments, in addition to the display generation component, the computer system also has one or more output devices, which include one or more haptic output generators and one or more audio output devices. In some embodiments, the computer system has a graphical user interface (GUI), one or more processors, a memory, and one or more modules, programs, or instruction sets stored in the memory for performing multiple functions. In some embodiments, the user interacts with the GUI through contact and gestures of a stylus and / or finger on a touch-sensitive surface, movements of the user's eyes and hands in space relative to the GUI or the user's body (as captured by cameras and other motion sensors), and voice input (as captured by one or more audio input devices). In some embodiments, the functions performed through the interaction optionally include image editing, drawing, presentation, word processing, spreadsheet creation, playing games, making phone calls, video conferencing, sending and receiving emails, instant messaging, fitness support, digital photography, digital video recording, web browsing, digital music playback, note-taking, and / or digital video playback. The executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program products configured to be executed by one or more processors.

[0009] An electronic device with improved methods and interfaces is needed to interact with a three-dimensional environment. Such methods and interfaces can supplement or replace conventional methods for interacting with a three-dimensional environment. Such methods and interfaces reduce the amount, degree, and / or nature of input from the user and produce a more efficient human-machine interface.

[0010] According to some embodiments, a method is performed at a computer system including a display generation component and one or more cameras. The method includes: displaying a view of a three-dimensional environment; when displaying the view of the three-dimensional environment, using the one or more cameras to detect movement of a user's thumb of a first hand above the user's index finger; in response to detecting, using the one or more cameras, movement of the user's thumb above the user's index finger: performing a first operation based on determining that the movement is a swipe of the thumb of the first hand above the index finger in a first direction; and performing a second operation different from the first operation based on determining that the movement is a tap of the thumb above the index finger at a first position on the index finger of the first hand.

[0011] According to some embodiments, a method is performed at a computing system including a display generation component and one or more input devices. The method includes: displaying a view of a three-dimensional environment; when the three-dimensional environment is displayed, detecting a hand at a first position corresponding to a portion of the three-dimensional environment; in response to detecting the hand at the first position corresponding to the portion of the three-dimensional environment: displaying a visual indication of a first operation context for gesture input using a gesture in the three-dimensional environment based on determining that the hand is not held in a first predefined configuration; and abandoning the display of the visual indication of the first operation context for gesture input using a gesture in the three-dimensional environment based on determining that the hand is not held in the first predefined configuration.

[0012] According to some embodiments, a method is performed at a computing system including a display generation component and one or more input devices. The method includes: displaying a three-dimensional environment, including displaying a representation of a physical environment; when displaying the representation of the physical environment, detecting a gesture; and in response to detecting the gesture: displaying a system user interface in the three-dimensional environment based on determining that the user's gaze points to a position corresponding to a predefined physical position in the physical environment; and performing an operation in the current context of the three-dimensional environment without displaying the system user interface based on determining that the user's gaze does not point to a position corresponding to a predefined physical position in the physical environment.

[0013] According to some embodiments, a method is performed at an electronic device including a display generation component and one or more input devices, the method including: displaying a three-dimensional environment that includes one or more virtual objects; detecting a gaze directed at a first object in the three-dimensional environment, where the gaze meets a first criterion and the first object responds to at least one gesture input; and in response to detecting a gaze that meets the first criterion and is directed at the first object that responds to at least one gesture input: displaying an indication of one or more interaction options available for the first object in the three-dimensional environment based on determining that the hand is in a predefined ready state for providing a gesture input; and foregoing displaying an indication of one or more interaction options available for the first object based on determining that the hand is not in a predefined ready state for providing a gesture input.

[0014] There is a need for electronic devices with improved methods and interfaces to facilitate user interaction with three-dimensional environments using the electronic devices. Such methods and interfaces can supplement or replace conventional methods for facilitating user interaction with three-dimensional environments using the electronic devices. Such methods and interfaces result in a more efficient human-machine interface and allow the user to have more control over the device, allowing the user to use a device that is safer, has reduced cognitive burden, and has an improved user experience.

[0015] In some embodiments, a method is performed at a computer system including a display generation component and one or more input devices, the method including: detecting the placement of the display generation component in a predefined position relative to a user of the electronic device; in response to detecting the placement of the display generation component in a predefined position relative to a user of the computer system, displaying, via the display generation component, a first view of a three-dimensional environment that includes a see-through portion, where the see-through portion includes a representation of at least a portion of the real world around the user; when displaying the first view of the three-dimensional environment that includes the see-through portion, detecting a change in the grip of the hand on a housing physically coupled to the display generation component; and in response to detecting a change in the grip of the hand on the housing physically coupled to the display generation component: replacing the first view of the three-dimensional environment with a second view of the three-dimensional environment based on determining that the change in the grip of the hand on the housing physically coupled to the display generation component meets a first criterion, where the second view replaces at least a portion of the see-through portion with virtual content.

[0016] In some embodiments, a method is performed at a computer system including a display generation component and one or more input devices, the method comprising: displaying a view of a virtual environment via the display generation component; detecting a first movement of a user in a physical environment when the view of the virtual environment is being displayed and when the view of the virtual environment does not include a visual representation of a first portion of a first physical object that exists in the physical environment in which the user is located; and in response to detecting the first movement of the user in the physical environment: changing the appearance of the view of the virtual environment in a first manner indicative of a physical characteristic of the first portion of the first physical object according to a determination that the user is within a threshold distance of the first portion of the first physical object, wherein the first physical object has a range potentially visible to the user based on the user's field of view for the virtual environment, without changing the appearance of the view of the virtual environment to be indicative of a second portion of the first physical object, the second portion being part of the range of the first physical object that is potentially visible to the user based on the user's field of view for the virtual environment; and refraining from changing the appearance of the view of the virtual environment in the first manner indicative of the physical characteristic of the first portion of the first physical object according to a determination that the user is not within the threshold distance of the first physical object that exists in the physical environment surrounding the user.

[0017] According to some embodiments, a computer system includes a display generation component (e.g., a display, a projector, a head-mounted display, etc.), one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), optionally one or more tactile output generators, one or more processors, and a memory storing one or more programs; the one or more programs are configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing the performance of the operations of any one of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), and optionally one or more tactile output generators, cause the device to perform the operations of any one of the methods described herein or cause the performance of the operations of any one of the methods described herein. According to some embodiments, the graphical user interface of a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), optionally one or more tactile output generators, a memory, and one or more processors for executing one or more programs stored in the memory includes one or more of the elements displayed in any one of the methods described herein, and the one or more elements are updated in response to input as described in any one of the methods described herein. According to some embodiments, a computer system includes: a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), and optionally one or more tactile output generators; and means for performing or causing the performance of the operations of any one of the methods described herein. According to some embodiments, an information processing apparatus in a computer system having a display generation component, one or more input devices (e.g., one or more cameras, a touch-sensitive surface, and optionally one or more sensors for detecting the intensity of contact with the touch-sensitive surface), and optionally one or more tactile output generators includes means for performing or causing the performance of the operations of any one of the methods described herein.

[0018] Accordingly, improved methods and interfaces are provided for a computer system having a display generation component for interacting with a three-dimensional environment and facilitating use of the computer system by a user when interacting with the three-dimensional environment, thereby enhancing the effectiveness, efficiency, user safety, and satisfaction of such computer systems. Such methods and interfaces may supplement or replace conventional methods for interacting with a three-dimensional environment and facilitating use of the computer system by a user when interacting with the three-dimensional environment.

[0019] Note that the various embodiments described above may be combined with any other embodiment described herein. The features and advantages described in this specification are not exhaustive, and in particular, many additional features and advantages will be apparent to those of ordinary skill in the art from the drawings, the specification, and the claims. Additionally, it should be noted that the language used in this specification has been selected for readability and guidance purposes and may not have been selected to delineate or circumscribe the subject matter of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the following drawings, in which like reference numerals indicate corresponding parts in all the drawings.

[0021] Figure 1 is a block diagram showing an operating environment of a computer system for providing a CGR experience according to some embodiments.

[0022] Figure 2 is a block diagram showing a controller of a computer system configured to manage and coordinate a user's CGR experience according to some embodiments.

[0023] Figure 3 is a block diagram showing a display generation component of a computer system configured to provide a visual component of a CGR experience to a user according to some embodiments.

[0024] Figure 4 is a block diagram showing a hand tracking unit of a computer system configured to capture a user's gesture input according to some embodiments.

[0025] Figure 5 is a block diagram showing an eye tracking unit of a computer system configured to capture a user's gaze input according to some embodiments.

[0026] Figure 6 is a flowchart showing a flash-assisted gaze tracking pipeline according to some embodiments.

[0027] 7A to 7J is a block diagram showing a user's interaction with a three-dimensional environment according to some embodiments.

[0028] Figure 7K to Figure 7P is a block diagram showing a method for facilitating a user to use a device in a physical environment to interact with a computer - generated three - dimensional environment according to some embodiments.

[0029] Figure 8 is a flowchart of a method for interacting with a three - dimensional environment according to some embodiments.

[0030] Fig. 9 is a flowchart of a method for interacting with a three - dimensional environment according to some embodiments.

[0031] Fig.10 is a flowchart of a method for interacting with a three - dimensional environment according to some embodiments.

[0032] Fig.11 is a flowchart of a method for interacting with a three - dimensional environment according to some embodiments.

[0033] Fig.12 is a flowchart of a method for facilitating a user to transition into and out of a three - dimensional environment according to some embodiments.

[0034] Fig.13 is a flowchart of a method for facilitating a user to transition into and out of a three - dimensional environment according to some embodiments. Detailed Description

[0035] According to some embodiments, the present disclosure relates to a user interface for providing a computer - generated reality (CGR) experience to a user.

[0036] The systems, methods, and GUIs described herein improve user - interface interactions with virtual / augmented reality environments in a variety of ways.

[0037] In some embodiments, a computer system allows a user to interact with a three - dimensional environment (e.g., a virtual or mixed - reality environment) using micro - gestures performed using small movements of a finger relative to other fingers or parts of the same hand. For example, as opposed to a touch - sensitive surface or other physical controller, a camera is used to detect the micro - gestures (e.g., a camera integrated with a head - mounted device or a camera mounted away from the user (e.g., in a CGR room)). Different movements and positions of the micro - gestures, as well as various movement parameters, are used to determine the operations performed in the three - dimensional environment. Using a camera to capture micro - gestures for interacting with the three - dimensional environment allows the user to move freely around the physical environment without being hindered by physical input equipment, which allows the user to explore the three - dimensional environment more naturally and effectively. Additionally, micro - gestures are discrete and unobtrusive, and are suitable for interactions that may occur in public settings and / or require decorum.

[0038] In some embodiments, a ready state configuration of a hand is defined. Detecting additional requirements of the hand at a location corresponding to a portion of the displayed three-dimensional environment ensures that the ready state configuration of the hand is not accidentally recognized by the computer system. The ready state configuration of the hand is used by the computer system as an indication that the user wants to interact with the computer system in a predefined operating context different from the currently displayed operating context. For example, the predefined operating context is one or more interactions with the device outside of the currently displayed application (e.g., game, communication session, media playback session, navigation, etc.). The predefined operating context is optionally a system interaction, such as displaying a main user interface or a start user interface from which other experiences and / or applications can be launched, a multitasking user interface from which recently displayed experiences and / or applications can be selected and restarted, and a control user interface for adjusting one or more device parameters of the computer system (e.g., brightness of the display, audio volume, network connection, etc.). Using a special gesture to trigger the display of a visual indication of a predefined operating context different from the currently displayed operating context for gesture input allows the user to easily access the predefined operating context without cluttering the three-dimensional environment with visual controls and without accidentally triggering an interaction in the predefined operating context.

[0039] In some embodiments, a physical object (e.g., a user's hand or a hardware device) or a portion of the physical object is selected by the user or the computer system to be associated with a system user interface (e.g., a control user interface for the device) that is not currently displayed in a three-dimensional environment (e.g., a mixed reality environment). When the user's gaze points to a location in the three-dimensional environment other than the location corresponding to the predefined physical object or a portion of the predefined physical object, a gesture performed by the user's hand causes an operation in the currently displayed context to be executed without causing the display of the system user interface; and when the user's gaze points to a location in the three-dimensional environment corresponding to the predefined physical object or an option of the predefined physical object, a gesture performed by the user's hand causes the display of the system user interface. Selectively performing an operation in the currently displayed operating context or displaying the system user interface in response to an input gesture based on whether the user's gaze points to a predefined physical object (e.g., the user's hand performing the gesture or the physical object the user wants to control using the gesture) allows the user to effectively interact with the three-dimensional environment in more than one context without visually cluttering the three-dimensional environment with multiple controls and improves the interaction efficiency of the user interface (e.g., reduces the amount of input required to achieve a desired result).

[0040] In some embodiments, a user's gaze directed at a virtual object in a three-dimensional environment in response to gesture input causes a visual indication of one or more interaction options available for the virtual object to be displayed only if the user's hand is also found to be in a predefined ready state for providing gesture input. If the user's hand is not found to be in a ready state for providing gesture input, the user's gaze directed at the virtual object does not trigger the display of the visual indication. Using a combination of the user's gaze and the ready state of the user's hand to determine whether to display a visual indication that the virtual object has associated interaction options for gesture input provides useful feedback to the user as the user explores the three-dimensional environment with his / her eyes, without unnecessarily bombarding the user with changing displays of the environment as the user shifts his / her gaze around the three-dimensional environment, thereby reducing user confusion when exploring the three-dimensional environment.

[0041] In some embodiments, when the display generation component of a computer system is placed in a predefined position relative to the user (e.g., placing a display in front of his / her eyes, or wearing a head-mounted device on his / her head), the user's view of the real world is blocked by the display generation component, and the content presented by the display generation component dominates the user's view. Sometimes, the user would benefit from a more gradual and controlled process for transitioning from the real world to a computer-generated experience. Thus, when presenting content to the user via the display generation component, the computer system displays a see-through portion (which includes a representation of at least a portion of the real world surrounding the user), and displays virtual content that replaces at least a portion of the see-through portion only in response to detecting a change in the user's hand grip on the housing of the display generation component. The change in the user's hand grip on the housing of the display generation component serves as an indication that the user is ready to transition into an experience that is more immersive than the experience currently presented by the display generation component. The staged transition into and out of the immersive environment, controlled by the user changing the hand grip on the housing of the display generation component, is intuitive and natural for the user, and improves the user's experience and comfort when using the computer system for a computer-generated immersive experience.

[0042] In some embodiments, when a computer system displays a virtual three-dimensional environment, the computer system applies a visual change to a portion of the virtual environment at a location corresponding to a portion of a physical object that is already within a threshold distance of a user and potentially within the user's field of view for the virtual environment (e.g., these portions of the physical object would be visible to the user if there were no display generation components blocking the user's view of the real world around the user). Further, rather than simply rendering all portions of the physical object that are potentially within the field, portions of the physical object that are not within the threshold distance of the user are not visually represented to the user (e.g., by changing the appearance of portions of the virtual environment corresponding to these portions of the physical object that are not within the threshold distance of the user). In some embodiments, the visual change applied to the portion of the virtual environment causes one or more physical characteristics of the portion of the physical object that is within the threshold distance of the user to be represented in the virtual environment, without completely stopping the display of those portions of the virtual environment or completely stopping providing the user with an immersive virtual experience. This technique allows a user to be alerted to physical obstacles approaching the user while the user is moving around in the physical environment while exploring an immersive virtual environment, without overly disturbing and disrupting the user's immersive virtual experience. Thus, a safer and smoother immersive virtual experience can be provided to the user.

[0043] Figures 1 to 6 A description of an exemplary computer system for providing a CGR experience to a user is provided. FIG. 7A to FIG. 7G An exemplary interaction with a three-dimensional environment using gesture input and / or gaze input is shown, according to some embodiments. Figures 7K to 7M An exemplary user interface displayed when a user transitions into and out of an interaction with a three-dimensional environment is shown, according to some embodiments. Figure 7N to Figure 7P An exemplary user interface displayed when a user moves around in a physical environment while interacting with a virtual environment is shown, according to some embodiments. Figures 8 to 11 A flowchart of a method for interacting with a three-dimensional environment, according to various embodiments. FIG. 7A to FIG. 7G The user interfaces in Figures 8 to 11 are respectively used to illustrate the Fig.12 A flowchart of a method that facilitates a user to use a computer system to interact with a three-dimensional environment, according to various embodiments. Figures 7K to 7M The user interfaces in Figure 12 to Figure 13 are respectively used to illustrate the

[0044] In some embodiments, as Figure 1As shown in [description], a CGR experience is provided to a user via an operating environment 100 that includes a computer system 101. The computer system 101 includes a controller 110 (e.g., a processor of a portable electronic device or a remote server), a display generation component 120 (e.g., a head-mounted device (HMD), a display, a projector, a touch screen, etc.), one or more input devices 125 (e.g., an eye tracking device 130, a hand tracking device 140, other input devices 150), one or more output devices 155 (e.g., a speaker 160, a haptic output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a haptic sensor, an orientation sensor, a proximity sensor, a temperature sensor, a position sensor, a motion sensor, a speed sensor, etc.), and optionally one or more peripheral devices 195 (e.g., home appliances, wearable devices, etc.). In some embodiments, one or more of the input device 125, the output device 155, the sensor 190, and the peripheral device 195 are integrated with the display generation component 120 (e.g., in a head-mounted device or a handheld device).

[0045] When describing a CGR experience, various terms are used to differently refer to several related but different environments that a user can sense and / or with which a user can interact (e.g., using inputs detected by the computer system 101 that generates the CGR experience, which inputs cause the computer system generating the CGR experience to generate audio, visual, and / or haptic feedback corresponding to the various inputs provided to the computer system 101). The following is a subset of these terms:

[0046] Physical environment: The physical environment refers to the physical world that people can sense and / or interact with without the help of an electronic system. Physical environments such as a physical park include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.

[0047] Computer-Generated Reality: In contrast, a computer-generated reality (CGR) environment is a fully or partially simulated environment in which people sense and / or interact via an electronic system. In CGR, a subset of a person's physical movements or their representations are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the CGR environment are adjusted in a manner consistent with at least one physical law. For example, a CGR system can detect a person's head rotation and, in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the characteristics of virtual objects in a CGR environment can be made in response to a representation of a physical movement (e.g., a voice command). A person can use any of their senses to sense and / or interact with CGR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of a point audio source in 3D space. As another example, an audio object can enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some CGR environments, a person can sense and / or interact only with audio objects.

[0048] Examples of CGR include virtual reality and mixed reality.

[0049] Virtual Reality: A virtual reality (VR) environment is a simulated environment that is designed to be completely computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects with which a person can sense and / or interact. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with the virtual objects in a VR environment by way of a simulation of the person's presence within the computer-generated environment and / or by way of a simulation of a subgroup of the person's physical movements within the computer-generated environment.

[0050] Mixed Reality: Compared to a VR environment that is designed to be based entirely on computer-generated sensory input, a mixed reality (MR) environment is an analog environment that is designed to include, in addition to computer-generated sensory input (e.g., virtual objects), sensory input or its representation from the physical environment. On the virtual continuum, an MR environment is any condition between a fully physical environment at one end and a virtual reality environment at the other end, but does not include these two ends. In some MR environments, the computer-generated sensory input can respond to changes in the sensory input from the physical environment. Additionally, some electronic systems for presenting an MR environment can track the position and / or orientation relative to the physical environment so that virtual objects can interact with real objects (i.e., physical items from the physical environment or their representations). For example, the system can cause movement so that a virtual tree appears stationary relative to the physical ground.

[0051] Examples of mixed reality include augmented reality and augmented virtuality.

[0052] Augmented Reality: An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are superimposed over a physical environment or a representation of a physical environment. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or translucent display such that the person, using the system, perceives the virtual objects superimposed over the physical environment. Alternatively, the system can have an opaque display and one or more imaging sensors that capture images or video of the physical environment, which are representations of the physical environment. The system combines the images or video with virtual objects and presents the combination on the opaque display. The person, using the system, indirectly views the physical environment via the images or video of the physical environment and perceives the virtual objects superimposed over the physical environment. As used herein, a video of a physical environment displayed on an opaque display is referred to as a "passthrough video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system can have a projection system that projects virtual objects into the physical environment, such as as a hologram or on a physical surface, such that the person, using the system, perceives the virtual objects superimposed over the physical environment. An augmented reality environment is also a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing a passthrough video, the system can transform one or more sensor images to impose an alternative perspective (e.g., viewpoint) different from the perspective captured by the imaging sensors. As another example, a representation of a physical environment can be transformed by graphically modifying (e.g., magnifying) portions thereof such that the modified portions can be a representative but not true version of the originally captured image. As yet another example, a representation of a physical environment can be transformed by graphically removing portions thereof or blurring portions thereof.

[0053] Augmented Virtuality: An augmented virtuality (AV) environment is a simulated environment in which a virtual environment or computer-generated environment incorporates one or more sensory inputs from a physical environment. The sensory inputs can be representations of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but a person's face is a realistic reproduction from an image of a physical person. As another example, a virtual object can adopt the shape or color of a physical item imaged by one or more imaging sensors. As yet another example, a virtual object can adopt a shadow that conforms to the positioning of the sun in the physical environment.

[0054] Hardware: There are many different types of electronic systems that enable a person to sense and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields integrated with display capabilities, windows integrated with display capabilities, displays formed as lenses designed to be placed on a person's eye (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smart phones, tablets, and desktop / laptop computers. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, the head-mounted system can be configured to accept an external opaque display (e.g., a smart phone). A head-mounted system can incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mounted system can have a transparent or translucent display instead of an opaque display. The transparent or translucent display can have a medium through which light representing an image is directed to a person's eye. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanned light sources, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display can be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphical image onto a person's retina. The projection system can also be configured to project virtual objects into the physical environment, for example as a hologram, or onto a physical surface. In some embodiments, controller 110 is configured to manage and coordinate a user's CGR experience. In some embodiments, controller 110 includes a suitable combination of software, firmware, and / or hardware. Reference is made below to Figure 2Controller 110 is described in more detail. In some embodiments, controller 110 is a computing device that is local or remote relative to scene 105 (e.g., a physical set / environment). For example, controller 110 is a local server located within scene 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside of scene 105. In some embodiments, controller 110 is communicatively coupled to display generation component 120 (e.g., an HMD, a display, a projector, a touch screen, etc.) via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE802.11x, IEEE802.16x, IEEE 802.3x, etc.). In another example, controller 110 is included within the housing (e.g., a physical enclosure) of one or more of display generation component 120 (e.g., an HMD or a portable electronic device including a display and one or more processors, etc.), one or more input devices of input device 125, one or more output devices of output device 155, one or more sensors of sensor 190, and / or one or more peripheral devices of peripheral device 195, or shares the same physical housing or support structure with one or more of the above devices.

[0055] In some embodiments, display generation component 120 is configured to provide a CGR experience (e.g., at least the visual component of the CGR experience) to a user. In some embodiments, display generation component 120 includes a suitable combination of software, firmware, and / or hardware. Display generation component 120 is described in more detail below with respect to Figure 3 In some embodiments, the functionality of controller 110 is provided by and / or combined with display generation component 120.

[0056] According to some embodiments, when a user is virtually and / or physically present within scene 105, display generation component 120 provides a CGR experience to the user.

[0057] In some embodiments, the display generation component is worn on a part of the user's body (e.g., on his / her head, on his / her hand, etc.). Accordingly, the display generation component 120 includes one or more CGR displays provided for displaying CGR content. For example, in various embodiments, the display generation component 120 surrounds the user's field of view. In some embodiments, the display generation component 120 is a handheld device (such as a smart phone or a tablet) configured to present CGR content, and the user holds the device having a display facing the user's field of view and a camera facing the scene 105. In some embodiments, the handheld device is optionally placed in a housing worn on the user's head. In some embodiments, the handheld device is optionally placed on a support (e.g., a tripod) in front of the user. In some embodiments, the display generation component 120 is a CGR chamber, housing, or room configured to present CGR content, where the user does not wear or hold the display generation component 120. Many user interfaces described with reference to one type of hardware for displaying CGR content (e.g., a handheld device or a device on a tripod) can be implemented on another type of hardware for displaying CGR content (e.g., an HMD or other wearable computing device). For example, a user interface showing an interaction with CGR content triggered based on an interaction occurring in the space in front of a handheld device or a tripod-mounted device can be similarly implemented with an HMD, where the interaction occurs in the space in front of the HMD and the response to the CGR content is displayed via the HMD. Similarly, a user interface showing an interaction with CRG content triggered based on the movement of a handheld device or a tripod-mounted device relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)) can be similarly implemented with an HMD, where the movement is caused by the movement of the HMD relative to the physical environment (e.g., the scene 105 or a part of the user's body (e.g., the user's eyes, head, or hand)).

[0058] Although relevant features of the operating environment 100 are shown in Figure 1 for the sake of brevity and to not obscure more relevant aspects of the exemplary embodiments disclosed herein, various other features are not shown.

[0059] Figure 2FIG. 0 is a block diagram of an example of controller 110 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein. For that purpose, by way of non-limiting example, in some embodiments, controller 110 includes one or more processing units 202 (e.g., microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, global system for mobile communications (GSM), code division multiple access (CDMA), time division multiple access (TDMA), global positioning system (GPS), infrared (IR), BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0060] In some embodiments, one or more communication buses 204 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0061] Memory 220 includes high-speed random access memory such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid state memory devices. In some embodiments, memory 220 includes non-volatile memory such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes non-transitory computer-readable storage media. In some embodiments, memory 220 or the non-transitory computer-readable storage media of memory 220 stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 230 and a CGR experience module 240.

[0062] The operating system 230 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, the CGR experience module 240 is configured to manage and coordinate single or multiple CGR experiences for one or more users (e.g., a single CGR experience for one or more users, or multiple CGR experiences for respective groups of one or more users). To this end, in various embodiments, the CGR experience module 240 includes a data acquisition unit 242, a tracking unit 244, a coordination unit 246, and a data transmission unit 248.

[0063] In some embodiments, the data acquisition unit 242 is configured to obtain data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the display generation component 120 of Figure 1 and optionally from one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For this purpose, in various embodiments, the data acquisition unit 242 includes instructions and / or logic for instructions and heuristics and metadata for heuristics.

[0064] In some embodiments, the tracking unit 244 is configured to map the scene 105 and track the position of at least the display generation component 120 relative to Figure 1 the scene 105, and optionally track the position of one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For this purpose, in various embodiments, the tracking unit 244 includes instructions and / or logic for instructions and heuristics and metadata for heuristics. In some embodiments, the tracking unit 244 includes a hand tracking unit 243 and / or an eye tracking unit 245. In some embodiments, the hand tracking unit 243 is configured to track the position of one or more parts of the user's hand and / or the movement of one or more parts of the user's hand relative to Figure 1 the scene 105, relative to the display generation component 120, and / or relative to a coordinate system (which is defined relative to the user's hand). The hand tracking unit 243 is described in more detail below with respect to Figure 4 . In some embodiments, the eye tracking unit 245 is configured to track the position or movement of the user's gaze (or more generally, the user's eyes, face, or head) relative to the scene 105 (e.g., relative to the physical environment and / or relative to the user (e.g., the user's hand)) or relative to the CGR content displayed via the display generation component 120. The eye tracking unit 245 is described in more detail below with respect to Figure 5 .

[0065] In some embodiments, the coordination unit 246 is configured to manage and coordinate the CGR experience presented to the user by the display generation component 120, and optionally by one or more of the output device 155 and / or the peripheral device 195. For this purpose, in various embodiments, the coordination unit 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0066] In some embodiments, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the display generation component 120, and optionally to one or more of the input device 125, the output device 155, the sensor 190, and / or the peripheral device 195. For this purpose, in various embodiments, the data transmission unit 248 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0067] Although the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 are shown as residing on a single device (e.g., the controller 110), it should be understood that in other embodiments, any combination of the data acquisition unit 242, the tracking unit 244 (e.g., including the eye tracking unit 243 and the hand tracking unit 244), the coordination unit 246, and the data transmission unit 248 may be located in separate computing devices.

[0068] Furthermore, Figure 2 It is intended to be more of a functional description of the various features that may be present in a particular implementation than a schematic diagram of the embodiments described herein. As will be recognized by those of ordinary skill in the art, items shown separately may be combined, and some items may be separated. For example, Figure 2 Some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions, as well as how the features are allocated therein, will vary depending on the implementation, and in some embodiments, will depend in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.

[0069] Figure 3FIG. is a block diagram of an example of a display generation component 120 according to some embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that, for the sake of brevity and to not obscure more relevant aspects of the embodiments disclosed herein, various other features are not shown. For that purpose, by way of non-limiting example, in some embodiments, the HMD 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more CGR displays 312, one or more optional internal- and / or external-facing image sensors 314, a memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0070] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time-of-flight, etc.), and the like.

[0071] In some embodiments, one or more CGR displays 312 are configured to provide a CGR experience to a user. In some embodiments, one or more CGR displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conduction electron-emitter display (SED), field-emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more CGR displays 312 correspond to diffractive, reflective, polarization, holographic, and other waveguide displays. For example, HMD 120 includes a single CGR display. As another example, HMD 120 includes CGR displays for each eye of the user. In some embodiments, one or more CGR displays 312 are capable of presenting MR and VR content. In some embodiments, one or more CGR displays 312 are capable of presenting AR or VR content.

[0072] In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's face including the user's eyes (and may be referred to as an eye-tracking camera). In some embodiments, one or more image sensors 314 are configured to acquire image data corresponding to at least a portion of the user's hand and optionally the user's arm (and may be referred to as a hand-tracking camera). In some embodiments, one or more image sensors 314 are configured to face forward to acquire image data corresponding to the scene that the user would see in the absence of HMD 120 (and may be referred to as a scene camera). One or more optional image sensors 314 may include one or more RGB cameras (e.g., having a complementary metal-oxide semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), one or more infrared (IR) cameras, and / or one or more event-based cameras, etc.

[0073] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores the following programs, modules, and data structures, or subsets thereof, including optional operating system 330 and CGR rendering module 340.

[0074] Operating system 330 includes instructions for handling various basic system services and for performing hardware-related tasks. In some embodiments, CGR rendering module 340 is configured to present CGR content to a user via one or more CGR displays 312. For this purpose, in various embodiments, CGR rendering module 340 includes data acquisition unit 342, CGR rendering unit 344, CGR mapping generation unit 346, and data transmission unit 348.

[0075] In some embodiments, data acquisition unit 342 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) at least from Figure 1 controller 110 of. For this purpose, in various embodiments, data acquisition unit 342 includes instructions and / or logic for these instructions and heuristics and metadata for these heuristics.

[0076] In some embodiments, CGR rendering unit 344 is configured to present CGR content via one or more CGR displays 312. For this purpose, in various embodiments, CGR rendering unit 344 includes instructions and / or logic for these instructions and heuristics and metadata for these heuristics.

[0077] In some embodiments, CGR mapping generation unit 346 is configured to generate a CGR map based on media content data (e.g., a 3D map of a mixed reality scene or a map of a physical environment in which computer-generated objects can be placed to generate computer-generated reality). For this purpose, in various embodiments, CGR mapping generation unit 346 includes instructions and / or logic for these instructions and heuristics and metadata for these heuristics.

[0078] In some embodiments, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the controller 110 and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various embodiments, the data transmission unit 348 includes instructions and / or logic for these instructions and heuristics and metadata for these heuristics.

[0079] Although the data acquisition unit 342, CGR presentation unit 344, CGR mapping generation unit 346, and data transmission unit 348 are shown as residing in a single device (e.g., Figure 1 the display generation component 120 of ), it should be understood that in other embodiments, any combination of the data acquisition unit 342, CGR presentation unit 344, CGR mapping generation unit 346, and data transmission unit 348 may be located in separate computing devices.

[0080] In addition, Figure 3 is intended more as a functional description of the various features that may be present in a particular embodiment rather than a structural schematic of the embodiments described herein. As will be recognized by those of ordinary skill in the art, items shown separately may be combined and some items may be separated. For example, Figure 3 some of the functional modules shown separately in may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various embodiments. The actual number of modules and the specific division of functions and how the features are allocated therein will vary according to the specific implementation and, in some embodiments, will depend in part on the particular combination of hardware, software, and / or firmware selected for the specific implementation.

[0081] Figure 4 is a schematic illustration of an exemplary embodiment of the hand tracking device 140. In some embodiments, the hand tracking device 140 ( Figure 1 ) is controlled by the hand tracking unit 243 ( Figure 2 ) to track the position of one or more parts of the user's hand and / or one or more parts of the user's hand relative to Figure 1Movement of the scene 105 (e.g., relative to a portion of the physical environment around the user, relative to the display generation component 120, or relative to a portion of the user (e.g., the user's face, eyes, or head), and / or relative to a coordinate system that is defined relative to the user's hand). In some embodiments, the hand tracking device 140 is part of the display generation component 120 (e.g., embedded in or attached to a head-mounted device). In some embodiments, the hand tracking device 140 is separate from the display generation component 120 (e.g., located in a separate housing or attached to a separate physical support structure).

[0082] In some embodiments, the hand tracking device 140 includes an image sensor 404 that captures three-dimensional scene information including at least the hand 406 of a human user (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras, etc.). The image sensor 404 captures hand images at a sufficient resolution to enable the fingers and their corresponding positions to be distinguished. The image sensor 404 typically captures images of other parts of the user's body, and may also or possibly capture images of all parts of the body, and may have zoom capabilities or a dedicated sensor with increased magnification to capture images of the hand at a desired resolution. In some embodiments, the image sensor 404 also captures 2D color video images of the hand 406 and other elements of the scene. In some embodiments, the image sensor 404 is used in combination with other image sensors to capture the physical environment of the scene 105, or serves as the image sensor for capturing the physical environment of the scene 105. In some embodiments, the image sensor 404 or a portion of its field of view is positioned relative to the user or the user's environment in a manner that defines an interaction space in which hand movements captured by the image sensor are treated as inputs to the controller 110.

[0083] In some embodiments, the image sensor 404 outputs a sequence of frames containing 3D mapping data (and in addition, possibly color image data) to the controller 110, which extracts high-level information from the mapping data. This high-level information is typically provided to an application running on the controller via an application programming interface (API), and the application drives the display generation component 120 accordingly. For example, a user can interact with software running on the controller 110 by moving his hand 408 and changing his hand posture.

[0084] In some embodiments, the image sensor 404 projects a speckle pattern onto a scene that includes the hand 406 and captures an image of the projected pattern. In some embodiments, the controller 110 calculates the 3D coordinates of points in the scene (including points on the surface of the user's hand) by triangulation based on the lateral offset of the speckles in the pattern. This method is advantageous because it does not require the user to hold or wear any kind of beacon, sensor, or other marker. The method gives the depth coordinates of points in the scene at a particular distance from the image sensor 404 relative to a pre-determined reference plane. In the present disclosure, it is assumed that the image sensor 404 defines an orthogonal set of x, y, and z axes such that the depth coordinate of a point in the scene corresponds to the z component measured by the image sensor. Alternatively, the hand tracking device 440 may use other 3D mapping methods, such as stereoscopic imaging or time-of-flight measurement, based on a single or multiple cameras or other types of sensors.

[0085] In some embodiments, the hand tracking device 140 captures and processes a time series of depth maps that include the user's hand as the user moves his hand (e.g., the entire hand or one or more fingers). Software running on a processor in the image sensor 404 and / or the controller 110 processes the 3D mapping data to extract image patch descriptors of the hand in these depth maps. The software may match these descriptors with image patch descriptors stored in the database 408 based on a previous learning process in order to estimate the pose of the hand in each frame. The pose generally includes the 3D positions of the user's hand joints and finger tips.

[0086] The software may also analyze the trajectories of the hand and / or fingers over multiple frames in the sequence to identify gestures. The pose estimation function described herein may alternate with the motion tracking function such that the image patch-based pose estimation is performed only once every two (or more) frames, while tracking is used to find the changes in pose that occur in the remaining frames. Pose, motion, and gesture information is provided to an application running on the controller 110 via the aforementioned API. The program may, for example, move and modify the image presented on the display generation component 120 in response to the pose and / or gesture information, or perform other functions.

[0087] In some embodiments, the software may be downloaded electronically to the controller 110, for example, via a network, or may alternatively be provided on a tangible non-transitory medium such as an optical, magnetic, or electronic memory medium. In some embodiments, the database 408 is likewise stored in a memory associated with the controller 110. Alternatively or in addition, some or all of the described functions of the computer may be implemented in dedicated hardware, such as a custom or semi-custom integrated circuit or a programmable digital signal processor (DSP). Although Figure 4Controller 110 is shown, but by way of example, some or all of the processing functions of the controller, as a unit separate from image sensor 440, may be performed by a suitable microprocessor and software or by dedicated circuitry within the housing of hand tracking device 402 or other devices associated with image sensor 404. In some embodiments, at least some of these processing functions may be performed by a suitable processor integrated with display generation component 120 (e.g., in a television receiver, a handheld device, or a head-mounted device) or integrated with any other suitable computerized device such as a game console or a media player. The sensing function of image sensor 404 may likewise be integrated into a computer or other computerized device to be controlled by the sensor output.

[0088] Figure 4 Also shown schematically is a depth map 410 captured by image sensor 404 according to some embodiments. As described above, the depth map includes a matrix of pixels having corresponding depth values. Pixels 412 corresponding to hand 406 have been segmented from the background and wrist in the map. The brightness of each pixel within depth map 410 is inversely proportional to its depth value (i.e., the measured z-distance from image sensor 404), where the shade of gray becomes darker with increasing depth. Controller 110 processes these depth values to identify and segment the components of the image having human hand characteristics (i.e., a group of adjacent pixels). These characteristics may include, for example, overall size, shape, and motion from frame to frame in a sequence of depth maps.

[0089] Figure 4 Also schematically shown is a hand skeleton 414 that controller 110 ultimately extracts from depth map 410 of hand 406 according to some embodiments. In Figure 4 this, skeleton 414 is superimposed on hand background 416 that has been segmented from the original depth map. In some embodiments, key feature points on the hand and optionally on the wrist or arm connected to the hand (e.g., points corresponding to knuckles, finger tips, palm center, the end of the hand connected to the wrist, etc.) are identified and located on hand skeleton 414. In some embodiments, controller 110 uses the positions and movements of these key feature points across multiple image frames to determine, according to some embodiments, the gesture being performed by the hand or the current state of the hand.

[0090] Figure 5 An exemplary embodiment of an eye tracking device 130 ( Figure 1 ) is shown. In some embodiments, eye tracking device 130 consists of eye tracking unit 245 ( Figure 2)Control is used to track the position and movement of the user's gaze relative to the scene 105 or relative to the CGR content displayed via the display generation component 120. In some embodiments, the eye tracking device 130 is integrated with the display generation component 120. For example, in some embodiments, when the display generation component 120 is a head-mounted device (such as, a head-mounted headset, helmet, goggles, or glasses) or a hand-held device placed in a wearable frame, the head-mounted device includes both components for generating CGR content for the user to view and components for tracking the user's gaze relative to the CGR content. In some embodiments, the eye tracking device 130 is separate from the display generation component 120. For example, when the display generation component is a hand-held device or a CGR room, the eye tracking device 130 is optionally a device separate from the hand-held device or the CGR room. In some embodiments, the eye tracking device 130 is a head-mounted device or part of a head-mounted device. In some embodiments, the head-mounted eye tracking device 130 is optionally used in combination with a display generation component that is also head-mounted or a display generation component that is not head-mounted. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally used in combination with a head-mounted display generation component. In some embodiments, the eye tracking device 130 is not a head-mounted device and is optionally part of a non-head-mounted display generation component.

[0091] In some embodiments, the display generation component 120 uses a display mechanism (e.g., a left near-eye display panel and a right near-eye display panel) to display a frame including a left image and a right image in front of the user's eyes, thereby providing the user with a 3D virtual view. For example, a head-mounted display generation component may include a left optical lens and a right optical lens (referred to herein as eye lenses) located between the display and the user's eyes. In some embodiments, the display generation component may include or be coupled to one or more external cameras that capture video of the user's environment for display. In some embodiments, a head-mounted display generation component may have a transparent or semi-transparent display, and virtual objects are displayed on the transparent or semi-transparent display, through which the user can directly view the physical environment. In some embodiments, the display generation component projects virtual objects into the physical environment. The virtual objects may be projected, for example, onto a physical surface or projected as a hologram such that an individual uses the system to observe the virtual objects superimposed over the physical environment. In this case, separate display panels and image frames for the left and right eyes may not be required.

[0092] As Figure 5As shown, in some embodiments, the gaze tracking device 130 includes at least one eye tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera), and an illumination source (e.g., an IR or NIR light source, such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye tracking camera may be pointed at the user's eyes to receive the IR or NIR light directly reflected from the eyes by the light source, or alternatively may be pointed at a "hot" mirror located between the user's eyes and the display panel, which reflects the IR or NIR light from the eyes to the eye tracking camera while allowing visible light to pass through. The gaze tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60 to 120 frames per second (fps)), analyzes the images to generate gaze tracking information, and transmits the gaze tracking information to the controller 110. In some embodiments, both of the user's eyes are tracked separately by corresponding eye tracking cameras and illumination sources. In some embodiments, only one of the user's eyes is tracked by corresponding eye tracking cameras and illumination sources.

[0093] In some embodiments, a device-specific calibration process is used to calibrate the eye tracking device 130 to determine the parameters of the eye tracking device for a particular operating environment 100, such as the 3D geometric relationships and parameters of the LEDs, cameras, hot mirrors (if any), eye lenses, and display screens. The device-specific calibration process can be performed at the factory or another facility before the AR / VR equipment is delivered to the end user. The device-specific calibration process can be an automatic calibration process or a manual calibration process. According to some embodiments, the user-specific calibration process may include an estimation of the eye parameters of a particular user, such as pupil position, fovea position, optical axis, visual axis, eye spacing, etc. According to some embodiments, once the device-specific parameters and user-specific parameters are determined for the eye tracking device 130, a flash-assisted method can be used to process the images captured by the eye tracking camera to determine the current visual axis and the user's fixation point relative to the display.

[0094] As Figure 5As shown, the eye tracking device 130 (e.g., 130A or 130B) includes an eye lens 520 and a gaze tracking system that includes at least one eye tracking camera 540 (e.g., an infrared (IR) or near-infrared (NIR) camera) positioned on the side of the user's face where eye tracking is to be performed, and an illumination source 530 (e.g., an IR or NIR light source, such as an array or ring of NIR light-emitting diodes (LEDs)) that emits light (e.g., IR or NIR light) toward the user's eye 592. The eye tracking camera 540 may be pointed at a mirror 550 (which reflects IR or NIR light from the eye 592 while allowing visible light to pass through) (e.g., as shown in the top portion of Figure 5 ), or alternatively may be pointed at the user's eye 592 to receive reflected IR or NIR light from the eye 592 (e.g., as shown in the bottom portion of Figure 5 ).

[0095] In some embodiments, the controller 110 renders AR or VR frames 562 (e.g., left and right frames for a left display panel and a right display panel) and provides the frames 562 to the display 510. The controller 110 uses the gaze tracking input 542 from the eye tracking camera 540 for various purposes, such as for processing the frames 562 for display. The controller 110 optionally estimates the user's gaze point on the display 510 based on the gaze tracking input 542 obtained using a flash assist method or other suitable method. The gaze point estimated based on the gaze tracking input 542 is optionally used to determine the direction in which the user is currently looking.

[0096] The following describes several possible use cases of the current gaze direction of a user and is not intended to be limiting. As an exemplary use case, the controller 110 may render virtual content differently based on the determined direction of the user's gaze. For example, the controller 110 may generate virtual content at a higher resolution in the foveal region determined according to the user's current gaze direction than in the peripheral region. As another example, the controller may position or move virtual content in the view at least in part based on the user's current gaze direction. As another example, the controller may display specific virtual content in the view at least in part based on the user's current gaze direction. As another exemplary use case in an AR application, the controller 110 may direct an external camera for capturing the physical environment for a CGR experience to focus in the determined direction. Then, the autofocus mechanism of the external camera may focus on an object or surface in the environment that the user is currently looking at on the display 510. As another exemplary use case, the eye lens 520 may be a focusable lens, and the controller uses the gaze tracking information to adjust the focus of the eye lens 520 such that the virtual object that the user is currently looking at has an appropriate vergence to match the convergence of the user's eyes 592. The controller 110 may utilize the gaze tracking information to direct the eye lens 520 to adjust the focus such that a nearby object that the user is looking at appears at the correct distance.

[0097] In some embodiments, the eye tracking device is part of a head-mounted device that includes a display (e.g., display 510) mounted in a wearable housing, two eye lenses (e.g., eye lenses 520), an eye tracking camera (e.g., eye tracking camera 540), and a light source (e.g., light source 530 (e.g., IR or NIR LED)). The light source emits light (e.g., IR or NIR light) towards the user's eyes 592. In some embodiments, the light source may be arranged in a ring or circle around each of the lenses, as Figure 5 shown. In some embodiments, for example, eight light sources 530 (e.g., LEDs) are arranged around each lens 520. However, more or fewer light sources 530 may be used, and other arrangements and positions of the light sources 530 may be used.

[0098] In some embodiments, the display 510 emits light in the visible light range and does not emit light in the IR or NIR range, and thus does not introduce noise into the gaze tracking system. Note that the position and angle of the eye tracking camera 540 are given by way of example and are not intended to be limiting. In some embodiments, a single eye tracking camera 540 is located on each side of the user's face. In some embodiments, two or more NIR cameras 540 may be used on each side of the user's face. In some embodiments, cameras 540 with a wider field of view (FOV) and cameras 540 with a narrower FOV may be used on each side of the user's face. In some embodiments, cameras 540 operating at one wavelength (e.g., 850 nm) and cameras 540 operating at a different wavelength (e.g., 940 nm) may be used on each side of the user's face.

[0099] As Figure 5 shown, embodiments of the gaze tracking system may be used, for example, in computer-generated reality (e.g., including virtual reality and / or mixed reality) applications to provide a computer-generated reality (e.g., including virtual reality, augmented reality, and / or augmented virtual) experience to a user.

[0100] Figure 6 A flash-assisted gaze tracking pipeline according to some embodiments is shown. In some embodiments, the gaze tracking pipeline is implemented by a flash-assisted gaze tracking system (e.g., an eye tracking device 130 as Figure 1 and Figure 5 shown). The flash-assisted gaze tracking system can maintain a tracking state. Initially, the tracking state is off or "no". When in the tracking state, when analyzing the current frame to track the pupil contour and flash in the current frame, the flash-assisted gaze tracking system uses previous information from the previous frame. When not in the tracking state, the flash-assisted gaze tracking system attempts to detect the pupil and flash in the current frame, and if successful, initializes the tracking state to "yes" and continues with the next frame in the tracking state.

[0101] As Figure 6 shown, the gaze tracking camera can capture left and right images of the user's left and right eyes. The captured images are then input into the gaze tracking pipeline for processing to begin at 610. As indicated by the arrow returning to element 600, the gaze tracking system can continue to capture images of the user's eyes, for example, at a rate of 60 to 120 frames per second. In some embodiments, each set of captured images can be input into the pipeline for processing. However, in some embodiments or under some conditions, not all captured frames are processed by the pipeline.

[0102] At 610, for the currently captured image, if the tracking status is yes, the method proceeds to element 640. At 610, if the tracking status is no, then as indicated at 620, the image is analyzed to detect the user's pupil and flash in the image. At 630, if the pupil and flash are successfully detected, the method proceeds to element 640. Otherwise, the method returns to element 610 to process the next image of the user's eye.

[0103] At 640, if proceeding from element 410, the current frame is analyzed to track the pupil and flash based in part on previous information from a previous frame. At 640, if proceeding from element 630, the tracking status is initialized based on the pupil and flash detected in the current frame. The processing result at element 640 is checked to verify that the result of the tracking or detection can be trusted. For example, the result can be checked to determine whether the pupil and a sufficient number of flashes for performing a gaze estimate are successfully tracked or detected in the current frame. At 650, if the result cannot be trusted, the tracking status is set to no, and the method returns to element 610 to process the next image of the user's eye. At 650, if the result is trusted, the method proceeds to element 670. At 670, the tracking status is set to yes (if it is not already yes), and the pupil and flash information is passed to element 680 to estimate the user's gaze point.

[0104] Figure 6 It is intended to be used as an example of an eye tracking technique that can be used for a particular specific implementation. As will be recognized by those of ordinary skill in the art, according to various embodiments, in the computer system 101 for providing a CGR experience to a user, other eye tracking techniques that currently exist or are developed in the future can be used to replace or be used in combination with the flash-assisted eye tracking technique described herein.

[0105] In this disclosure, various input methods are described in relation to interaction with a computer system. When one input device or input method is used to provide an example and another input device or input method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the input device or input method described in relation to the other example. Similarly, various output methods are described in relation to interaction with a computer system. When one output device or output method is used to provide an example and another output device or output method is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the output device or output method described in relation to the other example. Similarly, various methods are described in relation to interaction with a virtual environment or a mixed reality environment via a computer system. When interaction with a virtual environment is used to provide an example and a mixed reality environment is used to provide another example, it should be understood that each example can be compatible with and optionally utilize the methods described in relation to the other example. Accordingly, this disclosure discloses embodiments that are combinations of features of multiple examples without exhaustively listing all features of the embodiments in the description of each exemplary embodiment.

[0106] User interface and associated processes

[0107] Attention is now turned to embodiments of a user interface (“UI”) and associated processes that can be implemented on a computer system (such as a portable multifunctional device or a head-mounted device) having a display generation component, one or more input devices, and optionally one or more cameras.

[0108] 7A to 7C Examples of input gestures for interacting with a virtual or mixed reality environment are shown (e.g., discrete small movement gestures that are performed by moving a user's finger relative to other fingers or parts of the user's hand, optionally without the need to primarily move the user's entire hand or arm from its natural position and posture before or during the gesture to immediately perform an operation). Regarding 7A to 7C the input gestures described are used to illustrate the processes described below, including Figure 8 the processes in

[0109] In some embodiments, regarding 7A to 7C the input gestures described are analyzed by a sensor system (e.g., sensor 190, Figure 1 ; image sensor 314, Figure 3)Detected by the captured data or signals. In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras, such as a motion RGB camera, an infrared camera, a depth camera, etc.). For example, the one or more imaging sensors are components of a computer system (e.g., Figure 1 the computer system 101 in Figure 7C such as the portable electronic device 7100 or the HMD shown in Figure 1 , Figure 3 and Figure 4 the display generation component 120 in Figure 7CWhen the hand is on the hand (7200), an image of the hand under illumination by the light is captured by the one or more cameras, and the captured image is analyzed to determine the position and / or configuration of the hand. Using signals from an image sensor directed at the hand to determine an input gesture, rather than signals from a touch-sensitive surface or other direct contact mechanism or proximity-based mechanism, allows the user to freely choose whether to perform large movements or remain relatively stationary when providing an input gesture with his / her hand, without being subject to the limitations imposed by a particular input device or input area.

[0110] Fig. 7A Part (A) of shows a tap input where the thumb 7106 is above the index finger 7108 of the user's hand (e.g., above the side of the index finger 7108 adjacent to the thumb 7106). The thumb 7106 moves along the axis shown by arrow 7110, including moving from the raised position 7102 to the depressed position 7104 (e.g., where the thumb 7106 has contacted and remains resting on the index finger 7110), and optionally, after the thumb 7106 contacts the index finger 7110, moving again from the depressed position 7104 to the raised position 7102 within a threshold amount of time (e.g., a tap time threshold). In some embodiments, the tap input is detected without the need to lift the thumb off this side of the index finger. In some embodiments, the tap input is detected based on determining that an upward movement of the thumb follows a downward movement of the thumb, where the thumb contacts this side of the index finger for less than a threshold amount of time. In some embodiments, a tap-and-hold input is detected based on determining that the thumb moves from the raised position 7102 to the depressed position 7104 and remains in the depressed position 7104 for at least a first threshold amount of time (e.g., a tap time threshold or another time threshold longer than the tap time threshold). In some embodiments, the computer system requires the hand as a whole to remain substantially stationary in position for at least a first threshold amount of time in order to detect a tap-and-hold input of the thumb on the index finger. In some embodiments, a touch-and-hold input is detected without the need for the hand as a whole to remain substantially stationary (e.g., the hand as a whole can move while the thumb is resting on this side of the index finger). In some embodiments, a tap-and-hold-and-drag input is detected when the thumb is depressing this side of the index finger and the hand as a whole is moving while the thumb is resting on this side of the index finger.

[0111] Fig. 7APortion (B) shows a push or flick input performed by moving the thumb 7116 across the index finger 7118 (e.g., from the palmar side to the dorsal side of the index finger). The thumb 7116 moves from a retracted position 7112 to an extended position 7114 across the index finger 7118 (e.g., across the middle phalanx of the index finger 7118) along the axis shown by arrow 7120. In some embodiments, the extension movement of the thumb is accompanied by an upward movement away from that side of the index finger, e.g., as in an upward flick input performed by the thumb. In some embodiments, during the forward and upward movement of the thumb, the index finger moves in a direction opposite to that of the thumb. In some embodiments, a reverse flick input is performed by moving the thumb from the extended position 7114 to the retracted position 7112. In some embodiments, during the backward and downward movement of the thumb, the index finger moves in a direction opposite to that of the thumb.

[0112] Fig. 7A Portion (C) shows a swipe input performed by moving the thumb 7126 along the index finger 7128 (e.g., along the side of the index finger 7128 adjacent to or on the palmar side of the thumb 7126). The thumb 7126 moves along the axis shown by arrow 7130, along the length of the index finger 7128 from the proximal position 7122 of the index finger 7128 (e.g., at or near the proximal phalanx of the index finger 7118) to the distal position 7124 (e.g., at or near the distal phalanx of the index finger 7118) and / or from the distal position 7124 to the proximal position 7122. In some embodiments, the index finger is optionally in an extended state (e.g., substantially straight) or a curled state. In some embodiments, during the movement of the thumb in the swipe input gesture, the index finger changes between the extended state and the curled state.

[0113] Fig. 7A Portion (D) shows a tap input of the thumb 7106 above various phalanges of various fingers (e.g., the index finger, middle finger, ring finger, and optionally, the little finger). For example, as shown in portion (A), the thumb 7106 moves from a raised position 7102 to a pressed position, as shown at any of 7130 to 7148 shown in portion (D). In the pressed position 7130, the thumb 7106 is shown contacting the position 7150 on the proximal phalanx of the index finger 7108. In the pressed position 7134, the thumb 7106 contacts the position 7152 on the middle phalanx of the index finger 7108. In the pressed position 7136, the thumb 7106 contacts the position 7154 on the distal phalanx of the index finger 7108.

[0114] In the pressed positions shown at 7138, 7140, and 7142, the thumb 7106 contacts the positions 7156, 7158, and 7160 corresponding to the proximal phalanx, middle phalanx, and distal phalanx of the middle finger, respectively.

[0115] Among the touch positions shown at 7144, 7146, and 7150, the thumb 7106 contacts positions 7162, 7164, and 7166 corresponding to the proximal phalanx of the ring finger, the middle phalanx of the ring finger, and the distal phalanx of the ring finger, respectively.

[0116] In various embodiments, tap inputs performed by the thumb 7106 on different parts of another finger or on different parts of two adjacent fingers correspond to different inputs and trigger different operations in the corresponding user interface context. Similarly, in some embodiments, different push or click inputs can be performed by the thumb across different fingers and / or different parts of a finger to trigger different operations in the corresponding user interface context. Similarly, in some embodiments, different swipe inputs performed by the thumb along different fingers and / or in different directions (e.g., towards the distal or proximal end of a finger) trigger different operations in the corresponding user interface context.

[0117] In some embodiments, the computer system treats tap inputs, flick inputs, and swipe inputs as different types of inputs based on the type of movement of the thumb. In some embodiments, the computer system treats inputs with different finger positions tapped, touched, or swiped by the thumb as different sub-input types of a given input type (e.g., tap input type, flick input type, swipe input type, etc.) (e.g., proximal, middle, distal subtypes, or index finger, middle finger, ring finger, or little finger subtypes). In some embodiments, the amount of movement performed by moving a finger (e.g., the thumb) and / or other movement metrics associated with the movement of the finger (e.g., speed, initial speed, end speed, duration, direction, movement pattern, etc.) are used to quantitatively affect the operations triggered by the finger input.

[0118] In some embodiments, the computer system identifies combined input types that combine a series of movements performed by the thumb, such as tap-swipe inputs (e.g., a touch of the thumb on a finger followed by a swipe along that side of the finger), tap-flick inputs (e.g., a touch of the thumb above a finger followed immediately by a flick across the finger from the palm side to the dorsal side of the finger), double-tap inputs (e.g., two consecutive taps on the side of a finger at approximately the same position), etc.

[0119] In some embodiments, gesture input is performed by the index finger rather than the thumb (e.g., the index finger performs a tap or a swipe over the thumb, or the thumb and index finger move towards each other to perform a pinch gesture, etc.). In some embodiments, a wrist movement (e.g., a flick of the wrist in a horizontal or vertical direction) is performed immediately before, immediately after (e.g., within a threshold amount of time), or simultaneously with a finger movement input, compared to a finger movement input that does not have a modified input by a wrist movement, to trigger an additional operation, a different operation, or a modified operation in the current user interface context. In some embodiments, a finger input gesture performed with a user's palm facing the user's face is considered a different type of gesture than a finger input gesture performed with a user's palm facing away from the user's face. For example, compared to an operation (e.g., the same operation) performed in response to a tap gesture performed with a user's palm facing away from the user's face, the operation performed by a tap gesture performed with a user's palm facing the user has increased (or decreased) privacy protection.

[0120] Although in the examples provided in this disclosure, one type of finger input may be used to trigger a certain type of operation, in other embodiments, other types of finger inputs are optionally used to trigger the same type of operation.

[0121] Figure 7B An exemplary user interface context is shown that shows a menu 7170 including user interface objects 7172 to 7194 in some embodiments.

[0122] In some embodiments, the menu 7170 is displayed in a mixed reality environment (e.g., floating in the air or over a physical object in a three-dimensional environment, and corresponding to an operation associated with the mixed reality environment or an operation associated with the physical object). For example, the menu 7170 is displayed by the device (e.g., device 7100( Figure 7C) or the display of the HMD) (e.g., the menu is displayed as overlaying at least a portion thereof). In some embodiments, menu 7170 is displayed on a transparent or translucent display of the device (e.g., a head-up display or HMD), and the physical environment is visible through the transparent or translucent display. In some embodiments, menu 7170 is displayed in a user interface including a see-through portion surrounded by virtual content (e.g., a transparent or translucent portion through which the physical surrounding environment is visible, or a portion of a camera view showing the surrounding physical environment). In some embodiments, the hand of the user performing gesture input that causes an operation to be performed in a mixed reality environment is visible to the user on the display of the device. In some embodiments, the hand of the user performing gesture input that causes an operation to be performed in a mixed reality environment is not visible to the user on the display of the device (e.g., the camera providing a view of the physical world has a different field of view than the camera capturing the user's finger input).

[0123] In some embodiments, menu 7170 is displayed in a virtual reality environment (e.g., hovering in virtual space, or overlaying a virtual surface). In some embodiments, hand 7200 is visible in the virtual reality environment (e.g., an image of hand 7200 captured by one or more cameras is rendered in a virtual reality scene). In some embodiments, a representation of hand 7200 (e.g., a cartoon version of hand 7200) is rendered in a virtual reality scene. In some embodiments, hand 7200 is not visible in the virtual reality environment (e.g., omitted from the virtual reality environment). In some embodiments, device 7100( Figure 7C ) is not visible in the virtual reality environment (e.g., when device 7100 is an HMD). In some embodiments, an image of device 7100 or a representation of device 7100 is visible in the virtual reality environment.

[0124] In some embodiments, one or more of user interface objects 7172 to 7194 are application launch icons (e.g., for performing an operation to launch the corresponding application). In some embodiments, one or more of user interface objects 7172 to 7194 are controls for performing corresponding operations within an application (e.g., increasing volume, decreasing volume, playing, pausing, fast-forwarding, rewinding, initiating communication with a remote device, terminating communication with a remote device, transmitting communication to a remote device, starting a game, etc.). In some embodiments, one or more of user interface objects 7172 to 7194 are corresponding representations (e.g., avatars) of users of a remote device (e.g., for performing an operation to initiate communication with the corresponding user of the remote device). In some embodiments, one or more of user interface objects 7172 to 7194 are representations (e.g., thumbnails, two-dimensional images, or album covers) of media items (e.g., images, virtual objects, audio files, and / or video files). For example, activating a user interface object that is a representation of an image causes the image (e.g., at a location corresponding to a surface detected by one or more cameras) to be displayed and to be displayed in a computer-generated reality view (e.g., at a location corresponding to a surface in the physical environment or at a location corresponding to a surface displayed in a virtual space).

[0125] When the thumb of the hand 7200 performs the input gesture regarding Fig. 7A as described, an operation corresponding to the menu 7170 is performed according to the position and / or type of the detected input gesture. For example, in response to an input including the thumb moving along the y-axis (e.g., a movement from the proximal position on the index finger to the distal position on the index finger, as described in part (C) of Fig. 7A ), the current selection indicator 7198 (e.g., a selector object or a movable visual effect, such as highlighting the object by a change in the outline or appearance of the object) is iterated from the item 7190 to the subsequent user interface object 7192 to the right. In some embodiments, in response to an input including the thumb moving along the y-axis from the distal position on the index finger to the proximal position on the index finger, the current selection indicator 7198 is iterated from the item 7190 to the previous user interface object 7188 to the left. In some embodiments, in response to an input including a tap of the thumb above the index finger (e.g., a movement of the thumb along the z-axis, as described in Fig. 7AThe input as described in part (A) activates the currently selected user interface object 7190 and performs an operation corresponding to the currently selected user interface object 7190. For example, the user interface object 7190 is an application launch icon, and in response to a tap input when the user interface object 7190 is selected, the application corresponding to the user interface object 7190 is launched and displayed on the display. In some embodiments, in response to an input including a thumb moving from a retracted position to an extended position along the x-axis, the current selection indicator 7198 is moved upward from the item 7190 to the upper user interface object 7182. Other types of finger inputs provided relative to one or more of the user interface objects 7172 to 7194 are possible and optionally cause the execution of other types of operations corresponding to the user interface object to be subject to these inputs.

[0126] Figure 7C Shows a visual indication of a menu 7170 visible in a mixed reality view (e.g., an augmented reality view of a physical environment) displayed by a computer system (e.g., device 7100 or HMD). In some embodiments, a hand 7200 in the physical environment is visible in the displayed augmented reality view (e.g., as part of a view of the physical environment captured by a camera), as shown at 7200'. In some embodiments, the hand 7200 is visible through a transparent or translucent display surface on which the menu 7170 is displayed (e.g., the device 7100 is a head-up display or an HMD with a see-through portion).

[0127] In some embodiments, as Figure 7C shown, the menu 7170 is displayed in a mixed reality environment at a position corresponding to a predefined portion of the user's hand (e.g., the tip of the thumb) and having an orientation corresponding to the orientation of the user's hand. In some embodiments, when the user's hand moves (e.g., laterally or rotates) relative to the physical environment (e.g., a camera that captures the user's hand or the user's eyes or physical objects or walls around the user), the menu 7170 is shown to move with the user's hand in the mixed reality environment. In some embodiments, the menu 7170 moves according to the movement of the user's gaze directed into the mixed reality environment. In some embodiments, the menu 7170 is displayed at a fixed position on the display regardless of the view of the physical environment shown on the display.

[0128] In some embodiments, the menu 7170 is displayed on the display in response to detecting a ready position of the user's hand (e.g., the thumb resting on the side of the index finger). In some embodiments, the user interface object displayed in response to detecting the hand in a ready position is different depending on the current user interface context and / or the position of the user's gaze in the mixed reality environment.

[0129] FIG. 7D to FIG. 7E shows a hand 7200 in an exemplary non-ready state configuration (e.g., a rest configuration) ( Fig.7D ) and an exemplary ready state configuration ( Fig. 7E ). Regarding FIG. 7D to FIG. 7E the input gestures described are used to illustrate the processes described below, including Fig. 9 the processes in

[0130] In Fig.7D , the hand 7200 is shown in an exemplary non-ready state configuration (e.g., a rest configuration (e.g., the hand is in a relaxed or arbitrary state and the thumb 7202 is not resting on the index finger 7204)). In an exemplary user interface context, a container object 7206 (e.g., an application dock, folder, control panel, menu, record, preview, etc.) including user interface objects 7208, 7210, and 7212 (e.g., application icons, media objects, controls, menu items, etc.) is generated by a display component of a computer system (e.g., Figure 1 the computer system 101 in Figure 1 , Figure 2 and Figure 4The display generation component 120 (e.g., a touch screen display, a stereoscopic projector, a head-up display, an HMD, etc.) in displays in a three-dimensional environment (e.g., a virtual environment or a mixed reality environment). In some embodiments, a static configuration of the hand is an example of a hand configuration that is not a ready state configuration. For example, other hand configurations that are not necessarily relaxed and static but do not meet the criteria for detecting a ready state gesture (e.g., where the thumb is resting on the index finger (e.g., the middle phalanx of the index finger)) are also categorically recognized as being in a non-ready state configuration. For example, according to some embodiments, when the user waves their hand in the air or holds an object or makes a fist, etc., the computer system determines that the user's hand does not meet the criteria for detecting the ready state of the hand and determines that the user's hand is in a non-ready state configuration. In some embodiments, the criteria for detecting a ready state configuration of the hand include detecting that the user has changed his / her hand configuration and that the change results in the user's thumb resting on a predefined portion of the user's index finger (e.g., the middle phalanx of the index finger). In some embodiments, the criteria for detecting a ready state configuration require that the change in the gesture results in the user's thumb resting on a predefined portion of the user's index finger for at least a first threshold amount of time in order for the computer system to recognize that the hand is in a ready state configuration. In some embodiments, if the user has not changed his / her hand configuration and has not provided any valid input gesture for at least a second threshold amount of time after entering the ready state configuration, the computer system treats the current hand configuration as a non-ready state configuration. The computer system requires the user to change his / her current hand configuration and then return to the ready state configuration in order to recognize the ready state configuration again. In some embodiments, the ready state configuration is user-configurable and user-customizable, e.g., by the user demonstrating to the computer system the expected ready state configuration of the hand in a gesture settings environment provided by the computer system, and optionally an acceptable range of variations of the ready state configuration.

[0131] In some embodiments, the container 7206 is displayed in a mixed reality environment (e.g., as Fig.7D and Fig. 7Eas shown in). For example, container 7206 is displayed by a display of device 7100 together with at least a portion of a view of a physical environment captured by one or more rear cameras of device 7100. In some embodiments, hand 7200 in the physical environment is also visible in the displayed mixed reality environment, e.g., where the actual spatial relationship between the hand and the physical environment is represented in the view of the displayed mixed reality environment as shown at 7200b. In some embodiments, container 7206 and hand 7200 are displayed relative to a physical environment that is positioned away from the user and is displayed via a live feed of a camera juxtaposed with the remote physical environment. In some embodiments, container 7206 is displayed on a transparent or translucent display of the device, and the physical environment around the user (e.g., including hand 7200 as shown at 7200b) is visible through the transparent or translucent display.

[0132] In some embodiments, container 7206 is displayed in a virtual reality environment (e.g., hovering in virtual space). In some embodiments, hand 7200 is visible in a virtual reality scene (e.g., an image of hand 7200 captured by one or more cameras is rendered in the virtual reality environment). In some embodiments, a representation of hand 7200 is visible in the virtual reality environment. In some embodiments, hand 7200 is not visible in the virtual reality environment (e.g., omitted from the virtual reality environment). In some embodiments, device 7100 is not visible in the virtual reality environment. In some embodiments, an image or a representation of device 7100 is visible in the virtual reality environment.

[0133] In some embodiments, when hand 7200 is not in a ready configuration (e.g., in any non-ready configuration or stops remaining in a ready configuration (e.g., due to a change in hand posture or failure to provide a valid input gesture within a threshold amount of time to enter the ready configuration)), the computer system (in addition to determining whether the hand has entered a ready configuration) does not perform input gesture recognition to perform an operation within the current user interface context, and thus, does not respond to an input gesture performed by the hand (e.g., a tap of the thumb above the index finger that includes movement of the thumb along the axis shown by arrow 7110, as described in part (A) with respect to Fig. 7A ; movement of the thumb across the index finger along the axis indicated by arrow 7120, as described in part (B) with respect to Fig. 7A ; and / or movement of the thumb above the index finger along the axis indicated by arrow 7130, as described in part (C) with respect to Fig. 7Aperforms an operation as described in part (C). In other words, the computer system requires the user to place the hand in a ready state configuration (e.g., change from a non-ready state configuration to a ready state configuration), and then provide a valid input gesture for the current user interface context (e.g., within a threshold amount of time after the hand enters the ready state configuration) in order for the input gesture to be considered valid, and perform the corresponding operation in the current user interface context. In some embodiments, if a valid input gesture is detected without first detecting that the hand is in a ready state configuration, the computer system performs certain types of operations (e.g., interact with a currently displayed user interface object (e.g., scroll or activate a currently displayed user interface object)), and prohibits other types of operations (e.g., call a new user interface, trigger a system-level operation (e.g., navigate to a multitasking user interface or an application launch user interface, activate a device function control panel, etc.)). These safeguards help prevent and reduce the inadvertent and unintended triggering of operations, and avoid unnecessarily restricting the user's free hand movement when the user does not wish to perform an operation or certain types of operations within the current user interface context. Additionally, imposing a small and discrete motion requirement for the ready state configuration does not impose an excessive physical burden on the user (e.g., overly moving the user's arm or hand), and tends to reduce the embarrassment of the user when interacting with the user interface in a social setting.

[0134] The user interface objects 7208 to 7212 of the container 7206 include, for example, one or more application launch icons, one or more controls for performing operations within an application, one or more representations of users of a remote device, and / or one or more representations of media (e.g., as described above with respect to user interface objects 7172 to 7194). In some embodiments, when the computer system selects a user interface object without first detecting that the hand is in a ready state configuration and detects an input gesture, the computer system performs a first operation based on the input gesture relative to the selected user interface object (e.g., launching an application corresponding to the selected application icon, changing the control value of the selected control, initiating communication with the user of the selected representation of the user, and / or initiating playback of a media item corresponding to the selected representation of the media item); and when the computer system selects a user interface object after first detecting that the hand is in a ready state configuration and detects the same input gesture, the computer system performs a second operation different from the first operation (e.g., the second operation is a system operation not specific to the currently selected user interface object (e.g., system operations include displaying a system affordance representation in response to detecting that the hand is in a ready state configuration, and launching a system menu in response to the input gesture)). In some embodiments, placing the hand in a ready state configuration enables certain input gestures that are not paired with any function in the current user interface context (e.g., a thumb flick gesture), and detecting a newly enabled input gesture after detecting that the hand is in a ready state configuration causes the computer system to perform an additional function associated with the newly enabled input gesture. In some embodiments, the computer system optionally displays a user interface indication (e.g., an additional option, a system affordance representation, or a system menu) in response to detecting that the hand is in a ready state configuration, and allows the user to interact with the user interface indication or trigger an additional function using the newly enabled input gesture (e.g., a thumb flick gesture detected when a system affordance representation is displayed causes the system menu to be displayed, and a thumb flick gesture detected when the system menu is displayed causes navigation through the system menu or expansion of the system menu).

[0135] In Fig. 7EIn [the figure], the hand 7200 is shown in a ready state configuration (e.g., the thumb 7202 is resting on the index finger 7204). Based on determining that the hand 7200 has moved into the ready state configuration, the computer system (e.g., in the region corresponding to the fingertip of the thumb in a mixed reality environment) displays a system affordance representation icon 7214. The system affordance representation icon 7214 indicates an area from which one or more user interface objects (e.g., a menu of application icons corresponding to different applications, a menu of currently open applications, a device control user interface, etc.) can be displayed and / or accessed (e.g., in response to a thumb flick gesture or other predefined input gesture). In some embodiments, when the system affordance representation icon 7214 is displayed, the computer system performs an operation in response to an input gesture performed by the hand 7200 (e.g., as described below with respect to Figure 7F ).

[0136] In some embodiments, the movement of the hand 7200 from a non-ready state configuration to a ready state configuration is detected by analyzing data captured by a sensor system (e.g., an image sensor or other sensors (e.g., motion sensors, touch sensors, vibration sensors, etc.)), as described above with respect to 7A to 7C . In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras associated with a portable device or a head-mounted system), one or more depth sensors, and / or one or more light emitters.

[0137] In some embodiments, the system affordance representation icon 7214 is displayed in a mixed reality environment. For example, the system affordance representation icon 7214 is displayed by a display generation component (e.g., the display of the device 7100 or the HMD) together with at least a portion of a view of the physical environment captured by one or more cameras of the computer system (e.g., one or more rear cameras of the device 7100, or the forward-facing or downward-facing cameras of the HMD). In some embodiments, the system affordance representation icon 7214 is displayed on a transparent or semi-transparent display of the device (e.g., a heads-up display or an HMD with a see-through portion), through which the physical environment is visible. In some embodiments, the system affordance representation icon 7214 is displayed in a virtual reality environment (e.g., hovering in virtual space).

[0138] In some embodiments, in response to detecting that the hand has changed its pose without providing a valid input gesture for the current user interface context and is no longer in a ready state configuration, the computer system stops displaying the system affordance representation icon 7214. In some embodiments, in response to detecting that the hand has remained in a ready state pose without providing a valid input gesture for more than a threshold amount of time, the computer system stops displaying the system affordance representation icon 7214 and determines that the criteria for detecting the ready state configuration are no longer met. In some embodiments, after stopping the display of the system affordance representation icon 7214, in accordance with determining that a change in the hand pose of the user causes the criteria for detecting the ready state configuration to be met again, the computer system redisplay the system affordance representation icon (e.g., at the fingertip of the thumb at the new hand position).

[0139] In some embodiments, more than one ready state configuration of the hand is optionally defined and recognized by the computer system, and each ready state configuration of the hand causes the computer system to display a different type of affordance representation and enables different sets of input gestures and / or operations to be performed in the current user interface context. For example, a second ready state configuration is optionally all fingers pulled together into a fist with the thumb resting on a finger other than the index finger. When the computer system detects this second ready state configuration, the computer system displays a system affordance representation icon different from icon 7214, and subsequent input gestures (e.g., a thumb swipe across the index finger) cause the computer system to initiate a system shutdown operation or display a menu of power options (e.g., shut down, sleep, suspend, etc.).

[0140] In some embodiments, the system affordance representation icon displayed in response to the computer detecting that the hand is in a ready state configuration is a primary affordance representation that indicates that a selection user interface including a plurality of currently installed applications will be displayed in response to detecting a predefined input gesture (e.g., a thumb flick input, a thumb push input, or other input as described 7A to 7C above). In some embodiments, an application dock including a plurality of application icons for launching corresponding applications is displayed in response to the computer detecting that the hand is in a ready state configuration, and subsequent activation or selection input gestures made by the hand cause the computer system to launch the corresponding application.

[0141] FIG. 7F to FIG. 7G Various examples of operations performed in response to input gestures detected in the case of finding that the hand is in a ready state configuration are provided in accordance with various embodiments. Although FIG. 7F to FIG. 7GDescribes performing different operations based on user gaze. However, it should be understood that in some embodiments, gaze is not an essential component when detecting a readiness gesture and / or an input gesture to trigger the execution of those various operations. In some embodiments, the computer system uses a combination of user gaze and hand configuration in conjunction with the current user interface context to determine which operation will be performed.

[0142] Figure 7F Illustrates exemplary input gestures performed in the case of detecting that the hand is in a readiness configuration and exemplary responses of the displayed three-dimensional environment (e.g., virtual reality environment or mixed reality environment) according to some embodiments. In some embodiments, the computer system (e.g., device 7100 or HMD) displays a system affordance representation (e.g., system affordance representation icon 7214) at the tip of the thumb to indicate that a readiness configuration of the hand has been detected, and that, for example, in addition to other input gestures already available in the current user interface context (such as thumb swipe input for scrolling user interface objects, etc.), input of system gestures is enabled to trigger predefined system operations (e.g., a thumb flick gesture to display a task bar or system menu, a thumb tap gesture to activate a voice-based assistant). In some embodiments, when gesture input is provided, the system affordance representation moves according to the movement of the hand as a whole in space and / or the movement of the thumb, such that the position of the system affordance representation remains fixed relative to a predefined part of the hand (e.g., the tip of the thumb of hand 7200).

[0143] According to some embodiments, Figure 7F Parts (A) to (C) of [description] show a three-dimensional environment 7300 (e.g., virtual reality environment or mixed reality environment) displayed by a display generation component of a computer system (e.g., the touch screen display of device 7100 or a stereoscopic projector or the display of an HMD). In some embodiments, device 7100 is a handheld device (e.g., a cellular phone, a tablet computer, or other mobile electronic device) including a display, a touch-sensitive display, etc. In some embodiments, device 7100 represents a wearable head-mounted headset that includes a head-up display, a head-mounted display, etc.

[0144] In some embodiments, the three-dimensional environment 7300 is a virtual reality environment that includes virtual objects (e.g., user interface objects 7208, 7210, and 7212). In some embodiments, the virtual reality environment does not correspond to the physical environment in which the device 7100 is located. In some embodiments, the virtual reality environment corresponds to the physical environment (e.g., based on the position of physical objects in the physical environment determined using one or more cameras of the device 7100, at least some of the virtual objects are displayed at positions in the virtual reality environment corresponding to the positions of the physical objects in the corresponding physical environment). In some embodiments, the three-dimensional environment 7300 is a mixed reality environment. In some embodiments, the device 7100 includes one or more cameras configured to continuously provide a real-time view of at least a portion of the surrounding physical environment within the field of view of the one or more cameras of the device 7100, and the mixed reality environment corresponds to the portion of the surrounding physical environment within the field of view of the one or more cameras of the device 7100. In some embodiments, the mixed reality environment at least partially includes the real-time view of the one or more cameras of the device 7100. In some embodiments, the mixed reality environment includes one or more virtual objects that are displayed (e.g., at positions in the three-dimensional environment 7300 corresponding to the positions of physical objects in the physical environment based on the position of the physical objects in the physical environment determined using the one or more cameras of the device 7100 or the real-time view of the one or more cameras) to replace the real-time camera view (e.g., superimposed over, covering, or replacing the real-time camera view). In some embodiments, the display of the device 7100 includes at least a partially transparent heads-up display (e.g., having an opacity less than a threshold opacity such as less than 25%, 20%, 15%, 10%, or 5%, or having a see-through portion), such that the user can see at least a portion of the surrounding physical environment through the at least partially transparent region of the display. In some embodiments, the three-dimensional environment 7300 includes one or more virtual objects displayed on the display (e.g., container 7206 including user interface objects 7208, 7210, and 7212). In some embodiments, the three-dimensional environment 7300 includes one or more virtual objects displayed on the transparent region of the display so as to appear superimposed over the portion of the surrounding physical environment visible through the transparent region of the display. In some embodiments, one or more corresponding virtual objects are displayed at positions in the three-dimensional environment 7300 corresponding to the positions of physical objects in the physical environment (e.g., based on the position of physical objects in the physical environment determined using one or more cameras of the device 7100, the one or more cameras monitoring the portion of the physical environment visible through the transparent region of the display), such that the corresponding virtual objects are displayed to replace the corresponding physical objects (e.g., obscuring and replacing the view of the corresponding physical objects).

[0145] In some embodiments, a sensor system of a computer system (e.g., one or more cameras of device 7100 or the HMD) tracks the position and / or movement of one or more features of a user (such as the user's hand). In some embodiments, the position and / or movement of the user's hand (e.g., fingers) is used as an input to the computer system (e.g., device 7100 or the HMD). In some embodiments, although the user's hand is within the field of view of one or more cameras of the computer system (e.g., device 7100 or the HMD), and the position and / or movement of the user's hand is tracked by the sensor system of the computer system (e.g., device 7100 or the HMD) as an input to the control unit of the computer system (e.g., device 7100 or the HMD), the user's hand is not shown in the three-dimensional environment 7300 (e.g., the three-dimensional environment 7300 does not include a live view from the one or more cameras, the hand is removed from the live view of the one or more cameras, or the user's hand is within the field of view of the one or more cameras but outside of the portion of the field of view that is shown in the live view in the three-dimensional environment 7300). In some embodiments, as in the example shown in Figure 7F the hand 7200 (e.g., a representation of the user's hand or a portion of the hand within the field of view of one or more cameras of device 7100) is visible in the three-dimensional environment 7100 (e.g., is shown as a rendered representation, is shown as part of a live camera view, or is visible through a see-through portion of the display). In Figure 7F it is detected that the hand 7200 is in a ready-state configuration (e.g., the thumb resting on the middle phalanx of the index finger) for providing a gesture input (e.g., a gesture input). In some embodiments, the computer system (e.g., device 7100 or the HMD) determines that the hand 7200 is in the ready-state configuration by performing image analysis on the live view of the one or more cameras. More details regarding the ready-state configuration of the hand and input gestures are provided in at least 7A to 7E and the accompanying description, and for the sake of brevity, are not repeated here.

[0146] According to some embodiments, Figure 7F part (A) of shows a first type of input gesture (e.g., a thumb flick gesture or a thumb push gesture), which includes the movement of the thumb of the hand 7200 along an axis shown by arrow 7120 across a portion of the index finger of the hand 7200 (e.g., across the middle phalanx of the index finger from the palm side to the dorsal side) (e.g., as shown by the hand configuration transition from A(1) to A(2) of Figure 7F ). As described above, in some embodiments, the hand 7200 is tracked by the one or more cameras of device 7100 during the execution of a gesture such as Figure 7Fmovements (e.g., including the movement of the hand as a whole and the relative movement of each finger) when performing the gestures shown, and the position of the hand 7200 (e.g., including the position of the hand as a whole and the relative position of the fingers). In some embodiments, the device 7100 detects gestures performed by the hand 7200 by performing image analysis on the live view of the one or more cameras. In some embodiments, based on determining that the hand has provided a flick gesture starting from a ready-state configuration of the hand, the computer system performs a first operation corresponding to the flick gesture (e.g., a system operation such as displaying a menu 7170 including a plurality of application launch icons, or an operation corresponding to the current user interface context that would not be enabled if the hand had not first been found in the ready-state configuration). In some embodiments, the hand 7200 provides additional gestures to interact with the menu. For example, a subsequent flick gesture after the display of the menu 7170 causes the menu 7170 to be pushed into the three-dimensional environment and displayed in an enhanced state in virtual space (e.g., with a larger animated representation of the menu items). In some embodiments, a subsequent swipe gesture horizontally scrolls a selection indicator within the currently selected row of the menu, and a subsequent push or pull gesture with the thumb scrolls the selection indicator up and down across different rows of the menu. In some embodiments, a subsequent tap gesture with the thumb causes the activation of the currently selected menu item (e.g., application icon) to be launched in the three-dimensional environment.

[0147] According to some embodiments, Figure 7F Part (B) of shows a second type of gesture (e.g., a swipe gesture) that includes movement of the thumb of the hand 7200 along the axis shown by the arrow 7130, along the length of the index finger of the hand 7200 (e.g., as shown by the movement from Figure 7Fas shown by the hand configuration transition from B(1) to B(2). In some embodiments, gestures for interacting with virtual objects (e.g., virtual objects 7208, 7210, and 7212 in container object 7206) in the current user interface context are enabled regardless of whether the hand is first found in the ready state configuration. In some embodiments, different types of interactions with container object 7206 are enabled depending on whether the thumb swipe gesture starts from the ready state configuration. In some embodiments, based on determining that the hand has provided a thumb swipe gesture starting from the ready state configuration of the hand, the computer system performs a second operation corresponding to the thumb swipe gesture (e.g., scrolling the view of container object 7206 to reveal one or more virtual objects initially not visible to the user in container object 7206, or another operation corresponding to the current user interface context that would not be enabled if the hand had not first been found in the ready state configuration). In some embodiments, hand 7200 provides additional gestures to interact with the container object or perform system operations. For example, a subsequent thumb flick gesture causes menu 7170 to be displayed (e.g., as shown in part (A) of Figure 7F ). In some embodiments, subsequent thumb swipe gestures in different directions scroll the view of container object 7206 in opposite directions. In some embodiments, a subsequent thumb tap gesture causes activation of the currently selected virtual object in container 7206. In some embodiments, if the thumb swipe gesture does not start from the ready state configuration, in response to the thumb swipe gesture, the selection indicator shifts through the virtual objects in container 7206 in the direction of movement of the thumb of hand 7200.

[0148] According to some embodiments, Figure 7F part (C) of shows a third type of gesture input (e.g., a thumb tap gesture) (e.g., a tap of the thumb of hand 7200 on a predefined portion (e.g., the middle phalanx) of the index finger of hand 7200 (e.g., by moving the thumb downward along the axis shown by arrow 7110 from the raised position)) (e.g., as shown by the transition from Figure 7Fas shown by the hand configuration transition from C(1) to C(2). In some embodiments, completing the thumb tap gesture requires lifting the thumb away from the index finger. In some embodiments, gestures for interacting with virtual objects (e.g., virtual objects 7208, 7210, and 7212 in container object 7206) in the current user interface context are enabled regardless of whether the hand is first found in the ready configuration. In some embodiments, different types of interactions with container object 7206 are enabled depending on whether the thumb tap gesture starts from the ready configuration. In some embodiments, based on determining that the hand has provided a thumb tap gesture starting from the ready configuration of the hand (e.g., the gesture includes an upward movement of the thumb away from the index finger prior to the tap of the thumb on the index finger), the computer system performs a third operation corresponding to the thumb tap gesture (e.g., activation of a voice-based assistant 7302 or a communication channel (e.g., a voice communication application), or another operation corresponding to the current user interface context that would not be enabled if the hand had not first been found in the ready configuration). In some embodiments, the hand 7200 provides additional gestures for interacting with the voice-based assistant or the communication channel. For example, a subsequent thumb flick gesture causes the voice-based assistant or the voice communication user interface to be pushed to a more distant location in the space within the three-dimensional environment from a position next to the user's hand. In some embodiments, a subsequent thumb swipe gesture scrolls through different preset functions of the voice-based assistant or through a list of potential recipients of the voice communication channel. In some embodiments, a subsequent thumb tap gesture causes the voice-based assistant or the voice communication channel to be cleared. In some embodiments, if the thumb tap gesture does not start from the ready configuration, then in response to Fig. 7E the thumb tap gesture in part (C) of, the operation available in the current user interface context will be activated (e.g., the currently selected virtual object will be activated).

[0149] Figure 7F The examples shown in are merely illustrative. Providing additional and / or different operations in response to detecting an input gesture starting from the ready configuration of the hand allows the user to perform additional functions without cluttering the user interface with controls and reduces the amount of user input required to perform those functions, thereby making the user interface more efficient and saving time in the interaction between the user and the device.

[0150] The user interface interactions shown in are described without regard to the position of the user's gaze. 7A to 7F In some embodiments, the interaction is independent of the position of the user's gaze or the exact position of the user's gaze within a portion of the three-dimensional environment. However, in some embodiments, gaze is used to modify the response behavior of the system, and the operations and user interface feedback change depending on the different positions of the user's gaze when the user input is detected. Figure 7G shows exemplary gestures performed with a hand in a ready state and exemplary responses of the displayed three-dimensional environment according to user gaze. For example, the left column of the figure (e.g., parts A-0, A-1, A-2, and A-3) shows exemplary scenarios where, when the user's gaze is focused on the user's hand (e.g., as shown in part A-0 in Figure 7G ), the hand provides one or more gesture inputs starting from the ready state configuration (e.g., the thumb resting on the index finger) (e.g., relative to Fig. 7E and Figure 7F the ready state configuration described). The right column of the figure (e.g., parts B-0, B-1, B-2, and B-3) shows exemplary scenarios where, when the user's gaze is focused on the user interface environment (e.g., a user interface object or a physical object in the three-dimensional environment) other than the user's hand in the ready state configuration (e.g., as shown in part A-0 in Figure 7G ), the hand provides one or more gesture inputs, optionally also starting from the ready state configuration (e.g., the thumb resting on the index finger) (e.g., relative to Fig. 7E and Figure 7F the ready state configuration described). In some embodiments, for other physical objects (e.g., the top or front surface of the housing of a physical media player device, a physical window on a wall in a room, a physical controller device, etc., rather than the user's hand providing the input gesture) or a predefined portion of such other physical object, the user interface responses and interactions described relative to Figure 7G are implemented such that when the user's gaze is focused on the other physical object or the predefined portion of the other physical object, special interactions are enabled for the input gesture provided by the hand starting from the ready state configuration. Figure 7G The left column and right column figures in the same row (e.g., A-1 and B-1, A-2 and B-2, A-3 and B-3) of Figure 7G show different user interface responses according to whether the user's gaze is focused on the user's hand (or another physical object defined by the computer system as a controlled or controlling physical object) for the same input gesture provided in the same user interface context. Regarding the Fig.10 input gestures described are used to illustrate the processes described below, including the

[0151] processes in Figure 7GAs shown, the system affordance representation (e.g., affordance representation 7214) is optionally displayed at a predefined position corresponding to the hand in the ready state configuration in three dimensions. In some embodiments, whenever the hand is determined to be in the stable state configuration, the system affordance representation is always displayed at a predefined position (e.g., a static position or a dynamically determined position). In some embodiments, the dynamically determined position of the system affordance representation is fixed relative to the user's hand in the ready state configuration, e.g., when the user's hand moves as a whole while remaining in the ready state configuration. In some embodiments, in response to the user's gaze pointing to a predefined physical object (e.g., the user's hand in the ready state configuration, or other predefined physical object in the environment) while the user's hand remains in the ready state configuration, the system affordance representation is displayed (e.g., at a static position or near the tip of the thumb), and the display of the system affordance representation stops in response to the user's hand exiting the ready state configuration and / or the user's gaze moving away from the predefined physical object. In some embodiments, the computer system displays the system affordance representation in a first appearance (e.g., an enlarged and prominent appearance) in response to the user's gaze pointing to a predefined physical object (e.g., the user's hand in the ready state configuration, or other predefined physical object in the environment) while the user's hand remains in the ready state configuration, and displays the system affordance representation in a second appearance (e.g., a small and unobtrusive appearance) in response to the user's hand exiting the ready state configuration and / or the user's gaze moving away from the predefined physical object. In some embodiments, the system affordance representation is displayed (e.g., at a static position or near the tip of the thumb) only in response to an indication that the user is ready to provide input (e.g., when the user's hand, while remaining in the ready state configuration, is lifted relative to the body from a previous level) (e.g., regardless of whether the user's gaze is on the user's hand), and the display of the system affordance representation stops in response to the user's hand being lowered from the lifted state. In some embodiments, in some embodiments, the system affordance representation changes the appearance of the system affordance representation (e.g., from a simple indicator to an object menu) in response to an indication that the user is ready to provide input (e.g., when the user's hand, while remaining in the ready state configuration and / or the user's gaze is focused on the user's hand in the ready state configuration or a predefined physical object, is lifted relative to the body from a previous level), and restores the appearance of the system affordance representation (e.g., restores back to a simple indicator) in response to the cessation of the indication that the user is ready to provide input (e.g., the user's device is lowered from the lifted state, and / or the user's gaze moves away from the user's hand or the predefined physical object).In some embodiments, indications that the user is ready to provide input include one or more of the following: the user's finger is touching a physical controller or the user's hand (e.g., the index finger is resting on the controller, or the thumb is resting on the index finger); the user's hand is lifted from a lower level to a higher level relative to the user's body (e.g., there is an upward wrist rotation of the hand in a ready state configuration, or a bending movement of the elbow with the hand in a ready state configuration), thereby changing the hand configuration to a ready state configuration, etc. In some embodiments, the physical position of a predefined physical object is static relative to a three-dimensional environment compared to the position of the user's gaze in these embodiments (e.g., also referred to as "fixed to the world"). For example, system affordances are represented as being displayed on a wall in a three-dimensional environment. In some embodiments, the physical position of a predefined physical object is static relative to a display (e.g., a display generation component) compared to the position of the user's gaze in these embodiments (e.g., also referred to as "fixed to the user"). For example, system affordances are represented as being displayed at the bottom of the display or the user's field of view. In some embodiments, the physical position of a predefined physical object is static relative to a moving part of the user (e.g., the user's hand) or a moving part of the physical environment (e.g., a moving car on a highway) compared to the position of the user's gaze in these embodiments.

[0152] According to some embodiments, Figure 7G Part A-0 of shows user 7320 directing his / her gaze towards a predefined physical object in a three-dimensional environment (e.g., his / her hand in a ready state configuration) (e.g., user hand 7200 or a representation of the user's hand is visible within the field of view of one or more cameras of device 7100 or through the transmissive or transparent portion of an HMD or a head-up display). In some embodiments, device 7100 uses one or more cameras facing the user (e.g., a forward camera) to track the movement of the user's eyes (or the movement of both of the user's eyes) in order to determine the direction and / or object of the user's gaze. According to some embodiments, more details of exemplary gaze tracking techniques are provided with respect to Figures 1 to 6 (specifically with respect to Figure 5 and Figure 6 ). In Figure 7GIn part A-0, since when the hand 7200 is in the ready state configuration, the user's gaze is directed at the hand 7200 (e.g., as indicated by the dashed line connecting the user's eyeball 7512 to the user's hand 7200 or the representation 7200' of the user's hand (e.g., the actual hand or the representation of the hand presented via the display generation component)), system user interface operations are performed in response to gestures executed using the hand 7200 (e.g., user interface operations associated with the system affordance representation 7214 or with a system menu associated with the system affordance representation (e.g., menu 7170), rather than user interface operations associated with other areas or elements of the user interface or with a separate software application executed on the device 7100). Figure 7G Part B-0 shows that when the user's hand is in the ready state configuration, the user moves his / her gaze away from a predefined physical object in the three-dimensional environment (e.g., his / her hand in the ready state configuration or another predefined physical object) (e.g., the user's hand 7200 or the other predefined physical object visible within the field of view of one or more cameras of the device 7100 or through the transmissive or transparent portion of the HMD or head-up display, or the representation of the user's hand or the other predefined physical object). In some embodiments, when processing an input gesture of the hand to provide a system response corresponding to the input gesture, the computer system requires the user's gaze to remain on the predefined physical object (e.g., the user's hand or another predefined physical object). In some embodiments, the computer system requires the user's gaze to remain on the predefined physical object for more than a threshold amount of time and with a preset amount of stability (e.g., remaining substantially stationary or having less than a threshold amount of movement within the threshold amount of time) in order to provide a system response corresponding to the input gesture. For example, optionally, after meeting the time and stability requirements and before the input gesture is fully completed, the gaze can move away from the predefined input gesture.

[0153] In one example, Figure 7G Part A-1 shows a thumb flick gesture starting from the ready state configuration of the hand and including the forward movement of the thumb across the index finger of the hand 7200 along the axis indicated by the arrow 7120. In response to Figure 7G the thumb flick gesture starting from the ready state configuration in part A-1 and based on determining that the user's gaze is directed at a predefined physical object (e.g., the hand 7200, as seen in the real world or through the display generation component of the computer system), the computer system displays the system menu 7170 (e.g., a menu of application icons) in the three-dimensional environment (e.g., which replaces the system affordance representation 7214 at the tip of the user's thumb).

[0154] In another example, Figure 7GPart A-2 shows a thumb swipe gesture that starts from a ready state configuration and includes the movement of the thumb of the hand 7200 along the axis shown by arrow 7130 and along the length of the index finger of the hand 7200. In this example, the gesture in Part A-2 is performed when the computer system is displaying a system menu (e.g., menu 7170) Figure 7G (e.g., in response to the thumb flick gesture described in Part A-1 herein to display the system menu 7170). In response to Figure 7G the thumb swipe gesture starting from the ready state configuration in Part A-2 of, and based on determining that the user's gaze is directed at a predefined physical object (e.g., the hand 7200) (e.g., the gaze meets predefined position, duration, and stability requirements), the computer system moves the current selection indicator 7198 in the direction of movement of the thumb of the hand 7200 on the system menu (e.g., menu 7170) (e.g., moves it to an adjacent user interface object on the menu 7170). In some embodiments, the thumb swipe gesture is one input gesture in a sequence of two or more input gestures starting from a hand in a ready state configuration and represents a continuous series of user interactions with system user interface elements (e.g., a system menu or a system control object, etc.). Thus, in some embodiments, the requirement for the gaze to remain on a predefined physical object (e.g., the user's hand) is optionally applied only at the start of the first input gesture (e.g., Figure 7G the thumb flick gesture in Part A-1 of), and this requirement is not imposed on subsequent input gestures as long as the user's gaze is directed at a system user interface element during the subsequent input gestures. For example, based on determining that the user's gaze is on the user's hand or determining that the user's gaze is on a system menu placed at a position fixed relative to the user's hand (e.g., the tip of the thumb), the computer system performs a system operation (e.g., navigates within the system menu) in response to the thumb swipe gesture. Figure 7G the thumb flick gesture in Part A-1 of), and this requirement is not imposed on subsequent input gestures as long as the user's gaze is directed at a system user interface element during the subsequent input gestures. For example, based on determining that the user's gaze is on the user's hand or determining that the user's gaze is on a system menu placed at a position fixed relative to the user's hand (e.g., the tip of the thumb), the computer system performs a system operation (e.g., navigates within the system menu) in response to the thumb swipe gesture.

[0155] In yet another example, Figure 7G Part A-3 shows a thumb tap gesture that starts from the ready state configuration of the hand and includes the movement generated by the thumb of the hand 7200 tapping on the index finger of the hand 7200 (e.g., by the thumb moving upward from the index finger to a raised position and then moving downward along the axis shown by arrow 7110 from the raised position to a hand position where the thumb touches the index finger again). In this example, the thumb tap gesture in Part A-3 is performed when the computer system (e.g., in response to the thumb swipe gesture described in Part A-2 herein) is displaying a system menu (e.g., menu 7170), where the current selection indicator is displayed on the corresponding user interface object. In response to Figure 7G the thumb tap gesture in Part A-3 of when the computer system (e.g., in response to the thumb swipe gesture described in Part A-2 herein) is displaying a system menu (e.g., menu 7170), where the current selection indicator is displayed on the corresponding user interface object. In response to Figure 7G the thumb tap gesture in Part A-3 of, where the current selection indicator is displayed on the corresponding user interface object. In response to Figure 7Ga thumb tap gesture in portion A-3, and in response to the user's gaze being directed at a predefined physical object (e.g., the hand 7200), activate the currently selected user interface object 7190 and perform an operation corresponding to the user interface object 7190 (e.g., stop displaying the menu 7170, display a user interface object 7306 associated with the user interface object 7190 (e.g., a preview or control panel, etc.), and / or initiate an application corresponding to the user interface object 7190). In some embodiments, the user interface object 7306 is displayed at a position corresponding to the position of the hand 7200 in the three-dimensional environment. In some embodiments, the thumb tap gesture is one of a sequence of two or more input gestures starting from a hand in a ready-state configuration and represents a continuous series of user interactions with system user interface elements (e.g., a system menu or a system control object, etc.). Thus, in some embodiments, the requirement to maintain the gaze on a predefined physical object (e.g., the user's hand) is optionally only applied at the start of the first input gesture (e.g., Figure 7G the thumb flick gesture in portion A-1 of Figure 7G and is not imposed on subsequent input gestures (e.g., Figure 7G the thumb swipe gesture in portion A-2 and the thumb tap gesture in portion A-3) as long as the user's gaze is directed at a system user interface element during the subsequent input gestures. For example, in response to the thumb tap gesture, the computer system performs a system operation (e.g., activates the currently selected user interface object within the system menu) based on determining that the user's gaze is on the user's hand or on a system menu placed at a position fixed relative to the user's hand (e.g., the tip of the thumb).

[0156] Compared with Figure 7G portion A-0 of Figure 7G portion B-0 of Figure 7G shows the user moving his / her gaze away from a predefined physical object in the three-dimensional environment (e.g., his / her hand in a ready-state configuration) (e.g., the user's hand 7200 or a representation of the user's hand is within the field of view of one or more cameras of the device 7100, but the user's gaze is not (e.g., neither directly, nor through the pass-through or transparent portion of an HMD or a head-up display, nor through the camera view) on the hand 7100). Instead, the user's gaze is typically directed at a container 7206 of the user interface object (e.g., a menu row) (e.g., as described herein with reference to the container 7206, FIG. 7D to FIG. 7E) or the displayed user interface (e.g., the user's gaze does not meet the stability and duration requirements for a specific location or object in the three-dimensional environment). Based on determining that the user's gaze has left a predefined physical object (e.g., the hand 7200), the computer system abandons performing a system user interface operation (e.g., related to the system affordance representation 7214 or related to displaying a system menu associated with the system affordance representation (e.g., a menu of application icons) (e.g., as shown in part A-1 of Figure 7G )) in response to a thumb flick gesture performed using the hand 7200. Optionally, performing a user interface operation associated with other areas or elements of the user interface or associated with a separate software application executed on the computer system (e.g., device 100 or HMD) in response to a thumb flick gesture performed using the hand 7200 when the user's gaze has left the predefined physical object (e.g., the hand 7200 in a ready state) is not a system user interface operation. In one example, as shown in Figure 7G , the entire user interface including the container 7206 scrolls up according to an upward thumb flick gesture. In another example, performing a user interface operation associated with the container 7206 in response to a gesture performed using the hand 7200 when the user's gaze has left the hand 7200 and is pointing to the container 7206 is not a system user interface operation (e.g., displaying a system menu).

[0157] Similar to Figure 7G 's part A-1, Figure 7G 's part B-1 also shows a thumb flick gesture of the hand 7200 starting from the ready state configuration. Compared to the behavior shown in Figure 7G 's part A-1, in Figure 7G 's part B-1, based on determining that the user's gaze is not pointing to a predefined physical object (e.g., the user's hand in a ready state configuration), the computer system abandons displaying a system menu in response to the thumb flick gesture. Instead, the user interface scrolls up according to the thumb flick gesture. For example, the container 7206 moves up in the three-dimensional environment according to the movement of the thumb of the hand 7200 across the index finger of the hand 7200 and according to the user's gaze leaving the hand 7200 and pointing to the container 7206. In this example, although no system operation is performed, the system affordance representation 7214 is still displayed next to the user's thumb, indicating that the system operation is available (e.g., because the user's hand is in a ready state configuration), and the system affordance representation 7214 moves with the user's thumb during the input gesture and when the user interface scrolls up in response to the input gesture.

[0158] Similar to Figure 7G 's part A-2, Figure 7GPart B-2 also shows a thumb-swipe gesture of the hand 7200 starting from the ready-state configuration. As compared with Figure 7G Part A-2, the thumb-swipe gesture in Part B-2 is performed when the system menu (e.g., menu 7170) is not displayed and when the user's gaze is not focused on a predefined physical object (e.g., the user's hand). Figure 7G (e.g., because menu 7170 is not displayed in response to the thumb-flick gesture described in Part B-1 of this document). Figure 7G In response to the thumb-swipe gesture in Part B-2 of Figure 7G and based on determining that the user's gaze has left the predefined physical object (e.g., hand 7200) and is pointing to the container 7206, the current selection indicator in the container 7206 scrolls in the direction of movement of the thumb of the hand 7200.

[0159] Similar to Figure 7G Part A-3, Figure 7G Part B-3 also shows a thumb-tap gesture of the hand 7200 starting from the ready-state configuration. As compared with Figure 7G Part A-3, the thumb-tap gesture in Part B-3 is performed when the system menu (e.g., menu 7170) is not displayed and when the user's gaze is not focused on a predefined physical object (e.g., the user's hand). Figure 7G (e.g., because menu 7170 is not displayed in response to the thumb-flick gesture described in Part B-1 of this document). Figure 7G In response to the thumb-tap gesture in Part A-3 of Figure 7G and based on the user's gaze leaving the predefined physical object (e.g., hand 7200) and in the absence of any system user interface elements displayed in response to a previously received input gesture, the computer system abandons the execution of the system operation. Based on determining that the user's gaze is pointing to the container 7206, the currently selected user interface object in the container 7206 is activated, and an operation corresponding to the currently selected user interface object is performed (e.g., stopping the display of the container 7206 and displaying the user interface 7308 corresponding to the activated user interface object). In this example, although the system operation is not performed, the system enablement representation 7214 is still displayed next to the user's thumb, thereby indicating that the system operation is available (e.g., because the user's hand is in the ready-state configuration), and the system enablement representation 7214 moves with the user's thumb during the input gesture and when the user interface scrolls up in response to the input gesture. In some embodiments, because the user interface object 7308 is activated from the container 7206, the user interface object 7308 is displayed at a position that does not correspond to the position of the hand 7200 in the three-dimensional environment as compared with the user interface object 7306 in Part A-3 of Figure 7G .

[0160] It should be understood that, in the example shown in Figure 7G , the computer system treats the user's hand 7200 as a predefined physical object that uses its position (e.g., compared to the user's gaze) to determine whether a system operation should be performed in response to a predefined gesture input. Although the position of the hand 7200 appears different on the display in the examples shown in the left and right columns of Figure 7G , this only indicates that the position of the user's gaze has changed relative to the three-dimensional environment and does not necessarily impose a restriction on the position of the user's hand relative to the three-dimensional environment. In fact, in most cases, the user's hand as a whole is not typically fixed in position during the input gesture, and the gaze is compared to the moving physical position of the user's hand to determine whether the gaze is focused on the user's hand. In some embodiments, if another physical object in the user's environment other than the user's hand is used as the predefined physical object for determining whether a system operation should be performed in response to a predefined gesture input, then the gaze is compared to the physical position of that physical object even when the physical object can move relative to the environment or the user and / or when the user moves relative to the physical object.

[0161] In Figure 7G , according to some embodiments, a combination of whether the user's gaze is directed at a predefined physical location (e.g., focused on a predefined physical object (e.g., the user's hand is in a ready state configuration, or a system user interface object is displayed at a fixed position relative to the predefined physical object)) and whether the input gesture starts from a hand in a ready state configuration is used to determine whether to perform a system operation (e.g., display a system user interface or a system user interface object), or whether to perform an operation in the current context of the three-dimensional environment without performing a system operation. FIG. 7H to FIG. 7J shows an example behavior of the displayed three-dimensional environment (e.g., a virtual reality or mixed reality environment) according to some embodiments, which depends on whether the user is ready to provide a gesture input (e.g., whether the user's hand meets a predefined requirement (e.g., is lifted to a predefined level and settled in a ready state configuration for at least a threshold amount of time), and the user's gaze meets a predefined requirement (e.g., the gaze is focused on an activatable virtual object and meets stability and duration requirements)). Regarding the FIG. 7H to FIG. 7J input gesture described above is used to illustrate the processes described below, including the Fig.11 processes in

[0162] Figure 7H shows an exemplary computer-generated environment corresponding to a physical environment. As referred to herein Figure 7HAs described, the computer-generated environment can be a virtual reality environment, an augmented reality environment, or a computer-generated environment that is displayed on a display such that the computer-generated environment is superimposed over a view of the physical environment visible through a transparent portion of the display. As Figure 7H shown, user 7502 is standing in a physical environment (e.g., scene 105) and operating a computer system (e.g., computer system 101) (e.g., holding device 7100 or wearing an HMD). In some embodiments, as in the Figure 7H example shown in, device 7100 is a handheld device (e.g., a cellular phone, a tablet, or other mobile electronic device) that includes a display, a touch-sensitive display, etc. In some embodiments, device 7100 represents a wearable head-mounted headset and is optionally replaced by a wearable head-mounted headset that includes a heads-up display, a head-mounted display, etc. In some embodiments, the physical environment includes one or more physical surfaces and physical objects around user 7502 (e.g., the walls of a room, furniture (e.g., represented by shaded 3D cuboid 7504)).

[0163] In Figure 7HIn the example shown in part (B), a computer-generated three-dimensional environment corresponding to a physical environment (e.g., the part of the physical environment that is within the field of view of one or more cameras of device 7100 or visible through the transparent portion of the display of device 7100) is displayed on device 7100. The physical environment includes a physical object 7504, which is represented by an object 7504' in the computer-generated environment shown on the display (e.g., the computer-generated environment is a virtual reality environment that includes a virtual representation of physical object 7504, the computer-generated environment is an augmented reality environment that includes the representation 7504' of physical object 7504 as part of a live view of one or more cameras of device 7100, or physical object 7504 is visible through the transparent portion of the display of device 7100). Additionally, the computer-generated environment shown on the display includes virtual objects 7506, 7508, and 7510. Virtual object 7508 is shown as being attached to object 7504' (e.g., covering the flat front surface of physical object 7504). Virtual object 7506 is shown as being attached to the wall of the computer-generated environment (e.g., covering a part of the wall or the representation of the wall of the physical environment). Virtual object 7510 is shown as being attached to the floor of the computer-generated environment (e.g., covering a part of the floor or the representation of the floor of the physical environment). In some embodiments, virtual objects 7506, 7508, and 7510 are activatable user interface objects that, when activated by user input, cause object-specific operations to be performed. In some embodiments, the computer-generated environment also includes virtual objects that cannot be activated by user input and are displayed to enhance the aesthetic quality of the computer-generated environment and provide information to the user. Figure 7H Part (C) shows that the computer-generated environment shown on device 7100 is a three-dimensional environment: According to some embodiments, when the viewing angle of device 7100 relative to the physical environment changes (e.g., when the viewing angle of device 7100 or one or more cameras of device 7100 relative to the physical environment changes in response to the movement and / or rotation of device 7100 within the physical environment), the viewing angle of the computer-generated environment displayed on device 7100 is changed accordingly (e.g., including changing the viewing angles of physical surfaces and objects (e.g., walls, floors, physical object 7504) and virtual objects 7506, 7508, and 7510).

[0164] Fig.7I Shows an example behavior of the computer-generated environment in response to the user 7502 directing his / her gaze to a corresponding virtual object in the computer-generated environment when the user is not ready to provide gesture input (e.g., the user's hand is not in a ready state configuration). As Fig.7IAs shown in parts (A) to (C) thereof, the user holds his / her left hand 7200 in a state other than the state for providing gesture input preparation (e.g., holds it in a position other than the first predefined preparation state configuration). In some embodiments, the computer system determines that the user's hand is in a predefined preparation state for providing gesture input based on detecting that a predefined part of the user's finger is touching a physical control element (e.g., the thumb touches the middle phalanx of the index finger, or the index finger touches the physical controller, etc.). In some embodiments, the computer system determines that the user's hand is in a predefined preparation state for providing gesture input based on detecting that the user's hand is lifted relative to the user to a level higher than a predetermined level (e.g., the hand is lifted in response to the arm rotating around the elbow joint, or the wrist rotating around the wrist joint, or the finger being lifted relative to the hand, etc.). In some embodiments, the computer system determines that the user's hand is in a predefined preparation state for providing gesture input based on detecting that the posture of the user's hand changes to a predefined configuration (e.g., the thumb is resting on the middle phalanx of the index finger, the fingers are closed to form a fist, etc.). In some embodiments, multiple requirements among the above requirements are combined to determine whether the user's hand is in a preparation state for providing gesture input. In some embodiments, the computer system also requires that the user's hand as a whole is stationary (e.g., less movement than a threshold amount without reaching a threshold amount of time) in order to determine that the hand is ready to provide gesture input. According to some embodiments, when the user's hand is not found to be in a preparation state for providing gesture input and the user's gaze is focused on an activatable virtual object, subsequent movements of the user's hand (e.g., free movement or movement simulating a predefined gesture) are not considered and / or regarded as user input pointing to the virtual object (which is the focus of the user's gaze).

[0165] In this example, a representation of the hand 7200 is displayed in a computer-generated environment. The computer-generated environment does not include a representation of the user's right hand (e.g., because the right hand is not within the field of view of one or more cameras of the device 7100). Additionally, in some embodiments, such as in Fig.7IIn the example shown, where device 7100 is a handheld device, the user is able to see portions of the surrounding physical environment independent of any representation of the physical environment displayed on device 7100. For example, portions of the user's hand are visible to the user outside of the display of device 7100. In some embodiments, device 7100 in these examples represents and is replaced by a head-mounted headset having a display (e.g., a head-mounted display) that completely blocks the user's view of the surrounding physical environment. In some such embodiments, portions of the physical environment are not directly visible to the user; rather, the physical environment is visible to the user through a representation of a portion of the physical environment displayed by the device. In some embodiments, the user's hand is not visible to the user, either directly or via the display of device 7100, and the current state of the user's hand is continuously or periodically monitored by the device to determine whether the user's hand has entered a ready state to provide gesture input. In some embodiments, the device displays an indicator of whether the user's hand is in a ready state to provide an input gesture to provide feedback to the user and to alert the user to adjust his / her hand position if he / she wishes to provide an input gesture.

[0166] In Fig.7I Part (A) of, the user's gaze is directed at virtual object 7506 (e.g., as indicated by the dashed line connecting the representation of the user's eyeball 7512 and virtual object 7506). In some embodiments, device 7100 uses one or more cameras facing the user (e.g., a front-facing camera) to track the movement of the user's eyes (or the movement of both of the user's eyes) to determine the direction and / or object at which the user is gazing. Figures 1 to 6 (Specifically Figures 5 and 6 ) and the accompanying description provide more details of eye tracking or gaze tracking techniques. In Fig.7I Part (A) of, in accordance with determining that the user's hand is not in a ready state for providing gesture input (e.g., the left hand is not stationary and has not been held in a first predefined ready state configuration for more than a threshold amount of time), no operation is performed relative to virtual object 7506 in response to the user directing his / her gaze at virtual object 7506. Similarly, in Fig.7I Part (B) of, the user's gaze has left virtual object 7506 and is now directed at virtual object 7508 (e.g., as indicated by the dashed line connecting the representation of the user's eyeball 7512 and virtual object 7508), and in accordance with determining that the user's hand is not in a ready state for providing gesture input, no operation is performed relative to virtual object 7508 in response to the user directing his / her gaze at virtual object 7508. Likewise, in Fig.7IIn part (C), the user's gaze is directed towards virtual object 7510 (e.g., as indicated by the dashed line connecting the representation of the user's eyeball 7512 and virtual object 7510). Based on determining that the user's hand is not in a ready state for providing gesture input, no operation is performed relative to virtual object 7510 in response to the user directing his / her gaze towards virtual object 7510. In some embodiments, it is advantageous to require the user's hand to be in a ready state for providing gesture input in order to trigger a visual change indicating that the virtual object under the user's gaze can be activated by gesture input, as this will tend to prevent unnecessary visual changes in the displayed environment when the user only wishes to examine the environment (e.g., briefly or intently gaze at various virtual objects for a period of time) rather than interact with any particular virtual object in the environment. When experiencing a computer-generated three-dimensional environment using a computer system, this reduces the user's visual fatigue and distraction, and thus reduces user errors.

[0167] Compared with Fig.7I the exemplary scenario shown in Figure 7J FIG. shows an example behavior of a computer-generated environment in accordance with some embodiments in response to the user directing his / her gaze towards a corresponding virtual object in the computer-generated environment when the user is ready to provide gesture input. As Figure 7J shown in parts (A) to (C) of FIG., when virtual objects 7506, 7608, and 7510 are displayed in the three-dimensional environment, the user holds his / her left hand in a first ready-state configuration for providing gesture input (e.g., where the thumb is resting on the index finger and the hand is raised relative to the user's body to a level above a preset level).

[0168] In Figure 7J part (A) of FIG., the user's gaze is directed towards virtual object 7506 (e.g., as indicated by the dashed line connecting the representation of the user's eyeball 7512 and virtual object 7506). Based on determining that the user's left hand is in a ready state for providing gesture input, in response to the user directing his / her gaze towards virtual object 7506 (e.g., at virtual object 7506, the gaze meets the duration and stability requirements), the computer system provides visual feedback indicating that virtual object 7506 can be activated by gesture input (e.g., virtual object 7506 is highlighted, expanded, or enhanced with additional information or user interface details to indicate that virtual object 7506 is interactive (e.g., one or more operations associated with virtual object 7506 can be performed in response to the user's gesture input)). Similarly, in Figure 7JIn part (B), the user's gaze has moved away from the virtual object 7506 and is now directed at the virtual object 7508 (e.g., as indicated by the dashed line connecting the representation of the user's eyeball 7512 and the virtual object 7508). Based on determining that the user's hand is in a ready state for providing gesture input, in response to the user directing his / her gaze at the virtual object 7508 (e.g., at the virtual object 7508, the gaze meets the stability and duration requirements), the computer system provides visual feedback indicating that the virtual object 7506 can be activated by gesture input (e.g., the virtual object 7508 is highlighted, expanded, or enhanced with additional information or user interface details to indicate that the virtual object 7508 is interactive (e.g., one or more operations associated with the virtual object 7508 can be performed in response to the user's gesture input)). Similarly, in Figure 7J In part (C), the user's gaze has moved away from the virtual object 7508 and is now directed at the virtual object 7510 (e.g., as indicated by the dashed line connecting the representation of the user's eyeball 7512 and the virtual object 7510). Based on determining that the user's hand is in a ready state for providing gesture input, in response to detecting that the user directs his / her gaze at the virtual object 7510, the computer system provides visual feedback indicating that the virtual object 7506 can be activated by gesture input (e.g., the virtual object 7510 is highlighted, expanded, or enhanced with additional information or user interface details to indicate that the virtual object 7510 is interactive (e.g., one or more operations associated with the virtual object 7510 can be performed in response to gesture input)).

[0169] In some embodiments, when displaying visual feedback indicating that a virtual object can be activated by gesture input, and in response to detecting gesture input starting from the user's hand in a ready state, the computer system performs an operation corresponding to the virtual object that is the object of the user's gaze according to the user's gesture input. In some embodiments, in response to the user's gaze moving away from the corresponding virtual object and / or the user's hand ceasing to be in a ready state for providing gesture input without providing valid gesture input, the display of the visual feedback indicating that the corresponding virtual object can be activated by gesture input is stopped.

[0170] In some embodiments, the respective virtual objects (e.g., virtual objects 7506, 7508, or 7510) correspond to applications (e.g., the respective virtual objects are application icons), and the operations associated with the respective virtual objects available for execution include launching the corresponding application, performing one or more operations within the application, or displaying a menu of operations to be performed with respect to or within the application. For example, in the case where the respective virtual object corresponds to a media player application, the one or more operations include: increasing the output volume of the media (e.g., in response to a thumb swipe gesture or a pinch and twist gesture in a first direction), decreasing the output volume (e.g., in response to a thumb swipe gesture or a pinch and twist gesture in a second direction opposite the first direction), switching the playback of the media (e.g., playing or pausing the media) (e.g., in response to a thumb tap gesture), fast forwarding, rewinding, browsing the media for playback (e.g., in response to multiple consecutive thumb swipe gestures in the same direction), or otherwise controlling media playback (e.g., menu navigation in response to a thumb flick gesture followed by a thumb swipe gesture). In some embodiments, the respective virtual object is a simplified user interface for controlling a physical object (e.g., an electronic household appliance, a smart speaker, a smart light, etc.) underlying the respective virtual object, and a wrist flick gesture or a thumb flick gesture detected when a visual indication that the respective virtual object is interactive causes the computer system to display an enhanced user interface for controlling the physical object (e.g., displaying an on / off button and the currently playing media album, as well as additional playback controls and output adjustment controls, etc.).

[0171] In some embodiments, the visual feedback indicating that the virtual object is interactive (e.g., in response to user input, including gesture input and other types of input, such as audio input and touch input, etc.) includes displaying one or more user interface objects or information, prompts that were not displayed before the user's gaze input was on the virtual object. In one example, in the case where the respective virtual object is a virtual window overlaid on a physical wall represented in a three-dimensional environment, in response to the user directing his / her gaze to the virtual window when the user's hand is in a ready state for providing gesture input, the computer system displays the location and / or the time of day associated with the virtual scene visible through the displayed virtual window to indicate that the scene can be changed (e.g., by changing the location, the time of day, the season, etc. according to the user's subsequent gesture input). In another example, in the case where the respective virtual object includes a displayed still photo (e.g., the respective virtual object is a picture frame), in response to the user directing his / her gaze to the displayed photo when the user's hand is in a ready state for providing gesture input, the computer system displays a multi-frame photo or a video clip associated with the displayed still photo to indicate that the photo is interactive and optionally to indicate that the photo can be changed (e.g., browsing an album according to the user's subsequent gesture input).

[0172] Figures 7K to 7M shows an exemplary view of a three - dimensional environment (e.g., a virtual reality environment or a mixed reality environment) when a display generation component (e.g., the display generation component 120 in Figure 1 , Figure 3 and Figure 4 ) of a computer system is placed at a predefined position relative to a user of the device (e.g., when the user initially enters a computer - generated reality experience (e.g., when the user holds the device in front of his / her eyes or when the user wears an HMD on his / her head)), and the exemplary view of the three - dimensional environment changes in response to detecting a change in the grip of the user's hand on the housing of the display generation component of the computer system (e.g., the computer system 101 in Figure 1 (e.g., a handheld device or an HMD)). The change in the view of the three - dimensional environment forms an initial transition into a computer - generated reality experience that is controlled by the user rather than being determined entirely by the computer system without user input (e.g., by changing his / her grip on the device or the housing of the display generation component). Regarding Figure 7G the input gestures described are used to illustrate the processes described below, including Fig.12 the processes in

[0173] Figure 7K Part (A) shows a physical environment 7800 in which a user (e.g., user 7802) is using a computer system. The physical environment 7800 includes one or more physical surfaces (e.g., walls, floors, surfaces of physical objects, etc.) and physical objects (e.g., physical object 7504, the user's hand, body, etc.). Figure 7K Part (B) shows an exemplary view 7820 of a three - dimensional environment (also referred to as the "first view 7820 of the three - dimensional environment" or "first view 7820") displayed by a display generation component of a computer system (e.g., device 7100 or an HMD). In some embodiments, the first view 7820 is displayed when the display generation component (e.g., the display of device 7100 or an HMD) is placed at a predefined position relative to user 7802. For example, in Figure 7KIn this case, the display of device 7100 is placed in front of the user's eyes. In another example, the computer system determines that the display generation component (e.g., HMD) is placed on the user's head such that the user's field of view of the physical environment can only be obtained through the display generation component, and determines that the display generation component is placed at a predefined position relative to the user. In some embodiments, the computer system determines that the display generation component is placed at a predefined position relative to the user based on determining that the user has positioned himself in front of the head-up display of the computer system. In some embodiments, placing the display generation component at a predefined position relative to the user, or positioning the user at a predefined position relative to the display generation component, allows the user to view content (e.g., real content or virtual content) through the display generation component. In some embodiments, once the display generation component and the user are in a predefined relative position, the user's field of view of the physical environment may be at least partially (or completely) blocked by the display generation component.

[0174] In some embodiments, the placement of the display generation component of the computer system is determined based on an analysis of data captured by the sensor system. In some embodiments, the sensor system includes one or more sensors that are components of the computer system (e.g., internal components enclosed in the same housing as the display generation component of device 7100 or HMD). In some embodiments, the sensor system is an external system and is not enclosed in the same housing as the display generation component of the computer system (e.g., the sensor is an external camera that provides captured image data to the computer system for data analysis).

[0175] In some embodiments, the sensor system includes one or more imaging sensors (e.g., one or more cameras) that track the movement of the user and / or the display generation component of the computer system. In some embodiments, the one or more imaging sensors track the position and / or movement of one or more features of the user (such as the user's hand and / or the user's head) to detect the placement of the display generation component relative to the user or a predefined portion of the user (e.g., the head, eyes, etc.). For example, the image data is analyzed in real time to determine whether the user is holding the display of device 7100 in front of the user's eyes or whether the user is putting on a head-mounted display on the user's head. In some embodiments, the one or more imaging sensors track the user's eye gaze to determine where the user is looking (e.g., whether the user is looking at the display). In some embodiments, the sensor system includes one or more touch-based sensors (e.g., mounted on the display) to detect the user's hand grip on the display, such as holding the device with one or two hands and / or holding the edge of device 7100, or using two hands to hold the head-mounted display to put on the head-mounted display on the user's head. In some embodiments, the sensor system includes one or more motion sensors (e.g., accelerometers) and / or position sensors (e.g., gyroscopes, GPS sensors, and / or proximity sensors) that detect the motion and / or position information (e.g., position, height, and / or orientation) of the display of the electronic device to determine the placement of the display relative to the user. For example, the motion and / or position data is analyzed to determine whether the mobile device is being lifted and facing the user's eyes, or whether the head-mounted display is being lifted and put on the user's head. In some embodiments, the sensor system includes one or more infrared sensors that detect the positioning of the head-mounted display on the user's head. In some embodiments, the sensor system includes a combination of different types of sensors to provide data for determining the placement of the display generation component relative to the user. For example, the grip of the user's hand on the housing of the display generation component, the motion and / or orientation information of the display generation component, and the user's eye gaze information are analyzed in combination to determine the placement of the display generation component relative to the user.

[0176] In some embodiments, based on the analysis of the data captured by the sensor system, it is determined that the display of the electronic device is placed at a predefined position relative to the user. In some embodiments, the predefined position of the display relative to the user indicates that the user is about to use the computer system to initiate a virtual and immersive experience (e.g., start playing a 3D movie, enter a 3D virtual world, etc.). For example, when the user's eye gaze is directed at the display screen or the user is using two hands to hold and lift the head-mounted display to put it on the user's head, the sensor data indicates that the user is holding the mobile device in the two palms of the user (e.g., Figure 7Kthe hand configurations shown in). In some embodiments, the computer system allows a period of time for the user to adjust the position of the display generation component relative to the user (e.g., to shift the HMD so that it is comfortable to wear and the display is well-aligned with the eyes), and during this time, changes in hand grip and position do not trigger any changes in the first view being displayed. In some embodiments, the initial hand grip being monitored for change is not a grip for holding the display generation component, but rather a touch of the hand or finger on a specific part of the display generation component (e.g., a switch or control for turning on the HMD or starting the display of virtual content). In some embodiments, the combination of placing the HMD on the user's head and activating the control to start the immersive experience hand grip is the initial hand grip being monitored for change.

[0177] In some embodiments, as Figure 7K shown in part (B) of, in response to detecting that the display is in a predefined position relative to the user, the display generation component of the computer system displays a first view 7820 of a three-dimensional environment. In some embodiments, the first view 7820 of the three-dimensional environment is a welcome / onboarding user interface. In some embodiments, the first view 7820 includes a passthrough portion that includes a representation of at least a portion of the physical environment 7800 around the user 7802.

[0178] In some embodiments, the passthrough portion is a transparent or translucent (e.g., see-through) portion of the display generation component that shows at least a portion of the physical environment 7800 around or within the field of view of the user 7802. For example, the passthrough portion is made translucent (e.g., less than 50%, 40%, 30%, 20%, 15%, 10%, or 5% opacity) or transparent in a head-mounted display such that the user can see a portion of the real world around them through it without removing the display generation component. In some embodiments, when the welcome / onboarding user interface changes to an immersive virtual or mixed reality environment, for example in response to subsequent changes in the user's hand grip indicating that the user is ready to enter the fully immersive environment, the passthrough portion gradually transitions from translucent or transparent to fully opaque.

[0179] In some embodiments, the passthrough portion of the first view 7820 displays a live feed of an image or video of at least a portion of the physical environment 7800 captured by one or more cameras (e.g., a rear-facing camera of a mobile device associated with the head-mounted display, or other cameras that feed image data to the electronic device). For example, the passthrough portion includes all or a portion of a display screen that displays a live image or video of the physical environment 7800. In some embodiments, the one or more cameras are pointed at a portion of the physical environment that is directly in front of the user's eyes (e.g., behind the display generation component). In some embodiments, the one or more cameras are pointed at a portion of the physical environment that is not directly in front of the user's eyes (e.g., in a different physical environment, or to the side or rear of the user).

[0180] In some embodiments, the first view 7820 of the three-dimensional environment includes three-dimensional virtual reality (VR) content. In some embodiments, the VR content includes one or more virtual objects corresponding to one or more physical objects (e.g., a bookshelf and / or a wall) in the physical environment 7800. For example, at least some of the virtual objects are displayed at positions corresponding to the positions of the physical objects in the corresponding physical environment 7800 in the virtual reality environment (e.g., using one or more cameras to determine the positions of the physical objects in the physical environment). In some embodiments, the VR content does not correspond to the physical environment 7800 viewed through the passthrough portion and / or is displayed independently of the physical objects in the passthrough portion. For example, the VR content includes virtual user interface elements (e.g., a virtual task bar or a virtual menu including user interface objects) or other virtual objects that are not related to the physical environment 7800.

[0181] In some embodiments, the first view 7820 of the three-dimensional environment includes three-dimensional augmented reality (AR) content. In some embodiments, one or more cameras (e.g., a rear-facing camera of a mobile device or associated with a head-mounted display, or other cameras that feed image data to a computer system) continuously provide a real-time view of at least a portion of the surrounding physical environment 7800 within the field of view of the one or more cameras, and the AR content corresponds to the portion of the surrounding physical environment 7800 within the field of view of the one or more cameras. In some embodiments, the AR content at least partially includes the real-time view of the one or more cameras. In some embodiments, the AR content includes one or more virtual objects that are displayed to replace a portion of the real-time view (e.g., appear to be superimposed on a portion of the real-time view or block the portion). In some embodiments, the virtual object is displayed in the virtual environment 7820 at a position corresponding to the position of the corresponding object in the physical environment 7800. For example, the corresponding virtual object is displayed to replace the corresponding physical object in the physical environment 7800 (e.g., superimposed on the corresponding physical object, block the corresponding physical object and / or replace the view of the corresponding physical object).

[0182] In some embodiments, in the first view 7820 of the three-dimensional environment, the see-through portion (e.g., representing at least a portion of the physical environment 7800) is surrounded by virtual content (e.g., VR and / or AR content). For example, the see-through portion does not overlap with the virtual content on the display. In some embodiments, in the first view 7820 of the three-dimensional virtual environment, the VR and / or AR virtual content is displayed to replace the see-through portion (e.g., superimposed on top of the see-through portion or replacing the content displayed in the see-through portion). For example, virtual content (e.g., a virtual taskbar or virtual start menu listing multiple virtual user interface elements) is superimposed on top of or blocks a portion of the physical environment 7800 displayed through the translucent or transparent see-through portion. In some embodiments, the first view 7820 of the three-dimensional environment initially includes only the see-through portion without any virtual content. For example, when the user initially holds the device in the palm of the user's hand (e.g., as in Figure 7L ), or when the user initially puts the head mounted display on the user's head, the user sees a portion of the physical environment through the see-through portion that is within the field of view of the user's eyes or within the field of view of the live feed camera. Then, the virtual content (e.g., a welcome / getting started user interface with virtual menus / icons) gradually fades in to overlay or block the see-through portion or block the see-through portion for a period of time while the user's hand grip remains unchanged. In some embodiments, the welcome / getting started user interface remains displayed (e.g., in a stable state where both the virtual content and the see-through portion show the physical world) as long as the user's hand grip does not change.

[0183] In some embodiments, achieving a user's virtual immersive experience causes the user's current view of the surrounding real world to be temporarily blocked by a display generation component (e.g., by presenting a display directly in front of the user's eyes and providing a noise cancellation function for a head-mounted display). This occurs at a time point before the user's virtual immersive experience begins. By having a passthrough portion in the welcome / onboarding user interface, the transition from seeing the physical environment around the user to the user's virtual immersive experience benefits from better control and a smoother transition (e.g., a cognitively gentle transition). This allows the user to have more control over how much time is needed after seeing the welcome / onboarding user interface to be ready to enter a fully immersive experience, rather than having the computer system or content provider dictate the timing of the transition to a fully immersive experience for all users.

[0184] Figure 7M Another exemplary view 7920 (also referred to as the "second view 7920 of the three-dimensional environment" or "second view 7920") of a three-dimensional environment displayed by a display generation component of a computer system (e.g., on the display of device 7100) is shown. In some embodiments, the second view 7920 of the three-dimensional environment replaces the first view 7820 of the three-dimensional environment in response to detecting a change in the grip of the user's hand on the housing of the display generation component of the computer system (e.g., changing from the hand configuration in part (B) of Figure 7K (e.g., a two-handed grip) to the hand configuration in part (B) of Figure 7L (e.g., a one-handed grip)), where the change in grip meets a first predetermined criterion (e.g., a criterion corresponding to detecting a sufficient decrease in the user's control or alertness).

[0185] In some embodiments, the change in the grip of the user's hand is detected by a sensor system as discussed above with reference to Figure 7K . For example, one or more imaging sensors track the movement and / or position of the user's hand to detect the change in the grip of the user's hand. In another example, one or more touch-based sensors on the display detect the change in the grip of the user's hand.

[0186] In some embodiments, the first predetermined criterion for a change in the grip of the user's hand requires a change in the total number of hands detected on the display (e.g., changing from two hands to one hand, or from one hand to no hands, or from two hands to no hands), a change in the total number of fingers in contact with the display generating component (e.g., changing from eight fingers to six fingers, from four fingers to two fingers, from two fingers to no fingers, etc.), a change from having hand contact on the display generating component to having no hand contact, a change in the contact location (e.g., changing from the palm to the fingers) and / or a change in the contact intensity on the display (e.g., which is caused by changes in hand posture, orientation, relative gripping forces of different fingers on the display generating component). In some embodiments, a change in the grip of the hand on the display does not cause a change in the predefined position of the display relative to the user (e.g., a head-mounted display remains covering the user's eyes on the user's head). In some embodiments, a change in the grip of the hand indicates that the user (e.g., gradually or firmly) releases the display and is ready to immerse in a virtual immersive experience.

[0187] In some embodiments, the initial hand grip being monitored for change is not a grip for holding the display generating component, but rather a touch of the hand or fingers on a specific part of the display generating component (e.g., a switch or control for turning on the HMD or starting the display of virtual content), and the first predetermined criterion for a change in the grip of the user's hand requires the fingers touching the specific part of the display generating component (e.g., the fingers activating the switch or control for turning on the HMD or starting the display of virtual content) to stop touching the specific part of the display generating component.

[0188] In some embodiments, the second view 7920 of the three-dimensional environment replaces at least a portion of the passthrough portion in the first view 7820 with virtual content. In some embodiments, the virtual content in the second view 7920 of the three-dimensional environment includes VR content (e.g., virtual objects 7510 (e.g., virtual user interface elements or system affordance representations)), AR content (e.g., virtual objects 7506 (e.g., virtual windows overlaid on a live view of a wall captured by one or more cameras)), and / or virtual objects 7508 (e.g., shown as replacing a portion or all of the representation 7504' of a physical object 7504 in the physical environment or superimposed over a portion or all of the representation).

[0189] In some embodiments, replacing the first view 7820 with the second view 7920 includes increasing the opacity of the passthrough portion (e.g., when the passthrough portion is implemented with a translucent or transparent state of the display), such that the virtual content superimposed over the translucent or transparent portion of the display becomes more visible and has a more saturated color. In some embodiments, the virtual content in the second view 7920 provides a more immersive experience to the user than the virtual content in the first view 7820. For example, the virtual content in the first view 7820 is displayed in front of the user, while the virtual content in the second view 7920 includes a panoramic or 360-degree view of a three-dimensional world represented as the user turns his / her head and / or walks around. In some embodiments, compared to the first view 7820, the second view 7920 includes a smaller passthrough portion that shows a smaller or lesser portion of the physical environment 7800 surrounding the user. For example, the passthrough portion of the first view 7820 shows a real window on one of the walls in the room where the user is located, and the passthrough portion of the second view 7920 shows a window on one of the walls that has been replaced with a virtual window, such that the area of the passthrough portion is reduced in the second view 7920.

[0190] Figure 7M Another exemplary third view 7821 is shown (e.g., the first view 7820 of the three-dimensional environment or a modified version thereof, or a different view), and the display generation component of the computer system displays this exemplary third view in response to detecting the initial hand grip configuration again on the housing of the display generation component (e.g., after displaying the second view 7920 in response to detecting a desired hand grip change, as Figure 7L shown). In some embodiments, the third view 7821 re-establishes the passthrough portion in response to detecting another hand grip change of the user's hand on the display generation component of the computer system (e.g., a change in the hand configuration from Figure 7L or a change from no hand grip to Figure 7M ). In some embodiments, the hand grip change of the user's hand represents the re-establishment of the hand grip of the user on the housing of the display generation component and indicates that the user wants to (e.g., partially or fully, gradually or immediately) exit the virtual immersive experience.

[0191] In some embodiments, the sensor system detects changes in the total number of hands detected on the housing of the display generating component (e.g., changing from one hand to two hands, or from no hands to two hands), increases in the total number of fingers in contact with the housing of the display generating component, changes from no hand contact to hand contact on the housing of the display generating component, changes in the contact location (e.g., from a finger to a palm), and / or changes in the contact intensity on the housing of the display generating component. In some embodiments, the re - establishment of the user's hand grip causes a change in the position and / or orientation of the display generating component (e.g., a change in the position and angle of device 7100 relative to the environment in part (A) of Figure 7M compared to the angle in part (A) of Figure 7L ). In some embodiments, a change in the user's hand grip causes a change in the viewing angle of the user relative to the physical environment 7800 (e.g., a change in the viewing angle of device 7100 or one or more cameras of device 7100 relative to the physical environment 7800). Accordingly, the viewing angle of the displayed third view 7821 is changed (e.g., including changing the viewing angle of the passthrough portion and / or virtual objects on the display). Figure 7M the position and angle of device 7100 relative to the environment in part (A) (compared to the angle in part (A) of Figure 7L ) Figure 7L In some embodiments, a change in the user's hand grip causes a change in the viewing angle of the user relative to the physical environment 7800 (e.g., a change in the viewing angle of device 7100 or one or more cameras of device 7100 relative to the physical environment 7800). Accordingly, the viewing angle of the displayed third view 7821 is changed (e.g., including changing the viewing angle of the passthrough portion and / or virtual objects on the display).

[0192] In some embodiments, the passthrough portion in the third view 7821 is the same as the passthrough portion in the first view 7820, or is at least increased relative to the passthrough portion in the second view 7920 (if any). In some embodiments, the passthrough portion in the third view 7821 shows physical objects 7504 at different viewing angles in the physical environment 7800 compared to the passthrough portion in the first view 7820. In some embodiments, the passthrough portion in the third view 7821 is a transparent or translucent perspective portion of the display generating component. In some embodiments, the passthrough portion in the third view 7821 displays a live feed from one or more cameras configured to capture image data of at least a portion of the physical environment 7800. In some embodiments, there is no virtual content displayed together with the passthrough portion in the third view 7821. In some embodiments, the virtual content is paused or becomes translucent or less saturated in color in the third view 7821 and is displayed simultaneously with the passthrough portion in the third view 7821. When the third view is displayed, the user can restore a fully immersive experience by changing the hand grip again, as described with respect to Figures 7K to 7L . Figures 7K to 7L as described.

[0193] Figure 7N to Figure 7P An exemplary view of a three - dimensional virtual environment is shown according to some embodiments, the exemplary view of the three - dimensional virtual environment changing in response to detecting a change in the user's position relative to objects (e.g., obstacles or targets) in the physical environment surrounding the user. With respect to Figure 7N to Figure 7P Figure 7N to Figure 7PThe input gesture described above is used to illustrate the process described below, including Fig.13 the process in

[0194] In Figure 7N part (A) of , user 7802 holds device 7100 in physical environment 7800. The physical environment includes one or more physical surfaces and physical objects (e.g., walls, floors, physical object 7602). Device 7100 displays virtual three-dimensional environment 7610 without displaying a see-through portion showing the physical environment around the user. In some embodiments, device 7100 may be represented and replaced by an HMD or other computer system including a display generation component that blocks the user's view of the physical environment when displaying virtual environment 7610. In some embodiments, the HMD or display generation component of the computer system at least encloses the user's eyes, and the user's view of the physical environment is partially or completely blocked by the virtual content displayed by the display generation component and other physical barriers or portions of their enclosures formed by the display generation component.

[0195] Figure 7N Part (B) of shows a first view 7610 of the three-dimensional environment displayed by a display generation component (also referred to as the "display") of a computer system (e.g., device 7100 or HMD).

[0196] In some embodiments, first view 7610 is a three-dimensional virtual environment that provides an immersive virtual experience (e.g., a three-dimensional movie or game). In some embodiments, first view 7610 includes three-dimensional virtual reality (VR) content. In some embodiments, the VR content includes one or more virtual objects corresponding to one or more physical objects in a physical environment that does not correspond to the physical environment 7800 around the user. For example, at least some of the virtual objects are displayed at positions corresponding to the positions of the physical objects in a physical environment (which is remote from physical environment 7800) in the virtual reality environment. In some embodiments, the first view includes virtual user interface elements (e.g., a virtual task bar or virtual menu including user interface objects) or other virtual objects that are not related to physical environment 7800.

[0197] In some embodiments, the first view 7610 includes 100% virtual content (e.g., virtual objects 7612 and virtual surfaces 7614 (e.g., virtual walls and floors)) that does not include and is different from any representation of the physical environment 7800 surrounding the user 7802. In some embodiments, the virtual content (e.g., virtual objects 7612 and virtual surfaces 7614) in the first view 7610 does not correspond to or visually convey the presence, location, and / or physical structure of any physical object in the physical environment 7800. In some embodiments, the first view 7610 optionally includes a virtual representation that indicates the presence and location of a first physical object in the physical environment 7800 but does not visually convey the presence, location, and / or physical structure of a second physical object in the physical environment 7800, where both the first physical object and the second physical object would be within the user's field of view if the user's field of view were not blocked by the display generation component. In other words, the first view 7610 includes virtual content that replaces the display of at least some physical objects or portions thereof that would be present in the user's normal field of view (e.g., the user's field of view without a display generation component placed in front of the user's eyes).

[0198] Fig.7O Another exemplary view 7620 of a three-dimensional virtual environment (also referred to as the "second view 7620 of the three-dimensional environment," the "second view 7620 of the virtual environment," or the "second view 7620") displayed by a display generation component of a computer system is shown. In some embodiments, the sensor system detects that the user 7802 is moving toward a physical object 7602 in the physical environment 7800, and the sensor data acquired by the sensor system is analyzed to determine whether the distance between the user 7802 and the physical object 7602 is within a predefined threshold distance (e.g., within the length of the user's arm or the length of a normal gait). In some embodiments, when it is determined that a portion of the physical object 7602 is within the threshold distance from the user 7602, the appearance of the view of the virtual environment is changed to indicate the physical characteristics of a portion of the physical object 7602 (e.g., Fig.7OThe second view 7620 in part (B) shows the portion 7604 of the physical object 7602 that is within the threshold distance from the user 7802, and does not show other portions of the physical object 7602 that are also within the user's field of view but outside the user's threshold distance from the portion 7604. In some embodiments, rather than replacing a portion of the virtual content with a direct view or camera view of the portion 7604 of the physical object, the visual characteristics (e.g., opacity, color, texture, virtual material, etc.) of a portion of the virtual content at a location corresponding to the portion 7604 of the physical object are changed to indicate the physical characteristics (e.g., size, color, pattern, structure, contour, shape, surface, etc.) of the portion 7604 of the physical object. The change to the virtual content at the location corresponding to the portion 7604 of the physical object 7602 is not applied to other portions of the virtual content, including portions of the virtual content at locations corresponding to portions of the physical object 7602 outside of the portion 7604. In some embodiments, the computer system provides a blend (e.g., makes the visual transition smooth) between the portion of the virtual content at the location corresponding to the portion 7604 of the physical object and the portion of the virtual content immediately outside the location corresponding to the portion 7604 of the physical object.

[0199] In some embodiments, the physical object 7602 is a static object in the physical environment 7800, such as a wall, chair, or table. In some embodiments, the physical object 7602 is a moving object in the physical environment 7800, such as another person or a dog that moves relative to the user 7802 in the physical environment 7800 while the user 7802 is static relative to the physical environment 7800 (e.g., the user's pet moves around while the user is sitting on the couch watching a movie).

[0200] In some embodiments, when the user 7802 is enjoying a three - dimensional immersive virtual experience (e.g., including a panoramic three - dimensional display with surround sound effects and other virtual sensory perceptions), and real - time analysis of sensor data from a sensor system coupled to the computer system indicates that the user 7802 is close enough to the physical object 7602 (e.g., by the user moving towards the physical object, or the physical object moving towards the user), the user 7802 can benefit from receiving a warning that blends with the virtual environment in a smooth and less disruptive manner. This allows the user to make a more informed decision about whether to modify his / her movement and / or stop / continue the immersive experience without losing the immersive quality of the experience.

[0201] In some embodiments, a second view 7620 is displayed when analysis of sensor data shows that user 7802 is within a threshold distance of at least a portion of physical object 7602 in physical environment 7800 (e.g., physical object 7602 has a range that is potentially visible to the user based on the user's field of view for the virtual environment). In some embodiments, the computer system requires that, given the position of a portion of the physical object relative to the user in physical environment 7800, if the display has a see-through portion or the display generation component is not in front of the user's eyes, that portion of the physical object would be visible in the user's field of view.

[0202] In some embodiments, portion 7604 in the second view 7620 of the virtual environment includes a translucent visual representation of the corresponding portion of physical object 7602. For example, the translucent representation is overlaid on the virtual content. In some embodiments, portion 7604 in the second view 7620 of the virtual environment includes a vitreous appearance of the corresponding portion of physical object 7602. For example, when user 7802 moves closer to a table placed in a room while enjoying an immersive virtual experience, the portion of the table closest to the user is shown to have a shiny translucent perspective appearance (e.g., a virtual ball or virtual grass in the virtual view) overlaid on the virtual content, and the virtual content behind that portion of the table is visible through the vitreous appearance of the portion of the table. In some embodiments, the second view 7620 of the virtual environment shows a predefined distortion or other visual effect (e.g., blinking, rippling, glowing, dimming, blurring, rotational visual effect, or different text effects) applied to portion 7604 corresponding to the portion of physical object 7602 closest to user 7802.

[0203] In some embodiments, when the user moves towards the corresponding portion of physical object 7602 and enters within the threshold distance of that corresponding portion, the second view 7620 of the virtual environment immediately replaces the first view 7610 to provide the user with a timely warning. In some embodiments, for example, the second view 7620 of the virtual environment is gradually displayed with a fade-in / fade-out effect to provide a smoother transition and a less disruptive / intrusive user experience. In some embodiments, the computer system allows the user to navigate within the three-dimensional environment by moving in the physical environment and change the view of the three-dimensional environment presented to the user such that it reflects computer-generated movement within the three-dimensional environment. For example, as Figure 7N and 7OAs shown, when the user walks towards the physical object 7602, the user perceives his / her movement as moving towards the virtual object 7612 in the same direction in the three-dimensional virtual environment (e.g., seeing the virtual object 7612 getting closer and larger). In some embodiments, except when the user has reached within a threshold distance of the physical object in the physical environment, the virtual content presented to the user is independent of the user's movement in the physical environment and does not change based on the user's movement in the physical environment.

[0204] Figure 7P Another exemplary view 7630 of the three-dimensional environment (also referred to as the "third view 7630 of the three-dimensional environment", "third view 7630 of the virtual environment", or "third view 7630") is shown as displayed by the display generation component of the computer system (e.g., device 7100 or HMD). In some embodiments, when the user 7802 continues to move towards the physical object 7602 in the physical environment 7800 after viewing the second view 7620 of the virtual three-dimensional environment as discussed with reference to Fig.7O the analysis of the sensor data shows that the distance between the user 7802 and the portion 7606 of the physical object 7602 is below a predefined threshold distance. In response, the display transitions from the second view 7620 to the third view 7630. In some embodiments, depending on the structure (e.g., size, shape, length, width, etc.) and relative position of the user and the physical object 7602, the portion 7606 and the portion 7604 of the physical object 7602 (which are within the user's predefined threshold distance when the user is at different positions in the physical environment 7800) are optionally completely different and non-overlapping portions of the physical object, the portion 7606 optionally completely contains the portion 7604, the portion 7606 and the portion 7604 optionally only partially overlap, or the portion 7604 optionally completely contains the portion 7606. In some embodiments, when the user moves around the room relative to those physical objects, portions or the whole of one or more other physical objects may be visually represented or not represented in the current display view of the virtual three-dimensional environment, depending on whether those physical objects are within or outside the user's predefined threshold distance.

[0205] In some embodiments, the computer system optionally allows the user to pre-select a subset of physical objects in the physical environment 7800 in order to monitor the distance between the user and the pre-selected physical objects and apply visual changes to the virtual environment. For example, the user may wish to pre-select furniture and pets as a subset of physical objects and not select clothing, curtains, etc. as a subset of physical objects, and will not apply visual changes to the virtual environment so as not to alert the user to the presence of clothing and curtains even if the user walks into them. In some embodiments, whether or not the user is within a threshold distance of the physical object, the computer system allows the user to pre-specify one or more physical objects that are always visually represented in the virtual environment by applying visual effects (e.g., changes in transparency, opacity, emission, refractive index, etc.) to the portion of the virtual environment corresponding to the respective location of the physical object. These visual cues help the user to orient themselves relative to the real world even when immersed in the virtual world and feel safer and more secure when exploring the virtual world.

[0206] In some embodiments, as Figure 7P shown, as the user moves closer to the physical object, the third view 7630 includes a rendering of the portion 7606 of the physical object 7602 that is within the threshold distance from the user. In some embodiments, as the distance between the user and the portion of the physical object decreases, the computer system optionally further increases the value of the display characteristic of the visual effect applied to the portion of the virtual environment that indicates the corresponding portion of the physical object 7602. For example, as the user gradually moves closer to the portion of the physical object, the computer system optionally increases the refractive index, color saturation, visual effect, opacity, and / or sharpness of the portion of the virtual environment corresponding to the portion of the physical object in the third view 7630. In some embodiments, the spatial extent of the visual effect increases as the user 7802 moves closer to the physical object 7602, and the corresponding portion of the physical object 7602 appears larger in the user's field of view of the virtual environment. For example, as the user 7802 moves closer to the physical object 7602 in the physical environment 7800, for at least two reasons: (1) more portions of the physical object 7602 are within the user's predefined distance, and (2) the same portion of the physical object 7602 (e.g., portion 7604) occupies a larger portion of the user's field of view of the virtual environment because it is closer to the user's eyes, the portion 7606 in the third view 7630 appears to gradually increase in size compared to the portion 7604 in the second view 7620 and extends out from the virtual object 7612 and in the direction towards the user.

[0207] In some embodiments, a computer system defines a gesture input (e.g., a user raises one or both arms to a preset level relative to the user's body within a threshold amount of time (e.g., a sudden and unexpected movement of a muscle reflex to prevent a fall or collision with something)), the gesture input causing a portion of a physical object that is partially within a threshold distance of the user (e.g., all portions potentially visible within the user's field of view for the virtual environment) or all physical objects potentially visible in the user's field of view for the virtual environment to be visually represented in the virtual environment by modifying display characteristics of the virtual environment at locations corresponding to the physical object or those portions of all physical objects. This feature helps to allow a user to quickly reorient when feeling that his / her body position in the physical environment is unsafe without having to fully exit the immersive experience.

[0208] The following is referenced with respect to the Figures 8 to 13 methods 8000, 9000, 10000, 11000, 12000, and 13000 described below provide additional description regarding FIG. 7A to FIG. 7P of.

[0209] Figure 8 is a flowchart of an exemplary method 8000 for interacting with a three-dimensional environment using predefined input gestures according to some embodiments. In some embodiments, method 8000 is executed at a computer system (e.g., Figure 1 the computer system 101 in Figure 1 , Figure 3 and Figure 4 the display generation component 120 in Figure 1 ), the computer system including a display generation component (e.g.,

[0210] In method 8000, a computer system displays (8002) a view of a three-dimensional environment (e.g., a virtual or mixed reality environment). When displaying the view of the three-dimensional environment, the computer system uses one or more cameras (e.g., using one or more cameras positioned on the lower edge of the HMD, rather than using a touch-sensitive glove, or a touch-sensitive surface on a hand-held input device, or other non-image-based means (e.g., acoustic waves)) to detect (8004) movement of a user's thumb of a first hand (e.g., a left or right hand not wearing a glove or not covered with or attached to an input device / surface) above the user's index finger. For example, this is shown in Fig. 7A and the accompanying description (e.g., thumb tap gesture, thumb swipe gesture, and thumb flick gesture). In some embodiments, the user's hand or a graphical representation of the hand is displayed in the view of the three-dimensional environment (e.g., in the passthrough portion of the display generation component or as part of an augmented reality view of the physical environment surrounding the user). In some embodiments, the user's hand or a graphical representation of the hand is not shown in the view of the three-dimensional environment or is displayed in a portion of the display outside the view of the three-dimensional environment (e.g., in a separate or floating window). Benefits of using one or more cameras, especially cameras that are part of the HMD, include: the spatial position and size of the user's hand appear to the user as the hand's natural state in the physical or virtual environment with which the user is interacting, and the user gains an intuitive sense of scale, orientation, and anchor position to perceive the three-dimensional environment on the display without the need for additional calculations to match the space of the input device to the three-dimensional environment and / or without the need to otherwise scale, rotate, and translate the representation of the user's hand before placing it in the displayed three-dimensional environment. Referring back to Figure 8, in response to detecting movement of a user's thumb above the user's index finger using one or more cameras (e.g., rather than a more exaggerated gesture created by waving a finger or hand in the air or sliding on a touch-sensitive surface) (8006): Based on determining that the movement is a swipe of the thumb of a first hand above the index finger in a first direction (e.g., a movement along a first axis (e.g., the x-axis) of the x-axis and y-axis, where the movement along the x-axis is a movement along the length of the index finger and the movement along the y-axis is a movement in a direction across the index finger (e.g., substantially perpendicular to the movement along the length of the index finger)), the computer system performs a first operation (e.g., changing a selected user interface object in the displayed user interface (e.g., iteratively selecting items in a list of items in a first direction corresponding to the first direction (left and right in an item row); adjusting the position of the user interface object in the displayed user interface (e.g., moving the object in a direction corresponding to the first direction (e.g., left and right) in the user interface); and / or adjusting system settings of the device (e.g., adjusting volume, moving to a subsequent list item, moving to a previous list item, jumping forward (e.g., fast forwarding and / or advancing to the next chapter, audio track, and / or content item), jumping backward (e.g., rewinding and / or moving to the previous chapter, audio track, and / or content item))). In some embodiments, a swipe in a first sub-direction (e.g., toward the fingertip of the index finger) of the first direction (e.g., along the length of the index finger) corresponds to performing the first operation in one way, and a swipe in a second sub-direction (e.g., toward the base of the index finger) of the first direction corresponds to performing the first operation in another way. For example, this is shown in Figure 7B , Figure 7C , Figure 7F and the accompanying description. Referring back to Figure 8, in response to detecting, using one or more cameras, movement of a user's thumb above the user's index finger (e.g., rather than a more exaggerated gesture created by waving a finger or hand in the air or sliding on a touch-sensitive surface) (8006): Based on determining that the movement is a tap of the thumb of a first hand above the index finger at a first position on the index finger (e.g., at a first portion of the index finger such as the distal phalanx, middle phalanx, and / or proximal phalanx), the computer system performs a second operation different from a first operation (e.g., performs an operation corresponding to a currently selected user interface object, and / or changes the selected user interface object in the displayed user interface). In some embodiments, performing the first / second operation includes changing the view of a three-dimensional user interface, and the change depends on the current operation context. In other words, each gesture triggers different operations and corresponding changes in the view of the three-dimensional environment in a corresponding manner, depending on the current operation context (e.g., which object the user is looking at, which direction the user is facing, the last function performed immediately before the current gesture, and / or which object is currently selected). For example, this is shown in Figure 7B , Figure 7C , Figure 7F and the accompanying description.

[0211] In some embodiments, in response to detecting, using one or more cameras, movement of a user's thumb above the user's index finger, based on determining that the movement is a swipe of the thumb of a first hand above the index finger in a second direction substantially perpendicular to a first direction (e.g., a movement along a second axis (e.g., the y-axis) among the x-axis and the y-axis, where the movement along the x-axis is a movement along the length of the index finger and the movement along the y-axis is a movement in a direction across the index finger (e.g., substantially perpendicular to the movement along the length of the index finger)), the computer system performs a third operation different from the first operation and different from the second operation (e.g., changes the selected user interface object in the displayed user interface (e.g., iteratively selects in a second direction corresponding to the second direction (e.g., up and down among multiple item rows in a 2D menu, or up or down in a vertically arranged list)); adjusts the position of the user interface object in the displayed user interface (e.g., moves the object in a direction corresponding to the second direction (e.g., up and down) in the user interface); and / or adjusts the system settings of the device (e.g., volume)). In some embodiments, the third operation is different from the first operation and / or the second operation. In some embodiments, a swipe in a first sub-direction of the second direction (e.g., around the index finger away from the palm) corresponds to performing the third operation in one way, and a swipe in a second sub-direction of the second direction (e.g., around the index finger towards the palm) corresponds to performing the third operation in another way.

[0212] In some embodiments, in response to detecting, using one or more cameras, movement of a user's thumb above the user's index finger, and determining that the movement is a movement of the thumb above the index finger in a third direction different from a first direction (and a second direction) (and not a tap of the thumb above the index finger), the computer system performs a fourth operation different from the first operation and different from the second operation (and different from a third operation). In some embodiments, the third direction is a direction away from the index finger and upward from the index finger (e.g., opposite to tapping on that side of the index finger), and the gesture is a flick of the thumb away from the index finger and the palm from that side of the index finger. In some embodiments, such an upward flick gesture using the thumb across the middle phalanx of the index finger causes the currently selected user interface object to be pushed into a three-dimensional environment and initiates an immersive experience corresponding to the currently selected user interface object (e.g., movie icon, application icon, image, etc.) (e.g., 3D movie or 3D virtual experience, panoramic display mode, full-screen mode, etc.). In some embodiments, during the immersive experience, a downward swipe across the middle phalanx of the index finger toward the palm (e.g., a movement in one of the sub-directions of the second direction and not a tap on the middle phalanx of the index finger) causes the immersive experience to be paused, stopped, and / or reduced to a reduced immersion state (e.g., non-full screen, 2D mode, etc.).

[0213] In some embodiments, performing a first operation includes: increasing a value corresponding to the first operation (e.g., a value set by the system, a value indicating the position and / or selection of at least a portion of a user interface (e.g., a user interface object), and / or a value corresponding to selected content or a portion of the content). For example, increasing the value includes: moving in a first predefined portion of the index finger according to a determined swipe of the thumb above the index finger in a first direction (e.g., along the length of the index finger, or around the index finger), such as moving towards the tip of the index finger, or towards the dorsal side of the index finger, increasing the volume, moving an object in an increasing direction (e.g., up and / or right), and / or adjusting a position (e.g., in a list and / or content item) to a subsequent position or other later position. In some embodiments, performing the first operation further includes: moving away from a first predefined portion of the index finger (e.g., away from the tip of the index finger, or away from the dorsal side (on the dorsal side of the hand) of the index finger) and towards a second predefined portion of the index finger (e.g., towards the base of the index finger, or towards the front side (on the palmar side of the hand) of the index finger) according to a determined swipe of the thumb above the index finger in the first direction (e.g., along the length of the index finger, or around the index finger), and decreasing the value corresponding to the first operation (e.g., decreasing the value includes: decreasing the volume, moving an object in a decreasing direction (e.g., down and / or left), and / or adjusting a position (e.g., in a list and / or content item) to a previous position or other earlier position). In some embodiments, the swipe direction of the thumb above the index finger in a second direction also determines the direction of a third operation in a manner similar to how the swipe direction in the first direction determines the direction of the first operation.

[0214] In some embodiments, performing a first operation includes: adjusting a value corresponding to the first operation (e.g., a value set by the system, a value indicating the position and / or selection of at least a portion of a user interface (e.g., a user interface object), and / or a value corresponding to selected content or a portion of the content) by an amount corresponding to the movement of the thumb above the index finger. In some embodiments, the movement of the thumb is measured at a threshold position on the index finger, and the value corresponding to the first operation is adjusted between multiple discrete levels according to the reached threshold position. In some embodiments, the movement of the thumb is measured continuously, and the value corresponding to the first operation is adjusted continuously and dynamically based on the current position of the thumb on the index finger (e.g., along or around the index finger). In some embodiments, the speed of movement of the thumb is used to determine the magnitude of the operation and / or a threshold, which is used to determine when different discrete values of the operation are triggered.

[0215] In some embodiments, in response to detecting, using one or more cameras, movement of a user's thumb above a user's index finger, and based on determining that the movement is a tap of the thumb above the index finger at a second position on the index finger that is different from a first position (e.g., at a portion and / or phalanx of the index finger), the computer system performs a fifth operation that is different from a second operation (e.g., performs an operation corresponding to a currently selected user interface object, and / or changes the selected user interface object in the displayed user interface). In some embodiments, the fifth operation is different from the first operation, the third operation, and / or the fourth operation. In some embodiments, tapping on the middle phalanx portion of the index finger activates the currently selected object, and tapping on the tip of the index finger minimizes / pauses / closes the currently active application or experience. In some embodiments, detecting a tap of the thumb on the index finger does not require detecting the lifting of the thumb from the index finger, and when the thumb remains on the index finger, movement of the thumb or the entire hand can be considered a movement combined with a tap-and-hold input of the thumb (e.g., for dragging an object).

[0216] In some embodiments, the computer system detects, using one or more cameras, a swipe of a user's thumb above a user's middle finger (e.g., when detecting that the user's index finger extends away from the middle finger). In response to detecting the swipe of the user's thumb above the user's middle finger, the computer system performs a sixth operation. In some embodiments, the sixth operation is different from the first operation, the second operation, the third operation, the fourth operation, and / or the fifth operation. In some embodiments, the swipe of the user's thumb above the middle finger includes movement of the thumb along the length of the middle finger (e.g., from the base towards the tip of the middle finger, or vice versa), and one or more different operations are performed based on determining that the swipe of the user's thumb above the middle finger includes movement of the thumb along the length of the middle finger from the tip towards the base of the middle finger and / or movement across the middle finger from the palm side of the middle finger to the top of the middle finger.

[0217] In some embodiments, the computer system detects, using the one or more cameras, a tap of a user's thumb above a user's middle finger (e.g., when detecting that the user's index finger extends away from the middle finger). In response to detecting the tap of the user's thumb above the user's middle finger, the computer system performs a seventh operation. In some embodiments, the seventh operation is different from the first operation, the second operation, the third operation, the fourth operation, the fifth operation, and / or the sixth operation. In some embodiments, the tap of the user's thumb above the middle finger is at a first position on the middle finger, and different operations are performed based on determining that the tap of the user's thumb above the middle finger is at a second position that is different from the first position on the middle finger. In some embodiments, flicking upwards from the first position and / or the second position on the middle finger causes the device to perform other operations that are different from the first operation, the second operation,..., and / or the seventh operation.

[0218] In some embodiments, the computer system displays, in a three-dimensional environment, a visual indication of the operating context of a thumb gesture (e.g., the thumb swiping / tapping / flicking on other fingers of the hand) (e.g., displaying a menu of selectable options, a dial for adjusting a value, an avatar of a digital assistant, a selection indicator for a currently selected object, highlighting of an interactive object, etc.) (e.g., when the device uses one or more cameras to detect that the user's hand is in or enters a predefined ready state (e.g., the thumb is resting on or hovering over one side of the index finger, and / or flicking the wrist in the case where the back side of the thumb faces upward / the thumb is resting on one side of the index finger), or the thumb of the hand faces laterally toward the camera), and displays a plurality of user interface objects in the three-dimensional environment, where the user interface objects respond to swiping and tapping gestures of the thumb on other fingers of the hand), where performing a first operation (or a second operation, a third operation, etc.) includes: displaying, in the three-dimensional environment, a visual change corresponding to the execution of the first operation (or the second operation, the third operation, etc.) (e.g., displaying the visual change includes: activating the corresponding user interface object among the plurality of user interface objects and causing the operation associated with the corresponding user interface object to be performed).

[0219] In some embodiments, when displaying a visual indication of the operating context of a thumb gesture (e.g., when displaying the plurality of user interface objects in a three-dimensional environment in response to detecting that the user's hand is in a predefined ready state), the computer system uses one or more cameras to detect the movement of the user's first hand (e.g., the movement of the entire hand relative to the camera in the physical environment, rather than the internal movement of the fingers relative to each other) (e.g., detecting the movement of the hand while the hand remains in the ready state) (e.g., detecting the movement and / or rotation of the hand / wrist in a three-dimensional environment). In response to detecting the movement of the first hand, the computer system changes the display position in the three-dimensional environment of the visual indication of the operating context of the thumb gesture (e.g., the plurality of user interface objects) according to the detected change in the position of the hand (e.g., to display the plurality of user interface objects within a predefined distance of the hand during the movement of the hand (e.g., the menu of objects is pasted to the tip of the thumb)). In some embodiments, the visual indication is a system affordance representation (e.g., an indicator for an application launch user interface or task bar). In some embodiments, the visual indication is a task bar that includes a plurality of application launch icons. In some embodiments, the task bar changes as the hand configuration changes (e.g., the position of the thumb, the position of the index finger / middle finger). In some embodiments, the visual indication disappears when the hand moves out of the micro-gesture orientation (e.g., when the thumb is raised with the hand below the shoulder). In some embodiments, the visual indication reappears when the hand enters the micro-gesture orientation. In some embodiments, the visual indication appears in response to a gesture (e.g., an upward swipe of the thumb on the index finger when the user is looking at the hand). In some embodiments, the visual indication is reset (e.g., disappears) after a time threshold of inactivity of the hand (e.g., 8 seconds). For example, more details are described with respect to FIG. 7D to FIG. 7F , Fig. 9 and the accompanying description.

[0220] In some embodiments, a computer system uses one or more cameras to detect movement of a user's thumb of a second hand (e.g., different from the first hand) above the user's index finger (e.g., when movement of the user's thumb of the first hand above the user's index finger is detected (e.g., in a two-handed gesture scenario); or when movement of the user's thumb of the first hand above the user's index finger is not detected (e.g., in a one-handed gesture scenario)). In response to detecting movement of the user's thumb of the second hand above the user's index finger using the one or more cameras: Based on determining that the movement is a swipe of the thumb of the second hand above the index finger in a first direction (e.g., along the length of the index finger, or around the index finger, or upward away from a side of the index finger), the computer system performs an eighth operation different from a first operation; and based on determining that the movement is a tap of the thumb of the second hand above the index finger at a first position on the index finger of the second hand (e.g., at a first portion of the index finger such as the distal phalanx, middle phalanx, and / or proximal phalanx), the computer system performs a ninth operation different from a second operation (and the eighth operation). In some embodiments, the eighth operation and / or the ninth operation are different from the first operation, the second operation, the third operation, the fourth operation, the fifth operation, the sixth operation, and / or the seventh operation. In some embodiments, if two hands are used to perform a two-handed gesture, the movement of the thumbs on both hands is considered simultaneous input and is used together to determine what function to trigger. For example, if the thumbs on both hands move away from the fingertips of the index fingers and toward the bases of the index fingers (and the hands face each other), the device expands the currently selected object, and if the thumbs on both hands move from the bases of the index fingers toward the fingertips of the index fingers (and the hands face each other), the device minimizes the currently selected object. In some embodiments, if the thumbs tap down on the index fingers simultaneously on both hands, the device activates the currently selected object in a first manner (e.g., starts video recording using the camera application), and if the thumb taps down on the index finger on the left hand, the device activates the currently selected object in a second manner (e.g., performs autofocus using the camera application), and if the thumb taps down on the index finger on the right hand, the device activates the currently selected object in a third manner (e.g., takes a snapshot using the camera application).

[0221] In some embodiments, in response to detecting, using one or more cameras, movement of a user's thumb above the user's index finger (e.g., rather than a more exaggerated gesture created by waving a finger or hand in the air or sliding on a touch-sensitive surface), and based on determining that the movement includes a press of the thumb of a first hand on the index finger followed by a flick gesture of the wrist of the first hand (e.g., an upward movement of the first hand relative to the wrist of the first hand), the computer system performs a tenth operation that is different from a first operation (e.g., different from each operation or subset of the first through ninth operations corresponding to other types of movement patterns of the user's finger). For example, the tenth operation provides an input for operating a selected user interface object, provides an input for selecting an object (e.g., a virtual object selected and / or held by the user), and / or provides an input for discarding an object. In some embodiments, when the device detects that the user's gaze is directed at an optional object in a three-dimensional environment (e.g., a photo file icon, a movie file icon, a notification banner, etc.), and the device detects a press of the user's thumb on the index finger followed by an upward wrist flick gesture, the device initiates an experience corresponding to the object (e.g., opens the photo in the air, starts playing a 3D movie, opens an expanded notification, etc.).

[0222] In some embodiments, in response to detecting, using one or more cameras, movement of a user's thumb above the user's index finger (e.g., rather than a more exaggerated gesture created by waving a finger or hand in the air or sliding on a touch-sensitive surface), and based on determining that the movement includes a press of the thumb of a first hand on the index finger followed by a hand rotation gesture of the first hand (e.g., a rotation of at least a portion of the first hand relative to the wrist of the first hand), the computer system performs an eleventh operation that is different from a first operation (e.g., different from each operation or subset of the first through tenth operations corresponding to other types of movement patterns of the user's finger). For example, the eleventh operation adjusts a value by an amount corresponding to the amount of rotation of the hand. For example, the eleventh operation causes a virtual object (e.g., selected and / or held by the user (e.g., using gaze)) or a user interface object (e.g., a virtual dial control) to rotate based on the hand rotation gesture.

[0223] In some embodiments, when displaying a view of a three-dimensional environment, the computer system detects a movement of the palm of the user's first hand towards the user's face. Based on determining that the movement of the palm of the user's first hand towards the user's face meets a call criterion, the computer system performs a twelfth operation that is different from a first operation (e.g., different from each operation or subset of the first through eleventh operations corresponding to other types of movement patterns of the user's fingers), such as displaying a user interface object associated with a virtual assistant and / or displaying an image at a position corresponding to the palm of the first hand (e.g., a virtual representation of the user, a camera view of the user, a magnified view of the three-dimensional environment, and / or a magnified view of an object (e.g., a virtual object and / or a real object in the three-dimensional environment)). In some embodiments, the call criterion includes a criterion that is met based on determining that the distance between the user's palm and the user's face has decreased to below a threshold distance. In some embodiments, the call criterion includes a criterion that is met based on determining that the fingers of the hand are extended.

[0224] It should be understood that the specific order in which the operations are described Figure 8 is merely exemplary and is not intended to indicate that this is the only order in which the operations can be performed. Those of ordinary skill in the art will envision various ways to reorder the operations described herein. Additionally, it should be noted that the details of the other processes described herein with respect to other methods (e.g., methods 9000, 10000, 11000, 12000, and 13000) similarly apply in a like manner to the method 8000 described above with respect to Figure 8 For example, the gestures, gaze inputs, physical objects, user interface objects, and / or animations described above with respect to method 8000 optionally have one or more of the features of the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to other methods (e.g., methods 9000, 10000, 11000, 12000, and 13000). For the sake of brevity, these details are not repeated here.

[0225] Fig. 9 is a flowchart of an exemplary method 9000 for interacting with a three-dimensional environment using predefined input gestures according to some embodiments. In some embodiments, method 9000 is performed at a computer system (e.g., Figure 1 the computer system 101 in Figure 1 , Figure 3 and Figure 4a display generation component 120) (e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera pointing downward at the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 9000 is managed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of a computer system, such as one or more processors 202 of computer system 101 (e.g., Figure 1 the control unit 110) in

[0226] In method 9000, the computer system displays (9002) a view of a three-dimensional environment (e.g., a virtual environment or an augmented reality environment). When displaying the three-dimensional environment, the computer system detects a hand at a first position corresponding to a portion of the three-dimensional environment (e.g., detecting the hand at a position in the physical environment that is visible to the user according to the user's current field of view of the three-dimensional environment (e.g., a position where the user's hand has moved to intersect or be near the user's line of sight)). In some embodiments, in response to detecting the hand at the first position in the physical environment, a representation or image of the user's hand is displayed in that portion of the three-dimensional environment. In response to detecting the hand at the first position (9004) corresponding to the portion of the three-dimensional environment: According to determining that the hand remains in a first predefined configuration (e.g., a predefined ready state, such as detecting that the thumb is resting on the index finger using a camera or a touch-sensitive glove or a touch-sensitive finger attachment), a visual indication of a first operation context of a gesture input using the gesture is displayed in the three-dimensional environment (e.g., adjacent to the representation of the hand displayed in that portion of the three-dimensional environment) (e.g., a visual indication, such as a system affordance representation (e.g., Fig. 7E , Figure 7F and Figure 7G the system affordance representation 7214 in Fig.7D ), a task bar, a menu, an avatar of a voice-based virtual assistant; additional information about user interface elements that can be manipulated in response to gesture input, etc.) is displayed in the three-dimensional environment; and according to determining that the hand does not remain in the first predefined configuration, the visual indication of the first operation context of the gesture input using the gesture is abandoned in the three-dimensional environment (e.g., displaying a representation of the hand at that portion of the three-dimensional environment without displaying a visual indication adjacent to the representation of the hand, as shown in

[0227] In some embodiments, a visual indication of a first operational context of gesture input using a gesture is displayed at a location in the portion of the three-dimensional environment corresponding to the first location (e.g., the detected hand location). For example, the visual indication (e.g., a main affordance or a task bar, etc.) is displayed at the detected hand location and / or at a location within a predefined distance of the detected hand location. In some embodiments, the visual indication is displayed at a location corresponding to a particular portion of the hand (e.g., above the detected upper portion of the hand, below the detected lower portion of the hand, and / or covering the hand).

[0228] In some embodiments, when the visual indication is displayed in the portion of the three-dimensional environment, the computer system detects a change in the position of the hand from the first position to the second position (e.g., detects the movement and / or rotation of the hand in the three-dimensional environment when the hand is in a first predefined configuration or some other predefined configuration that also indicates the readiness of the hand). In response to detecting the change in the position of the hand from the first position to the second position, the computer system changes the display position of the visual indication according to the detected change in the position of the hand (e.g., to maintain the visual indication being displayed within a predefined distance of the hand in the three-dimensional environment).

[0229] In some embodiments, the visual indication includes one or more user interface objects. In some embodiments, the visual indicator is a system affordance icon (e.g., Fig. 7E the system affordance 7120 in Figure 7F ), which indicates an area from which one or more user interface objects can be displayed and / or accessed. For example, as shown in part (A) of

[0230] In some embodiments, the one or more user interface objects include a plurality of application launch icons (e.g., the one or more user interface objects is a task bar including a row of application launch icons for a plurality of frequently used applications or experiences), wherein activation of a corresponding application launch icon among the application launch icons causes an operation associated with the corresponding application to be performed (e.g., causes the corresponding application to be launched).

[0231] In some embodiments, when a visual indication is displayed, the computer system detects a change in the configuration of the hand from a first predefined configuration to a second predefined configuration (e.g., detects a change in the position of the thumb (e.g., relative to another finger, such as a movement across another finger)). In response to detecting the detected change in the configuration of the hand from the first predefined configuration to the second predefined configuration, the computer system (e.g., in addition to or instead of displaying the visual indication) displays a first set of user interface objects (e.g., a main area or an application launch user interface), wherein activation of a corresponding user interface object in the first set of user interface objects causes an operation associated with the corresponding user interface object to be performed. In some embodiments, the visual indicator is a system affordance representation icon (e.g., Fig. 7E and Figure 7F the system affordance representation 7214 in), the system affordance representation icon indicating an area from which the main area or the application launch user interface can be displayed and / or accessed. For example, as shown in part (A) of Figure 7F , when the thumb of the hand 7200 moves across the index finger in the direction 7120, the display of the visual indicator 7214 is replaced by the display of a set of user interface objects 7170. In some embodiments, at least some of the user interface objects in the first set of user interface objects include application launch icons, wherein activation of an application launch icon causes the corresponding application to be launched.

[0232] In some embodiments, when a visual indication is displayed, the computer system determines whether the movement of the hand during a time window (e.g., a time window of 5 seconds, 8 seconds, 15 seconds, etc. starting from the time when the visual indication is displayed in response to detecting a hand in a ready state at a first position) meets an interaction criterion (e.g., an interaction criterion met according to determining that at least one finger and / or thumb of the hand moves a distance greater than a threshold distance and / or according to a predefined gesture movement). Based on determining that the movement of the hand during the time window does not meet the interaction criterion, the computer system stops displaying the visual indication. In some embodiments, when the hand is detected again in the user's field of view in a first predefined configuration after the user's hand exits the user's field of view or the user's hand changes to another configuration that is not the first predefined configuration or other predefined configuration corresponding to the ready state of the hand, the device redisplay the visual indication.

[0233] In some embodiments, when displaying a visual indication, the computer system detects a change in the hand configuration from a first predefined configuration to a second predefined configuration that meets the input criteria (e.g., the configuration of the hand has changed, but the hand remains within the user's field of view). For example, the detected change is a change in the position of the thumb (e.g., relative to another finger, such as contacting and / or releasing contact with another finger; moving along the length of another finger and / or moving across another finger) and / or a change in the position of the index finger and / or middle finger of the hand (e.g., extension of the finger and / or other movement of the finger relative to the rest of the hand). In response to detecting a change in the hand configuration from the first predefined configuration to the second predefined configuration that meets the input criteria (e.g., based on determining that the user's hand has changed from the configuration that is the starting state of a first acceptance gesture to the starting state of a second acceptance gesture), the computer system adjusts the visual indication (e.g., adjusts the selected corresponding user interface object in a set of one or more user interface objects from a first corresponding user interface object to a second corresponding user interface object; changes the display position of the one or more user interface objects; and / or displays and / or stops displaying the corresponding user interface object among the one or more user interface objects).

[0234] In some embodiments, when displaying a visual indication, the computer system detects a change in the hand configuration from a first predefined configuration to a third configuration that does not meet the input criteria (e.g., based on determining that at least a portion of the hand is outside the user's field of view, the configuration does not meet the input criteria). In some embodiments, the device determines that the third configuration does not meet the input criteria based on determining that the user's hand has changed from the configuration that is the starting state of a first acceptance gesture to a state that does not correspond to the starting state of any acceptance gesture. In response to detecting the detected change in the configuration of the hand from the first predefined configuration to the third configuration that does not meet the input criteria, the computer system stops displaying the visual indication.

[0235] In some embodiments, after stopping the display of the visual indication, the computer system detects a change in the hand configuration to the first predefined configuration (and detects that the hand is within the user's field of view). In response to detecting the detected change in the hand configuration to the first predefined configuration, the computer system redisplay the visual indication.

[0236] In some embodiments, in response to detecting a hand at a first position corresponding to the portion of the three-dimensional environment, based on determining that the hand is not maintained in the first predefined configuration, the computer system performs an operation different from displaying a visual indication of a first operational context of gesture input using a gesture (e.g., displays a representation of the hand without a visual indication and / or provides a cue for indicating that the hand is not maintained in the first predefined configuration).

[0237] It should be understood that Fig. 9 the particular order in which the operations in Fig. 9 are described is merely exemplary and is not intended to indicate that the order is the only order in which the operations can be performed. Those of ordinary skill in the art will think of various ways to reorder the operations described herein. Additionally, it should be noted that the details of the other processes described herein with respect to other methods described herein (e.g., methods 8000, 10000, 11000, 12000, and 13000) apply to the method 9000 described above in a similar manner. For example, the gestures, gaze inputs, physical objects, user interface objects, and / or animations referred to above with respect to method 9000 optionally have one or more of the features of the gestures, gaze inputs, physical objects, user interface objects, and / or animations described herein with respect to other methods described herein (e.g., methods 8000, 10000, 11000, 12000, and 13000). For the sake of brevity, these details are not repeated here.

[0238] Fig.10 is a flowchart of an exemplary method 10000 for interacting with a three-dimensional environment using predefined input gestures. In some embodiments, method 10000 is executed at a computer system (e.g., Figure 1 the computer system 101 in Figure 1 , Figure 3 and Figure 4 the display generation component 120 in Figure 1 e.g., a head-up display, a display, a touch screen, a projector, etc.) and one or more cameras (e.g., a camera pointing down at the user's hand (e.g., a color sensor, an infrared sensor, and other depth-sensing cameras) or a camera pointing forward from the user's head). In some embodiments, method 10000 is managed by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of the computer system, such as one or more processors 202 of the computer system 101 (e.g., Figure 1 the control unit 110 in

[0239] In method 10000, a computer system displays (10002) a three-dimensional environment (e.g., an augmented reality environment), including displaying a representation of the physical environment (e.g., displaying a camera view of the physical environment around the user, or including a see-through portion in a displayed user interface or virtual environment that shows the physical environment around the user). When displaying the representation of the physical environment, the computer system (e.g., using a camera or one or more motion sensors) detects (10004) a gesture (e.g., a gesture involving a predefined movement of the user's hand, finger, wrist, or arm, or a predefined static pose of the hand that is different from the natural resting pose of the hand). In response to detecting the gesture (10006): Based on determining that the user's gaze points to a position corresponding to a predefined physical position (e.g., the user's hand) in the physical environment (e.g., in the three-dimensional environment) (e.g., based on determining that during the time when the gesture is initiated and completed, the gaze points to and remains at that position; or based on determining that when the hand is in the final state of the gesture (e.g., the ready state of the hand (e.g., the predefined static pose of the hand)), the gaze points to the hand), the computer system displays a system user interface in the three-dimensional environment (e.g., a user interface including visual indications and / or selectable options for interaction options available in the three-dimensional environment, the user interface being displayed in response to the gesture and not being displayed before the gesture is detected (e.g., when the gaze points)). For example, this is shown in Figure 7G Sections A-1, A-2, and A-3 of, where input gestures made by hand 7200 cause interactions with system user interface elements such as system affordance representations 7214, system menu 7170, and application icons 7190. In some embodiments, the position corresponding to the predefined physical position is a representation (e.g., a video image or a graphical abstraction) of the predefined physical position (e.g., the user's hand that can move within the physical environment, or a physical object that is stationary in the physical environment) in the three-dimensional environment. In some embodiments, the system user interface includes one or more application icons (e.g., the corresponding application icon in the one or more application icons launches the corresponding application when activated). In response to detecting the gesture (10006): Based on determining that the user's gaze does not point to a position corresponding to a predefined physical position in the physical environment (e.g., in the three-dimensional environment) (e.g., based on determining that during the time when the gesture is initiated and completed, the gaze points to and / or remains at another position or does not point to that position; or based on determining that when the hand is in the first state of the gesture (e.g., the ready state of the hand (e.g., the predefined static pose of the hand)), the gaze does not point to the hand), an operation is performed in the current context of the three-dimensional environment without displaying the system user interface. For example, this is shown in Figure 7Gare shown in portions B-1, B-2, and B-3. In some embodiments, the operation includes a first operation that changes the state of the electronic device (e.g., changes the output volume of the device) and that does not cause a visual change in the three-dimensional environment. In some embodiments, the operation includes a second operation that displays a hand gesture and that does not cause further interaction with the three-dimensional environment. In some embodiments, the operation includes an operation of changing the state of the virtual object that the gaze is currently pointing to. In some embodiments, the operation includes an operation of changing the state of the virtual object that the user last interacted with in the three-dimensional environment. In some embodiments, the operation includes an operation of changing the state of the currently selected virtual object that has an input focus.

[0240] In some embodiments, the computer system displays a system affordance representation at a predefined location relative to a location corresponding to a predefined physical location (e.g., a primary affordance representation indicating that the device is ready to detect one or more system gestures for displaying a user interface for system-level (rather than application-level) operations). In some embodiments, the location corresponding to the predefined physical location is a location in the three-dimensional environment. In some embodiments, the location corresponding to the predefined physical location is a location on the display. In some embodiments, as long as the predefined location of the system affordance representation is a location in the displayed three-dimensional environment, the system affordance representation is maintained on display even if the location corresponding to the predefined physical location is no longer visible in the displayed portion of the three-dimensional environment (e.g., the system affordance representation continues to be displayed even if the predefined physical location moves out of the field of view of one or more cameras of the electronic device). In some embodiments, the system affordance representation is displayed at a predefined fixed location relative to the user's hand, wrist, or finger or relative to the representation of the user's hand, wrist, or finger in the three-dimensional environment (e.g., superimposed on or replacing a portion of the display of the user's hand, wrist, or finger, or at a fixed location offset from the user's hand, wrist, or finger). In some embodiments, the system affordance representation is displayed at a predefined location relative to the location corresponding to the predefined physical location regardless of whether the user's gaze remains pointed at a location in the three-dimensional environment (e.g., the system affordance representation remains on display for a predefined timeout period even after the user's gaze has moved away from the user's hand in the ready state or after a gesture has been completed).

[0241] In some embodiments, displaying a system affordance representation at a predefined location relative to a location corresponding to a predefined physical location includes: detecting movement of the location corresponding to the predefined physical location in a three-dimensional environment (e.g., detecting that the location of a user's hand shown in the three-dimensional environment has changed as the user's head or hand has moved); and in response to detecting the movement of the location corresponding to the predefined physical location in the three-dimensional environment, moving the system affordance representation in the three-dimensional environment such that the relative position of the system affordance representation relative to the location corresponding to the predefined physical location remains unchanged in the three-dimensional environment (e.g., when the location of the user's hand changes in the three-dimensional environment, the system affordance representation follows the location of the user's hand (e.g., the system affordance representation is shown in the display view of the three-dimensional environment at a location corresponding to the top of the user's thumb)).

[0242] In some embodiments, a system affordance representation is displayed at a predefined location (e.g., sometimes referred to as a "predefined relative location") relative to a location corresponding to a predefined physical location based on determining that a user's gaze is directed at the location corresponding to the predefined physical location. In some embodiments, a system affordance representation is displayed at the predefined relative location based on determining that the user's gaze is directed at a location near the predefined physical location (e.g., within a predefined threshold distance of the predefined physical location). In some embodiments, the system affordance representation is not displayed when the user's gaze is not directed at the predefined physical location (e.g., when the user's gaze has moved away from the predefined physical location or is at least a predefined distance away from the predefined physical location). In some embodiments, when a system affordance representation is displayed at a predefined location relative to a location corresponding to a predefined physical location in a three-dimensional environment, the device detects that the user's gaze has moved away from the location corresponding to the predefined physical location, and in response to detecting that the user's gaze has moved away from the location corresponding to the predefined physical location in the three-dimensional environment, the device stops displaying the system affordance representation at the predefined location in the three-dimensional environment.

[0243] In some embodiments, displaying a system affordance representation at a predefined location relative to a location corresponding to a predefined physical location in a three-dimensional environment includes: displaying the system affordance representation in a first appearance (e.g., shape, size, color, etc.) based on determining that the user's gaze does not point to the location corresponding to the predefined physical location; and displaying the system affordance representation in a second appearance different from the first appearance based on determining that the user's gaze points to the location corresponding to the predefined physical location. In some embodiments, when the user's gaze leaves the location corresponding to the predefined physical location, the system affordance representation has the first appearance. In some embodiments, when the user's gaze points to the location corresponding to the predefined physical location, the system affordance representation has the second appearance. In some embodiments, when the user's gaze shifts to the location corresponding to the predefined physical location (e.g., within a threshold distance of the location), the system affordance representation changes from the first appearance to the second appearance, and when the user's gaze shifts away from the location corresponding to the predefined physical location (e.g., at least a threshold distance away from the location), the system affordance representation changes from the second appearance to the first appearance.

[0244] In some embodiments, displaying a system affordance representation at a predefined location relative to a location corresponding to a predefined physical location includes: based on determining that the user is ready to perform a gesture, displaying the system affordance representation at the predefined location. In some embodiments, determining that the user is ready to perform a gesture includes detecting an indication that the user is ready to perform a gesture, such as by detecting that a predefined physical location (e.g., the user's hand, wrist, or finger) is in (or has entered) a predefined configuration (e.g., a predefined pose relative to a device in the physical environment). In one example, when in addition to detecting that the gaze is on a hand in a ready state, the device also detects that the user has placed their hand in a predefined ready state (e.g., a particular position and / or orientation of the hand) in the physical environment, the system affordance representation is displayed at a predefined location relative to the display representation of the user's hand in the three-dimensional environment. In some embodiments, the predefined configuration requires the predefined physical location (e.g., the user's hand) to have a particular position relative to the electronic device or one or more input devices of the electronic device, such as within the field of view of one or more cameras.

[0245] In some embodiments, displaying a system affordance representation at a predefined location relative to a location corresponding to a predefined physical location includes: based on determining that the user is not ready to perform a gesture, displaying the system affordance representation in a first appearance. In some embodiments, determining that the user is not ready to perform a gesture includes detecting an indication that the user is not ready to perform a gesture (e.g., detecting that the user's hand is not in a predefined ready state). In some embodiments, determining that the user is not ready includes failing to detect an indication that the user is ready (e.g., failing or being unable to detect that the user's hand is in a predefined ready state, such as when the user's hand is outside the field of view of one or more cameras of the electronic device). Reference is made herein Fig. 7E and the associated description further describes in detail an indication of the user's readiness to perform a gesture. In some embodiments, displaying a system affordance at a predefined location relative to a location corresponding to a predefined physical location further includes: based on determining that the user is ready to perform a gesture (e.g., based on detecting an indication that the user is ready to perform a gesture, as referred to herein with reference to Fig. 7Eand, as described in the accompanying description, display a system affordance representation in a second appearance that is different from the first appearance. One of ordinary skill in the art will recognize that the presence or absence of a system affordance representation and the particular appearance of the system affordance representation can be modified based on what information is to be conveyed to the user in a particular context (e.g., what operation is to be performed in response to a gesture, and / or whether additional criteria need to be met for the gesture to invoke the system user interface). In some embodiments, when a system affordance representation is displayed at a predefined location in a three-dimensional environment relative to a location corresponding to a predefined physical position, the device detects a change of the user's hand from a first state to a second state, and in response to detecting the change from the first state to the second state: based on determining that the first state is a ready state and the second state is not a ready state, the device displays the system affordance representation in a second appearance (changing it from the first appearance to the second appearance); and based on determining that the first state is not a ready state and the second state is a ready state, the device displays the system affordance representation in a first appearance (e.g., changing it from the second appearance to the first appearance). In some embodiments, if the computer system does not detect that the user is gazing at the user's hand and the user's hand is not in a ready state configuration, the computer system does not display the system affordance representation, or optionally, displays the system affordance representation in a first appearance. If a subsequent input gesture is detected (e.g., when the system affordance representation is not displayed or is displayed in a first appearance), the computer system does not perform the system operation corresponding to the input gesture, or optionally, performs an operation in the current user interface context corresponding to the input gesture. In some embodiments, if the computer system does detect that the user is gazing at the user's hand but the hand is not in a ready state configuration, the computer system does not display the system affordance representation, or optionally displays the system affordance representation in a first appearance or a second appearance. If a subsequent gesture input is detected (e.g., when the system affordance representation is not displayed or is displayed in a first appearance or a second appearance), the computer system does not perform the system operation corresponding to the input gesture, or optionally, displays the system user interface (e.g., a task bar or a system menu). In some embodiments, if the computer system does not detect that the user is gazing at the user's hand but the hand is in a ready state configuration, the computer system does not display the system affordance representation, or optionally, displays the system affordance representation in a first appearance or a second appearance. If a subsequent gesture input is detected (e.g., when the system affordance representation is not displayed or is displayed in a first appearance or a second appearance), the computer system does not perform the system operation corresponding to the input gesture, or optionally, performs an operation in the current user interface context. In some embodiments, if the computer system detects that the user is gazing at the user's hand and the hand is in a ready state configuration, the computer system displays the system affordance representation in a second appearance or a third appearance.If a subsequent gesture input is detected (e.g., when the system is enabled to be displayed in a second or third appearance), the computer system performs a system operation (e.g., displays a system user interface). In some embodiments, multiple of the above operations are combined in the same implementation.

[0246] In some embodiments, the predefined physical location is the user's hand, and determining that the user is ready to perform a gesture (e.g., the hand is currently in a predefined ready state, or an initiation gesture has just been detected) includes determining that a predefined portion of the hand (e.g., a specified finger) is in contact with a physical control element. In some embodiments, the physical control element is a controller separate from the user (e.g., a corresponding input device) (e.g., the ready state is the user's thumb in contact with a touch-sensitive strip or ring attached to the user's index finger). In some embodiments, the physical control element is a different part of the user's hand (e.g., the ready state is the thumb in contact with the upper side of the index finger (e.g., near the second knuckle)). In some embodiments, the device uses a camera to detect whether the hand is in a predefined ready state and, optionally, displays the ready hand in a view of the three-dimensional environment. In some embodiments, the device uses a physical control element that is touch-sensitive and communicatively coupled to the electronic device to transmit touch inputs to the electronic device to detect whether the hand is in a predefined ready state.

[0247] In some embodiments, the predefined physical location is the user's hand, and determining that the user is ready to perform a gesture includes determining that the hand is lifted relative to the user to a level above a predefined level. In some embodiments, determining that the hand is lifted includes determining that the hand is positioned relative to the user at a location above a specific lateral plane (e.g., above the user's waist, i.e., closer to the user's head than to the user's feet). In some embodiments, determining that the hand is lifted includes determining that the user's wrist or elbow is bent by at least a certain amount (e.g., within a 90-degree angle). In some embodiments, the device uses a camera to detect whether the hand is in a predefined ready state and, optionally, displays the ready hand in a view of the three-dimensional environment. In some embodiments, the device uses one or more sensors (e.g., motion sensors) attached to the user's hand, wrist, or arm and communicatively coupled to the electronic device to transmit movement inputs to the electronic device to detect whether the hand is in a predefined ready state.

[0248] In some embodiments, the predefined physical location is the user's hand, and determining that the user is ready to perform a gesture includes determining that the hand is in a predefined configuration. In some embodiments, the predefined configuration requires a respective finger of the hand (e.g., the thumb) to contact a different part of the user's hand (e.g., an opposing finger, such as the index finger, or a predefined portion of the opposing finger, such as the middle phalanx or middle knuckle of the index finger). In some embodiments, as described above, the predefined configuration requires the hand to be above a particular transverse plane (e.g., above the user's waist). In some embodiments, the predefined configuration requires the wrist to be bent towards the thumb side and away from the little finger side (e.g., radially flexed) (e.g., without axially rotating the arm). In some embodiments, when the hand is in the predefined configuration, one or more fingers are in a natural rest position (e.g., curled), and the entire hand is tilted or displaced from the natural rest position of the hand, wrist, or arm to indicate the user's readiness to perform a gesture. Those of ordinary skill in the art will recognize that the particular predefined readiness state used can be selected to have an intuitive and natural user interaction, and may require any combination of the foregoing criteria. In some embodiments, when the user only wishes to view the three-dimensional environment rather than provide input to and interact with the three-dimensional environment, the predefined configuration is different from the natural rest pose of the user's hand (e.g., a relaxed and stationary pose on the person's thigh, a tabletop, or a side of the body). The change from the natural rest pose to the predefined configuration is meaningful and requires the user to intentionally move the hand into the predefined configuration.

[0249] In some embodiments, the location corresponding to the predefined physical location is a fixed location within the three-dimensional environment (e.g., the corresponding predefined physical location is a fixed location in the physical environment). In some embodiments, the physical environment is the user's frame of reference. That is, those of ordinary skill in the art will recognize that a location referred to as a "fixed" location in the physical environment may not be an absolute location in space, but is fixed relative to the user's frame of reference. In some examples, if the user is in a room of a building, the location is a fixed location within the three-dimensional environment corresponding to a fixed location in the room (e.g., on a wall, floor, or ceiling of the room). In some examples, if the user is inside a moving vehicle, the location is a fixed location within the three-dimensional environment corresponding to a fixed location along the interior of the vehicle (e.g., a representation of a fixed location along the interior of the vehicle). In some embodiments, the location is fixed relative to the content displayed within the three-dimensional environment, where the displayed content corresponds to the fixed predefined physical location in the physical environment.

[0250] In some embodiments, the position corresponding to the predefined physical location is a fixed position relative to the display of the three-dimensional environment (e.g., relative to the display generation component). In some embodiments, the position is fixed relative to the user's perspective of the three-dimensional environment (e.g., is a position fixed relative to the display of the three-dimensional environment by the display generation component), regardless of the specific content displayed within the three-dimensional environment, which typically updates in response to or in conjunction with a change in the user's perspective. In some examples, the position is a fixed position along the edge of the display of the three-dimensional environment (e.g., within a predefined distance of the edge). In some examples, the position is centered relative to the display of the three-dimensional environment (e.g., centered within a display area along the bottom edge, top edge, left edge, or right edge of the display of the three-dimensional environment).

[0251] In some embodiments, the predefined physical location is a fixed position on the user. In some examples, the predefined physical location is the user's hand or finger. In some such examples, the position corresponding to the predefined physical location includes the display representation of the user's hand or finger within the three-dimensional environment.

[0252] In some embodiments, after displaying a system user interface in a three-dimensional environment, the computer system detects a second gesture (e.g., a second gesture performed by the user's hand, wrist, finger, or arm) (e.g., when the system user interface is displayed after detecting a first gesture and a gaze directed to a location corresponding to a predefined physical location). In response to detecting the second gesture, the system user interface (e.g., an application launch user interface) is displayed. In some embodiments, the second gesture is a continuation of the first gesture. For example, the first gesture is a swipe gesture (e.g., by moving the user's thumb above the user's index finger on the same hand), and the second gesture is a continuation of the swipe gesture (e.g., continued movement of the thumb above the index finger) (e.g., the second gesture starts from the end position of the first gesture without resetting the start position of the second gesture to the start position of the first gesture). In some embodiments, the second gesture is a repetition of the first gesture (e.g., after performing the first gesture, the start position of the second gesture is reset to within a predefined distance of the start position of the first gesture, and the second gesture moves along the movement of the first gesture within a predefined tolerance). In some embodiments, displaying the main user interface includes expanding a system affordance representation from a predefined position relative to the position corresponding to the predefined physical location to occupy a larger portion of the displayed three-dimensional environment and showing additional user interface objects and options. In some embodiments, the system affordance representation is an indicator without corresponding content, and the corresponding content (e.g., a task bar with a row of application icons for recently used or frequently used applications) replaces the indicator in response to a first swipe gesture made by the hand, a two-dimensional grid of application icons for all installed applications replaces the task bar in response to a second swipe gesture made by the hand; and a three-dimensional working environment in which interactive application icons float at different depths and positions in the three-dimensional working environment replaces the two-dimensional grid in response to a third swipe gesture made by the hand.

[0253] In some embodiments, the current context of the three-dimensional environment includes an indication of a received notification being displayed (e.g., initially displaying a subset of information about the received notification), and performing an operation in the current context of the three-dimensional environment includes displaying an expanded notification that includes additional information about the received notification (e.g., displaying information in addition to the initially displayed subset). In some embodiments, the current context of the three-dimensional environment is determined based on the location currently pointed to by the gaze. In some embodiments, when a notification is received and indicated in the three-dimensional environment and the user's gaze is detected towards the notification (and not at a location corresponding to a predefined physical location (e.g., the user's hand)), the device determines that the current context is an interaction with the notification, and in response to detecting a gesture by the user (e.g., an upward flick gesture made by the thumb or wrist), the expanded notification content is displayed in the three-dimensional environment.

[0254] In some embodiments, the current context of the three-dimensional environment includes displaying an indication of one or more photos (e.g., one or more corresponding thumbnails of the one or more photos), and performing an operation in the current context of the three-dimensional environment includes displaying at least one of the one or more photos in the three-dimensional environment (e.g., displaying the photo in an enhanced (e.g., expanded, animated, improved, 3D, etc.) manner). In some embodiments, the current context of the three-dimensional environment is determined based on the location where the gaze is currently directed. In some embodiments, when an image is displayed in the three-dimensional environment and the user's gaze is detected to be toward the image (and not at a location corresponding to a predefined physical location (e.g., the user's hand)), the device determines that the current context is an interaction with the image, and displays the image in the three-dimensional environment in an enhanced manner in response to detecting a gesture of the user (e.g., an upward flick gesture by a thumb or wrist).

[0255] It should be understood that Fig.10 The specific order in which the operations in the description are described is merely exemplary and is not intended to indicate that the order is the only order in which the operations may be performed. A person of ordinary skill in the art will recognize many ways to reorder the operations described herein. In addition, it...

Claims

1. A method for interacting with a three-dimensional environment, comprising: At a computing system including a display generation component and one or more input devices: Display a view of the three-dimensional environment; When displaying the view of the three-dimensional environment, detect a hand at a first position corresponding to a portion of the three-dimensional environment; In response to detecting the hand at the first position corresponding to the portion of the three-dimensional environment: Based on determining that the hand is held in a first predefined configuration at the first position, display a visual indication at a first display position indicating that a corresponding user interface object can be displayed in response to detecting gesture input using a gesture in the three-dimensional environment starting from the first predefined configuration; And Based on determining that the hand is not held in the first predefined configuration at the first position, abandon displaying the visual indication; When displaying the visual indication, detect movement of at least a portion of the hand; And In response to detecting the movement of at least a portion of the hand: Based on determining that the movement of at least a portion of the hand causes a change in the position of the hand as a whole, and as a result of the movement of at least a portion of the hand, the hand is held in the first predefined configuration at a second position different from the first position, Display the visual indication at a second display position different from the first display position; And Based on determining that the movement of at least a portion of the hand causes the configuration of the hand to change from the first predefined configuration to a second predefined configuration different from the first predefined configuration as a result of the movement of at least a portion of the hand at the first position, Display the corresponding user interface object.

2. The method according to claim 1, wherein the first display position is a position in the portion of the three-dimensional environment.

3. The method according to any one of claims 1 to 2, wherein the visual indication includes one or more user interface objects.

4. The method according to claim 3, wherein the one or more user interface objects include a plurality of application launch icons, and activation of a corresponding one of the plurality of application launch icons causes an operation associated with the corresponding application to be executed.

5. The method according to claim 1, comprising: When displaying the visual indication, determine whether the movement of the hand satisfies an interaction criterion during a time window; And Based on determining that the movement of the hand does not satisfy the interaction criterion during the time window, stop displaying the visual indication.

6. The method according to claim 1, comprising: When displaying the visual indication, detect a change in the configuration of the hand from the first predefined configuration to a third configuration that does not satisfy an input criterion; And In response to detecting the detected change in the configuration of the hand from the first predefined configuration to the third configuration that does not satisfy the input criterion, stop displaying the visual indication.

7. The method according to claim 6, comprising: After stopping displaying the visual indication, detect a change in the configuration of the hand to the first predefined configuration; And In response to detecting the change in the configuration of the detected hand to the first predefined configuration, redisplay the visual indication.

8. A computer-readable storage medium storing one or more programs including executable instructions that, when executed by a computer system having one or more processors and a display generation component, cause the computer system to perform the following operations: Display a view of a three-dimensional environment; When displaying the view of the three-dimensional environment, detect a hand at a first position corresponding to a portion of the three-dimensional environment; In response to detecting the hand at the first position corresponding to the portion of the three-dimensional environment: Based on determining that the hand is held in a first predefined configuration at the first position, display a visual indication at a first display position indicating that a corresponding user interface object can be displayed in response to detecting gesture input using a gesture in the three-dimensional environment starting from the first predefined configuration; And Based on determining that the hand is not held in the first predefined configuration at the first position, abandon displaying the visual indication; When displaying the visual indication, detect movement of at least a portion of the hand; And In response to detecting the movement of at least a portion of the hand: Based on determining that the movement of at least a portion of the hand causes a change in the position of the hand as a whole, and as a result of the movement of at least a portion of the hand, the hand is held in the first predefined configuration at a second position different from the first position, Display the visual indication at a second display position different from the first display position; And Based on determining that the movement of at least a portion of the hand causes the configuration of the hand to change from the first predefined configuration to a second predefined configuration different from the first predefined configuration as a result of the movement of at least a portion of the hand at the first position, Display the corresponding user interface object.

9. The computer-readable storage medium according to claim 8, wherein the one or more programs include instructions that, when executed by the computer system, cause the computer system to perform the method according to any one of claims 2 to 7.

10. A computer system comprising: One or more processors; A display generation component; And A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: Display a view of a three-dimensional environment; When displaying the view of the three-dimensional environment, detect a hand at a first position corresponding to a portion of the three-dimensional environment; In response to detecting the hand at the first position corresponding to the portion of the three-dimensional environment: Based on determining that the hand is held in a first predefined configuration at the first position, display a visual indication at a first display position indicating that a corresponding user interface object can be invoked in response to detecting gesture input using a gesture in the three-dimensional environment starting from the first predefined configuration; and Based on determining that the hand is not held in the first predefined configuration at the first position, abandon displaying the visual indication; When the visual indication is being displayed, detect movement of at least a portion of the hand; and In response to detecting the movement of at least a portion of the hand: Based on determining that the movement of at least a portion of the hand causes a change in the position of the hand as a whole, and as a result of the movement of at least a portion of the hand, the hand is held in the first predefined configuration at a second position different from the first position, display the visual indication at a second display position different from the first display position; and Based on determining that the movement of at least a portion of the hand causes the configuration of the hand at the first position to change from the first predefined configuration to a second predefined configuration different from the first predefined configuration as a result of the movement of at least a portion of the hand, display the corresponding user interface object.

11. The computer system according to claim 10, wherein the one or more programs include instructions for performing the method according to any one of claims 2 - 7.

12. A computer program product, the computer program product including one or more programs, the one or more programs including executable instructions that, when executed by a computer system having one or more processors and a display generation component, cause the computer system to perform the method according to any one of claims 1 - 7.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    CN105074625A

  • Methods and systems for creating virtual and augmented reality

    CN106937531A